humanish 0.86.1 → 0.88.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/actor-contract.d.ts +7 -0
- package/dist/actor-contract.js.map +1 -1
- package/dist/actor-stop-cause.d.ts +9 -0
- package/dist/actor-stop-cause.js +41 -0
- package/dist/actor-stop-cause.js.map +1 -0
- package/dist/computer-use.d.ts +4 -1
- package/dist/computer-use.js +45 -12
- package/dist/computer-use.js.map +1 -1
- package/dist/cua-actor-lab.d.ts +10 -3
- package/dist/cua-actor-lab.js +29 -4
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/cua-admission-limit.d.ts +12 -0
- package/dist/cua-admission-limit.js +20 -0
- package/dist/cua-admission-limit.js.map +1 -0
- package/dist/cua-diagnostics.d.ts +38 -0
- package/dist/cua-diagnostics.js +68 -0
- package/dist/cua-diagnostics.js.map +1 -0
- package/dist/feedback.js +11 -1
- package/dist/feedback.js.map +1 -1
- package/dist/index.d.ts +2 -1
- package/dist/index.js +1 -0
- package/dist/index.js.map +1 -1
- package/dist/observer-app.html +5 -5
- package/dist/observer-data.d.ts +5 -0
- package/dist/observer-data.js +27 -4
- package/dist/observer-data.js.map +1 -1
- package/dist/observer.js +3 -3
- package/dist/observer.js.map +1 -1
- package/dist/openai-responses-cu.js +14 -3
- package/dist/openai-responses-cu.js.map +1 -1
- package/dist/program.d.ts +2 -0
- package/dist/program.js +9 -4
- package/dist/program.js.map +1 -1
- package/dist/run.d.ts +6 -3
- package/dist/run.js +22 -7
- package/dist/run.js.map +1 -1
- package/dist/telemetry.d.ts +3 -0
- package/dist/telemetry.js +15 -1
- package/dist/telemetry.js.map +1 -1
- package/docs/contracts/adapter-admission.md +54 -0
- package/docs/contracts/run-bundle.md +24 -0
- package/docs/contracts/schemas.md +1 -1
- package/docs/goals/current.md +151 -719
- package/docs/ramp/README.md +11 -3
- package/docs/release/0.87.0-participant-endings-and-phone-review.md +64 -0
- package/docs/release/0.88.0-study-diagnostics.md +43 -0
- package/package.json +1 -1
package/docs/ramp/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Status: public-safe contributor and agent ramp.
|
|
4
4
|
|
|
5
|
-
Package/source version in this tree: `0.
|
|
5
|
+
Package/source version in this tree: `0.88.0` (2026-09-11). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
|
|
6
6
|
`0.15.1` remains in force: managed run and output paths bind to validated
|
|
7
7
|
physical filesystem identities, and stored provider IDs are evidence, not
|
|
8
8
|
cleanup authority. The bundled OSS meta-lab is dry-run only until
|
|
@@ -47,6 +47,14 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
|
|
|
47
47
|
|
|
48
48
|
## Current State
|
|
49
49
|
|
|
50
|
+
The [0.88.0 release note](../release/0.88.0-study-diagnostics.md) describes
|
|
51
|
+
computer-use CLI diagnostics, explicit local admission limits and retained
|
|
52
|
+
uncertainty when an earlier provider request did not report usage.
|
|
53
|
+
|
|
54
|
+
The [0.87.0 release note](../release/0.87.0-participant-endings-and-phone-review.md)
|
|
55
|
+
describes recorded interruption causes, phone playback controls beside fitted
|
|
56
|
+
captures, and credential and evidence-sharing guidance by execution route.
|
|
57
|
+
|
|
50
58
|
The [0.86.1 release note](../release/0.86.1-task-preflight-saved-recordings.md)
|
|
51
59
|
describes task preflight on unsupported execution paths and saved HTML recordings
|
|
52
60
|
that remain snapshots when served over HTTP.
|
|
@@ -79,7 +87,7 @@ Implemented:
|
|
|
79
87
|
- first-class lab manifest resolution through `humanish/labs/*.yaml` and
|
|
80
88
|
ignored `.humanish/labs/*.yaml` overlays — `humanish.lab.v2` compositions
|
|
81
89
|
(`src/lab-config.ts`), one engine, no hardcoded lab kinds;
|
|
82
|
-
- a first-party actor registry with
|
|
90
|
+
- a first-party actor registry with seven registered descriptors
|
|
83
91
|
(`src/actor-registry.ts`); `actors[0].type` is a real dispatch key on the
|
|
84
92
|
computer-use, scripted-browser, and terminal-product routes;
|
|
85
93
|
- a computer-use route and clone subject provider: `subject.source: app-url`
|
|
@@ -130,7 +138,7 @@ without changing its success standard.
|
|
|
130
138
|
- capability receipts are not adopter replacement: no first-party deletion
|
|
131
139
|
branch has yet removed a bespoke generic harness while preserving
|
|
132
140
|
decision-equivalent proof;
|
|
133
|
-
- the
|
|
141
|
+
- the seven actor descriptors are a closed first-party union, not a supported
|
|
134
142
|
out-of-tree actor-registration API;
|
|
135
143
|
- run storage and provider-resource lifecycle logic still spans several routes
|
|
136
144
|
instead of one `RunStore` and `ResourceLease` boundary;
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# Humanish 0.87.0: see why a participant stopped
|
|
2
|
+
|
|
3
|
+
The Observer now distinguishes a provider response limit from elapsed time or
|
|
4
|
+
estimated spend. Participant cards, summaries and reports show the recorded
|
|
5
|
+
cause, including for participants with no screenshot frames. New computer-use
|
|
6
|
+
traces retain an optional `stopCause` alongside the existing status and
|
|
7
|
+
completion reason.
|
|
8
|
+
|
|
9
|
+
New computer-use sessions that hit their estimated spend cap before any action
|
|
10
|
+
are classified as interrupted, with no target-app feedback inferred from that
|
|
11
|
+
harness stop. The cap timing and recorded usage are unchanged.
|
|
12
|
+
|
|
13
|
+
Older recordings use exact machine-generated notices where available. A recorded
|
|
14
|
+
“provider token limit reached” notice remains **provider token limit**: it does
|
|
15
|
+
not establish whether the provider exhausted its output or context allowance.
|
|
16
|
+
When an older run records only a generic limit, the Observer says **limit
|
|
17
|
+
reached**. Updating a viewer preserves the original evidence and outcome counts;
|
|
18
|
+
it does not infer a cause from participant prose.
|
|
19
|
+
|
|
20
|
+
On phones, a saved landscape capture in Fit mode keeps playback controls beside
|
|
21
|
+
the fitted image. Portrait captures, zoomed views and live views retain their
|
|
22
|
+
bounded viewing area. The run banner states **Saved recording** once; the player
|
|
23
|
+
keeps the participant status and selected capture timestamp. Exact action links
|
|
24
|
+
and saved-moment recall retain their existing behavior.
|
|
25
|
+
|
|
26
|
+
Expanded assignments show the recorded lane focus before the mission, so a long
|
|
27
|
+
mission does not put those instructions below it in the phone panel. The full
|
|
28
|
+
mission and participant tasks remain available.
|
|
29
|
+
|
|
30
|
+
The website also explains credential placement by execution route. Computer-use
|
|
31
|
+
model keys remain on the host. Terminal actors receive a command-scoped runtime
|
|
32
|
+
key by default; optional `openai-egress` keeps the raw key outside the sandbox
|
|
33
|
+
while still allowing sandbox processes to make paid requests through its proxy.
|
|
34
|
+
Evidence stays local by default, and exposing Observer or sharing an export makes
|
|
35
|
+
it accessible to others. See the [credential and sharing guide](https://humanish.dev/docs/budgets-and-privacy#store-credentials).
|
|
36
|
+
|
|
37
|
+
Install with `npm install -g humanish@0.87.0`. Reopen a run with
|
|
38
|
+
`humanish observe --run RUN`, or regenerate its HTML export, to use the updated
|
|
39
|
+
viewer. Existing HTML files retain the renderer that produced them.
|
|
40
|
+
|
|
41
|
+
## Verification and limits
|
|
42
|
+
|
|
43
|
+
[Interruption-cause checks](https://github.com/danielgwilson/humanish/pull/747)
|
|
44
|
+
cover recorded provider output limits, legacy token-limit notices, elapsed time,
|
|
45
|
+
spend and incomplete provider responses. Tests preserve old outcome counts and
|
|
46
|
+
source evidence through served and exported projections. A retained real
|
|
47
|
+
recording was inspected at desktop and phone widths, including its participant
|
|
48
|
+
whose provider stopped before any app action.
|
|
49
|
+
|
|
50
|
+
[Zero-action spend-cap checks](https://github.com/danielgwilson/humanish/pull/752)
|
|
51
|
+
verify one provider turn, zero executed actions, retained usage, an interrupted
|
|
52
|
+
outcome and no inferred target-app feedback. Older bundle counts are unchanged.
|
|
53
|
+
|
|
54
|
+
[Phone review checks](https://github.com/danielgwilson/humanish/pull/746) cover
|
|
55
|
+
landscape and portrait capture/control spacing, keyboard seeking, focus, zoom,
|
|
56
|
+
action links, refresh and saved-moment recall against retained recordings.
|
|
57
|
+
These browser checks do not establish physical-phone behavior or new live-stream
|
|
58
|
+
coverage. Screenshot replay remains sparse, and no participant success rate or
|
|
59
|
+
incremental persona benefit is inferred from these changes.
|
|
60
|
+
|
|
61
|
+
The [assignment ordering change](https://github.com/danielgwilson/humanish/pull/754)
|
|
62
|
+
preserves recorded text while moving lane instructions to the top of expanded
|
|
63
|
+
assignments. This addresses visibility; it does not establish that every reader
|
|
64
|
+
will interpret the instructions correctly.
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
# Humanish 0.88.0: see what ended a study
|
|
2
|
+
|
|
3
|
+
When a computer-use study stops before the task is done, `humanish lab run` now
|
|
4
|
+
shows its diagnostic category and recorded stop cause. Fan-out results retain
|
|
5
|
+
each participant's ending. Different endings remain mixed. Older generic budget
|
|
6
|
+
stops retain an unspecified limit. Successful dry-runs are reported as contract
|
|
7
|
+
previews.
|
|
8
|
+
|
|
9
|
+
New review packets and feedback participant summaries use recorded interruption
|
|
10
|
+
labels. Previously saved review packets keep their original wording.
|
|
11
|
+
|
|
12
|
+
An adapter can use the new `CuaAdmissionLimitError` to declare that its local
|
|
13
|
+
limit refused a request before dispatch. Humanish records **adapter admission
|
|
14
|
+
limit** and ends the interaction without another retry, action or closing
|
|
15
|
+
report. Earlier actions and measured usage remain available. A refusal of an
|
|
16
|
+
optional closing report preserves the already observed task outcome. See the
|
|
17
|
+
[adapter contract](../contracts/adapter-admission.md) for when this signal applies.
|
|
18
|
+
|
|
19
|
+
If an earlier request stalled or had an ambiguous transport outcome, a later
|
|
20
|
+
success or local refusal does not make its missing usage known. The cost summary
|
|
21
|
+
keeps measured token estimates and an explicit unmeasured line. With a spending
|
|
22
|
+
cap, Humanish skips the optional closing report while that usage is uncertain.
|
|
23
|
+
|
|
24
|
+
Install with `npm install -g humanish@0.88.0`. CLI JSON adds optional diagnostics
|
|
25
|
+
without changing the existing status or exit-code contract. Usage telemetry,
|
|
26
|
+
enabled by default, adds two fields with fixed allowed values: diagnostic
|
|
27
|
+
category and stop cause. See [TELEMETRY.md](https://github.com/danielgwilson/humanish/blob/main/TELEMETRY.md)
|
|
28
|
+
for their scope and opt-out controls.
|
|
29
|
+
|
|
30
|
+
## Verification and limits
|
|
31
|
+
|
|
32
|
+
[Admission-limit checks](https://github.com/danielgwilson/humanish/pull/759)
|
|
33
|
+
cover refusal before and after activity, sanitized error handling, preserved
|
|
34
|
+
outcomes and unknown usage after both successful and refused retries.
|
|
35
|
+
[CLI diagnostic checks](https://github.com/danielgwilson/humanish/pull/758)
|
|
36
|
+
cover previews, mixed endings, missing causes, invalid evidence and telemetry
|
|
37
|
+
allowlists. The required release checks also cover package installation,
|
|
38
|
+
Observer rendering and the website.
|
|
39
|
+
|
|
40
|
+
The admission signal depends on an accurate adapter declaration; it is not a
|
|
41
|
+
provider billing receipt. Existing generic failures keep their recorded
|
|
42
|
+
outcomes. This release does not infer causes from participant prose or establish
|
|
43
|
+
why an older study failed.
|
package/package.json
CHANGED