humanish 0.86.1 → 0.88.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/dist/actor-contract.d.ts +7 -0
  2. package/dist/actor-contract.js.map +1 -1
  3. package/dist/actor-stop-cause.d.ts +9 -0
  4. package/dist/actor-stop-cause.js +41 -0
  5. package/dist/actor-stop-cause.js.map +1 -0
  6. package/dist/computer-use.d.ts +4 -1
  7. package/dist/computer-use.js +45 -12
  8. package/dist/computer-use.js.map +1 -1
  9. package/dist/cua-actor-lab.d.ts +10 -3
  10. package/dist/cua-actor-lab.js +29 -4
  11. package/dist/cua-actor-lab.js.map +1 -1
  12. package/dist/cua-admission-limit.d.ts +12 -0
  13. package/dist/cua-admission-limit.js +20 -0
  14. package/dist/cua-admission-limit.js.map +1 -0
  15. package/dist/cua-diagnostics.d.ts +38 -0
  16. package/dist/cua-diagnostics.js +68 -0
  17. package/dist/cua-diagnostics.js.map +1 -0
  18. package/dist/feedback.js +11 -1
  19. package/dist/feedback.js.map +1 -1
  20. package/dist/index.d.ts +2 -1
  21. package/dist/index.js +1 -0
  22. package/dist/index.js.map +1 -1
  23. package/dist/observer-app.html +5 -5
  24. package/dist/observer-data.d.ts +5 -0
  25. package/dist/observer-data.js +27 -4
  26. package/dist/observer-data.js.map +1 -1
  27. package/dist/observer.js +3 -3
  28. package/dist/observer.js.map +1 -1
  29. package/dist/openai-responses-cu.js +14 -3
  30. package/dist/openai-responses-cu.js.map +1 -1
  31. package/dist/program.d.ts +2 -0
  32. package/dist/program.js +9 -4
  33. package/dist/program.js.map +1 -1
  34. package/dist/run.d.ts +6 -3
  35. package/dist/run.js +22 -7
  36. package/dist/run.js.map +1 -1
  37. package/dist/telemetry.d.ts +3 -0
  38. package/dist/telemetry.js +15 -1
  39. package/dist/telemetry.js.map +1 -1
  40. package/docs/contracts/adapter-admission.md +54 -0
  41. package/docs/contracts/run-bundle.md +24 -0
  42. package/docs/contracts/schemas.md +1 -1
  43. package/docs/goals/current.md +151 -719
  44. package/docs/ramp/README.md +11 -3
  45. package/docs/release/0.87.0-participant-endings-and-phone-review.md +64 -0
  46. package/docs/release/0.88.0-study-diagnostics.md +43 -0
  47. package/package.json +1 -1
@@ -2,7 +2,7 @@
2
2
 
3
3
  Status: public-safe contributor and agent ramp.
4
4
 
5
- Package/source version in this tree: `0.86.1` (2026-09-09). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
5
+ Package/source version in this tree: `0.88.0` (2026-09-11). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
6
6
  `0.15.1` remains in force: managed run and output paths bind to validated
7
7
  physical filesystem identities, and stored provider IDs are evidence, not
8
8
  cleanup authority. The bundled OSS meta-lab is dry-run only until
@@ -47,6 +47,14 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
47
47
 
48
48
  ## Current State
49
49
 
50
+ The [0.88.0 release note](../release/0.88.0-study-diagnostics.md) describes
51
+ computer-use CLI diagnostics, explicit local admission limits and retained
52
+ uncertainty when an earlier provider request did not report usage.
53
+
54
+ The [0.87.0 release note](../release/0.87.0-participant-endings-and-phone-review.md)
55
+ describes recorded interruption causes, phone playback controls beside fitted
56
+ captures, and credential and evidence-sharing guidance by execution route.
57
+
50
58
  The [0.86.1 release note](../release/0.86.1-task-preflight-saved-recordings.md)
51
59
  describes task preflight on unsupported execution paths and saved HTML recordings
52
60
  that remain snapshots when served over HTTP.
@@ -79,7 +87,7 @@ Implemented:
79
87
  - first-class lab manifest resolution through `humanish/labs/*.yaml` and
80
88
  ignored `.humanish/labs/*.yaml` overlays — `humanish.lab.v2` compositions
81
89
  (`src/lab-config.ts`), one engine, no hardcoded lab kinds;
82
- - a first-party actor registry with six registered descriptors
90
+ - a first-party actor registry with seven registered descriptors
83
91
  (`src/actor-registry.ts`); `actors[0].type` is a real dispatch key on the
84
92
  computer-use, scripted-browser, and terminal-product routes;
85
93
  - a computer-use route and clone subject provider: `subject.source: app-url`
@@ -130,7 +138,7 @@ without changing its success standard.
130
138
  - capability receipts are not adopter replacement: no first-party deletion
131
139
  branch has yet removed a bespoke generic harness while preserving
132
140
  decision-equivalent proof;
133
- - the six actor descriptors are a closed first-party union, not a supported
141
+ - the seven actor descriptors are a closed first-party union, not a supported
134
142
  out-of-tree actor-registration API;
135
143
  - run storage and provider-resource lifecycle logic still spans several routes
136
144
  instead of one `RunStore` and `ResourceLease` boundary;
@@ -0,0 +1,64 @@
1
+ # Humanish 0.87.0: see why a participant stopped
2
+
3
+ The Observer now distinguishes a provider response limit from elapsed time or
4
+ estimated spend. Participant cards, summaries and reports show the recorded
5
+ cause, including for participants with no screenshot frames. New computer-use
6
+ traces retain an optional `stopCause` alongside the existing status and
7
+ completion reason.
8
+
9
+ New computer-use sessions that hit their estimated spend cap before any action
10
+ are classified as interrupted, with no target-app feedback inferred from that
11
+ harness stop. The cap timing and recorded usage are unchanged.
12
+
13
+ Older recordings use exact machine-generated notices where available. A recorded
14
+ “provider token limit reached” notice remains **provider token limit**: it does
15
+ not establish whether the provider exhausted its output or context allowance.
16
+ When an older run records only a generic limit, the Observer says **limit
17
+ reached**. Updating a viewer preserves the original evidence and outcome counts;
18
+ it does not infer a cause from participant prose.
19
+
20
+ On phones, a saved landscape capture in Fit mode keeps playback controls beside
21
+ the fitted image. Portrait captures, zoomed views and live views retain their
22
+ bounded viewing area. The run banner states **Saved recording** once; the player
23
+ keeps the participant status and selected capture timestamp. Exact action links
24
+ and saved-moment recall retain their existing behavior.
25
+
26
+ Expanded assignments show the recorded lane focus before the mission, so a long
27
+ mission does not put those instructions below it in the phone panel. The full
28
+ mission and participant tasks remain available.
29
+
30
+ The website also explains credential placement by execution route. Computer-use
31
+ model keys remain on the host. Terminal actors receive a command-scoped runtime
32
+ key by default; optional `openai-egress` keeps the raw key outside the sandbox
33
+ while still allowing sandbox processes to make paid requests through its proxy.
34
+ Evidence stays local by default, and exposing Observer or sharing an export makes
35
+ it accessible to others. See the [credential and sharing guide](https://humanish.dev/docs/budgets-and-privacy#store-credentials).
36
+
37
+ Install with `npm install -g humanish@0.87.0`. Reopen a run with
38
+ `humanish observe --run RUN`, or regenerate its HTML export, to use the updated
39
+ viewer. Existing HTML files retain the renderer that produced them.
40
+
41
+ ## Verification and limits
42
+
43
+ [Interruption-cause checks](https://github.com/danielgwilson/humanish/pull/747)
44
+ cover recorded provider output limits, legacy token-limit notices, elapsed time,
45
+ spend and incomplete provider responses. Tests preserve old outcome counts and
46
+ source evidence through served and exported projections. A retained real
47
+ recording was inspected at desktop and phone widths, including its participant
48
+ whose provider stopped before any app action.
49
+
50
+ [Zero-action spend-cap checks](https://github.com/danielgwilson/humanish/pull/752)
51
+ verify one provider turn, zero executed actions, retained usage, an interrupted
52
+ outcome and no inferred target-app feedback. Older bundle counts are unchanged.
53
+
54
+ [Phone review checks](https://github.com/danielgwilson/humanish/pull/746) cover
55
+ landscape and portrait capture/control spacing, keyboard seeking, focus, zoom,
56
+ action links, refresh and saved-moment recall against retained recordings.
57
+ These browser checks do not establish physical-phone behavior or new live-stream
58
+ coverage. Screenshot replay remains sparse, and no participant success rate or
59
+ incremental persona benefit is inferred from these changes.
60
+
61
+ The [assignment ordering change](https://github.com/danielgwilson/humanish/pull/754)
62
+ preserves recorded text while moving lane instructions to the top of expanded
63
+ assignments. This addresses visibility; it does not establish that every reader
64
+ will interpret the instructions correctly.
@@ -0,0 +1,43 @@
1
+ # Humanish 0.88.0: see what ended a study
2
+
3
+ When a computer-use study stops before the task is done, `humanish lab run` now
4
+ shows its diagnostic category and recorded stop cause. Fan-out results retain
5
+ each participant's ending. Different endings remain mixed. Older generic budget
6
+ stops retain an unspecified limit. Successful dry-runs are reported as contract
7
+ previews.
8
+
9
+ New review packets and feedback participant summaries use recorded interruption
10
+ labels. Previously saved review packets keep their original wording.
11
+
12
+ An adapter can use the new `CuaAdmissionLimitError` to declare that its local
13
+ limit refused a request before dispatch. Humanish records **adapter admission
14
+ limit** and ends the interaction without another retry, action or closing
15
+ report. Earlier actions and measured usage remain available. A refusal of an
16
+ optional closing report preserves the already observed task outcome. See the
17
+ [adapter contract](../contracts/adapter-admission.md) for when this signal applies.
18
+
19
+ If an earlier request stalled or had an ambiguous transport outcome, a later
20
+ success or local refusal does not make its missing usage known. The cost summary
21
+ keeps measured token estimates and an explicit unmeasured line. With a spending
22
+ cap, Humanish skips the optional closing report while that usage is uncertain.
23
+
24
+ Install with `npm install -g humanish@0.88.0`. CLI JSON adds optional diagnostics
25
+ without changing the existing status or exit-code contract. Usage telemetry,
26
+ enabled by default, adds two fields with fixed allowed values: diagnostic
27
+ category and stop cause. See [TELEMETRY.md](https://github.com/danielgwilson/humanish/blob/main/TELEMETRY.md)
28
+ for their scope and opt-out controls.
29
+
30
+ ## Verification and limits
31
+
32
+ [Admission-limit checks](https://github.com/danielgwilson/humanish/pull/759)
33
+ cover refusal before and after activity, sanitized error handling, preserved
34
+ outcomes and unknown usage after both successful and refused retries.
35
+ [CLI diagnostic checks](https://github.com/danielgwilson/humanish/pull/758)
36
+ cover previews, mixed endings, missing causes, invalid evidence and telemetry
37
+ allowlists. The required release checks also cover package installation,
38
+ Observer rendering and the website.
39
+
40
+ The admission signal depends on an accurate adapter declaration; it is not a
41
+ provider billing receipt. Existing generic failures keep their recorded
42
+ outcomes. This release does not infer causes from participant prose or establish
43
+ why an older study failed.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.86.1",
3
+ "version": "0.88.0",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [