humanish 0.86.1 → 0.87.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Status: public-safe contributor and agent ramp.
4
4
 
5
- Package/source version in this tree: `0.86.1` (2026-09-09). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
5
+ Package/source version in this tree: `0.87.0` (2026-09-10). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
6
6
  `0.15.1` remains in force: managed run and output paths bind to validated
7
7
  physical filesystem identities, and stored provider IDs are evidence, not
8
8
  cleanup authority. The bundled OSS meta-lab is dry-run only until
@@ -47,6 +47,10 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
47
47
 
48
48
  ## Current State
49
49
 
50
+ The [0.87.0 release note](../release/0.87.0-participant-endings-and-phone-review.md)
51
+ describes recorded interruption causes, phone playback controls beside fitted
52
+ captures, and credential and evidence-sharing guidance by execution route.
53
+
50
54
  The [0.86.1 release note](../release/0.86.1-task-preflight-saved-recordings.md)
51
55
  describes task preflight on unsupported execution paths and saved HTML recordings
52
56
  that remain snapshots when served over HTTP.
@@ -79,7 +83,7 @@ Implemented:
79
83
  - first-class lab manifest resolution through `humanish/labs/*.yaml` and
80
84
  ignored `.humanish/labs/*.yaml` overlays — `humanish.lab.v2` compositions
81
85
  (`src/lab-config.ts`), one engine, no hardcoded lab kinds;
82
- - a first-party actor registry with six registered descriptors
86
+ - a first-party actor registry with seven registered descriptors
83
87
  (`src/actor-registry.ts`); `actors[0].type` is a real dispatch key on the
84
88
  computer-use, scripted-browser, and terminal-product routes;
85
89
  - a computer-use route and clone subject provider: `subject.source: app-url`
@@ -130,7 +134,7 @@ without changing its success standard.
130
134
  - capability receipts are not adopter replacement: no first-party deletion
131
135
  branch has yet removed a bespoke generic harness while preserving
132
136
  decision-equivalent proof;
133
- - the six actor descriptors are a closed first-party union, not a supported
137
+ - the seven actor descriptors are a closed first-party union, not a supported
134
138
  out-of-tree actor-registration API;
135
139
  - run storage and provider-resource lifecycle logic still spans several routes
136
140
  instead of one `RunStore` and `ResourceLease` boundary;
@@ -0,0 +1,64 @@
1
+ # Humanish 0.87.0: see why a participant stopped
2
+
3
+ The Observer now distinguishes a provider response limit from elapsed time or
4
+ estimated spend. Participant cards, summaries and reports show the recorded
5
+ cause, including for participants with no screenshot frames. New computer-use
6
+ traces retain an optional `stopCause` alongside the existing status and
7
+ completion reason.
8
+
9
+ New computer-use sessions that hit their estimated spend cap before any action
10
+ are classified as interrupted, with no target-app feedback inferred from that
11
+ harness stop. The cap timing and recorded usage are unchanged.
12
+
13
+ Older recordings use exact machine-generated notices where available. A recorded
14
+ “provider token limit reached” notice remains **provider token limit**: it does
15
+ not establish whether the provider exhausted its output or context allowance.
16
+ When an older run records only a generic limit, the Observer says **limit
17
+ reached**. Updating a viewer preserves the original evidence and outcome counts;
18
+ it does not infer a cause from participant prose.
19
+
20
+ On phones, a saved landscape capture in Fit mode keeps playback controls beside
21
+ the fitted image. Portrait captures, zoomed views and live views retain their
22
+ bounded viewing area. The run banner states **Saved recording** once; the player
23
+ keeps the participant status and selected capture timestamp. Exact action links
24
+ and saved-moment recall retain their existing behavior.
25
+
26
+ Expanded assignments show the recorded lane focus before the mission, so a long
27
+ mission does not put those instructions below it in the phone panel. The full
28
+ mission and participant tasks remain available.
29
+
30
+ The website also explains credential placement by execution route. Computer-use
31
+ model keys remain on the host. Terminal actors receive a command-scoped runtime
32
+ key by default; optional `openai-egress` keeps the raw key outside the sandbox
33
+ while still allowing sandbox processes to make paid requests through its proxy.
34
+ Evidence stays local by default, and exposing Observer or sharing an export makes
35
+ it accessible to others. See the [credential and sharing guide](https://humanish.dev/docs/budgets-and-privacy#store-credentials).
36
+
37
+ Install with `npm install -g humanish@0.87.0`. Reopen a run with
38
+ `humanish observe --run RUN`, or regenerate its HTML export, to use the updated
39
+ viewer. Existing HTML files retain the renderer that produced them.
40
+
41
+ ## Verification and limits
42
+
43
+ [Interruption-cause checks](https://github.com/danielgwilson/humanish/pull/747)
44
+ cover recorded provider output limits, legacy token-limit notices, elapsed time,
45
+ spend and incomplete provider responses. Tests preserve old outcome counts and
46
+ source evidence through served and exported projections. A retained real
47
+ recording was inspected at desktop and phone widths, including its participant
48
+ whose provider stopped before any app action.
49
+
50
+ [Zero-action spend-cap checks](https://github.com/danielgwilson/humanish/pull/752)
51
+ verify one provider turn, zero executed actions, retained usage, an interrupted
52
+ outcome and no inferred target-app feedback. Older bundle counts are unchanged.
53
+
54
+ [Phone review checks](https://github.com/danielgwilson/humanish/pull/746) cover
55
+ landscape and portrait capture/control spacing, keyboard seeking, focus, zoom,
56
+ action links, refresh and saved-moment recall against retained recordings.
57
+ These browser checks do not establish physical-phone behavior or new live-stream
58
+ coverage. Screenshot replay remains sparse, and no participant success rate or
59
+ incremental persona benefit is inferred from these changes.
60
+
61
+ The [assignment ordering change](https://github.com/danielgwilson/humanish/pull/754)
62
+ preserves recorded text while moving lane instructions to the top of expanded
63
+ assignments. This addresses visibility; it does not establish that every reader
64
+ will interpret the instructions correctly.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.86.1",
3
+ "version": "0.87.0",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [