humanish 0.86.1 → 0.88.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/dist/actor-contract.d.ts +7 -0
  2. package/dist/actor-contract.js.map +1 -1
  3. package/dist/actor-stop-cause.d.ts +9 -0
  4. package/dist/actor-stop-cause.js +41 -0
  5. package/dist/actor-stop-cause.js.map +1 -0
  6. package/dist/computer-use.d.ts +4 -1
  7. package/dist/computer-use.js +45 -12
  8. package/dist/computer-use.js.map +1 -1
  9. package/dist/cua-actor-lab.d.ts +10 -3
  10. package/dist/cua-actor-lab.js +29 -4
  11. package/dist/cua-actor-lab.js.map +1 -1
  12. package/dist/cua-admission-limit.d.ts +12 -0
  13. package/dist/cua-admission-limit.js +20 -0
  14. package/dist/cua-admission-limit.js.map +1 -0
  15. package/dist/cua-diagnostics.d.ts +38 -0
  16. package/dist/cua-diagnostics.js +68 -0
  17. package/dist/cua-diagnostics.js.map +1 -0
  18. package/dist/feedback.js +11 -1
  19. package/dist/feedback.js.map +1 -1
  20. package/dist/index.d.ts +2 -1
  21. package/dist/index.js +1 -0
  22. package/dist/index.js.map +1 -1
  23. package/dist/observer-app.html +5 -5
  24. package/dist/observer-data.d.ts +5 -0
  25. package/dist/observer-data.js +27 -4
  26. package/dist/observer-data.js.map +1 -1
  27. package/dist/observer.js +3 -3
  28. package/dist/observer.js.map +1 -1
  29. package/dist/openai-responses-cu.js +14 -3
  30. package/dist/openai-responses-cu.js.map +1 -1
  31. package/dist/program.d.ts +2 -0
  32. package/dist/program.js +9 -4
  33. package/dist/program.js.map +1 -1
  34. package/dist/run.d.ts +6 -3
  35. package/dist/run.js +22 -7
  36. package/dist/run.js.map +1 -1
  37. package/dist/telemetry.d.ts +3 -0
  38. package/dist/telemetry.js +15 -1
  39. package/dist/telemetry.js.map +1 -1
  40. package/docs/contracts/adapter-admission.md +54 -0
  41. package/docs/contracts/run-bundle.md +24 -0
  42. package/docs/contracts/schemas.md +1 -1
  43. package/docs/goals/current.md +151 -719
  44. package/docs/ramp/README.md +11 -3
  45. package/docs/release/0.87.0-participant-endings-and-phone-review.md +64 -0
  46. package/docs/release/0.88.0-study-diagnostics.md +43 -0
  47. package/package.json +1 -1
@@ -1,754 +1,186 @@
1
1
  # Current Goals
2
2
 
3
- Status date: 2026-09-09 (rev 24)
3
+ Status date: 2026-09-11. Published baseline: `0.88.0`.
4
4
 
5
- This page is the current public-safe operating goal for `humanish`. Keep it
6
- short enough to reread before a coding session and concrete enough that future
7
- agents can choose useful work without private context.
5
+ This page guides work on current merged source. Published behavior is described
6
+ in the [release notes](../release/0.88.0-study-diagnostics.md).
7
+ The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
8
+ preserves the former status log; its queues do not supersede this page.
8
9
 
9
10
  ## North Star
10
11
 
11
- Humanish should be the open-source CLI that lets a maintainer ask:
12
+ Humanish should let a maintainer ask:
12
13
 
13
- > What happens when realistic synthetic personas try to use this app, CLI, or
14
+ > What happens when synthetic personas try to use this app, CLI, or
14
15
  > agent-facing workflow?
15
16
 
16
- The answer should be observable, verifiable, public-safe, and easy to turn into
17
- actionable feedback.
17
+ The answer should be observable, verifiable, public-safe, and useful for
18
+ making a repair. The working loop is:
18
19
 
19
- What humanish runs is synthetic user research, and every surface is checked
20
- against the three people a study involves — researcher, stakeholder,
21
- participant ([docs/principles/three-roles.md](../principles/three-roles.md)).
22
- The operational consequences: a study declares `tasks` with success criteria
23
- the participant never sees and gets a per-task completion funnel back; budgets
24
- are study-level recruiting decisions (`execution.caps.maxTotalUsd`) with
25
- per-lane caps as backstops; sessions end because the participant finished, not
26
- because a timer fired; abandonment and reported friction are findings that
27
- become feedback candidates, not failures. Receipts: the email-gated signup
28
- study completed, reproduced, and produced a real accessibility finding via a
29
- keyboard-first participant
30
- ([docs/goals/email-gated-signup/receipts/](email-gated-signup/receipts/)).
31
-
32
- ## Current Program Truth (source `0.86.1`)
33
-
34
- The package source and repository implementation in this tree agree on these
35
- points:
36
-
37
- **Declared tasks and saved recordings, 2026-09-09 (#737, #738).** Task declarations now fail preflight when the execution path cannot consume them, before hooks, processes or paid allocation. Portable HTML exports stay saved recordings over HTTP, including snapshots captured while a participant was running. Ordinary served runs still report real update failures. See the [0.86.1 release note](../release/0.86.1-task-preflight-saved-recordings.md).
38
-
39
- **Participant evidence context, 2026-09-09 (#733).** New runs preserve each participant's authored mission and lane focus. The Observer distinguishes individual actions within a capture interval, including links and saved moments, and labels the capture's time relative to the action. Supported desktop SDK startup cleanup shares a bounded result between SDK-internal cleanup and Humanish's fallback (#734). See the [0.86.0 release note](../release/0.86.0-participant-evidence.md).
40
-
41
- **Observer review continuity, 2026-09-09 (#731).** The run library opens during player/comparison review. Comparison frames link into the player and return to the saved alignment and cursor; participant names stay consistent. Explicit thinking filters reveal narration, and run/setup notices remain inspectable separately from frame-linked findings. See the [0.85.1 release note](../release/0.85.1-observer-continuity.md).
42
-
43
- **Observer watching and review, 2026-09-09 (#723, #724, #726, #728).** Grid previews preserve complete screens at a shared height, with compact captions and controls below the evidence. Run activity, live viewing, recorded replay and update freshness are distinct. Seeking, refresh and incoming captures preserve the selected viewing intent. The player adds elapsed-time review, zoom, saved moments and participant/cross-run comparison. The built TUI and observe serve existing evidence; exports contain no runtime desktop grants. See the [0.85.0 release note](../release/0.85.0-observer-review.md) for capabilities, acceptance evidence and limits.
20
+ ```text
21
+ study → recorded evidence → verify → feedback → repair → comparable rerun
22
+ ```
44
23
 
45
- **Portable feedback acceptance commands, 2026-09-07 (#720).** Generated proof commands use the installed CLI from the evidence workspace, including standalone exports without a package manifest. Redrafting recognized first-party candidates projects exact legacy command templates without rewriting source candidates or receipts. Custom instructions remain unchanged. See the [0.84.1 release note](../release/0.84.1-portable-feedback.md).
24
+ The run bundle is the source of truth; Observer is its review surface. Keep
25
+ product-specific tasks, state checks and vocabulary in adapters, with generic
26
+ execution, evidence and lifecycle primitives in core.
46
27
 
47
- **Readable evidence and shareable feedback, 2026-09-07 (#136).** `export --format bundle --redact-screenshots` creates a separate verified workspace while retaining the readable original. Four real operator-run recordings preserved all measured findings and costs through export and ordinary feedback commands; all 31 frames played in both bundle and HTML views. Static Observer grades now reflect their actual verification result. See the [retained workflow receipt](computer-use-actor/receipts/redacted-evidence-workflow-2026-09-07.md). External maintainer use and acceptance remain unproven.
28
+ ## The Three Roles
48
29
 
49
- **Declared evidence verification (#715).** Ordinary verification checks final/partial actor frame references and feedback candidate files. Explicit raw frame declarations keep partial traces `local_only`; valid zero-event terminal logs remain supported. Missing or unknown per-frame redaction metadata retains existing compatibility behavior, and unreferenced files are outside this check.
30
+ Check changes against the [researcher, stakeholder and participant](../principles/three-roles.md):
50
31
 
51
- **Narrow desktop and shared-study correction, 2026-09-07.** A minimum-width Chrome window can remain outside a 500px desktop after move/resize. A state-checked fullscreen fallback now proves physical containment before participant entry. Concurrent shared-world actors share the declared model-study budget, and a host ending before lobby handoff retains its actual failure rather than being reported as a timeout. See [bounded desktop receipt](computer-use-actor/receipts/narrow-desktop-containment-2026-09-07.md). This does not establish physical mobile fidelity or external adoption.
32
+ - The researcher declares the study, tasks, hidden success criteria, panel and
33
+ budgets. Supported routes return observation-backed task outcomes.
34
+ - The stakeholder needs inspectable moments, findings and denominators, with
35
+ evidence quality separate from participant success.
36
+ - The participant receives a goal and behavioral constraints. They may finish,
37
+ abandon, encounter a blocker or be interrupted; these outcomes must remain
38
+ distinct from failures of the harness.
52
39
 
53
- **Mobile input correction, 2026-09-05 (#676).** The historical 4/4 TodoMVC phone-lane rename
54
- failures describe the shipped desktop pointer-to-touch conversion path. In two new hosted
55
- conformance probes, SDK double click emitted two single clicks; direct touch opened the same
56
- original editor in both. A separate local native-X control reproduced the difference by toggling
57
- conversion. Mobile viewport and touch flags do not certify gesture equivalence or establish a
58
- physical-device app defect. Mobile lanes using touch conversion now carry that advisory in run
59
- warnings. [Method, traces summary and limits](computer-use-actor/receipts/mobile-input-conformance-2026-09-05.md).
40
+ A participant does not see their hidden success criteria. Missing observation
41
+ inputs remain unmeasured. Provider and session limits are safeguards, and a
42
+ limit ending a session must not be presented as natural completion.
60
43
 
61
- The immutable 2026-06-10 proof-roadmap packet is paired with a
62
- [current implementation checkpoint](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/README.md).
44
+ ## Best Next Work
63
45
 
64
- | Surface | Shipped | Still unproven or unbuilt |
46
+ An independent maintainer uses the published package on their own app,
47
+ adjudicates a useful finding, makes a repair and retains a comparable rerun.
48
+ Voluntary repeat use is the next adoption signal to establish.
49
+
50
+ Measure the effort and decisions in that loop:
51
+
52
+ - setup effort and interventions needed to obtain a usable recording;
53
+ - findings confirmed, dismissed or left uncertain after evidence review;
54
+ - accepted repair decisions and the results of comparable reruns;
55
+ - whether the maintainer chooses to use Humanish again.
56
+
57
+ Prioritize work that changes one of those outcomes. Fix a reproduced first-use
58
+ failure or evidence-review obstacle before expanding the platform. Use actual
59
+ recordings to check whether people can identify what happened, find the
60
+ relevant capture and understand the limits of the evidence.
61
+
62
+ The [own-app guide](https://humanish.dev/docs/your-app) and
63
+ [TodoMVC repair study](https://humanish.dev/docs/todomvc-edit-study) provide
64
+ starting points. Named public applications in study receipts are subjects,
65
+ not adopters or endorsements. An invitation is not activation.
66
+
67
+ ## What The Evidence Establishes
68
+
69
+ - A keyless `first-run` is a contract and review preview. It does not establish
70
+ live execution, app findings or independent use.
71
+ - Kept live studies establish the behavior observed on their named subjects,
72
+ versions, tasks and actors. Preserve interruptions and unsuccessful attempts
73
+ alongside successful ones.
74
+ - The [TodoMVC repair comparison](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/todomvc-edit-confirmation-2026-09-05.md)
75
+ supports one synthetic keyboard repair result; it does not estimate human
76
+ completion rates or general defect recall.
77
+ - The [evidence/export workflow receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/redacted-evidence-workflow-2026-09-07.md)
78
+ establishes readable originals and verified shareable derivatives on retained
79
+ runs. External maintainer acceptance remains a separate gate.
80
+ - Explicit keyboard/pointer contrasts do not isolate the value of rich persona
81
+ prompting. Matched comparisons need the same tasks, model, budgets and source
82
+ conditions, with validated findings and adjudication effort measured. See
83
+ [actor fidelity](../principles/actor-fidelity.md) for claim boundaries.
84
+
85
+ Capability proof also differs from replacing an adopter's bespoke harness.
86
+ The [proof roadmap](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/README.md)
87
+ requires decision-equivalent retained evidence and a real deletion branch.
88
+ No first-party deletion branch has met that gate. Public demonstrations do not
89
+ substitute for it.
90
+
91
+ ## Current Program Truth (source `0.88.0`)
92
+
93
+ | Surface | Available in merged source | Remaining boundary |
65
94
  | --- | --- | --- |
66
- | Actor execution | Six first-party registry descriptors; computer-use, scripted-browser, and terminal-product dispatch paths. `0.55.0` makes reasoning effort declarable per actor and per lane and records the resolved value on the trace (`humanish.model-settings.v1`) — it had been reachable in the provider and unreachable from a lab, so every prior run silently took the provider default | Public out-of-tree actor registration and conformance certification |
67
- | Persona scale | Bounded per-lane-world fan-out, including differentiated lanes and roster expansion; kept deterministic and live receipts | A completed first-party deletion branch that replaces a bespoke generic harness |
68
- | Shared state | Sequential and concurrent single-origin shared-world execution; sequential has deterministic proof, concurrent has deterministic and kept live proof | Multi-origin shared-world runtime/schema support; real-adopter deletion proof |
69
- | Subject sources/routes | Seven declared sources: `this-repo`, `clone`, `app-url`, `local-app`, `terminal-product`, `desktop-cli`, and `local-tree`; support is route-specific and `this-repo` remains dry-run-only | One centralized run/resource lifecycle boundary across all routes |
70
- | Public proof | A legible four-persona Observer hero from a verified real public-application study (commit-pinned drawDB) shipped in the npm payload (`0.16.0`) | Coverage beyond a single studied subject; the stratified breadth panel remains unbuilt |
71
- | OSS meta-lab | Dry-run contract and separate disposable smoke harness | Live meta-lab execution; disabled until repository instructions and actor credentials have an isolated boundary |
72
- | Observer serving | `watch`/`observe` loopback servers plus `serve` — the run-library surface with loopback default, capability-link exposure, `share_ready`-gated open mode, and optional operator-run tunnel; streams never served remotely | A remote live-stream (`--live-streams`) design; a persistent capability-link store |
73
- | Stakeholder terminal surface | `humanish tui` (`0.50.0`, redesigned to the reviewed spec in `0.51.0`, reworked again in `0.56.0` from stakeholder feedback: two explicit start rows instead of a hidden mode toggle, a description line so a list of studies says what they are, `?` keys, and ←/→ back to meaning back and open): labs -> lab -> run, arrow-key navigation, and starting a dry or live run from the lab screen. The run is DETACHED and outlives the terminal — live-proven by killing the terminal 42s into a real run that then ran on for ~3.5 minutes and finished `pass` at $0.639751. The run screen leads with the participant and their recorded thinking, live-proven mid-flight against a real computer-use run. Ships as one bundled file loaded on demand; refuses a non-interactive stdin/stdout naming the JSON commands instead. `0.51.0` builds the reviewed rev-8 design: wordmark and project context, content capped at 96 columns, live rows naming the PARTICIPANT and their elapsed clock, spinners and verdict glyphs, breadcrumbs, the lab's subject/model/caps/keys line, and one Start with a dry-run/live toggle. `0.52.0` completes the reviewed screen set: the run outcome card (denominator first, the participant's closing words, then Open in Observer / Run again), the interrupted card with Reclaim — live-proven by killing a real run and stopping its orphaned sandbox — and All runs, the cross-lab peer where participants lead and one thought line follows the cursor. `0.53.0` gives it the reviewed palette rather than the terminal's theme: exact hex on a truecolor terminal, downsampled where not, and never colour alone. `0.54.0` closes the last two gaps: stopping a running run (armed, signals the process GROUP, and says plainly that sandboxes are separate), and pricing a run WHILE it runs — the running usage now travels with the trace into the mid-run flush | Cancelling a run from the surface; per-user persisted config (#470); the TUI views over `export` (#471) and `stats` (#472), both shipped as CLI commands in `0.68.0` and `0.69.0`; agent-authored labs (#473); a reusable persona panel (#474) |
74
- | Off-app comms | Vendor-neutral in-sandbox email/SMS catch, a minimal persona inbox surface, and digest-only `humanish.comms-thread.v1` evidence; wired into the computer-use and shared-world routes over both HTTP and SMTP; live-proven end to end on 2026-08-08 — a persona signed up for a public app, read the emailed link in its inbox, and reached the signed-in product (`docs/goals/email-gated-signup/receipts/signup-verify-live-2026-08-08.md`); the adopter-hosted / app-url ingress plane is wired on the CUA and concurrent external-public routes (#387/#380, 2026-08-11) | Real-provider delivery; a live adopter-hosted receipt |
75
-
76
- Capability proof and adopter replacement are different gates. A deterministic
77
- test or kept live receipt proves that a Humanish mechanism works. The depth-axis
78
- goal is met only when an adopter produces decision-equivalent evidence on a
79
- green branch that deletes its bespoke generic harness and retains at most a
80
- thin product-specific extension. No first-party deletion branch had met that
81
- bar as of this status date.
82
-
83
- Multi-origin shared-world has a ratified core-design direction, but the
84
- implementation gate is still closed. A real adopter must first show a concrete
85
- cross-origin need that the single-origin path or a downstream facade cannot
86
- serve cleanly; the implementation packet then requires maintainer review before
87
- build work starts. The current amendment is
88
- [`docs/goals/multi-origin-shared-world/README.md`](https://github.com/danielgwilson/humanish/blob/main/docs/goals/multi-origin-shared-world/README.md);
89
- the dated design packet remains unchanged.
90
-
91
- ## Current Safety Boundary
92
-
93
- - Managed run, Observer, feedback, lab, actor-output, and source-archive paths
94
- bind to validated physical filesystem identities and fail closed on unsafe
95
- traversal, link, special-file, or retargeting states.
96
- - Provider IDs stored in `run.json` are mutable evidence, not cleanup
97
- authority. `humanish cleanup` writes an inspection receipt; same-process
98
- teardown continues to use the provider handles that created the resources.
99
- - The bundled `oss` manifest defaults to dry-run. Live OSS meta-lab execution
100
- fails with `HUMANISH_OSS_META_LIVE_ISOLATION_REQUIRED` before side effects
101
- until repository-derived instructions have an isolated credential boundary.
102
- - Ordinary Git repositories and verified linked worktrees remain supported.
103
- Git metadata that cannot pass containment validation is recorded as
104
- unavailable rather than followed.
105
-
106
- ## Definition Of Awesome
107
-
108
- A world-class Humanish run should eventually provide:
109
-
110
- - one human-friendly command that starts simulations and opens Observer;
111
- - multiple synthetic personas with different goals, patience, and skill levels;
112
- - UI, CLI, TUI, and code-agent lanes in one mission-control Observer;
113
- - real evidence: screenshots, terminal transcripts, lifecycle events, traces,
114
- filesystem setup-quality snapshots, artifacts, and verifier output;
115
- - clear pass, fail, blocked, and gap states;
116
- - public-safe feedback issue drafts that do not mutate GitHub by default;
117
- - first-class `.yaml` lab manifests for reusable simulation runs;
118
- - adapter contracts that let projects customize behavior without forking core;
119
- - release gates that prevent PII, PHI, secrets, private artifacts, and stale
120
- internal residue from reaching the public repo or package.
121
-
122
- ## Current Objective
123
-
124
- Make the public package and repo credible enough that an external maintainer can:
125
-
126
- 1. install the skill;
127
- 2. install `humanish`;
128
- 3. run `humanish init`;
129
- 4. run `humanish watch`;
130
- 5. run `humanish watch first-run` or another lab manifest;
131
- 6. inspect Observer evidence;
132
- 7. verify the bundle;
133
- 8. produce a public-safe feedback draft;
134
- 9. understand the next live-adapter path without reading chat history.
135
-
136
- ## Near-Term Goals
137
-
138
- ### 1. Public Readiness
139
-
140
- Keep the repository clean and public-safe.
141
-
142
- Acceptance:
95
+ | Study authoring | YAML labs, lane/roster composition, model settings and study/per-lane caps | Route support varies; declarations are not promises of engine parity |
96
+ | Actors | Seven first-party descriptors; computer-use, scripted-browser and terminal-product dispatch | No supported public out-of-tree actor-registration API |
97
+ | Subjects | `this-repo`, `clone`, `app-url`, `local-app`, `terminal-product`, `desktop-cli`, `local-tree` | `this-repo` is dry-run-only; `local-app` needs a caller-supplied executor/provider |
98
+ | Task protocol | Hidden criteria and per-task outcomes on supported per-lane CUA paths, including local-agent and desktop-cli | Shared-world, terminal-product, scripted and synthetic routes reject `tasks` before execution |
99
+ | Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
100
+ | Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
101
+ | Review and feedback | Verification grades, feedback drafts, portable HTML and redacted bundle derivatives | Sharing requires the appropriate grade; generated findings still need adjudication |
102
+ | TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
103
+ | Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
104
+ | Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
105
+
106
+ Use the [task support matrix](../architecture/task-protocol-support.md),
107
+ [actor registry](https://github.com/danielgwilson/humanish/blob/main/src/actor-registry.ts)
108
+ and [CLI reference](https://humanish.dev/docs/cli) when choosing a concrete path.
109
+ Source behavior and required tests outrank stale status prose.
110
+
111
+ ## Gates And Deferred Work
112
+
113
+ - Live OSS meta-lab execution remains disabled until repository-derived
114
+ instructions have an isolated credential boundary. Its dry-run and separate
115
+ disposable smoke harness do not open that gate.
116
+ - [Multi-origin shared-world work](https://github.com/danielgwilson/humanish/issues/239)
117
+ needs a real adopter's cross-origin requirement and a reviewed implementation
118
+ packet. It has a ratified core-design direction, but the implementation gate is still closed.
119
+ - [Nested provider grants](https://github.com/danielgwilson/humanish/pull/534)
120
+ remain unmerged. Do not assume a nested provider-credential channel exists.
121
+ - Paused adopter-deletion work stays paused until its current readiness and
122
+ maintainer decisions permit it. Do not resume it from the historical queue.
123
+ - Broader actor/plugin APIs, additional media, registry promotion and TUI
124
+ expansion are follow-ups, not substitutes for a useful maintainer workflow.
125
+ A concrete use case and current issue readiness determine when to take them up.
126
+
127
+ ## Safety And Autonomous Work
128
+
129
+ Follow [AGENTS.md](../../AGENTS.md), the [invariants](../principles/invariants-and-defaults.md)
130
+ and the [public-readiness standard](../release/public-readiness-standard.md).
131
+
132
+ - Keep `main` clean and work on scoped branches/worktrees. Substantial work
133
+ needs an issue with scope, authority, required proof and stop conditions.
134
+ - Existing explicit shipping authority governs implementation and merge;
135
+ otherwise issue readiness does not create authority by itself.
136
+ - Never commit secrets, private transcripts/screenshots, customer data or
137
+ private project context. Keep generated proof in ignored `.humanish/` and
138
+ retain needed evidence before removing a worktree.
139
+ - Managed paths bind to validated filesystem identities. Stored provider IDs
140
+ are evidence, not cleanup authority; reclaim only resources the operation is
141
+ authorized to own, and keep unknown cleanup explicitly unresolved.
142
+ - Verification distinguishes `share_ready`, `local_only` and `blocked`.
143
+ Feedback drafts do not mutate GitHub by default. Live spend, publishing,
144
+ external mutation and broader credential access are explicit choices.
145
+ - Provider credential placement is route-specific. The computer-use model key stays
146
+ on the host; the default terminal runtime uses command-scoped credentials. Do not infer safety from an unqualified “keys stay outside” claim.
147
+
148
+ ## Proof Before Shipping
149
+
150
+ From a clean contributor worktree:
143
151
 
144
152
  ```bash
153
+ pnpm install --frozen-lockfile
145
154
  pnpm release:check
155
+ pnpm docs:check
146
156
  git diff --check
147
157
  ```
148
158
 
149
- Fresh clone release checks should pass before public visibility changes.
150
-
151
- ### 2. Future-Agent Ramp
159
+ For Observer changes, build before its tests and inspect real bundle data at
160
+ 390px and desktop width. Run `pnpm --filter humanish-observer test` and
161
+ `pnpm observer:browser:proof`; preserve the source and limits of each receipt.
162
+ For website changes, run site typecheck, build and `registry:check`.
163
+ Required CI remains the merge gate.
152
164
 
153
- Maintain a durable ramp that tells future contributors and coding agents where
154
- to start, what exists, what remains, and what proof is required.
165
+ Before a release, follow the [release procedure](../release/open-source-readiness.md),
166
+ including the candidate-package `pnpm release:dogfood` study with authorized
167
+ credentials and bounded paid spend. A successful build alone is not participant
168
+ proof. Verify the published package and any changed publishing surface.
155
169
 
156
- Acceptance:
157
-
158
- - [`docs/ramp/README.md`](../ramp/README.md) stays current;
159
- - this page stays current;
160
- - README links both;
161
- - release package includes both docs directories.
162
-
163
- ### 3. Fresh-Agent Install Proof
164
-
165
- Prove the skill and package setup flow from a disposable target app with no chat
166
- context.
167
-
168
- Target proof:
170
+ A disposable, keyless consumer check is:
169
171
 
170
172
  ```bash
171
173
  npm i -D humanish
172
174
  npx humanish init --yes
173
- npx humanish watch --json --no-open
175
+ npx humanish run first-run --json
174
176
  npx humanish verify --run latest --json
175
177
  npx humanish feedback issue --run latest --repo owner/repo --format markdown
176
178
  ```
177
179
 
178
- The proof target must use synthetic personas and no real user data.
179
-
180
- ### 4. Live Browser Adapter
181
-
182
- Graduate from synthetic UI lanes to a real browser journey against a local app.
183
-
184
- Minimum acceptance:
185
-
186
- - local app target detection;
187
- - browser launch;
188
- - route/state capture;
189
- - screenshot artifact;
190
- - run bundle references screenshot evidence;
191
- - Observer renders the screenshot;
192
- - `verify` fails closed if required evidence is missing;
193
- - bounded desktop/mobile two-step browser persona proof with per-step traces and
194
- screenshots. `done`
195
- - LLM-driven browser lane: the registered `openai-computer-use` actor dispatches
196
- from a lab config (`subject.source: app-url`, loopback entry only) into a hosted
197
- E2B desktop, fills the provider-neutral `stream.actor` trace seam, and persists
198
- a verified redacted bundle (0.3.0 registered the actor; 0.4.0 made
199
- `actors[].type` a real dispatch key). `done`
200
- - Clone subject provider: `subject.source: clone` + `serve` clones a repo INTO the
201
- sandbox, installs/builds/starts it from config, probes readiness, and records
202
- provenance (repo, commit, env names) in the bundle — config-only computer-use
203
- labs against real apps (0.5.0; see `docs/goals/proof-roadmap/goal.md` —
204
- repo-only, not shipped in the npm package — and
205
- `docs/principles/invariants-and-defaults.md`, which ships in the package).
206
- `done`
207
- - De-paranoia (0.6.0): the redaction redesign + demoted defaults. Screenshots are
208
- full-fidelity by default (redaction binds the publish boundary, not capture —
209
- `policies.redactScreenshots` opts back in); `policies.allowPublicTargets` lets an
210
- owner drive a declared deployment/preview; `subject.clone.keep` is honored on
211
- failure for debugging; `serve.installTimeoutMs`/`buildTimeoutMs` are configurable
212
- for monorepo-scale builds. Doctrine updated with the capture-vs-publish rule. This
213
- re-sequences the proof roadmap: a redaction redesign and an overridable
214
- public-target policy are prerequisites for any decision-grade depth evidence, so
215
- they land BEFORE the consumer-web-app / agent-skill depth phases. `done`
216
- - Device presets (0.6.1): screen/device is a real dimension, with LITERAL values copied
217
- from the in-house sims (mobile 414×896 … wide 1920×1080; default `desktop` 1440×950) —
218
- not guessed. `execution.desktop.device` picks the per-run hosted screen; the guessed 1280×800
219
- is gone. Honest fidelity: on the E2B route only width/height render (real mobile *layout*)
220
- + the model is told its device, matching the sims' organic lanes; true touch/DPR/UA needs
221
- the CDP actor. Per-*persona* device (N×devices) rides fan-out. `done`
222
-
223
- ### 5. Live Terminal And Codex Lanes
224
-
225
- Make local PTY and Codex-style lanes reliable enough that Observer can show
226
- running, passed, failed, blocked, and timed-out states without human inference.
227
-
228
- Minimum acceptance:
229
-
230
- - sanitized transcript persistence;
231
- - explicit completion reason;
232
- - verifier checks redaction status;
233
- - Observer polling reflects lane completion;
234
- - no raw private transcript or credential values.
235
-
236
- Terminal-product real-agent lane (0.8.0; depth-axis layer 6, so an adopter can delete a
237
- bespoke real-agent sim for humanish + a thin adapter — see
238
- `docs/goals/terminal-product-lane/goal.md`):
239
-
240
- - `subject.source: terminal-product` + `execution.target: e2b-terminal` + the registered
241
- `codex-exec` terminal actor route a config to a real Codex agent studying a product from
242
- public surfaces inside an E2B shell. `done`
243
- - The credential-placement inversion is enforced by construction AND by verifier: the runtime
244
- key is injected ONLY command-scoped into the `codex` invocation, never sandbox-global; a
245
- deny-by-default allowlist excludes GitHub/payment/deploy/db creds; metadata is a positive
246
- allowlist; stdin is disabled with an always-present interventions ledger; cleanup is proven
247
- or the run fails closed. `done`
248
- - Cost/no-spend ledger with the null-vs-known-zero-vs-absent discipline (unknowns are `null`,
249
- never guessed); the no-spend proof is DERIVED from the ledger, never asserted; `maxUsd`/
250
- `maxJobs`/`maxMinutes` caps enforced fail-closed. `done`
251
- - Product-adapter extension seam: exported contract types + a scorer/feedback DI hook +
252
- adapter-namespaced product nouns, so an adopter attaches scoring/feedback as a thin
253
- in-repo extension without forking core. `done`
254
- - Cleanup is proven BY EXACT CREATED ID: `Sandbox.kill(id)`, confirmed further by
255
- `Sandbox.getInfo(id)` where the SDK exposes it, and humanish never calls `Sandbox.list`. A live
256
- rung never needs a dedicated or isolated E2B key; the SAME shared operator key used everywhere
257
- else in this repo is safe, because humanish only ever reaches a sandbox it created (see
258
- "The placement rule" corollary in `docs/principles/invariants-and-defaults.md`).
259
- - LIVE-PROVEN (2026-07-09): a real Codex agent, bootstrapped in a stock E2B shell (Node
260
- installed in-sandbox, run via `npx -y @openai/codex@latest exec`), studied a public
261
- agent-CLI product from its declared public surfaces and ran the product's free zero-spend
262
- guide within `$0` no-spend caps; verdict nonce-verified, cleanup proven BY EXACT ID
263
- (`getInfo(id)` SandboxNotFoundError, never `Sandbox.list`), verify 15/15, share_ready
264
- (`docs/goals/terminal-product-lane/receipts/terminal-live-rung-2026-07-09.md`). This closes
265
- the #159 live-receipt gap. Optional follow-up: a custom image with the agent runtime baked in
266
- to drop the per-run npx bootstrap. Duplex-PTY/xterm replay is a deferred SLICE 5.
267
-
268
- Multi-lane fan-out for the computer-use lab (0.9.0; proof-roadmap layer 2, the prerequisite
269
- for multi-actor shared-state work — #163, see `docs/goals/multi-lane-fanout/goal.md`):
270
-
271
- - `actors[0].lanes[]` (differentiated roster: per-lane persona/device/starting-surface) XOR
272
- `actors[0].count` (homogeneous) fan out N independent E2B desktops in ONE run bundle;
273
- `per-lane worlds` is the only topology this slice (shared-world is #164). `done`
274
- - `execution.concurrency` bounds in-flight paid desktops (default min(N,3); env may only
275
- lower it); a pre-flight spend/lane plan prints before any sandbox/provider call and at $0
276
- in dry-run; per-lane teardown reclaims ONLY each lane's own sandbox by id (never
277
- account-wide). `done`
278
- - Proven deterministically (fake substrate: bounded concurrency, by-id teardown, fail-fast,
279
- hollow-lane caught) AND with a kept live rung (2 lanes, two distinct desktops, both
280
- reclaimed by id, bundle verifies — `docs/goals/multi-lane-fanout/receipts/`). `done`
281
- - Deferred: seed-fork provisioning (PR-2), in-process-route fan-out, shared-world topology
282
- (#164).
283
-
284
- Shared-world topology — multi-actor against ONE shared mutable service (0.10.0; proof-roadmap
285
- layer 7; #164; `docs/goals/shared-world-topology/`). The north-star sim leverage: MANY personas,
286
- ONE shared world.
287
-
288
- - Sequential (`topology: shared-world`, concurrency 1): one sandbox, N role seats take turns
289
- against the shared DB; a checkpoint timeline proves role B acted on a world already containing
290
- role A's mutation. `done`
291
- - **Concurrent (`topology: shared-world` + `concurrency > 1`): one subject sandbox served +
292
- `getHost`-exposed, N actor desktop sandboxes drive that one URL SIMULTANEOUSLY** (reuses fan-out
293
- orchestration; all N+1 reclaimed by id). Honest attribution under concurrency: per-persona
294
- outcomes + harness-clocked `laneWindows` proving real overlap + a `stateSeries` of the shared
295
- world under load; causation is structurally inexpressible (independent series, no
296
- per-delta→actor field). `done`
297
- - A new `attributionClass: isolated | shared-world` honesty axis + verify FAIL-CLOSED on the
298
- required/forbidden `attributionLimits` sets + a concurrency-on-pass gate (a passed concurrent
299
- run must show real overlap AND a state delta coincident with it). `getHost` URLs are
300
- internet-reachable → the route is gated (verify) to synthetic+seeded subjects; the raw URL is
301
- digest-only in evidence. `done`
302
- - **LIVE-PROVEN (0.10.1):** a kept live receipt ran 3 personas concurrently against ONE
303
- getHost-exposed synthetic plane — all 3 passed, all 3 lane-windows overlapped on the real clock,
304
- the shared stateSeries evolved under load, N+1=4 sandboxes reclaimed by id, verify ok
305
- (`docs/goals/shared-world-topology/receipts/concurrent-live-rung-2026-06-17.md`). One trial =
306
- phase-change proof, not scale. The next step is the real downstream sim migration (a
307
- synthetic-seeded multi-role app in the adopter's domain). Per-action causation,
308
- cross-sandbox concurrency beyond getHost, and #108 PII/PHI remain out of scope.
309
- - shared-world (sequential AND concurrent) now also accepts `subject.source: local-tree` alongside
310
- `clone`: the ONE subject sandbox packs the operator's own working tree instead of cloning,
311
- reusing `provisionLocalTreeSubject` from the local-tree keystone (0.14.0). Provenance carries
312
- `archiveSha256` (the pin - one archive per run, so no per-lane unanimity math applies) plus
313
- host-side commit/dirty when the packed root is a git work tree; local-tree has no repo/publicRepo
314
- field. The N actor desktops on the concurrent route still drive the harness-minted getHost URL
315
- exactly as before; only the subject's provisioning + provenance source changed. The multi-origin
316
- design (`docs/goals/multi-origin-shared-world/design.md`) remains a separate,
317
- ratified but implementation-gated downstream slice. It is not part of
318
- `0.15.3`.
319
-
320
- Adopter-driven engine features (0.11.0; surfaced by real bespoke-sim migrations):
321
-
322
- - `execution.desktop.template` — run a lab on a CUSTOM E2B desktop image (name/ID) instead of the
323
- stock `desktop` template, threaded to `Sandbox.create(template, opts)` via one
324
- `createDesktopSandbox` seam across every desktop route (cua single+fan-out, sequential +
325
- concurrent shared-world subject+actors). Absent == the byte-stable stock-template call; recorded
326
- as `RunBundle.desktopTemplate`. Lets a Node/bun/DB-bearing adopter image run without
327
- installing the runtime per lane. `done`
328
- - `humanish observe --run <id>` — serves a run's Observer over `http://127.0.0.1:<port>` (loopback
329
- only, path-traversal-guarded to the run dir, `/`->`/observer/index.html`) instead of `file://`,
330
- so browsers/automation can open it and artifact links resolve. `done`
331
-
332
- Patch hardening (0.11.1):
333
-
334
- - concurrent shared-world review now fails closed when any actor lane records a failed terminal
335
- trace; a lane can remain evidence without making the aggregate review green. `done`
336
- - scripted-browser labs can provision a single cloned synthetic subject, expose it through a
337
- tokenless sandbox host, and drive deterministic scripted steps while persisting only public-safe
338
- provenance and host digests. `done`
339
-
340
- Adopter-driven roster/readback ergonomics (0.12.0):
341
-
342
- - Lane grouping metadata (`actorType`, `surface`, `caseGroup`) is adapter-owned and projected into
343
- Observer `laneGroups[]` plus stream labels, so downstream projects can group simulated users
344
- without teaching Humanish private role names. `done`
345
- - `actors[0].roster[]` is compact authoring sugar for repeated lane groups. The parser expands it
346
- into deterministic `lanes[]` (`<group.id>-01`, `<group.id>-02`, ...) before the engine runs, so
347
- the runtime and run bundle keep one normalized lane shape. `done`
348
-
349
- Provenance hardening (0.12.1):
350
-
351
- - Clone-subject provenance now refreshes after successful provisioning phases, so `subject.commit`
352
- records the served subject HEAD rather than only the initial clone HEAD. This preserves truthful
353
- run-bundle provenance when an adopter's install/provisioning step checks out the exact revision to
354
- test. `done`
355
-
356
- Adapter artifact evidence (0.12.15):
357
-
358
- - Browser/shared-world adapter hooks may now write product/state proof files under the ignored run
359
- directory and return namespaced `humanish.adapter-artifact.v1` references. Core validates only the
360
- generic reference shape and local-path safety, Observer links the artifacts, and `verify` fails
361
- closed if a referenced file disappears. The payload schema and product nouns stay in the adapter's
362
- namespace. `done`
363
-
364
- Evidence hygiene and readback polish (0.12.16):
365
-
366
- - Browser-backed lanes launch Chromium with shared evidence-hygiene defaults (first-run/update
367
- background surfaces suppressed, extensions/sync/component update disabled, password/autofill
368
- profile prompts disabled) so screenshots prefer product pixels over browser chrome. `done`
369
- - Run-bundle producers now use percent-scale simulation progress consistently: terminal states
370
- serialize as `100`, and only true in-progress shared-world snapshots serialize partial progress.
371
- This keeps Observer status pills from rendering completed runs as low-percentage complete states.
372
- `done`
373
- - Verify results now separate valid local evidence from public-promotable evidence with
374
- `shareSafety.status`. Raw full-fidelity screenshot runs remain valid local proof
375
- (`local_only`), while feedback draft/issue commands require `share_ready` and fail
376
- closed with structured reasons. `done`
377
-
378
- Attached CUA live Observer (shipped):
379
-
380
- - Plain computer-use labs now honor the same attached `onObserverReady` lifecycle as shared-world
381
- labs: a live CUA run writes an in-progress bundle before actor sessions complete, loopback
382
- `serveObserver` can hydrate desktop stream iframes while actors are still running, and stream auth
383
- URLs remain runtime-only through the Observer WeakMap rather than persisted into `run.json` or
384
- `observer-data.json`. `done`
385
-
386
- Local working-tree subject + operator observability (0.14.0):
387
-
388
- - `subject.source: local-tree` packs the lab resolution cwd on the host (git-aware
389
- enumeration honoring `.gitignore` and including uncommitted work, an always-on
390
- non-overridable secrets denylist, symlinks stored never dereferenced, one
391
- enumeration driving both the tar file list and the digest), uploads the
392
- once-per-run archive into each lane's desktop sandbox, extracts into the subject
393
- dir, and reuses the clone route's install/build/state/start/probe pipeline
394
- unchanged. Provenance pins the tree by `archiveSha256` (a dirty tree cannot be
395
- commit-pinned) plus host-side commit/dirty; `verify` fails closed on a live
396
- local-tree bundle without a well-formed pin. Live-proven twice with kept
397
- receipts: a dirty synthetic fixture and this repo packing itself
398
- (`docs/goals/local-tree-subject/receipts/`). `done`
399
- - Subject provisioning phase events: started/completed boundaries for
400
- clone/upload/extract/install/build/serve/readiness/seed-step phases stream to
401
- stderr by default (injectable via `CuaActorLabHooks.onPhase` /
402
- `SharedWorldLabHooks.onPhase`) and the completed trail persists into
403
- `bundle.events`; a single-lane provisioned boot is never silent again. `done`
404
- - Truthful CLI envelopes at the command boundary: any uncaught action error emits
405
- one structured `humanish.cli-response.v1` envelope (never a raw stack trace under
406
- `--json`, never a second stdout document after a flushed envelope), `humanish runs`
407
- gained a real failure branch, and `doctor` failure now exits 2 like every other
408
- structured command (behavioral change). `done`
409
-
410
- ### 6. Lab Manifest Shape
411
-
412
- Make reusable simulations feel like source artifacts, not hardcoded command
413
- branches.
414
-
415
- Minimum acceptance:
416
-
417
- - `humanish/labs/*.yaml` is the committed lab source convention;
418
- - `.humanish/labs/*.yaml` and `.humanish/local/labs/*.yaml` are ignored local
419
- overlays;
420
- - `humanish watch [lab]`, `humanish lab list`, `humanish lab inspect <lab>`, and
421
- `humanish lab run <lab>` are supported;
422
- - `--env-file <path>` loads local values for the current command without
423
- persisting values into artifacts;
424
- - maintainer dogfood labs such as `oss` are examples, not the canonical
425
- consumer taxonomy.
426
-
427
- ### 7. OSS Lab Health Readback
428
-
429
- Make the maintainer `oss` lab report nested lane health back into the
430
- top-level Observer instead of relying on a human watching the desktops.
431
-
432
- The current safety boundary above governs this lane. The completed bullets below
433
- record prior capability and evidence shape; they do not mean the live
434
- entrypoint is currently enabled.
435
-
436
- Minimum acceptance:
437
-
438
- - each lane records setup status; `done`
439
- - each lane records target app status/URL or blocker; `done`
440
- - each lane records nested Observer presence; `done`
441
- - each lane records nested verification status or blocker; `done`
442
- - each lane records setup-quality filesystem evidence and Observer can inspect
443
- it; `done`
444
- - top-level Observer updates lane verdicts from evidence; `done`
445
- - feedback candidates are derived from setup-quality/actor evidence; `done`
446
- - Codex app-server actor telemetry is persisted as redacted trace, event, and
447
- transcript artifacts; `done`
448
- - each lane receives a meaningful-use score over setup, filesystem, nested
449
- Humanish proof, actor activity, product surface, and feedback; `done`
450
- - provider-backed nested app-url proof now drives a bounded two-step
451
- desktop/mobile browser persona journey in a headed E2B lane; `done`
452
- - app-specific executable browser steps can now be authored under
453
- `humanish/scenarios/*.yaml` and are summarized into top-level nested proof
454
- evidence; `done`
455
- - repeated public app/tool headed proofs with app-specific manifests have passed
456
- against two public targets; `done`
457
- - next gap: richer multi-step product journeys and broader multi-persona
458
- matrices.
459
-
460
- ## Non-Goals
461
-
462
- Do not make these default behavior:
463
-
464
- - live provider spend;
465
- - GitHub API mutation;
466
- - hosted queues, databases, or webhooks;
467
- - production deploys;
468
- - real customer/user/patient data;
469
- - private screenshots or raw transcripts;
470
- - private upstream artifacts.
471
-
472
- Maintainer-only tooling can exist later, but it must be opt-in, token-explicit,
473
- and dry-run-first.
474
-
475
- ## Drift Alarms
476
-
477
- Stop and correct course if:
478
-
479
- - docs start depending on chat memory;
480
- - Observer gets prettier without stronger evidence;
481
- - feedback drafts imply product proof from synthetic contract proof;
482
- - tests pass while generated artifacts are not inspectable;
483
- - actor setup/use trials produce findings that never become feedback candidates;
484
- - live labs require private infrastructure to look impressive;
485
- - package docs link to files that are not shipped;
486
- - public-safety gates become optional.
487
-
488
- ## Best Next Work
489
-
490
- **2026-09-06 (0.83.1).** Desktop CLI studies without a declared product install
491
- now prepare Node/npm before participant entry, while leaving the product
492
- uninstalled. Two real stock desktops began without the runtime and reached the
493
- entry hook with Node/npm/npx available in ordinary and sudo shells. The hook
494
- then stopped deliberately without starting a participant or model. See the
495
- [runtime conformance receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/desktop-cli-runtime-2026-09-06.md).
496
-
497
- Terminal startup failures preserve confirmed cleanup from the desktop startup
498
- guard. An existing empty terminal event log is valid only when both embedded
499
- and retained terminal traces declare zero events. Unknown cleanup stays
500
- unproven, and a failed startup remains a failed run. The compiled CLI regression
501
- checks verified failure evidence and natural process exit using the installed
502
- SDK's debug constructor with network access blocked. Two hosted startup-fault
503
- controls also verified failed bundles, prompt natural exit and exact sandbox
504
- absence with real provider transports. See the
505
- [startup evidence receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/terminal-startup-evidence-2026-09-06.md).
506
-
507
- **2026-09-06 (0.83.0).** First-party OpenAI computer-use actors accept
508
- an optional per-response `maxOutputTokens` setting. Two independent lanes and two
509
- sequential shared-world roles forwarded the declared limit in real requests; all
510
- four preserved provider truncation as incomplete, with zero actions or closing
511
- requests. The same scoped integration confirmed sequential actor-level reasoning
512
- effort and a lane override on the request and trace. This proves configuration
513
- propagation, not participant success or persona efficacy. See the
514
- [four-role live receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/output-token-limit-2026-09-06.md).
515
-
516
- Hosted browser setup checks physical X window bounds before participant actions.
517
- A clipped window gets one correction and a measured readback; a window that
518
- remains clipped stops the lane. Two hosted fault-injection probes restored a
519
- bottom-edge button and clicked it successfully. This checks desktop containment,
520
- not mobile gesture fidelity. See the [desktop geometry receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/desktop-geometry-2026-09-06.md).
521
-
522
- Recovery hints now allow task-directed waiting and suggest recovery without
523
- asking a participant to abandon early. Deterministic delayed-start tests cover
524
- the prompt change and unchanged hard idle limits; live multiplayer patience
525
- has not been measured. Lane and roster-group objects reject unknown fields
526
- before allocation, so misspelled controls cannot silently become defaults.
527
-
528
- **2026-09-05 (0.82.1).** Computer-use actors now preserve a provider output-limit
529
- interruption as incomplete instead of treating a response without actions as
530
- success. Two captured live Responses API shapes reproduce the old false pass
531
- and verify the correction; other explicit non-completed statuses also fail
532
- closed. Usage and partial text remain in the run, with no execution of actions
533
- from an interrupted response. See the [provider-limit receipt](computer-use-actor/receipts/provider-token-limit-2026-09-05.md).
534
-
535
- **2026-09-05 (0.82.0).** Computer-use desktop estimates now use observed CPU and RAM,
536
- with missing resources and incomplete lifetimes left unknown (#687). Repeated clicks avoid
537
- the redundant cursor move when a fresh position check confirms the pointer is already there
538
- (#685). Terminal studies accept an exact Codex package version and forward the declared model
539
- and reasoning effort; evidence records the executed version without presenting runtime defaults
540
- as observed model usage (#688). Local-path redaction preserves nested terminal JSON framing,
541
- and the release report reader exposes skipped malformed lines (#686). Observer cards and reports
542
- show typed participant blockers as Blocked while retaining the original protocol trace (#691).
543
-
544
- The new [TodoMVC comparison](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/todomvc-edit-confirmation-2026-09-05.md)
545
- connects a keyboard blocker to a reproducible local patch: uninterrupted keyboard completion was
546
- 0/2 before and 2/2 after, with one provider interruption in each version retained in the twelve
547
- attempts. This descriptive synthetic comparison does not establish human completion rates.
548
- The public field notes and shorter reference-linked README make the method easier to inspect
549
- (#673, #682, #689).
550
-
551
- **2026-09-05 (0.81.0).** First use now distinguishes the free evidence preview
552
- from a live participant study, with concise successful setup output and complete JSON details
553
- (#660). The website has runnable docs and a generated CLI reference (#661, #668). Retained participant
554
- reports survive automatic stops. Supported providers with retained history can add one closing
555
- report when time and known budget remain; clean reports such as “no confusion or hesitation”
556
- stay clean (#658, #670, #671). Desktop startup cleanup retains acquired handles and records
557
- confirmed or unknown reclamation. The SDK's detached screenshot-file cleanup rejection is
558
- handled (#665, #666).
559
- Terminal output reconciles SDK callbacks with returned aggregates, preserving legitimate
560
- repeated lines and usage turns while avoiding doubled capture (#672). Opt-in `openai-egress`
561
- auth keeps the raw OpenAI runtime key outside the sandbox; every sandbox process can still
562
- spend through the proxy, so this is not a provider spending limit (#663). Stock terminal
563
- startup installs a checksum-verified Node archive without refreshing unrelated package mirrors
564
- and gives newly installed npm a default global prefix on the standard PATH (#677, #680). Mobile studies warn that desktop pointer-to-touch conversion can change repeated-tap
565
- behavior; gesture failures require direct or native touch confirmation before app attribution
566
- (#678).
567
-
568
- **2026-09-04, night (0.80.0).** The observation window reaches every desktop route: the
569
- sequential and concurrent shared-world seats forward `dwell` the way they forward `stopWhen`
570
- (#645), with a plumbing test per route and two live receipts, three participants holding together
571
- on one shared board whose checkpoint digest stood still until they acted, and two seats holding
572
- in turn on one sandbox (`receipts/dwell-window-2026-09-04.md`). The sequential route refuses a
573
- sandbox request over the provider's 60-minute cap before any call and names the per-role ceiling
574
- (#649); the provider's refusal had surfaced as a bundle that failed verification. The planted-defect
575
- benchmark ran with a second brain, the operator's own Claude Code as the participant: 14 of 15
576
- reported, nothing invented in three clean runs, 72 of 75 across both brains
577
- (`bench/RESULTS-2026-09-04-claude.md`). A two-minute window and a camera feed of the adopter's own
578
- are live receipts too.
579
-
580
- **2026-09-04, evening (0.79.0).** Two of the operator's own asks, each with a live receipt. A
581
- declared observation window (#510): `dwell` on an actor or a lane holds the page once its condition
582
- matches (or from the start), captures a frame on a cadence, takes no action and requests no model
583
- turn, then hands control back with a hint or ends the session; the window is recorded in the trace
584
- as deliberate and never outlasts the session budget. On TodoMVC the window opened when "1 item
585
- left" matched, held 60.6 s, captured six frames with no model turn inside, and the participant then
586
- added the second task (`receipts/dwell-window-2026-09-04.md`). The first live attempt found the
587
- option dropped between the lane and the loop by a spread the type checker cannot see; a plumbing
588
- test now pins it. A participant with a camera (#509, first slice): `execution.desktop.media.camera`
589
- gives a hosted Chrome lane a capture device (ffmpeg's test pattern generated in the sandbox, or a
590
- `.y4m` of the adopter's) behind the browser's own permission dialog by default;
591
- `policies.mediaPermission: granted` bypasses it; the bundle records the feed and the flags under
592
- `desktopBrowser.media`; a microphone is refused without an image that has an audio stack, before
593
- any spend. Live, Chrome raised its real dialog with a preview of the feed, the participant chose
594
- "Allow this time" and read back 640x480 (`receipts/participant-camera-2026-09-04.md`). Under
595
- the then-shipped mobile-emulation input path, TodoMVC rename had stopped 3 of 3 phone
596
- participants at this checkpoint; Excalidraw read 12 of 12. The 2026-09-05 input-conformance
597
- correction above qualifies attribution of those mobile failures.
598
-
599
- **2026-09-04, later (0.78.0).** `@e2b/desktop` moved from 2.2.3 to 2.3.3 (#638): its 2.3.1
600
- changelog names the socket #581 found today, a background command's event stream the SDK kept
601
- open after `Xvfb` and `startxfce4` were launched, which held the CLI process alive twelve minutes
602
- past a run's written result. `doctor` now prints the installed desktop SDK version on its row and
603
- attaches an advisory with the fix below 2.3.1 (#639). A live try-live on the new SDK reached its
604
- goal in 104 s and settled with no active resources.
605
-
606
- **2026-09-04 (0.77.0).** Four fixes found by runs and one by a scanner. A phone-emulated lane now
607
- follows the participant into a tab that opens later: the hold-mode applier attaches to every page
608
- target Chrome creates, paused before its first navigation, so the new tab lays out at the phone
609
- width, and the state observer reads each new tab's own fidelity report, recording a covered tab
610
- under `desktopGeometry.fidelity.laterTargets` (width, DPR, touch points) and a drifted one as a
611
- lane warning with the number the page gave; the applier's own log travels as `holderLog`. The
612
- first live proof failed in the participant's hands (a paused popup never resumed) and the shipped
613
- design is the one four live runs settled on: never pause a later tab, send the overrides the
614
- moment it exists, reload it once after its first navigation commits
615
- (`docs/goals/computer-use-actor/receipts/later-tab-emulation-2026-09-04.md`, #623, #636). A sandbox create
616
- or archive upload that fails on a transient provider error is retried once and named on the lane
617
- (#630): six lanes started 20 s apart lost five to E2B SDK errors before any turn, a probe of the
618
- same SDK a minute later worked in 6 s. The provider's `Retry-After` now governs a retry (capped
619
- at 60 s) and a `403 misalignment_policy_violation` ends the lane named and unresumed (#633). The
620
- provisioned-clone CI flake runs on an injected clock (#276 closed). The fourth benchmark run on
621
- this build read 15 of 15 planted defects with nothing invented in three clean runs, cumulative 58
622
- of 60 (`bench/RESULTS-2026-09-04-0.76.0.md`). humanish.dev negotiates `Accept: text/markdown`
623
- and llms.txt says when to use humanish (#634).
624
-
625
- ## Best Next Work
180
+ This checks installation, preview evidence and feedback generation. A live
181
+ claim needs actual participant execution, retained evidence and verified
182
+ cleanup. Use N > 1 where the claim needs replication; report cost estimates,
183
+ unknowns and failures without turning them into zeros or successes.
626
184
 
627
- **2026-09-03, later (0.76.0).** A hosted Chromium lane on a mobile preset can now be a
628
- mobile-emulated browser: `execution.desktop.fidelity.mobileEmulation: true` applies a real 414 px
629
- CSS viewport, the preset's device pixel ratio, touch events and a mobile user agent before the
630
- participant's first observation, and the bundle records what the page then reported about itself
631
- under `desktopGeometry.fidelity` (#221, `docs/goals/computer-use-actor/receipts/mobile-emulation-2026-09-03.md`).
632
- On the published 0.76.0, neither phone participant could rename through the emulated input path,
633
- where the 500 px runs without touch had finished. Those are historical instrument observations:
634
- the 2026-09-05 conformance check found SDK double click and direct touch differ on the original
635
- editor, so the earlier outcome does not establish a touch-device app defect.
636
- The persona axis was replicated the same evening on drawDB and TodoMVC and given a third app,
637
- Excalidraw, as the clean control (`persona-axis-phone-2026-09-03.md`); a multi-lane study's
638
- second and third findings reach `feedback draft` through `--candidate` (#609); a negated report
639
- word no longer counts as friction (#614); and the adopter metric excludes this project's own
640
- checkout wherever its CLI is spawned from (#611). Three host-timing test flakes were made
641
- load-independent.
642
-
643
- **2026-09-03 (0.75.0).** Three fixes found by runs, each with a receipt. The DevTools probe behind
644
- every url, page-text and CSS-viewport observation ran on Node inside the sandbox, and the stock
645
- desktop has none, so on the app-url route and on any subject served by something other than Node
646
- every `urlIncludes` / `textIncludes` stop condition and task criterion had been silently blind; it
647
- runs on python3 now, a dark channel is named in the lane warnings, and the same lab reads
648
- "never measured in 2" on 0.74.0 and 2/2 at turn 0 on 0.75.0
649
- (`docs/goals/computer-use-actor/receipts/task-observation-514-2026-09-03.md`, #514). A subject
650
- install or runtime bootstrap that exits non-zero is retried once, and the error leads with a line a
651
- person can act on (#602). The gpt-5.6-sol rates were re-pinned at the live promotional sheet
652
- (4 / 0.40 / 5 / 20 per 1M through at least 2026-11-21), so estimates stop running 25-33% high, and
653
- gpt-6-astra is priced. #548 closed on the benchmark receipts.
654
-
655
- **2026-09-01 (0.66.0 through 0.70.0, one session).** The product got its first efficacy numbers and
656
- they are in the repo with run ids: planted-defect recall 43 of 45 over three benchmark runs and
657
- 12 clean runs with nothing invented (`bench/`; 58 of 60 and 15 clean runs after the 09-04 rerun), 5 of 6 findings confirmed on TodoMVC and 11 of 12
658
- on drawDB against the source (20 participants on apps we did not write, 0 false reports), and the
659
- comparison of keyboard and pointer use on two apps: every keyboard-first participant hit a defect no
660
- mouse-driving participant met (drawDB's mouse-only database modal, TodoMVC's double-click-only
661
- rename), receipts under `docs/goals/computer-use-actor/receipts/`. These observations show
662
- sensitivity to explicit keyboard-versus-pointer instructions; they do not establish incremental
663
- benefit from persona prompting over a matched generic tester. Five cold installs of two
664
- published versions reached the goal in under three minutes each. On the way the runs found and
665
- the releases fixed: telemetry that never said what a study was (and stored IPs), a refused lane
666
- written up as a pass (#476), a blocker regex that refused five of five finished runs (#565, then
667
- the structural #570: the participant declares its own outcome, 12 of 12 adherence), a
668
- study-participant marker nothing set (#546), hung turns that ate a lane's budget (#469, #480),
669
- and a Claude participant with no memory across turns (#520). `humanish stats` (#472) and
670
- `humanish export` (#471) exist as CLI commands.
671
-
672
- The next outcome to establish is an external maintainer using Humanish on their
673
- own app, finding a useful problem, repairing it and retaining a comparable rerun.
674
- The published studies establish synthetic capability on their named subjects;
675
- they do not establish external adoption or incremental persona value.
676
- [Runnable documentation](https://humanish.dev/docs) and the repair-study recipe
677
- are available; the documentation gap from #513 is closed.
678
-
679
- Remaining engineering work includes:
680
-
681
- - #581: #665 already reclaims acquired desktop instances when startup fails,
682
- and #666 handles detached screenshot cleanup errors. The compiled CLI now has
683
- deterministic exit proof and two hosted startup-fault receipts. Remaining
684
- boundaries are allocation ambiguity before instance acquisition, unconfirmed
685
- reclamation, and exit behavior after other provider/network failures.
686
- The 2.3.3 SDK update fixed the identified background stream holder; older
687
- optional peers still receive an advisory, so a universal no-linger claim
688
- would be too broad.
689
- - #509: microphone support requires a desktop image with an audio stack and
690
- the matching launch path. The unsupported declaration remains rejected.
691
- - #221 and #676: real-device and touch-input fidelity beyond viewport emulation;
692
- #623: scripted-browser support beyond its current model-free route.
693
- - TUI views over `stats` and `export` (#455), registry promotions (#431), and
694
- the remaining shared-world evidence work (#365, #446).
695
-
696
- Earlier state, kept for the record:
697
-
698
- The bounded public-proof side task is done: a verified, legible four-persona
699
- Observer hero from a commit-pinned public application (drawDB) shipped in
700
- `0.16.0`. That run treats the application as a study subject, not a Humanish
701
- adopter, and does not imply endorsement; it does not satisfy the depth gate.
702
-
703
- The operator-pinned train-of-thought thread (#427, both stages) and the
704
- #441 queue head are done as of `0.47.0`, live-proven the same day they
705
- shipped: trace items carry recording stamps and structured click
706
- coordinates; the live path flushes `liveActor` partials mid-run so the
707
- attached Observer's timeline grows while a study runs; playback runs at
708
- the participant's recorded pace when stamps exist (and the transport says
709
- which clock it is on); the participant card's decide-line ticks the
710
- newest reported thought during a live lane; and a live-shaped golden cut
711
- from a kept receipt run pins the whole shape (`?fixture=live`). On the way
712
- the capture exposed and closed two evidence-quality defects: the persona
713
- identity leak (#452) and the credible-pass guard's incentive inversion
714
- against post-success defect reports (#453 — the verdict scan and the
715
- friction tally are now separate reads of the same narrative). Shared-world
716
- roll-up honesty (#364) closed via an external contribution that separates
717
- credibility, mission endpoint, and convergence into three explicit claims.
718
-
719
- Deep links landed in `0.48.0` (#464): every participant and frame is
720
- addressable (`#/lane/<id>/f/<n>`), Back/Forward restore the view, and a
721
- reload or shared link lands on the exact moment — #441 closed entirely.
722
-
723
- The npx-first-try adoption cluster closed in `0.49.0`, operator-prompted and
724
- adversarially red-teamed before merge (both arcs): provider keys now resolve
725
- through each vendor's native chain (#436 — the documented project overlay,
726
- `e2b auth login`'s store, `gh auth token`, and a `humanish keys set` user
727
- store; fills announced by name and source, never value; `HUMANISH_STRICT_KEYS=1`
728
- opts out) and #346 closed on its receipts. The computer-use default moved to
729
- `gpt-5.6-sol` with the whole 5.6 family priced, and the cost estimate now
730
- models the two billing mechanics 5.6 introduced — cache writes at 1.25x and
731
- long-context re-tiering — exactly, from a new per-request usage ledger on the
732
- trace (#334). Both spend caps price through the same tier-aware estimator.
733
-
734
- The standing queue, in rough order:
735
-
736
- 1. registry promotions: the wordmark (#431), the participant card, and the two
737
- vendored Base UI wrappers (drawer, popover) once a second surface consumes
738
- them;
739
- 2. shared-world honesty, remaining half: per-action evidence (#365) and
740
- exposure-flag coverage (#446);
741
- 3. the stakeholder TUI (#455): research + token-translated mocks on the
742
- operator review surface, design sign-off gated before any code.
743
-
744
- The depth-axis deletion front (an adopter's bespoke terminal-product sim,
745
- comparator contract posted on the adopter's tracker) is paused awaiting the
746
- operator's decision-ledger sign-off, deliberately: it resumes on that
747
- sign-off, not by default. The multi-actor shared-state adopter gate (#166)
748
- needs one live rerun of its replacement-critical family on current humanish;
749
- the recon (family, runner pin) is recorded and it runs as its own arc.
750
-
751
- The version-pinned README hero is the drawDB real-application study: a live
752
- four-persona capture that proves package/Observer rendering, public-safe asset
753
- delivery, and real-application evidence against a studied subject. It is not
754
- adopter deletion evidence, which is a separate and higher gate.
185
+ Start from the [ramp](../ramp/README.md), choose one changed user outcome, and
186
+ close work with what changed, what was checked and what remains uncertain.