humanish 0.86.0 → 0.87.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/actor-contract.d.ts +4 -0
- package/dist/actor-contract.js.map +1 -1
- package/dist/actor-stop-cause.d.ts +7 -0
- package/dist/actor-stop-cause.js +36 -0
- package/dist/actor-stop-cause.js.map +1 -0
- package/dist/computer-use.d.ts +1 -1
- package/dist/computer-use.js +15 -8
- package/dist/computer-use.js.map +1 -1
- package/dist/concurrent-shared-world-lab.d.ts +1 -1
- package/dist/concurrent-shared-world-lab.js +4 -0
- package/dist/concurrent-shared-world-lab.js.map +1 -1
- package/dist/cua-actor-lab.d.ts +1 -1
- package/dist/cua-actor-lab.js +9 -0
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/e2b-terminal-lab.d.ts +1 -1
- package/dist/e2b-terminal-lab.js +4 -0
- package/dist/e2b-terminal-lab.js.map +1 -1
- package/dist/export.js +1 -1
- package/dist/export.js.map +1 -1
- package/dist/index.d.ts +1 -1
- package/dist/index.js.map +1 -1
- package/dist/lab-config.d.ts +4 -0
- package/dist/lab-config.js +18 -0
- package/dist/lab-config.js.map +1 -1
- package/dist/lab-engine.js +24 -1
- package/dist/lab-engine.js.map +1 -1
- package/dist/observer-app.html +9 -9
- package/dist/observer-data.d.ts +5 -0
- package/dist/observer-data.js +27 -4
- package/dist/observer-data.js.map +1 -1
- package/dist/observer.d.ts +3 -1
- package/dist/observer.js +9 -6
- package/dist/observer.js.map +1 -1
- package/dist/openai-responses-cu.js +1 -1
- package/dist/openai-responses-cu.js.map +1 -1
- package/dist/oss-lab.d.ts +1 -1
- package/dist/oss-meta-lab.d.ts +1 -1
- package/dist/oss-meta-lab.js.map +1 -1
- package/dist/run.d.ts +6 -3
- package/dist/run.js +22 -7
- package/dist/run.js.map +1 -1
- package/dist/scripted-browser-lab.d.ts +1 -1
- package/dist/scripted-browser-lab.js +9 -0
- package/dist/scripted-browser-lab.js.map +1 -1
- package/dist/shared-world-lab.d.ts +1 -1
- package/dist/shared-world-lab.js +4 -0
- package/dist/shared-world-lab.js.map +1 -1
- package/docs/architecture/task-protocol-support.md +35 -0
- package/docs/contracts/run-bundle.md +17 -0
- package/docs/contracts/schemas.md +1 -1
- package/docs/goals/current.md +151 -717
- package/docs/ramp/README.md +11 -3
- package/docs/release/0.86.1-task-preflight-saved-recordings.md +40 -0
- package/docs/release/0.87.0-participant-endings-and-phone-review.md +64 -0
- package/package.json +3 -2
package/docs/goals/current.md
CHANGED
|
@@ -1,752 +1,186 @@
|
|
|
1
1
|
# Current Goals
|
|
2
2
|
|
|
3
|
-
Status date: 2026-09-
|
|
3
|
+
Status date: 2026-09-10. Published baseline: `0.87.0`.
|
|
4
4
|
|
|
5
|
-
This page
|
|
6
|
-
|
|
7
|
-
|
|
5
|
+
This page guides work on current merged source. Published behavior is described
|
|
6
|
+
in the [release notes](../release/0.87.0-participant-endings-and-phone-review.md).
|
|
7
|
+
The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
|
|
8
|
+
preserves the former status log; its queues do not supersede this page.
|
|
8
9
|
|
|
9
10
|
## North Star
|
|
10
11
|
|
|
11
|
-
Humanish should
|
|
12
|
+
Humanish should let a maintainer ask:
|
|
12
13
|
|
|
13
|
-
> What happens when
|
|
14
|
+
> What happens when synthetic personas try to use this app, CLI, or
|
|
14
15
|
> agent-facing workflow?
|
|
15
16
|
|
|
16
|
-
The answer should be observable, verifiable, public-safe, and
|
|
17
|
-
|
|
17
|
+
The answer should be observable, verifiable, public-safe, and useful for
|
|
18
|
+
making a repair. The working loop is:
|
|
18
19
|
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
The operational consequences: a study declares `tasks` with success criteria
|
|
23
|
-
the participant never sees and gets a per-task completion funnel back; budgets
|
|
24
|
-
are study-level recruiting decisions (`execution.caps.maxTotalUsd`) with
|
|
25
|
-
per-lane caps as backstops; sessions end because the participant finished, not
|
|
26
|
-
because a timer fired; abandonment and reported friction are findings that
|
|
27
|
-
become feedback candidates, not failures. Receipts: the email-gated signup
|
|
28
|
-
study completed, reproduced, and produced a real accessibility finding via a
|
|
29
|
-
keyboard-first participant
|
|
30
|
-
([docs/goals/email-gated-signup/receipts/](email-gated-signup/receipts/)).
|
|
31
|
-
|
|
32
|
-
## Current Program Truth (source `0.86.0`)
|
|
33
|
-
|
|
34
|
-
The package source and repository implementation in this tree agree on these
|
|
35
|
-
points:
|
|
36
|
-
|
|
37
|
-
**Participant evidence context, 2026-09-09 (#733).** New runs preserve each participant's authored mission and lane focus. The Observer distinguishes individual actions within a capture interval, including links and saved moments, and labels the capture's time relative to the action. Supported desktop SDK startup cleanup shares a bounded result between SDK-internal cleanup and Humanish's fallback (#734). See the [0.86.0 release note](../release/0.86.0-participant-evidence.md).
|
|
38
|
-
|
|
39
|
-
**Observer review continuity, 2026-09-09 (#731).** The run library opens during player/comparison review. Comparison frames link into the player and return to the saved alignment and cursor; participant names stay consistent. Explicit thinking filters reveal narration, and run/setup notices remain inspectable separately from frame-linked findings. See the [0.85.1 release note](../release/0.85.1-observer-continuity.md).
|
|
40
|
-
|
|
41
|
-
**Observer watching and review, 2026-09-09 (#723, #724, #726, #728).** Grid previews preserve complete screens at a shared height, with compact captions and controls below the evidence. Run activity, live viewing, recorded replay and update freshness are distinct. Seeking, refresh and incoming captures preserve the selected viewing intent. The player adds elapsed-time review, zoom, saved moments and participant/cross-run comparison. The built TUI and observe serve existing evidence; exports contain no runtime desktop grants. See the [0.85.0 release note](../release/0.85.0-observer-review.md) for capabilities, acceptance evidence and limits.
|
|
20
|
+
```text
|
|
21
|
+
study → recorded evidence → verify → feedback → repair → comparable rerun
|
|
22
|
+
```
|
|
42
23
|
|
|
43
|
-
|
|
24
|
+
The run bundle is the source of truth; Observer is its review surface. Keep
|
|
25
|
+
product-specific tasks, state checks and vocabulary in adapters, with generic
|
|
26
|
+
execution, evidence and lifecycle primitives in core.
|
|
44
27
|
|
|
45
|
-
|
|
28
|
+
## The Three Roles
|
|
46
29
|
|
|
47
|
-
|
|
30
|
+
Check changes against the [researcher, stakeholder and participant](../principles/three-roles.md):
|
|
48
31
|
|
|
49
|
-
|
|
32
|
+
- The researcher declares the study, tasks, hidden success criteria, panel and
|
|
33
|
+
budgets. Supported routes return observation-backed task outcomes.
|
|
34
|
+
- The stakeholder needs inspectable moments, findings and denominators, with
|
|
35
|
+
evidence quality separate from participant success.
|
|
36
|
+
- The participant receives a goal and behavioral constraints. They may finish,
|
|
37
|
+
abandon, encounter a blocker or be interrupted; these outcomes must remain
|
|
38
|
+
distinct from failures of the harness.
|
|
50
39
|
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
original editor in both. A separate local native-X control reproduced the difference by toggling
|
|
55
|
-
conversion. Mobile viewport and touch flags do not certify gesture equivalence or establish a
|
|
56
|
-
physical-device app defect. Mobile lanes using touch conversion now carry that advisory in run
|
|
57
|
-
warnings. [Method, traces summary and limits](computer-use-actor/receipts/mobile-input-conformance-2026-09-05.md).
|
|
40
|
+
A participant does not see their hidden success criteria. Missing observation
|
|
41
|
+
inputs remain unmeasured. Provider and session limits are safeguards, and a
|
|
42
|
+
limit ending a session must not be presented as natural completion.
|
|
58
43
|
|
|
59
|
-
|
|
60
|
-
[current implementation checkpoint](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/README.md).
|
|
44
|
+
## Best Next Work
|
|
61
45
|
|
|
62
|
-
|
|
46
|
+
An independent maintainer uses the published package on their own app,
|
|
47
|
+
adjudicates a useful finding, makes a repair and retains a comparable rerun.
|
|
48
|
+
Voluntary repeat use is the next adoption signal to establish.
|
|
49
|
+
|
|
50
|
+
Measure the effort and decisions in that loop:
|
|
51
|
+
|
|
52
|
+
- setup effort and interventions needed to obtain a usable recording;
|
|
53
|
+
- findings confirmed, dismissed or left uncertain after evidence review;
|
|
54
|
+
- accepted repair decisions and the results of comparable reruns;
|
|
55
|
+
- whether the maintainer chooses to use Humanish again.
|
|
56
|
+
|
|
57
|
+
Prioritize work that changes one of those outcomes. Fix a reproduced first-use
|
|
58
|
+
failure or evidence-review obstacle before expanding the platform. Use actual
|
|
59
|
+
recordings to check whether people can identify what happened, find the
|
|
60
|
+
relevant capture and understand the limits of the evidence.
|
|
61
|
+
|
|
62
|
+
The [own-app guide](https://humanish.dev/docs/your-app) and
|
|
63
|
+
[TodoMVC repair study](https://humanish.dev/docs/todomvc-edit-study) provide
|
|
64
|
+
starting points. Named public applications in study receipts are subjects,
|
|
65
|
+
not adopters or endorsements. An invitation is not activation.
|
|
66
|
+
|
|
67
|
+
## What The Evidence Establishes
|
|
68
|
+
|
|
69
|
+
- A keyless `first-run` is a contract and review preview. It does not establish
|
|
70
|
+
live execution, app findings or independent use.
|
|
71
|
+
- Kept live studies establish the behavior observed on their named subjects,
|
|
72
|
+
versions, tasks and actors. Preserve interruptions and unsuccessful attempts
|
|
73
|
+
alongside successful ones.
|
|
74
|
+
- The [TodoMVC repair comparison](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/todomvc-edit-confirmation-2026-09-05.md)
|
|
75
|
+
supports one synthetic keyboard repair result; it does not estimate human
|
|
76
|
+
completion rates or general defect recall.
|
|
77
|
+
- The [evidence/export workflow receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/redacted-evidence-workflow-2026-09-07.md)
|
|
78
|
+
establishes readable originals and verified shareable derivatives on retained
|
|
79
|
+
runs. External maintainer acceptance remains a separate gate.
|
|
80
|
+
- Explicit keyboard/pointer contrasts do not isolate the value of rich persona
|
|
81
|
+
prompting. Matched comparisons need the same tasks, model, budgets and source
|
|
82
|
+
conditions, with validated findings and adjudication effort measured. See
|
|
83
|
+
[actor fidelity](../principles/actor-fidelity.md) for claim boundaries.
|
|
84
|
+
|
|
85
|
+
Capability proof also differs from replacing an adopter's bespoke harness.
|
|
86
|
+
The [proof roadmap](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/README.md)
|
|
87
|
+
requires decision-equivalent retained evidence and a real deletion branch.
|
|
88
|
+
No first-party deletion branch has met that gate. Public demonstrations do not
|
|
89
|
+
substitute for it.
|
|
90
|
+
|
|
91
|
+
## Current Program Truth (source `0.87.0`)
|
|
92
|
+
|
|
93
|
+
| Surface | Available in merged source | Remaining boundary |
|
|
63
94
|
| --- | --- | --- |
|
|
64
|
-
|
|
|
65
|
-
|
|
|
66
|
-
|
|
|
67
|
-
|
|
|
68
|
-
|
|
|
69
|
-
|
|
|
70
|
-
|
|
|
71
|
-
|
|
|
72
|
-
| Off-app
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
## Current Objective
|
|
121
|
-
|
|
122
|
-
Make the public package and repo credible enough that an external maintainer can:
|
|
123
|
-
|
|
124
|
-
1. install the skill;
|
|
125
|
-
2. install `humanish`;
|
|
126
|
-
3. run `humanish init`;
|
|
127
|
-
4. run `humanish watch`;
|
|
128
|
-
5. run `humanish watch first-run` or another lab manifest;
|
|
129
|
-
6. inspect Observer evidence;
|
|
130
|
-
7. verify the bundle;
|
|
131
|
-
8. produce a public-safe feedback draft;
|
|
132
|
-
9. understand the next live-adapter path without reading chat history.
|
|
133
|
-
|
|
134
|
-
## Near-Term Goals
|
|
135
|
-
|
|
136
|
-
### 1. Public Readiness
|
|
137
|
-
|
|
138
|
-
Keep the repository clean and public-safe.
|
|
139
|
-
|
|
140
|
-
Acceptance:
|
|
95
|
+
| Study authoring | YAML labs, lane/roster composition, model settings and study/per-lane caps | Route support varies; declarations are not promises of engine parity |
|
|
96
|
+
| Actors | Seven first-party descriptors; computer-use, scripted-browser and terminal-product dispatch | No supported public out-of-tree actor-registration API |
|
|
97
|
+
| Subjects | `this-repo`, `clone`, `app-url`, `local-app`, `terminal-product`, `desktop-cli`, `local-tree` | `this-repo` is dry-run-only; `local-app` needs a caller-supplied executor/provider |
|
|
98
|
+
| Task protocol | Hidden criteria and per-task outcomes on supported per-lane CUA paths, including local-agent and desktop-cli | Shared-world, terminal-product, scripted and synthetic routes reject `tasks` before execution |
|
|
99
|
+
| Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
|
|
100
|
+
| Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
|
|
101
|
+
| Review and feedback | Verification grades, feedback drafts, portable HTML and redacted bundle derivatives | Sharing requires the appropriate grade; generated findings still need adjudication |
|
|
102
|
+
| TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
|
|
103
|
+
| Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
|
|
104
|
+
| Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
|
|
105
|
+
|
|
106
|
+
Use the [task support matrix](../architecture/task-protocol-support.md),
|
|
107
|
+
[actor registry](https://github.com/danielgwilson/humanish/blob/main/src/actor-registry.ts)
|
|
108
|
+
and [CLI reference](https://humanish.dev/docs/cli) when choosing a concrete path.
|
|
109
|
+
Source behavior and required tests outrank stale status prose.
|
|
110
|
+
|
|
111
|
+
## Gates And Deferred Work
|
|
112
|
+
|
|
113
|
+
- Live OSS meta-lab execution remains disabled until repository-derived
|
|
114
|
+
instructions have an isolated credential boundary. Its dry-run and separate
|
|
115
|
+
disposable smoke harness do not open that gate.
|
|
116
|
+
- [Multi-origin shared-world work](https://github.com/danielgwilson/humanish/issues/239)
|
|
117
|
+
needs a real adopter's cross-origin requirement and a reviewed implementation
|
|
118
|
+
packet. It has a ratified core-design direction, but the implementation gate is still closed.
|
|
119
|
+
- [Nested provider grants](https://github.com/danielgwilson/humanish/pull/534)
|
|
120
|
+
remain unmerged. Do not assume a nested provider-credential channel exists.
|
|
121
|
+
- Paused adopter-deletion work stays paused until its current readiness and
|
|
122
|
+
maintainer decisions permit it. Do not resume it from the historical queue.
|
|
123
|
+
- Broader actor/plugin APIs, additional media, registry promotion and TUI
|
|
124
|
+
expansion are follow-ups, not substitutes for a useful maintainer workflow.
|
|
125
|
+
A concrete use case and current issue readiness determine when to take them up.
|
|
126
|
+
|
|
127
|
+
## Safety And Autonomous Work
|
|
128
|
+
|
|
129
|
+
Follow [AGENTS.md](../../AGENTS.md), the [invariants](../principles/invariants-and-defaults.md)
|
|
130
|
+
and the [public-readiness standard](../release/public-readiness-standard.md).
|
|
131
|
+
|
|
132
|
+
- Keep `main` clean and work on scoped branches/worktrees. Substantial work
|
|
133
|
+
needs an issue with scope, authority, required proof and stop conditions.
|
|
134
|
+
- Existing explicit shipping authority governs implementation and merge;
|
|
135
|
+
otherwise issue readiness does not create authority by itself.
|
|
136
|
+
- Never commit secrets, private transcripts/screenshots, customer data or
|
|
137
|
+
private project context. Keep generated proof in ignored `.humanish/` and
|
|
138
|
+
retain needed evidence before removing a worktree.
|
|
139
|
+
- Managed paths bind to validated filesystem identities. Stored provider IDs
|
|
140
|
+
are evidence, not cleanup authority; reclaim only resources the operation is
|
|
141
|
+
authorized to own, and keep unknown cleanup explicitly unresolved.
|
|
142
|
+
- Verification distinguishes `share_ready`, `local_only` and `blocked`.
|
|
143
|
+
Feedback drafts do not mutate GitHub by default. Live spend, publishing,
|
|
144
|
+
external mutation and broader credential access are explicit choices.
|
|
145
|
+
- Provider credential placement is route-specific. The computer-use model key stays
|
|
146
|
+
on the host; the default terminal runtime uses command-scoped credentials. Do not infer safety from an unqualified “keys stay outside” claim.
|
|
147
|
+
|
|
148
|
+
## Proof Before Shipping
|
|
149
|
+
|
|
150
|
+
From a clean contributor worktree:
|
|
141
151
|
|
|
142
152
|
```bash
|
|
153
|
+
pnpm install --frozen-lockfile
|
|
143
154
|
pnpm release:check
|
|
155
|
+
pnpm docs:check
|
|
144
156
|
git diff --check
|
|
145
157
|
```
|
|
146
158
|
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
to start, what exists, what remains, and what proof is required.
|
|
159
|
+
For Observer changes, build before its tests and inspect real bundle data at
|
|
160
|
+
390px and desktop width. Run `pnpm --filter humanish-observer test` and
|
|
161
|
+
`pnpm observer:browser:proof`; preserve the source and limits of each receipt.
|
|
162
|
+
For website changes, run site typecheck, build and `registry:check`.
|
|
163
|
+
Required CI remains the merge gate.
|
|
153
164
|
|
|
154
|
-
|
|
165
|
+
Before a release, follow the [release procedure](../release/open-source-readiness.md),
|
|
166
|
+
including the candidate-package `pnpm release:dogfood` study with authorized
|
|
167
|
+
credentials and bounded paid spend. A successful build alone is not participant
|
|
168
|
+
proof. Verify the published package and any changed publishing surface.
|
|
155
169
|
|
|
156
|
-
|
|
157
|
-
- this page stays current;
|
|
158
|
-
- README links both;
|
|
159
|
-
- release package includes both docs directories.
|
|
160
|
-
|
|
161
|
-
### 3. Fresh-Agent Install Proof
|
|
162
|
-
|
|
163
|
-
Prove the skill and package setup flow from a disposable target app with no chat
|
|
164
|
-
context.
|
|
165
|
-
|
|
166
|
-
Target proof:
|
|
170
|
+
A disposable, keyless consumer check is:
|
|
167
171
|
|
|
168
172
|
```bash
|
|
169
173
|
npm i -D humanish
|
|
170
174
|
npx humanish init --yes
|
|
171
|
-
npx humanish
|
|
175
|
+
npx humanish run first-run --json
|
|
172
176
|
npx humanish verify --run latest --json
|
|
173
177
|
npx humanish feedback issue --run latest --repo owner/repo --format markdown
|
|
174
178
|
```
|
|
175
179
|
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
Graduate from synthetic UI lanes to a real browser journey against a local app.
|
|
181
|
-
|
|
182
|
-
Minimum acceptance:
|
|
183
|
-
|
|
184
|
-
- local app target detection;
|
|
185
|
-
- browser launch;
|
|
186
|
-
- route/state capture;
|
|
187
|
-
- screenshot artifact;
|
|
188
|
-
- run bundle references screenshot evidence;
|
|
189
|
-
- Observer renders the screenshot;
|
|
190
|
-
- `verify` fails closed if required evidence is missing;
|
|
191
|
-
- bounded desktop/mobile two-step browser persona proof with per-step traces and
|
|
192
|
-
screenshots. `done`
|
|
193
|
-
- LLM-driven browser lane: the registered `openai-computer-use` actor dispatches
|
|
194
|
-
from a lab config (`subject.source: app-url`, loopback entry only) into a hosted
|
|
195
|
-
E2B desktop, fills the provider-neutral `stream.actor` trace seam, and persists
|
|
196
|
-
a verified redacted bundle (0.3.0 registered the actor; 0.4.0 made
|
|
197
|
-
`actors[].type` a real dispatch key). `done`
|
|
198
|
-
- Clone subject provider: `subject.source: clone` + `serve` clones a repo INTO the
|
|
199
|
-
sandbox, installs/builds/starts it from config, probes readiness, and records
|
|
200
|
-
provenance (repo, commit, env names) in the bundle — config-only computer-use
|
|
201
|
-
labs against real apps (0.5.0; see `docs/goals/proof-roadmap/goal.md` —
|
|
202
|
-
repo-only, not shipped in the npm package — and
|
|
203
|
-
`docs/principles/invariants-and-defaults.md`, which ships in the package).
|
|
204
|
-
`done`
|
|
205
|
-
- De-paranoia (0.6.0): the redaction redesign + demoted defaults. Screenshots are
|
|
206
|
-
full-fidelity by default (redaction binds the publish boundary, not capture —
|
|
207
|
-
`policies.redactScreenshots` opts back in); `policies.allowPublicTargets` lets an
|
|
208
|
-
owner drive a declared deployment/preview; `subject.clone.keep` is honored on
|
|
209
|
-
failure for debugging; `serve.installTimeoutMs`/`buildTimeoutMs` are configurable
|
|
210
|
-
for monorepo-scale builds. Doctrine updated with the capture-vs-publish rule. This
|
|
211
|
-
re-sequences the proof roadmap: a redaction redesign and an overridable
|
|
212
|
-
public-target policy are prerequisites for any decision-grade depth evidence, so
|
|
213
|
-
they land BEFORE the consumer-web-app / agent-skill depth phases. `done`
|
|
214
|
-
- Device presets (0.6.1): screen/device is a real dimension, with LITERAL values copied
|
|
215
|
-
from the in-house sims (mobile 414×896 … wide 1920×1080; default `desktop` 1440×950) —
|
|
216
|
-
not guessed. `execution.desktop.device` picks the per-run hosted screen; the guessed 1280×800
|
|
217
|
-
is gone. Honest fidelity: on the E2B route only width/height render (real mobile *layout*)
|
|
218
|
-
+ the model is told its device, matching the sims' organic lanes; true touch/DPR/UA needs
|
|
219
|
-
the CDP actor. Per-*persona* device (N×devices) rides fan-out. `done`
|
|
220
|
-
|
|
221
|
-
### 5. Live Terminal And Codex Lanes
|
|
222
|
-
|
|
223
|
-
Make local PTY and Codex-style lanes reliable enough that Observer can show
|
|
224
|
-
running, passed, failed, blocked, and timed-out states without human inference.
|
|
225
|
-
|
|
226
|
-
Minimum acceptance:
|
|
227
|
-
|
|
228
|
-
- sanitized transcript persistence;
|
|
229
|
-
- explicit completion reason;
|
|
230
|
-
- verifier checks redaction status;
|
|
231
|
-
- Observer polling reflects lane completion;
|
|
232
|
-
- no raw private transcript or credential values.
|
|
233
|
-
|
|
234
|
-
Terminal-product real-agent lane (0.8.0; depth-axis layer 6, so an adopter can delete a
|
|
235
|
-
bespoke real-agent sim for humanish + a thin adapter — see
|
|
236
|
-
`docs/goals/terminal-product-lane/goal.md`):
|
|
237
|
-
|
|
238
|
-
- `subject.source: terminal-product` + `execution.target: e2b-terminal` + the registered
|
|
239
|
-
`codex-exec` terminal actor route a config to a real Codex agent studying a product from
|
|
240
|
-
public surfaces inside an E2B shell. `done`
|
|
241
|
-
- The credential-placement inversion is enforced by construction AND by verifier: the runtime
|
|
242
|
-
key is injected ONLY command-scoped into the `codex` invocation, never sandbox-global; a
|
|
243
|
-
deny-by-default allowlist excludes GitHub/payment/deploy/db creds; metadata is a positive
|
|
244
|
-
allowlist; stdin is disabled with an always-present interventions ledger; cleanup is proven
|
|
245
|
-
or the run fails closed. `done`
|
|
246
|
-
- Cost/no-spend ledger with the null-vs-known-zero-vs-absent discipline (unknowns are `null`,
|
|
247
|
-
never guessed); the no-spend proof is DERIVED from the ledger, never asserted; `maxUsd`/
|
|
248
|
-
`maxJobs`/`maxMinutes` caps enforced fail-closed. `done`
|
|
249
|
-
- Product-adapter extension seam: exported contract types + a scorer/feedback DI hook +
|
|
250
|
-
adapter-namespaced product nouns, so an adopter attaches scoring/feedback as a thin
|
|
251
|
-
in-repo extension without forking core. `done`
|
|
252
|
-
- Cleanup is proven BY EXACT CREATED ID: `Sandbox.kill(id)`, confirmed further by
|
|
253
|
-
`Sandbox.getInfo(id)` where the SDK exposes it, and humanish never calls `Sandbox.list`. A live
|
|
254
|
-
rung never needs a dedicated or isolated E2B key; the SAME shared operator key used everywhere
|
|
255
|
-
else in this repo is safe, because humanish only ever reaches a sandbox it created (see
|
|
256
|
-
"The placement rule" corollary in `docs/principles/invariants-and-defaults.md`).
|
|
257
|
-
- LIVE-PROVEN (2026-07-09): a real Codex agent, bootstrapped in a stock E2B shell (Node
|
|
258
|
-
installed in-sandbox, run via `npx -y @openai/codex@latest exec`), studied a public
|
|
259
|
-
agent-CLI product from its declared public surfaces and ran the product's free zero-spend
|
|
260
|
-
guide within `$0` no-spend caps; verdict nonce-verified, cleanup proven BY EXACT ID
|
|
261
|
-
(`getInfo(id)` SandboxNotFoundError, never `Sandbox.list`), verify 15/15, share_ready
|
|
262
|
-
(`docs/goals/terminal-product-lane/receipts/terminal-live-rung-2026-07-09.md`). This closes
|
|
263
|
-
the #159 live-receipt gap. Optional follow-up: a custom image with the agent runtime baked in
|
|
264
|
-
to drop the per-run npx bootstrap. Duplex-PTY/xterm replay is a deferred SLICE 5.
|
|
265
|
-
|
|
266
|
-
Multi-lane fan-out for the computer-use lab (0.9.0; proof-roadmap layer 2, the prerequisite
|
|
267
|
-
for multi-actor shared-state work — #163, see `docs/goals/multi-lane-fanout/goal.md`):
|
|
268
|
-
|
|
269
|
-
- `actors[0].lanes[]` (differentiated roster: per-lane persona/device/starting-surface) XOR
|
|
270
|
-
`actors[0].count` (homogeneous) fan out N independent E2B desktops in ONE run bundle;
|
|
271
|
-
`per-lane worlds` is the only topology this slice (shared-world is #164). `done`
|
|
272
|
-
- `execution.concurrency` bounds in-flight paid desktops (default min(N,3); env may only
|
|
273
|
-
lower it); a pre-flight spend/lane plan prints before any sandbox/provider call and at $0
|
|
274
|
-
in dry-run; per-lane teardown reclaims ONLY each lane's own sandbox by id (never
|
|
275
|
-
account-wide). `done`
|
|
276
|
-
- Proven deterministically (fake substrate: bounded concurrency, by-id teardown, fail-fast,
|
|
277
|
-
hollow-lane caught) AND with a kept live rung (2 lanes, two distinct desktops, both
|
|
278
|
-
reclaimed by id, bundle verifies — `docs/goals/multi-lane-fanout/receipts/`). `done`
|
|
279
|
-
- Deferred: seed-fork provisioning (PR-2), in-process-route fan-out, shared-world topology
|
|
280
|
-
(#164).
|
|
281
|
-
|
|
282
|
-
Shared-world topology — multi-actor against ONE shared mutable service (0.10.0; proof-roadmap
|
|
283
|
-
layer 7; #164; `docs/goals/shared-world-topology/`). The north-star sim leverage: MANY personas,
|
|
284
|
-
ONE shared world.
|
|
285
|
-
|
|
286
|
-
- Sequential (`topology: shared-world`, concurrency 1): one sandbox, N role seats take turns
|
|
287
|
-
against the shared DB; a checkpoint timeline proves role B acted on a world already containing
|
|
288
|
-
role A's mutation. `done`
|
|
289
|
-
- **Concurrent (`topology: shared-world` + `concurrency > 1`): one subject sandbox served +
|
|
290
|
-
`getHost`-exposed, N actor desktop sandboxes drive that one URL SIMULTANEOUSLY** (reuses fan-out
|
|
291
|
-
orchestration; all N+1 reclaimed by id). Honest attribution under concurrency: per-persona
|
|
292
|
-
outcomes + harness-clocked `laneWindows` proving real overlap + a `stateSeries` of the shared
|
|
293
|
-
world under load; causation is structurally inexpressible (independent series, no
|
|
294
|
-
per-delta→actor field). `done`
|
|
295
|
-
- A new `attributionClass: isolated | shared-world` honesty axis + verify FAIL-CLOSED on the
|
|
296
|
-
required/forbidden `attributionLimits` sets + a concurrency-on-pass gate (a passed concurrent
|
|
297
|
-
run must show real overlap AND a state delta coincident with it). `getHost` URLs are
|
|
298
|
-
internet-reachable → the route is gated (verify) to synthetic+seeded subjects; the raw URL is
|
|
299
|
-
digest-only in evidence. `done`
|
|
300
|
-
- **LIVE-PROVEN (0.10.1):** a kept live receipt ran 3 personas concurrently against ONE
|
|
301
|
-
getHost-exposed synthetic plane — all 3 passed, all 3 lane-windows overlapped on the real clock,
|
|
302
|
-
the shared stateSeries evolved under load, N+1=4 sandboxes reclaimed by id, verify ok
|
|
303
|
-
(`docs/goals/shared-world-topology/receipts/concurrent-live-rung-2026-06-17.md`). One trial =
|
|
304
|
-
phase-change proof, not scale. The next step is the real downstream sim migration (a
|
|
305
|
-
synthetic-seeded multi-role app in the adopter's domain). Per-action causation,
|
|
306
|
-
cross-sandbox concurrency beyond getHost, and #108 PII/PHI remain out of scope.
|
|
307
|
-
- shared-world (sequential AND concurrent) now also accepts `subject.source: local-tree` alongside
|
|
308
|
-
`clone`: the ONE subject sandbox packs the operator's own working tree instead of cloning,
|
|
309
|
-
reusing `provisionLocalTreeSubject` from the local-tree keystone (0.14.0). Provenance carries
|
|
310
|
-
`archiveSha256` (the pin - one archive per run, so no per-lane unanimity math applies) plus
|
|
311
|
-
host-side commit/dirty when the packed root is a git work tree; local-tree has no repo/publicRepo
|
|
312
|
-
field. The N actor desktops on the concurrent route still drive the harness-minted getHost URL
|
|
313
|
-
exactly as before; only the subject's provisioning + provenance source changed. The multi-origin
|
|
314
|
-
design (`docs/goals/multi-origin-shared-world/design.md`) remains a separate,
|
|
315
|
-
ratified but implementation-gated downstream slice. It is not part of
|
|
316
|
-
`0.15.3`.
|
|
317
|
-
|
|
318
|
-
Adopter-driven engine features (0.11.0; surfaced by real bespoke-sim migrations):
|
|
319
|
-
|
|
320
|
-
- `execution.desktop.template` — run a lab on a CUSTOM E2B desktop image (name/ID) instead of the
|
|
321
|
-
stock `desktop` template, threaded to `Sandbox.create(template, opts)` via one
|
|
322
|
-
`createDesktopSandbox` seam across every desktop route (cua single+fan-out, sequential +
|
|
323
|
-
concurrent shared-world subject+actors). Absent == the byte-stable stock-template call; recorded
|
|
324
|
-
as `RunBundle.desktopTemplate`. Lets a Node/bun/DB-bearing adopter image run without
|
|
325
|
-
installing the runtime per lane. `done`
|
|
326
|
-
- `humanish observe --run <id>` — serves a run's Observer over `http://127.0.0.1:<port>` (loopback
|
|
327
|
-
only, path-traversal-guarded to the run dir, `/`->`/observer/index.html`) instead of `file://`,
|
|
328
|
-
so browsers/automation can open it and artifact links resolve. `done`
|
|
329
|
-
|
|
330
|
-
Patch hardening (0.11.1):
|
|
331
|
-
|
|
332
|
-
- concurrent shared-world review now fails closed when any actor lane records a failed terminal
|
|
333
|
-
trace; a lane can remain evidence without making the aggregate review green. `done`
|
|
334
|
-
- scripted-browser labs can provision a single cloned synthetic subject, expose it through a
|
|
335
|
-
tokenless sandbox host, and drive deterministic scripted steps while persisting only public-safe
|
|
336
|
-
provenance and host digests. `done`
|
|
337
|
-
|
|
338
|
-
Adopter-driven roster/readback ergonomics (0.12.0):
|
|
339
|
-
|
|
340
|
-
- Lane grouping metadata (`actorType`, `surface`, `caseGroup`) is adapter-owned and projected into
|
|
341
|
-
Observer `laneGroups[]` plus stream labels, so downstream projects can group simulated users
|
|
342
|
-
without teaching Humanish private role names. `done`
|
|
343
|
-
- `actors[0].roster[]` is compact authoring sugar for repeated lane groups. The parser expands it
|
|
344
|
-
into deterministic `lanes[]` (`<group.id>-01`, `<group.id>-02`, ...) before the engine runs, so
|
|
345
|
-
the runtime and run bundle keep one normalized lane shape. `done`
|
|
346
|
-
|
|
347
|
-
Provenance hardening (0.12.1):
|
|
348
|
-
|
|
349
|
-
- Clone-subject provenance now refreshes after successful provisioning phases, so `subject.commit`
|
|
350
|
-
records the served subject HEAD rather than only the initial clone HEAD. This preserves truthful
|
|
351
|
-
run-bundle provenance when an adopter's install/provisioning step checks out the exact revision to
|
|
352
|
-
test. `done`
|
|
353
|
-
|
|
354
|
-
Adapter artifact evidence (0.12.15):
|
|
355
|
-
|
|
356
|
-
- Browser/shared-world adapter hooks may now write product/state proof files under the ignored run
|
|
357
|
-
directory and return namespaced `humanish.adapter-artifact.v1` references. Core validates only the
|
|
358
|
-
generic reference shape and local-path safety, Observer links the artifacts, and `verify` fails
|
|
359
|
-
closed if a referenced file disappears. The payload schema and product nouns stay in the adapter's
|
|
360
|
-
namespace. `done`
|
|
361
|
-
|
|
362
|
-
Evidence hygiene and readback polish (0.12.16):
|
|
363
|
-
|
|
364
|
-
- Browser-backed lanes launch Chromium with shared evidence-hygiene defaults (first-run/update
|
|
365
|
-
background surfaces suppressed, extensions/sync/component update disabled, password/autofill
|
|
366
|
-
profile prompts disabled) so screenshots prefer product pixels over browser chrome. `done`
|
|
367
|
-
- Run-bundle producers now use percent-scale simulation progress consistently: terminal states
|
|
368
|
-
serialize as `100`, and only true in-progress shared-world snapshots serialize partial progress.
|
|
369
|
-
This keeps Observer status pills from rendering completed runs as low-percentage complete states.
|
|
370
|
-
`done`
|
|
371
|
-
- Verify results now separate valid local evidence from public-promotable evidence with
|
|
372
|
-
`shareSafety.status`. Raw full-fidelity screenshot runs remain valid local proof
|
|
373
|
-
(`local_only`), while feedback draft/issue commands require `share_ready` and fail
|
|
374
|
-
closed with structured reasons. `done`
|
|
375
|
-
|
|
376
|
-
Attached CUA live Observer (shipped):
|
|
377
|
-
|
|
378
|
-
- Plain computer-use labs now honor the same attached `onObserverReady` lifecycle as shared-world
|
|
379
|
-
labs: a live CUA run writes an in-progress bundle before actor sessions complete, loopback
|
|
380
|
-
`serveObserver` can hydrate desktop stream iframes while actors are still running, and stream auth
|
|
381
|
-
URLs remain runtime-only through the Observer WeakMap rather than persisted into `run.json` or
|
|
382
|
-
`observer-data.json`. `done`
|
|
383
|
-
|
|
384
|
-
Local working-tree subject + operator observability (0.14.0):
|
|
385
|
-
|
|
386
|
-
- `subject.source: local-tree` packs the lab resolution cwd on the host (git-aware
|
|
387
|
-
enumeration honoring `.gitignore` and including uncommitted work, an always-on
|
|
388
|
-
non-overridable secrets denylist, symlinks stored never dereferenced, one
|
|
389
|
-
enumeration driving both the tar file list and the digest), uploads the
|
|
390
|
-
once-per-run archive into each lane's desktop sandbox, extracts into the subject
|
|
391
|
-
dir, and reuses the clone route's install/build/state/start/probe pipeline
|
|
392
|
-
unchanged. Provenance pins the tree by `archiveSha256` (a dirty tree cannot be
|
|
393
|
-
commit-pinned) plus host-side commit/dirty; `verify` fails closed on a live
|
|
394
|
-
local-tree bundle without a well-formed pin. Live-proven twice with kept
|
|
395
|
-
receipts: a dirty synthetic fixture and this repo packing itself
|
|
396
|
-
(`docs/goals/local-tree-subject/receipts/`). `done`
|
|
397
|
-
- Subject provisioning phase events: started/completed boundaries for
|
|
398
|
-
clone/upload/extract/install/build/serve/readiness/seed-step phases stream to
|
|
399
|
-
stderr by default (injectable via `CuaActorLabHooks.onPhase` /
|
|
400
|
-
`SharedWorldLabHooks.onPhase`) and the completed trail persists into
|
|
401
|
-
`bundle.events`; a single-lane provisioned boot is never silent again. `done`
|
|
402
|
-
- Truthful CLI envelopes at the command boundary: any uncaught action error emits
|
|
403
|
-
one structured `humanish.cli-response.v1` envelope (never a raw stack trace under
|
|
404
|
-
`--json`, never a second stdout document after a flushed envelope), `humanish runs`
|
|
405
|
-
gained a real failure branch, and `doctor` failure now exits 2 like every other
|
|
406
|
-
structured command (behavioral change). `done`
|
|
407
|
-
|
|
408
|
-
### 6. Lab Manifest Shape
|
|
409
|
-
|
|
410
|
-
Make reusable simulations feel like source artifacts, not hardcoded command
|
|
411
|
-
branches.
|
|
412
|
-
|
|
413
|
-
Minimum acceptance:
|
|
414
|
-
|
|
415
|
-
- `humanish/labs/*.yaml` is the committed lab source convention;
|
|
416
|
-
- `.humanish/labs/*.yaml` and `.humanish/local/labs/*.yaml` are ignored local
|
|
417
|
-
overlays;
|
|
418
|
-
- `humanish watch [lab]`, `humanish lab list`, `humanish lab inspect <lab>`, and
|
|
419
|
-
`humanish lab run <lab>` are supported;
|
|
420
|
-
- `--env-file <path>` loads local values for the current command without
|
|
421
|
-
persisting values into artifacts;
|
|
422
|
-
- maintainer dogfood labs such as `oss` are examples, not the canonical
|
|
423
|
-
consumer taxonomy.
|
|
424
|
-
|
|
425
|
-
### 7. OSS Lab Health Readback
|
|
426
|
-
|
|
427
|
-
Make the maintainer `oss` lab report nested lane health back into the
|
|
428
|
-
top-level Observer instead of relying on a human watching the desktops.
|
|
429
|
-
|
|
430
|
-
The current safety boundary above governs this lane. The completed bullets below
|
|
431
|
-
record prior capability and evidence shape; they do not mean the live
|
|
432
|
-
entrypoint is currently enabled.
|
|
433
|
-
|
|
434
|
-
Minimum acceptance:
|
|
435
|
-
|
|
436
|
-
- each lane records setup status; `done`
|
|
437
|
-
- each lane records target app status/URL or blocker; `done`
|
|
438
|
-
- each lane records nested Observer presence; `done`
|
|
439
|
-
- each lane records nested verification status or blocker; `done`
|
|
440
|
-
- each lane records setup-quality filesystem evidence and Observer can inspect
|
|
441
|
-
it; `done`
|
|
442
|
-
- top-level Observer updates lane verdicts from evidence; `done`
|
|
443
|
-
- feedback candidates are derived from setup-quality/actor evidence; `done`
|
|
444
|
-
- Codex app-server actor telemetry is persisted as redacted trace, event, and
|
|
445
|
-
transcript artifacts; `done`
|
|
446
|
-
- each lane receives a meaningful-use score over setup, filesystem, nested
|
|
447
|
-
Humanish proof, actor activity, product surface, and feedback; `done`
|
|
448
|
-
- provider-backed nested app-url proof now drives a bounded two-step
|
|
449
|
-
desktop/mobile browser persona journey in a headed E2B lane; `done`
|
|
450
|
-
- app-specific executable browser steps can now be authored under
|
|
451
|
-
`humanish/scenarios/*.yaml` and are summarized into top-level nested proof
|
|
452
|
-
evidence; `done`
|
|
453
|
-
- repeated public app/tool headed proofs with app-specific manifests have passed
|
|
454
|
-
against two public targets; `done`
|
|
455
|
-
- next gap: richer multi-step product journeys and broader multi-persona
|
|
456
|
-
matrices.
|
|
457
|
-
|
|
458
|
-
## Non-Goals
|
|
459
|
-
|
|
460
|
-
Do not make these default behavior:
|
|
461
|
-
|
|
462
|
-
- live provider spend;
|
|
463
|
-
- GitHub API mutation;
|
|
464
|
-
- hosted queues, databases, or webhooks;
|
|
465
|
-
- production deploys;
|
|
466
|
-
- real customer/user/patient data;
|
|
467
|
-
- private screenshots or raw transcripts;
|
|
468
|
-
- private upstream artifacts.
|
|
469
|
-
|
|
470
|
-
Maintainer-only tooling can exist later, but it must be opt-in, token-explicit,
|
|
471
|
-
and dry-run-first.
|
|
472
|
-
|
|
473
|
-
## Drift Alarms
|
|
474
|
-
|
|
475
|
-
Stop and correct course if:
|
|
476
|
-
|
|
477
|
-
- docs start depending on chat memory;
|
|
478
|
-
- Observer gets prettier without stronger evidence;
|
|
479
|
-
- feedback drafts imply product proof from synthetic contract proof;
|
|
480
|
-
- tests pass while generated artifacts are not inspectable;
|
|
481
|
-
- actor setup/use trials produce findings that never become feedback candidates;
|
|
482
|
-
- live labs require private infrastructure to look impressive;
|
|
483
|
-
- package docs link to files that are not shipped;
|
|
484
|
-
- public-safety gates become optional.
|
|
485
|
-
|
|
486
|
-
## Best Next Work
|
|
487
|
-
|
|
488
|
-
**2026-09-06 (0.83.1).** Desktop CLI studies without a declared product install
|
|
489
|
-
now prepare Node/npm before participant entry, while leaving the product
|
|
490
|
-
uninstalled. Two real stock desktops began without the runtime and reached the
|
|
491
|
-
entry hook with Node/npm/npx available in ordinary and sudo shells. The hook
|
|
492
|
-
then stopped deliberately without starting a participant or model. See the
|
|
493
|
-
[runtime conformance receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/desktop-cli-runtime-2026-09-06.md).
|
|
494
|
-
|
|
495
|
-
Terminal startup failures preserve confirmed cleanup from the desktop startup
|
|
496
|
-
guard. An existing empty terminal event log is valid only when both embedded
|
|
497
|
-
and retained terminal traces declare zero events. Unknown cleanup stays
|
|
498
|
-
unproven, and a failed startup remains a failed run. The compiled CLI regression
|
|
499
|
-
checks verified failure evidence and natural process exit using the installed
|
|
500
|
-
SDK's debug constructor with network access blocked. Two hosted startup-fault
|
|
501
|
-
controls also verified failed bundles, prompt natural exit and exact sandbox
|
|
502
|
-
absence with real provider transports. See the
|
|
503
|
-
[startup evidence receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/terminal-startup-evidence-2026-09-06.md).
|
|
504
|
-
|
|
505
|
-
**2026-09-06 (0.83.0).** First-party OpenAI computer-use actors accept
|
|
506
|
-
an optional per-response `maxOutputTokens` setting. Two independent lanes and two
|
|
507
|
-
sequential shared-world roles forwarded the declared limit in real requests; all
|
|
508
|
-
four preserved provider truncation as incomplete, with zero actions or closing
|
|
509
|
-
requests. The same scoped integration confirmed sequential actor-level reasoning
|
|
510
|
-
effort and a lane override on the request and trace. This proves configuration
|
|
511
|
-
propagation, not participant success or persona efficacy. See the
|
|
512
|
-
[four-role live receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/output-token-limit-2026-09-06.md).
|
|
513
|
-
|
|
514
|
-
Hosted browser setup checks physical X window bounds before participant actions.
|
|
515
|
-
A clipped window gets one correction and a measured readback; a window that
|
|
516
|
-
remains clipped stops the lane. Two hosted fault-injection probes restored a
|
|
517
|
-
bottom-edge button and clicked it successfully. This checks desktop containment,
|
|
518
|
-
not mobile gesture fidelity. See the [desktop geometry receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/desktop-geometry-2026-09-06.md).
|
|
519
|
-
|
|
520
|
-
Recovery hints now allow task-directed waiting and suggest recovery without
|
|
521
|
-
asking a participant to abandon early. Deterministic delayed-start tests cover
|
|
522
|
-
the prompt change and unchanged hard idle limits; live multiplayer patience
|
|
523
|
-
has not been measured. Lane and roster-group objects reject unknown fields
|
|
524
|
-
before allocation, so misspelled controls cannot silently become defaults.
|
|
525
|
-
|
|
526
|
-
**2026-09-05 (0.82.1).** Computer-use actors now preserve a provider output-limit
|
|
527
|
-
interruption as incomplete instead of treating a response without actions as
|
|
528
|
-
success. Two captured live Responses API shapes reproduce the old false pass
|
|
529
|
-
and verify the correction; other explicit non-completed statuses also fail
|
|
530
|
-
closed. Usage and partial text remain in the run, with no execution of actions
|
|
531
|
-
from an interrupted response. See the [provider-limit receipt](computer-use-actor/receipts/provider-token-limit-2026-09-05.md).
|
|
532
|
-
|
|
533
|
-
**2026-09-05 (0.82.0).** Computer-use desktop estimates now use observed CPU and RAM,
|
|
534
|
-
with missing resources and incomplete lifetimes left unknown (#687). Repeated clicks avoid
|
|
535
|
-
the redundant cursor move when a fresh position check confirms the pointer is already there
|
|
536
|
-
(#685). Terminal studies accept an exact Codex package version and forward the declared model
|
|
537
|
-
and reasoning effort; evidence records the executed version without presenting runtime defaults
|
|
538
|
-
as observed model usage (#688). Local-path redaction preserves nested terminal JSON framing,
|
|
539
|
-
and the release report reader exposes skipped malformed lines (#686). Observer cards and reports
|
|
540
|
-
show typed participant blockers as Blocked while retaining the original protocol trace (#691).
|
|
541
|
-
|
|
542
|
-
The new [TodoMVC comparison](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/todomvc-edit-confirmation-2026-09-05.md)
|
|
543
|
-
connects a keyboard blocker to a reproducible local patch: uninterrupted keyboard completion was
|
|
544
|
-
0/2 before and 2/2 after, with one provider interruption in each version retained in the twelve
|
|
545
|
-
attempts. This descriptive synthetic comparison does not establish human completion rates.
|
|
546
|
-
The public field notes and shorter reference-linked README make the method easier to inspect
|
|
547
|
-
(#673, #682, #689).
|
|
548
|
-
|
|
549
|
-
**2026-09-05 (0.81.0).** First use now distinguishes the free evidence preview
|
|
550
|
-
from a live participant study, with concise successful setup output and complete JSON details
|
|
551
|
-
(#660). The website has runnable docs and a generated CLI reference (#661, #668). Retained participant
|
|
552
|
-
reports survive automatic stops. Supported providers with retained history can add one closing
|
|
553
|
-
report when time and known budget remain; clean reports such as “no confusion or hesitation”
|
|
554
|
-
stay clean (#658, #670, #671). Desktop startup cleanup retains acquired handles and records
|
|
555
|
-
confirmed or unknown reclamation. The SDK's detached screenshot-file cleanup rejection is
|
|
556
|
-
handled (#665, #666).
|
|
557
|
-
Terminal output reconciles SDK callbacks with returned aggregates, preserving legitimate
|
|
558
|
-
repeated lines and usage turns while avoiding doubled capture (#672). Opt-in `openai-egress`
|
|
559
|
-
auth keeps the raw OpenAI runtime key outside the sandbox; every sandbox process can still
|
|
560
|
-
spend through the proxy, so this is not a provider spending limit (#663). Stock terminal
|
|
561
|
-
startup installs a checksum-verified Node archive without refreshing unrelated package mirrors
|
|
562
|
-
and gives newly installed npm a default global prefix on the standard PATH (#677, #680). Mobile studies warn that desktop pointer-to-touch conversion can change repeated-tap
|
|
563
|
-
behavior; gesture failures require direct or native touch confirmation before app attribution
|
|
564
|
-
(#678).
|
|
565
|
-
|
|
566
|
-
**2026-09-04, night (0.80.0).** The observation window reaches every desktop route: the
|
|
567
|
-
sequential and concurrent shared-world seats forward `dwell` the way they forward `stopWhen`
|
|
568
|
-
(#645), with a plumbing test per route and two live receipts, three participants holding together
|
|
569
|
-
on one shared board whose checkpoint digest stood still until they acted, and two seats holding
|
|
570
|
-
in turn on one sandbox (`receipts/dwell-window-2026-09-04.md`). The sequential route refuses a
|
|
571
|
-
sandbox request over the provider's 60-minute cap before any call and names the per-role ceiling
|
|
572
|
-
(#649); the provider's refusal had surfaced as a bundle that failed verification. The planted-defect
|
|
573
|
-
benchmark ran with a second brain, the operator's own Claude Code as the participant: 14 of 15
|
|
574
|
-
reported, nothing invented in three clean runs, 72 of 75 across both brains
|
|
575
|
-
(`bench/RESULTS-2026-09-04-claude.md`). A two-minute window and a camera feed of the adopter's own
|
|
576
|
-
are live receipts too.
|
|
577
|
-
|
|
578
|
-
**2026-09-04, evening (0.79.0).** Two of the operator's own asks, each with a live receipt. A
|
|
579
|
-
declared observation window (#510): `dwell` on an actor or a lane holds the page once its condition
|
|
580
|
-
matches (or from the start), captures a frame on a cadence, takes no action and requests no model
|
|
581
|
-
turn, then hands control back with a hint or ends the session; the window is recorded in the trace
|
|
582
|
-
as deliberate and never outlasts the session budget. On TodoMVC the window opened when "1 item
|
|
583
|
-
left" matched, held 60.6 s, captured six frames with no model turn inside, and the participant then
|
|
584
|
-
added the second task (`receipts/dwell-window-2026-09-04.md`). The first live attempt found the
|
|
585
|
-
option dropped between the lane and the loop by a spread the type checker cannot see; a plumbing
|
|
586
|
-
test now pins it. A participant with a camera (#509, first slice): `execution.desktop.media.camera`
|
|
587
|
-
gives a hosted Chrome lane a capture device (ffmpeg's test pattern generated in the sandbox, or a
|
|
588
|
-
`.y4m` of the adopter's) behind the browser's own permission dialog by default;
|
|
589
|
-
`policies.mediaPermission: granted` bypasses it; the bundle records the feed and the flags under
|
|
590
|
-
`desktopBrowser.media`; a microphone is refused without an image that has an audio stack, before
|
|
591
|
-
any spend. Live, Chrome raised its real dialog with a preview of the feed, the participant chose
|
|
592
|
-
"Allow this time" and read back 640x480 (`receipts/participant-camera-2026-09-04.md`). Under
|
|
593
|
-
the then-shipped mobile-emulation input path, TodoMVC rename had stopped 3 of 3 phone
|
|
594
|
-
participants at this checkpoint; Excalidraw read 12 of 12. The 2026-09-05 input-conformance
|
|
595
|
-
correction above qualifies attribution of those mobile failures.
|
|
596
|
-
|
|
597
|
-
**2026-09-04, later (0.78.0).** `@e2b/desktop` moved from 2.2.3 to 2.3.3 (#638): its 2.3.1
|
|
598
|
-
changelog names the socket #581 found today, a background command's event stream the SDK kept
|
|
599
|
-
open after `Xvfb` and `startxfce4` were launched, which held the CLI process alive twelve minutes
|
|
600
|
-
past a run's written result. `doctor` now prints the installed desktop SDK version on its row and
|
|
601
|
-
attaches an advisory with the fix below 2.3.1 (#639). A live try-live on the new SDK reached its
|
|
602
|
-
goal in 104 s and settled with no active resources.
|
|
603
|
-
|
|
604
|
-
**2026-09-04 (0.77.0).** Four fixes found by runs and one by a scanner. A phone-emulated lane now
|
|
605
|
-
follows the participant into a tab that opens later: the hold-mode applier attaches to every page
|
|
606
|
-
target Chrome creates, paused before its first navigation, so the new tab lays out at the phone
|
|
607
|
-
width, and the state observer reads each new tab's own fidelity report, recording a covered tab
|
|
608
|
-
under `desktopGeometry.fidelity.laterTargets` (width, DPR, touch points) and a drifted one as a
|
|
609
|
-
lane warning with the number the page gave; the applier's own log travels as `holderLog`. The
|
|
610
|
-
first live proof failed in the participant's hands (a paused popup never resumed) and the shipped
|
|
611
|
-
design is the one four live runs settled on: never pause a later tab, send the overrides the
|
|
612
|
-
moment it exists, reload it once after its first navigation commits
|
|
613
|
-
(`docs/goals/computer-use-actor/receipts/later-tab-emulation-2026-09-04.md`, #623, #636). A sandbox create
|
|
614
|
-
or archive upload that fails on a transient provider error is retried once and named on the lane
|
|
615
|
-
(#630): six lanes started 20 s apart lost five to E2B SDK errors before any turn, a probe of the
|
|
616
|
-
same SDK a minute later worked in 6 s. The provider's `Retry-After` now governs a retry (capped
|
|
617
|
-
at 60 s) and a `403 misalignment_policy_violation` ends the lane named and unresumed (#633). The
|
|
618
|
-
provisioned-clone CI flake runs on an injected clock (#276 closed). The fourth benchmark run on
|
|
619
|
-
this build read 15 of 15 planted defects with nothing invented in three clean runs, cumulative 58
|
|
620
|
-
of 60 (`bench/RESULTS-2026-09-04-0.76.0.md`). humanish.dev negotiates `Accept: text/markdown`
|
|
621
|
-
and llms.txt says when to use humanish (#634).
|
|
622
|
-
|
|
623
|
-
## Best Next Work
|
|
180
|
+
This checks installation, preview evidence and feedback generation. A live
|
|
181
|
+
claim needs actual participant execution, retained evidence and verified
|
|
182
|
+
cleanup. Use N > 1 where the claim needs replication; report cost estimates,
|
|
183
|
+
unknowns and failures without turning them into zeros or successes.
|
|
624
184
|
|
|
625
|
-
|
|
626
|
-
|
|
627
|
-
CSS viewport, the preset's device pixel ratio, touch events and a mobile user agent before the
|
|
628
|
-
participant's first observation, and the bundle records what the page then reported about itself
|
|
629
|
-
under `desktopGeometry.fidelity` (#221, `docs/goals/computer-use-actor/receipts/mobile-emulation-2026-09-03.md`).
|
|
630
|
-
On the published 0.76.0, neither phone participant could rename through the emulated input path,
|
|
631
|
-
where the 500 px runs without touch had finished. Those are historical instrument observations:
|
|
632
|
-
the 2026-09-05 conformance check found SDK double click and direct touch differ on the original
|
|
633
|
-
editor, so the earlier outcome does not establish a touch-device app defect.
|
|
634
|
-
The persona axis was replicated the same evening on drawDB and TodoMVC and given a third app,
|
|
635
|
-
Excalidraw, as the clean control (`persona-axis-phone-2026-09-03.md`); a multi-lane study's
|
|
636
|
-
second and third findings reach `feedback draft` through `--candidate` (#609); a negated report
|
|
637
|
-
word no longer counts as friction (#614); and the adopter metric excludes this project's own
|
|
638
|
-
checkout wherever its CLI is spawned from (#611). Three host-timing test flakes were made
|
|
639
|
-
load-independent.
|
|
640
|
-
|
|
641
|
-
**2026-09-03 (0.75.0).** Three fixes found by runs, each with a receipt. The DevTools probe behind
|
|
642
|
-
every url, page-text and CSS-viewport observation ran on Node inside the sandbox, and the stock
|
|
643
|
-
desktop has none, so on the app-url route and on any subject served by something other than Node
|
|
644
|
-
every `urlIncludes` / `textIncludes` stop condition and task criterion had been silently blind; it
|
|
645
|
-
runs on python3 now, a dark channel is named in the lane warnings, and the same lab reads
|
|
646
|
-
"never measured in 2" on 0.74.0 and 2/2 at turn 0 on 0.75.0
|
|
647
|
-
(`docs/goals/computer-use-actor/receipts/task-observation-514-2026-09-03.md`, #514). A subject
|
|
648
|
-
install or runtime bootstrap that exits non-zero is retried once, and the error leads with a line a
|
|
649
|
-
person can act on (#602). The gpt-5.6-sol rates were re-pinned at the live promotional sheet
|
|
650
|
-
(4 / 0.40 / 5 / 20 per 1M through at least 2026-11-21), so estimates stop running 25-33% high, and
|
|
651
|
-
gpt-6-astra is priced. #548 closed on the benchmark receipts.
|
|
652
|
-
|
|
653
|
-
**2026-09-01 (0.66.0 through 0.70.0, one session).** The product got its first efficacy numbers and
|
|
654
|
-
they are in the repo with run ids: planted-defect recall 43 of 45 over three benchmark runs and
|
|
655
|
-
12 clean runs with nothing invented (`bench/`; 58 of 60 and 15 clean runs after the 09-04 rerun), 5 of 6 findings confirmed on TodoMVC and 11 of 12
|
|
656
|
-
on drawDB against the source (20 participants on apps we did not write, 0 false reports), and the
|
|
657
|
-
comparison of keyboard and pointer use on two apps: every keyboard-first participant hit a defect no
|
|
658
|
-
mouse-driving participant met (drawDB's mouse-only database modal, TodoMVC's double-click-only
|
|
659
|
-
rename), receipts under `docs/goals/computer-use-actor/receipts/`. These observations show
|
|
660
|
-
sensitivity to explicit keyboard-versus-pointer instructions; they do not establish incremental
|
|
661
|
-
benefit from persona prompting over a matched generic tester. Five cold installs of two
|
|
662
|
-
published versions reached the goal in under three minutes each. On the way the runs found and
|
|
663
|
-
the releases fixed: telemetry that never said what a study was (and stored IPs), a refused lane
|
|
664
|
-
written up as a pass (#476), a blocker regex that refused five of five finished runs (#565, then
|
|
665
|
-
the structural #570: the participant declares its own outcome, 12 of 12 adherence), a
|
|
666
|
-
study-participant marker nothing set (#546), hung turns that ate a lane's budget (#469, #480),
|
|
667
|
-
and a Claude participant with no memory across turns (#520). `humanish stats` (#472) and
|
|
668
|
-
`humanish export` (#471) exist as CLI commands.
|
|
669
|
-
|
|
670
|
-
The next outcome to establish is an external maintainer using Humanish on their
|
|
671
|
-
own app, finding a useful problem, repairing it and retaining a comparable rerun.
|
|
672
|
-
The published studies establish synthetic capability on their named subjects;
|
|
673
|
-
they do not establish external adoption or incremental persona value.
|
|
674
|
-
[Runnable documentation](https://humanish.dev/docs) and the repair-study recipe
|
|
675
|
-
are available; the documentation gap from #513 is closed.
|
|
676
|
-
|
|
677
|
-
Remaining engineering work includes:
|
|
678
|
-
|
|
679
|
-
- #581: #665 already reclaims acquired desktop instances when startup fails,
|
|
680
|
-
and #666 handles detached screenshot cleanup errors. The compiled CLI now has
|
|
681
|
-
deterministic exit proof and two hosted startup-fault receipts. Remaining
|
|
682
|
-
boundaries are allocation ambiguity before instance acquisition, unconfirmed
|
|
683
|
-
reclamation, and exit behavior after other provider/network failures.
|
|
684
|
-
The 2.3.3 SDK update fixed the identified background stream holder; older
|
|
685
|
-
optional peers still receive an advisory, so a universal no-linger claim
|
|
686
|
-
would be too broad.
|
|
687
|
-
- #509: microphone support requires a desktop image with an audio stack and
|
|
688
|
-
the matching launch path. The unsupported declaration remains rejected.
|
|
689
|
-
- #221 and #676: real-device and touch-input fidelity beyond viewport emulation;
|
|
690
|
-
#623: scripted-browser support beyond its current model-free route.
|
|
691
|
-
- TUI views over `stats` and `export` (#455), registry promotions (#431), and
|
|
692
|
-
the remaining shared-world evidence work (#365, #446).
|
|
693
|
-
|
|
694
|
-
Earlier state, kept for the record:
|
|
695
|
-
|
|
696
|
-
The bounded public-proof side task is done: a verified, legible four-persona
|
|
697
|
-
Observer hero from a commit-pinned public application (drawDB) shipped in
|
|
698
|
-
`0.16.0`. That run treats the application as a study subject, not a Humanish
|
|
699
|
-
adopter, and does not imply endorsement; it does not satisfy the depth gate.
|
|
700
|
-
|
|
701
|
-
The operator-pinned train-of-thought thread (#427, both stages) and the
|
|
702
|
-
#441 queue head are done as of `0.47.0`, live-proven the same day they
|
|
703
|
-
shipped: trace items carry recording stamps and structured click
|
|
704
|
-
coordinates; the live path flushes `liveActor` partials mid-run so the
|
|
705
|
-
attached Observer's timeline grows while a study runs; playback runs at
|
|
706
|
-
the participant's recorded pace when stamps exist (and the transport says
|
|
707
|
-
which clock it is on); the participant card's decide-line ticks the
|
|
708
|
-
newest reported thought during a live lane; and a live-shaped golden cut
|
|
709
|
-
from a kept receipt run pins the whole shape (`?fixture=live`). On the way
|
|
710
|
-
the capture exposed and closed two evidence-quality defects: the persona
|
|
711
|
-
identity leak (#452) and the credible-pass guard's incentive inversion
|
|
712
|
-
against post-success defect reports (#453 — the verdict scan and the
|
|
713
|
-
friction tally are now separate reads of the same narrative). Shared-world
|
|
714
|
-
roll-up honesty (#364) closed via an external contribution that separates
|
|
715
|
-
credibility, mission endpoint, and convergence into three explicit claims.
|
|
716
|
-
|
|
717
|
-
Deep links landed in `0.48.0` (#464): every participant and frame is
|
|
718
|
-
addressable (`#/lane/<id>/f/<n>`), Back/Forward restore the view, and a
|
|
719
|
-
reload or shared link lands on the exact moment — #441 closed entirely.
|
|
720
|
-
|
|
721
|
-
The npx-first-try adoption cluster closed in `0.49.0`, operator-prompted and
|
|
722
|
-
adversarially red-teamed before merge (both arcs): provider keys now resolve
|
|
723
|
-
through each vendor's native chain (#436 — the documented project overlay,
|
|
724
|
-
`e2b auth login`'s store, `gh auth token`, and a `humanish keys set` user
|
|
725
|
-
store; fills announced by name and source, never value; `HUMANISH_STRICT_KEYS=1`
|
|
726
|
-
opts out) and #346 closed on its receipts. The computer-use default moved to
|
|
727
|
-
`gpt-5.6-sol` with the whole 5.6 family priced, and the cost estimate now
|
|
728
|
-
models the two billing mechanics 5.6 introduced — cache writes at 1.25x and
|
|
729
|
-
long-context re-tiering — exactly, from a new per-request usage ledger on the
|
|
730
|
-
trace (#334). Both spend caps price through the same tier-aware estimator.
|
|
731
|
-
|
|
732
|
-
The standing queue, in rough order:
|
|
733
|
-
|
|
734
|
-
1. registry promotions: the wordmark (#431), the participant card, and the two
|
|
735
|
-
vendored Base UI wrappers (drawer, popover) once a second surface consumes
|
|
736
|
-
them;
|
|
737
|
-
2. shared-world honesty, remaining half: per-action evidence (#365) and
|
|
738
|
-
exposure-flag coverage (#446);
|
|
739
|
-
3. the stakeholder TUI (#455): research + token-translated mocks on the
|
|
740
|
-
operator review surface, design sign-off gated before any code.
|
|
741
|
-
|
|
742
|
-
The depth-axis deletion front (an adopter's bespoke terminal-product sim,
|
|
743
|
-
comparator contract posted on the adopter's tracker) is paused awaiting the
|
|
744
|
-
operator's decision-ledger sign-off, deliberately: it resumes on that
|
|
745
|
-
sign-off, not by default. The multi-actor shared-state adopter gate (#166)
|
|
746
|
-
needs one live rerun of its replacement-critical family on current humanish;
|
|
747
|
-
the recon (family, runner pin) is recorded and it runs as its own arc.
|
|
748
|
-
|
|
749
|
-
The version-pinned README hero is the drawDB real-application study: a live
|
|
750
|
-
four-persona capture that proves package/Observer rendering, public-safe asset
|
|
751
|
-
delivery, and real-application evidence against a studied subject. It is not
|
|
752
|
-
adopter deletion evidence, which is a separate and higher gate.
|
|
185
|
+
Start from the [ramp](../ramp/README.md), choose one changed user outcome, and
|
|
186
|
+
close work with what changed, what was checked and what remains uncertain.
|