humanish 0.96.1 → 0.98.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (193) hide show
  1. package/AGENTS.md +86 -79
  2. package/CONTRIBUTING.md +7 -2
  3. package/README.md +11 -2
  4. package/dist/actor-contract.d.ts +35 -1
  5. package/dist/actor-contract.js +38 -0
  6. package/dist/actor-contract.js.map +1 -1
  7. package/dist/adapter-extension.js +1 -0
  8. package/dist/adapter-extension.js.map +1 -1
  9. package/dist/automatic-analysis-config.d.ts +15 -5
  10. package/dist/automatic-analysis-config.js +25 -4
  11. package/dist/automatic-analysis-config.js.map +1 -1
  12. package/dist/automatic-study-analysis.js +4 -2
  13. package/dist/automatic-study-analysis.js.map +1 -1
  14. package/dist/browser-control-client.d.ts +14 -0
  15. package/dist/browser-control-client.js +134 -0
  16. package/dist/browser-control-client.js.map +1 -0
  17. package/dist/browser-control-dispatcher.d.ts +14 -0
  18. package/dist/browser-control-dispatcher.js +109 -0
  19. package/dist/browser-control-dispatcher.js.map +1 -0
  20. package/dist/browser-control-protocol.d.ts +371 -0
  21. package/dist/browser-control-protocol.js +155 -0
  22. package/dist/browser-control-protocol.js.map +1 -0
  23. package/dist/browser-control-transport.d.ts +24 -0
  24. package/dist/browser-control-transport.js +156 -0
  25. package/dist/browser-control-transport.js.map +1 -0
  26. package/dist/comms-lease-store.d.ts +1 -0
  27. package/dist/comms-lease-store.js +9 -3
  28. package/dist/comms-lease-store.js.map +1 -1
  29. package/dist/computer-use-actor.d.ts +2 -2
  30. package/dist/computer-use-actor.js +6 -1
  31. package/dist/computer-use-actor.js.map +1 -1
  32. package/dist/computer-use.d.ts +23 -1
  33. package/dist/computer-use.js +253 -70
  34. package/dist/computer-use.js.map +1 -1
  35. package/dist/cua-actor-lab.d.ts +24 -289
  36. package/dist/cua-actor-lab.js +203 -1983
  37. package/dist/cua-actor-lab.js.map +1 -1
  38. package/dist/cua-desktop-lane.d.ts +35 -0
  39. package/dist/cua-desktop-lane.js +13 -0
  40. package/dist/cua-desktop-lane.js.map +1 -0
  41. package/dist/cua-executor-error.d.ts +31 -0
  42. package/dist/cua-executor-error.js +48 -0
  43. package/dist/cua-executor-error.js.map +1 -0
  44. package/dist/cua-provider-error.d.ts +12 -0
  45. package/dist/cua-provider-error.js +28 -0
  46. package/dist/cua-provider-error.js.map +1 -0
  47. package/dist/desktop-session.d.ts +41 -0
  48. package/dist/desktop-session.js +46 -0
  49. package/dist/desktop-session.js.map +1 -0
  50. package/dist/doctor-lab.d.ts +8 -1
  51. package/dist/doctor-lab.js +40 -8
  52. package/dist/doctor-lab.js.map +1 -1
  53. package/dist/e2b-cua-desktop.d.ts +3 -0
  54. package/dist/e2b-cua-desktop.js +675 -0
  55. package/dist/e2b-cua-desktop.js.map +1 -0
  56. package/dist/e2b-cua-provisioning.d.ts +311 -0
  57. package/dist/e2b-cua-provisioning.js +1213 -0
  58. package/dist/e2b-cua-provisioning.js.map +1 -0
  59. package/dist/e2b-desktop-executor.d.ts +1 -24
  60. package/dist/e2b-desktop-executor.js +2 -127
  61. package/dist/e2b-desktop-executor.js.map +1 -1
  62. package/dist/e2b-desktop-session.d.ts +7 -0
  63. package/dist/e2b-desktop-session.js +29 -0
  64. package/dist/e2b-desktop-session.js.map +1 -0
  65. package/dist/e2b-terminal-lab.js +1 -0
  66. package/dist/e2b-terminal-lab.js.map +1 -1
  67. package/dist/frame-signature.d.ts +24 -0
  68. package/dist/frame-signature.js +128 -0
  69. package/dist/frame-signature.js.map +1 -0
  70. package/dist/guest-bootstrap.d.ts +43 -0
  71. package/dist/guest-bootstrap.js +240 -0
  72. package/dist/guest-bootstrap.js.map +1 -0
  73. package/dist/guest-browser-tools.d.ts +8 -0
  74. package/dist/guest-browser-tools.js +66 -0
  75. package/dist/guest-browser-tools.js.map +1 -0
  76. package/dist/guest-chromium-text.d.ts +27 -0
  77. package/dist/guest-chromium-text.js +281 -0
  78. package/dist/guest-chromium-text.js.map +1 -0
  79. package/dist/guest-desktop-executor.d.ts +22 -0
  80. package/dist/guest-desktop-executor.js +177 -0
  81. package/dist/guest-desktop-executor.js.map +1 -0
  82. package/dist/guest-desktop-native.d.ts +14 -0
  83. package/dist/guest-desktop-native.js +131 -0
  84. package/dist/guest-desktop-native.js.map +1 -0
  85. package/dist/guest-runtime-desktop.d.ts +35 -0
  86. package/dist/guest-runtime-desktop.js +231 -0
  87. package/dist/guest-runtime-desktop.js.map +1 -0
  88. package/dist/guest-runtime-main.d.ts +1 -0
  89. package/dist/guest-runtime-main.js +31 -0
  90. package/dist/guest-runtime-main.js.map +1 -0
  91. package/dist/guest-runtime-revision.d.ts +1 -0
  92. package/dist/guest-runtime-revision.js +3 -0
  93. package/dist/guest-runtime-revision.js.map +1 -0
  94. package/dist/guest-runtime.d.ts +25 -0
  95. package/dist/guest-runtime.js +96 -0
  96. package/dist/guest-runtime.js.map +1 -0
  97. package/dist/index.d.ts +1 -1
  98. package/dist/lab-config.js +10 -3
  99. package/dist/lab-config.js.map +1 -1
  100. package/dist/lab-engine.js +6 -0
  101. package/dist/lab-engine.js.map +1 -1
  102. package/dist/lab-summary.d.ts +5 -1
  103. package/dist/lab-summary.js +5 -0
  104. package/dist/lab-summary.js.map +1 -1
  105. package/dist/local-agent-cli.js +1 -1
  106. package/dist/local-agent-cli.js.map +1 -1
  107. package/dist/local-firecracker-desktop.d.ts +13 -0
  108. package/dist/local-firecracker-desktop.js +150 -0
  109. package/dist/local-firecracker-desktop.js.map +1 -0
  110. package/dist/local-firecracker-study.d.ts +9 -0
  111. package/dist/local-firecracker-study.js +93 -0
  112. package/dist/local-firecracker-study.js.map +1 -0
  113. package/dist/local-runtime-config.d.ts +6 -0
  114. package/dist/local-runtime-config.js +56 -0
  115. package/dist/local-runtime-config.js.map +1 -0
  116. package/dist/local-runtime-release.d.ts +3 -0
  117. package/dist/local-runtime-release.js +8 -0
  118. package/dist/local-runtime-release.js.map +1 -0
  119. package/dist/local-runtime.d.ts +25 -0
  120. package/dist/local-runtime.js +113 -0
  121. package/dist/local-runtime.js.map +1 -0
  122. package/dist/observer-app.html +4 -4
  123. package/dist/pricing.d.ts +22 -1
  124. package/dist/pricing.js +22 -0
  125. package/dist/pricing.js.map +1 -1
  126. package/dist/program.js +50 -9
  127. package/dist/program.js.map +1 -1
  128. package/dist/restricted-codex-analysis.d.ts +15 -0
  129. package/dist/restricted-codex-analysis.js +13 -0
  130. package/dist/restricted-codex-analysis.js.map +1 -0
  131. package/dist/restricted-codex-participant-policy.d.ts +39 -0
  132. package/dist/restricted-codex-participant-policy.js +69 -0
  133. package/dist/restricted-codex-participant-policy.js.map +1 -0
  134. package/dist/restricted-codex-participant-run.d.ts +20 -0
  135. package/dist/restricted-codex-participant-run.js +78 -0
  136. package/dist/restricted-codex-participant-run.js.map +1 -0
  137. package/dist/restricted-codex-participant.d.ts +14 -0
  138. package/dist/restricted-codex-participant.js +178 -0
  139. package/dist/restricted-codex-participant.js.map +1 -0
  140. package/dist/restricted-codex-policy.d.ts +56 -0
  141. package/dist/restricted-codex-policy.js +151 -0
  142. package/dist/restricted-codex-policy.js.map +1 -0
  143. package/dist/restricted-codex-session.d.ts +19 -0
  144. package/dist/restricted-codex-session.js +413 -0
  145. package/dist/restricted-codex-session.js.map +1 -0
  146. package/dist/restricted-codex-transport.d.ts +58 -0
  147. package/dist/restricted-codex-transport.js +233 -0
  148. package/dist/restricted-codex-transport.js.map +1 -0
  149. package/dist/run-detail.js +4 -2
  150. package/dist/run-detail.js.map +1 -1
  151. package/dist/run.d.ts +12 -5
  152. package/dist/run.js +17 -1
  153. package/dist/run.js.map +1 -1
  154. package/dist/shared-world-lab.js +2 -2
  155. package/dist/shared-world-lab.js.map +1 -1
  156. package/dist/study-analysis-codex-config.d.ts +11 -0
  157. package/dist/study-analysis-codex-config.js +34 -0
  158. package/dist/study-analysis-codex-config.js.map +1 -0
  159. package/dist/study-analysis-engine.d.ts +6 -3
  160. package/dist/study-analysis-engine.js +19 -10
  161. package/dist/study-analysis-engine.js.map +1 -1
  162. package/dist/study-analysis-job.d.ts +3 -2
  163. package/dist/study-analysis-job.js +1 -1
  164. package/dist/study-analysis-job.js.map +1 -1
  165. package/dist/study-analysis-provider.d.ts +4 -2
  166. package/dist/study-analysis-provider.js +1 -1
  167. package/dist/study-analysis-provider.js.map +1 -1
  168. package/dist/study-analysis-service.d.ts +3 -0
  169. package/dist/study-analysis-service.js +24 -6
  170. package/dist/study-analysis-service.js.map +1 -1
  171. package/dist/study-analysis-validation.d.ts +43 -19
  172. package/dist/study-analysis-validation.js +23 -10
  173. package/dist/study-analysis-validation.js.map +1 -1
  174. package/dist/study-analysis.d.ts +27 -2
  175. package/dist/study-costs.js +6 -0
  176. package/dist/study-costs.js.map +1 -1
  177. package/dist/tui-app.js +102 -102
  178. package/docs/architecture/browser-control.md +117 -0
  179. package/docs/architecture/desktop-sessions.md +80 -0
  180. package/docs/architecture/guest-desktop.md +87 -0
  181. package/docs/architecture/local-browser-runtime.md +100 -0
  182. package/docs/architecture/restricted-codex-analysis.md +91 -0
  183. package/docs/architecture/runtime-broker-core.md +30 -0
  184. package/docs/contracts/schemas.md +1 -1
  185. package/docs/contracts/study-analysis.md +44 -2
  186. package/docs/goals/current.md +25 -9
  187. package/docs/product/automatic-analysis.md +24 -3
  188. package/docs/product/open-source-install-experience.md +7 -0
  189. package/docs/ramp/README.md +30 -13
  190. package/docs/release/0.97.0-codex-account-analysis.md +45 -0
  191. package/docs/release/0.98.0-local-browser-studies.md +31 -0
  192. package/package.json +4 -2
  193. package/skills/humanish/SKILL.md +39 -0
@@ -1,9 +1,9 @@
1
1
  # Current Goals
2
2
 
3
- Status date: 2026-09-21. Release baseline: `0.96.1`.
3
+ Status date: 2026-09-24. Release baseline: `0.98.0`.
4
4
 
5
5
  This page guides work on current merged source. Published behavior is described
6
- in the [release notes](../release/0.96.1-browser-navigation.md).
6
+ in the [release notes](../release/0.98.0-local-browser-studies.md).
7
7
  The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
8
8
  preserves the former status log; its queues do not supersede this page.
9
9
 
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
88
88
  No first-party deletion branch has met that gate. Public demonstrations do not
89
89
  substitute for it.
90
90
 
91
- ## Current Program Truth (source `0.96.1`)
91
+ ## Current Program Truth (source `0.98.0`)
92
92
 
93
93
  | Surface | Available in merged source | Remaining boundary |
94
94
  | --- | --- | --- |
@@ -99,7 +99,7 @@ substitute for it.
99
99
  | Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
100
100
  | Observer | Live/recorded views, shared grid and participant playback, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
101
101
  | Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
102
- | Study findings | Default post-run analysis on supported live routes with a separate disclosed $3 admission estimate limit and opt-out; explicit `analyze`, fairer evidence selection, concern review and versioned findings with exact source links | Model interpretation needs review; bounded selection and source truncation limit coverage; opening Observer never dispatches analysis |
102
+ | Study findings | Default post-run analysis on supported live routes with a separate disclosed $3 admission estimate limit and opt-out; explicit `analyze`, fairer evidence selection, concern review and versioned findings with exact source links; explicit restricted Codex account analysis on the qualified Linux profile | Account dollars/output-token caps are unavailable; Mac/keychain/other CLI profiles are unqualified. Model interpretation needs review; selection limits coverage; opening Observer never dispatches analysis |
103
103
  | TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving, run library and AgentMail setup, authentication and lab configuration | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
104
104
  | Off-app communication | Recipient-scoped local capture and fresh real AgentMail receiving, supported inline raster images, bounded collection and host-owned recovery | Real mail uses isolated participant surfaces and remains local-only for publication. Hosted mail/model processing, bounded fidelity and interrupted-run recovery are explicit; local-agent, borrowed inboxes and SMS are unsupported |
105
105
  | Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed; a synthetic video-only call with separate hosted peers is proven | Audio, TURN, provider-specific rooms, physical-device and touch fidelity remain unproven; unsupported media declarations are rejected |
@@ -107,13 +107,24 @@ substitute for it.
107
107
  Use the [task support matrix](../architecture/task-protocol-support.md),
108
108
  [actor registry](https://github.com/danielgwilson/humanish/blob/main/src/actor-registry.ts)
109
109
  and [CLI reference](https://humanish.dev/docs/cli) when choosing a concrete path.
110
- Source behavior and required tests outrank stale status prose.
110
+ Source and tests establish observed behavior. Resolve conflicts with requirements
111
+ explicitly; neither stale status prose nor a passing test makes a bug correct.
111
112
 
112
113
  The library-assisted `local-app` route now includes a
113
114
  [runnable npm example](../architecture/examples/state-driven-local-app/README.md).
114
115
  Its deterministic provider demonstrates the integration with a real loopback
115
116
  app; it does not establish persona effectiveness or independent adoption.
116
117
 
118
+ The [local browser runtime](../architecture/local-browser-runtime.md)
119
+ runs isolated Linux browser participants through the same scheduler, recordings
120
+ and automatic analysis as hosted studies. It uses Docker-owned resources and
121
+ ordinary TAP/NAT networking. Continue managed-local work from this complete study
122
+ path; the earlier offline owner/service qualification experiments are historical
123
+ fixtures, not an installation architecture or a prerequisite queue. Explicit
124
+ Linux local labs now use the installed CLI/TUI, with a verified runtime download
125
+ before the first live run. Mac support, inbox integration and optional media
126
+ remain unfinished. Existing hosted labs retain their behavior.
127
+
117
128
  ## Gates And Deferred Work
118
129
 
119
130
  - Live OSS meta-lab execution remains disabled until repository-derived
@@ -135,10 +146,12 @@ app; it does not establish persona effectiveness or independent adoption.
135
146
  Follow [AGENTS.md](../../AGENTS.md), the [invariants](../principles/invariants-and-defaults.md)
136
147
  and the [public-readiness standard](../release/public-readiness-standard.md).
137
148
 
138
- - Keep `main` clean and work on scoped branches/worktrees. Substantial work
139
- needs an issue with scope, authority, required proof and stop conditions.
149
+ - Keep `main` clean and work on scoped branches/worktrees. Keep the task's scope,
150
+ authority, relevant checks and material failure boundaries in its issue, PR or
151
+ current handoff; do not create a separate packet for routine work.
140
152
  - Existing explicit shipping authority governs implementation and merge;
141
- otherwise issue readiness does not create authority by itself.
153
+ otherwise issue readiness does not create authority by itself. Machine-readiness
154
+ fields gate automated queue pickup, not directly assigned interactive work.
142
155
  - Never commit secrets, private transcripts/screenshots, customer data or
143
156
  private project context. Keep generated proof in ignored `.humanish/` and
144
157
  retain needed evidence before removing a worktree.
@@ -153,7 +166,10 @@ and the [public-readiness standard](../release/public-readiness-standard.md).
153
166
 
154
167
  ## Proof Before Shipping
155
168
 
156
- From a clean contributor worktree:
169
+ Use the [verification guidance](../../AGENTS.md#verification): check the changed
170
+ behavior and material risks, then stop unless new evidence warrants more work.
171
+ Required CI remains the merge gate. For a release, run the full release gates
172
+ from a clean contributor worktree:
157
173
 
158
174
  ```bash
159
175
  pnpm install --frozen-lockfile
@@ -22,7 +22,7 @@ review:
22
22
 
23
23
  Omitting `review.analysis` uses these defaults. Set `review.analysis: false` to
24
24
  run participants without the additional analysis request. An explicit analysis
25
- mapping requires `maxCostUsd`. This limits an admission estimate, not the
25
+ mapping using the default OpenAI API provider requires `maxCostUsd`. This limits an admission estimate, not the
26
26
  provider's final bill, and is separate from participant spending limits. Analysis
27
27
  can decline a large study before dispatch when its conservative estimate exceeds
28
28
  that limit. Use `analyze --dry-run --max-cost <usd>` on retained evidence to inspect
@@ -37,6 +37,27 @@ its existing authentication boundary. Review the separate analysis budget before
37
37
  zero-dollar cap does not cap post-run analysis. The bundled first-contact
38
38
  zero-spend product fixture explicitly disables analysis.
39
39
 
40
+ To explicitly use your Codex ChatGPT account for the separate analyst:
41
+
42
+ ```yaml
43
+ review:
44
+ analysis:
45
+ provider: codex
46
+ model: gpt-6-astra
47
+ timeoutMs: 600000
48
+ ```
49
+
50
+ This requires Linux x64, qualified Codex CLI `0.154.0`, and a file-backed ChatGPT account login; the analyst
51
+ uses low reasoning effort and remote inference. Dollar cost and a provider
52
+ enforced output-token ceiling are unknown, so omit `maxCostUsd` and
53
+ `maxOutputTokens`. Numeric values are rejected before participant resources are
54
+ allocated. There is no API fallback. Missing or unsupported account setup leaves
55
+ an explicit failed analysis state and the original recording intact. Use
56
+ `humanish doctor --lab <lab>` for setup checks; account allowance and model access
57
+ remain untested until a request. An omitted provider still means OpenAI, including
58
+ hosted studies whose participant uses a local Codex or Claude login. This setting
59
+ does not enable managed local desktops.
60
+
40
61
  The same configuration works through `humanish run <lab>`, `lab run <lab>`,
41
62
  `watch <lab>`, and TUI live starts. Direct library calls to the five recording
42
63
  producers honor it too. Supported routes are computer-use, scripted-browser,
@@ -53,7 +74,7 @@ Default analysis also skips recordings containing only setup or failure records
53
74
  with no retained participant activity. A desktop startup failure does not start
54
75
  an analysis request. The original failure remains visible.
55
76
 
56
- CLI live starts disclose the separate admission estimate limit before execution.
77
+ CLI live starts disclose the selected analyst and its separate admission estimate limit or unknown account dollars before execution.
57
78
  `humanish lab preflight <lab> --json` and the TUI lab screen also expose the
58
79
  resolved budget without dispatching analysis. Library callers can inspect
59
80
  `resolveAutomaticAnalysis` or `automaticAnalysisBudget` before running.
@@ -64,7 +85,7 @@ existing job and does not start another request. Concurrent or repeated automati
64
85
  invocations cannot silently retry a paid attempt. If a process disappears while
65
86
  an attempt is in flight, its state can be unknown rather than falsely complete.
66
87
  Use manual `humanish analyze --run <exact-run-id> --max-cost 3` for an intentional
67
- follow-up after inspecting the existing attempt and its accounting.
88
+ follow-up after inspecting the existing attempt and its accounting. For the account branch, use `humanish analyze --run <exact-run-id> --provider codex --rerun` without a dollar limit.
68
89
 
69
90
  Stopping participant execution does not start a fresh automatic analysis. A
70
91
  recorded harness cancellation is skipped; ordinary time limits and participant
@@ -235,3 +235,10 @@ them should install `@e2b/desktop` explicitly instead of receiving that
235
235
  substrate as part of the default Humanish package install. When a GitHub token is
236
236
  present, repo labels are redacted in durable artifacts by default; live stream
237
237
  auth URLs are used only by the attached watch server and are not persisted.
238
+
239
+ An explicitly selected [local browser lab](../architecture/local-browser-runtime.md)
240
+ can instead use Linux x64, Docker/KVM and an existing Codex ChatGPT login for
241
+ participants and findings. The installed CLI prepares its pinned image before
242
+ the first live run. It does not install host prerequisites, and Mac, inbox and
243
+ media integration remain separate follow-ups. Existing hosted lab configuration
244
+ is preserved.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Status: public-safe contributor and agent ramp.
4
4
 
5
- Package/source version in this tree: `0.96.1` (2026-09-21). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
5
+ Package/source version in this tree: `0.98.0` (2026-09-24). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
6
6
  `0.15.1` remains in force: managed run and output paths bind to validated
7
7
  physical filesystem identities, and stored provider IDs are evidence, not
8
8
  cleanup authority. The bundled OSS meta-lab is dry-run only until
@@ -14,19 +14,26 @@ context.
14
14
 
15
15
  ## First Read
16
16
 
17
- Read these in order:
17
+ Start with three things:
18
18
 
19
- 1. [`AGENTS.md`](../../AGENTS.md) for public boundary and engineering rules.
20
- 2. [`docs/principles/invariants-and-defaults.md`](../principles/invariants-and-defaults.md) — which rules are invariants and which are overridable defaults, each with the reason it exists and the check that enforces it. (The enforcement is what makes a rule real here — when a doc sentence and a test disagree, trust the test and say so.)
21
- 3. [`README.md`](../../README.md) for install, commands, and package shape.
22
- 4. [`docs/goals/current.md`](../goals/current.md) for the active product goal.
23
- 5. [`docs/goals/proof-roadmap/goal.md`](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/goal.md) for the ratified proof architecture (repo-only; not shipped in the npm package, hence the absolute link).
24
- 6. [`docs/product/open-source-install-experience.md`](../product/open-source-install-experience.md) for first-run UX.
25
- 7. [`docs/roadmap/world-class-open-source-v0.md`](../roadmap/world-class-open-source-v0.md) for staged delivery history (historical; see its status banner).
26
- 8. [`docs/architecture/observer.md`](../architecture/observer.md) for Observer architecture.
27
- 9. [`docs/contracts/run-bundle.md`](../contracts/run-bundle.md) and [`docs/contracts/policy.md`](../contracts/policy.md) for proof contracts.
28
- 10. [`docs/release/public-readiness-standard.md`](../release/public-readiness-standard.md) before deciding what must be scrubbed.
29
- 11. [`docs/release/open-source-readiness.md`](../release/open-source-readiness.md) before touching public packaging or repository visibility.
19
+ 1. [`AGENTS.md`](../../AGENTS.md) for engineering judgment and public boundaries.
20
+ 2. The current task and [`docs/goals/current.md`](../goals/current.md) for current
21
+ product status. Explicit task direction takes precedence over historical queues.
22
+ 3. Instructions in the component being changed, then its relevant contracts.
23
+
24
+ Use the references below as needed. Historical plans are context, not a backlog
25
+ to resume automatically. Keep one concise current task handoff with the requested
26
+ outcome, demonstrated behavior, next complete result, constraints and rejected or
27
+ deferred approaches; link evidence rather than repeating its chronology.
28
+
29
+ | When working on | Reference |
30
+ | --- | --- |
31
+ | Install, commands or first-run UX | [`README.md`](../../README.md), [install experience](../product/open-source-install-experience.md) |
32
+ | Security, evidence handling or defaults | [Invariants and defaults](../principles/invariants-and-defaults.md) |
33
+ | Observer | [Observer architecture](../architecture/observer.md) and its component instructions |
34
+ | Bundle formats or policy | [Run bundle](../contracts/run-bundle.md), [policy](../contracts/policy.md) |
35
+ | Public artifacts or packaging | [Public-readiness standard](../release/public-readiness-standard.md), [release procedure](../release/open-source-readiness.md) |
36
+ | Proof architecture or historical decisions | [Proof roadmap](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/goal.md), [historical delivery roadmap](../roadmap/world-class-open-source-v0.md) |
30
37
 
31
38
  ## Mental Model
32
39
 
@@ -47,6 +54,16 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
47
54
 
48
55
  ## Current State
49
56
 
57
+ The [0.98.0 release note](../release/0.98.0-local-browser-studies.md) describes
58
+ installed Linux browser studies with managed runtime images, Codex account
59
+ participants and automatic analysis, and shared CLI/TUI setup checks.
60
+
61
+ The [0.97.0 release note](../release/0.97.0-codex-account-analysis.md) describes
62
+ explicit Codex account analysis on a qualified Linux CLI/login profile, with
63
+ separate analyst authority, evidence-linked reports and unknown-dollar accounting.
64
+ Existing API defaults remain unchanged; managed local desktops and Mac account
65
+ analysis are not qualified by this release.
66
+
50
67
  The [0.96.1 release note](../release/0.96.1-browser-navigation.md) describes
51
68
  accurate physical browser measurements and bounded fitting that preserves normal
52
69
  browser controls when they fit. Narrow screens retain the fullscreen fallback.
@@ -0,0 +1,45 @@
1
+ # 0.97.0 — Codex account reports on Linux
2
+
3
+ Saved studies can use an existing Codex ChatGPT login for their separate
4
+ findings report. Select it explicitly:
5
+
6
+ ```bash
7
+ humanish analyze --run latest --provider codex --dry-run
8
+ humanish analyze --run latest --provider codex
9
+ ```
10
+
11
+ For an automatic report after a live study, set:
12
+
13
+ ```yaml
14
+ review:
15
+ analysis:
16
+ provider: codex
17
+ ```
18
+
19
+ The initial qualification is Linux x64, Codex CLI 0.154.0, a file-backed
20
+ ChatGPT login, and `gpt-6-astra` with low reasoning effort. Other platforms,
21
+ CLI versions and keychain-only logins are refused. Selected evidence still goes
22
+ to remote inference. Model access and account allowance depend on the account.
23
+
24
+ Account dollars and the generated-token ceiling are unknown; numeric dollar or
25
+ output-token caps are rejected. Time and byte limits still apply. Token counts
26
+ are retained when available, and interrupted usage remains incomplete. Existing
27
+ OpenAI API analysis defaults and participant setup requirements are unchanged.
28
+
29
+ The analyst has its own conversation and restricted tool configuration.
30
+ Evidence admission, exact source references, secret scrubbing, immutable attempt
31
+ accounting and source-change checks apply to both providers. Cancellation
32
+ preserves recordings and earlier valid findings. Unconfirmed process cleanup or
33
+ unexpected replacement credentials produce a failed attempt and private recovery
34
+ guidance. Historical reports remain readable as execution profiles evolve.
35
+
36
+ Doctor checks the selected analyst's setup without a model turn. Dry-run checks
37
+ local evidence and configuration only; it does not test CLI/login/model access.
38
+ The CLI, TUI and Observer disclose account limits consistently. This release
39
+ does not introduce managed local desktops or qualify Mac account analysis.
40
+
41
+ The [qualification receipt](https://github.com/danielgwilson/humanish/blob/46330116726f74080fa18947c36da4fb4b333805/docs/goals/computer-use-actor/receipts/codex-account-analysis-2026-09-23.md)
42
+ records real screenshot analysis, cache reuse, cancellation, automatic reports
43
+ after desktop cleanup, desktop/phone evidence review, failures and proof limits.
44
+ See [configuration](../product/automatic-analysis.md) and the
45
+ [restricted launcher contract](../architecture/restricted-codex-analysis.md).
@@ -0,0 +1,31 @@
1
+ # 0.98.0 — Local browser studies on Linux
2
+
3
+ Explicit local browser labs now run from the installed CLI and TUI. On Linux
4
+ x64 with Docker/KVM and a supported Codex ChatGPT login, isolated Firecracker
5
+ participants use the normal scheduler, recordings, Observer and automatic
6
+ analysis without E2B or OpenAI API keys. Codex inference remains remote and
7
+ consumes account quota; dollar cost is unknown.
8
+
9
+ `humanish runtime setup` prepares a pinned, verified runtime image. A live local
10
+ lab also prepares it automatically before starting participants. `runtime status`
11
+ and `doctor --lab` inspect readiness without downloading or launching a browser.
12
+ The TUI shows that same runtime status. Existing hosted labs retain their
13
+ configuration, and a missing local prerequisite never selects another provider.
14
+
15
+ The [setup guide](../architecture/local-browser-runtime.md) includes a complete
16
+ manifest and the current limits. This release supports loopback apps, fixed
17
+ 960×720 Chromium desktops, and optional OpenAI API participants. Mac/Lima,
18
+ local inboxes and local camera/microphone integration remain follow-ups.
19
+
20
+ Private state uses a Docker-owned anonymous volume, removed with its container
21
+ on normal close or controller death. Runtime image downloads are checked for
22
+ exact size and SHA-256 before loading; matching source archives and notices
23
+ are distributed separately.
24
+
25
+ Validation includes a packed installation outside the checkout: two concurrent
26
+ Codex participants saved distinct notes in a local app, the app confirmed both
27
+ saves, the recordings verified, and automatic analysis completed without model
28
+ or desktop API keys. Separate checks cover controller death and state-volume
29
+ removal, download integrity/cancellation, missing prerequisites and preserving
30
+ existing hosted configurations. This is a small Linux integration proof, not
31
+ a claim about Mac readiness, high concurrency or real conferencing apps.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.96.1",
3
+ "version": "0.98.0",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [
@@ -76,13 +76,15 @@
76
76
  "tui:smoke": "node scripts/tui-smoke.mjs",
77
77
  "tui:test": "pnpm --filter humanish-tui test",
78
78
  "release:dogfood": "node scripts/release-dogfood.mjs",
79
+ "browser-control:proof": "node scripts/browser-control-proof.mjs",
79
80
  "docs:generate": "tsx scripts/generate-cli-docs.ts",
80
81
  "docs:check": "tsx scripts/generate-cli-docs.ts --check",
81
82
  "observer:browser:proof": "node scripts/observer-browser-proof.mjs",
82
83
  "observer:iframe:proof": "node scripts/observer-iframe-proof.mjs",
83
84
  "observer:chrome:proof": "node scripts/observer-chrome-proof.mjs",
84
85
  "observer:reliability:proof": "node scripts/observer-reliability-proof.mjs",
85
- "tui:connections:proof": "python3 scripts/tui-connections-proof.py"
86
+ "tui:connections:proof": "python3 scripts/tui-connections-proof.py",
87
+ "guest-desktop:proof": "node scripts/guest-desktop-proof.mjs"
86
88
  },
87
89
  "repository": {
88
90
  "type": "git",
@@ -93,6 +93,45 @@ exact returned path, not a basename that could resolve to another manifest.
93
93
  - keep `.env.example` commit-safe and value-free;
94
94
  - never commit generated run bundles.
95
95
 
96
+ ## Choosing a findings analyst
97
+
98
+ Analysis is separate from the participant. Hosted and manual defaults use the
99
+ OpenAI API with its own admission budget. An explicitly local browser study with
100
+ a Codex participant defaults to a separate Codex account analyst. To select that
101
+ restricted Codex ChatGPT account analyst on other supported studies,
102
+ set `review.analysis.provider: codex` or pass `analyze --provider codex` on a
103
+ completed recording. This uses remote inference and the qualified CLI/login,
104
+ not local inference or the participant's existing conversation. See
105
+ [the analysis contract](../../docs/contracts/study-analysis.md) for the current
106
+ CLI/model qualification and setup limits.
107
+
108
+ Do not pass numeric `maxCostUsd`/`maxOutputTokens` or their CLI flags to the
109
+ account branch. It cannot enforce those ceilings and rejects them. Account dollar
110
+ cost remains unknown even when token usage is reported. There is no fallback to
111
+ an API key or another provider. `analyze --dry-run --provider codex` validates
112
+ local evidence/configuration only; `doctor --lab` checks setup without a model
113
+ request. Inspect a failed attempt before explicitly retrying `--provider codex
114
+ --rerun`. Opening Observer never starts analysis.
115
+
116
+ ## Local browser setup
117
+
118
+ On Linux x64 with a local rootful Docker Engine, KVM and TUN, an `app-url` lab can
119
+ set `execution.target: local` and `actors[0].type: local-agent` with
120
+ `localAgent: codex`. It uses the supported Codex ChatGPT login, not E2B or an
121
+ OpenAI API key. Inference is remote and consumes account quota. Existing hosted
122
+ labs stay hosted; never silently change their execution or billing provider.
123
+
124
+ Use `humanish runtime status --json` and `humanish doctor --lab <path> --json`
125
+ for read-only setup inspection. `humanish runtime setup` downloads and verifies
126
+ the pinned runtime; a live local run also prepares it automatically. The normal
127
+ `humanish lab run <path>` command and TUI use the same study runner and Observer.
128
+ See [the complete example and limits](../../docs/architecture/local-browser-runtime.md).
129
+
130
+ Local browsers currently require a loopback app URL with an explicit port above
131
+ 1023, use a 960×720 Chromium desktop, and reject inbox/media declarations. Mac
132
+ setup is not integrated. Do not claim that installing the CLI also installs
133
+ Docker or makes these prerequisites available.
134
+
96
135
  ## Format Stack
97
136
 
98
137
  When creating or editing Humanish files: