humanish 0.96.1 → 0.98.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +86 -79
- package/CONTRIBUTING.md +7 -2
- package/README.md +11 -2
- package/dist/actor-contract.d.ts +35 -1
- package/dist/actor-contract.js +38 -0
- package/dist/actor-contract.js.map +1 -1
- package/dist/adapter-extension.js +1 -0
- package/dist/adapter-extension.js.map +1 -1
- package/dist/automatic-analysis-config.d.ts +15 -5
- package/dist/automatic-analysis-config.js +25 -4
- package/dist/automatic-analysis-config.js.map +1 -1
- package/dist/automatic-study-analysis.js +4 -2
- package/dist/automatic-study-analysis.js.map +1 -1
- package/dist/browser-control-client.d.ts +14 -0
- package/dist/browser-control-client.js +134 -0
- package/dist/browser-control-client.js.map +1 -0
- package/dist/browser-control-dispatcher.d.ts +14 -0
- package/dist/browser-control-dispatcher.js +109 -0
- package/dist/browser-control-dispatcher.js.map +1 -0
- package/dist/browser-control-protocol.d.ts +371 -0
- package/dist/browser-control-protocol.js +155 -0
- package/dist/browser-control-protocol.js.map +1 -0
- package/dist/browser-control-transport.d.ts +24 -0
- package/dist/browser-control-transport.js +156 -0
- package/dist/browser-control-transport.js.map +1 -0
- package/dist/comms-lease-store.d.ts +1 -0
- package/dist/comms-lease-store.js +9 -3
- package/dist/comms-lease-store.js.map +1 -1
- package/dist/computer-use-actor.d.ts +2 -2
- package/dist/computer-use-actor.js +6 -1
- package/dist/computer-use-actor.js.map +1 -1
- package/dist/computer-use.d.ts +23 -1
- package/dist/computer-use.js +253 -70
- package/dist/computer-use.js.map +1 -1
- package/dist/cua-actor-lab.d.ts +24 -289
- package/dist/cua-actor-lab.js +203 -1983
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/cua-desktop-lane.d.ts +35 -0
- package/dist/cua-desktop-lane.js +13 -0
- package/dist/cua-desktop-lane.js.map +1 -0
- package/dist/cua-executor-error.d.ts +31 -0
- package/dist/cua-executor-error.js +48 -0
- package/dist/cua-executor-error.js.map +1 -0
- package/dist/cua-provider-error.d.ts +12 -0
- package/dist/cua-provider-error.js +28 -0
- package/dist/cua-provider-error.js.map +1 -0
- package/dist/desktop-session.d.ts +41 -0
- package/dist/desktop-session.js +46 -0
- package/dist/desktop-session.js.map +1 -0
- package/dist/doctor-lab.d.ts +8 -1
- package/dist/doctor-lab.js +40 -8
- package/dist/doctor-lab.js.map +1 -1
- package/dist/e2b-cua-desktop.d.ts +3 -0
- package/dist/e2b-cua-desktop.js +675 -0
- package/dist/e2b-cua-desktop.js.map +1 -0
- package/dist/e2b-cua-provisioning.d.ts +311 -0
- package/dist/e2b-cua-provisioning.js +1213 -0
- package/dist/e2b-cua-provisioning.js.map +1 -0
- package/dist/e2b-desktop-executor.d.ts +1 -24
- package/dist/e2b-desktop-executor.js +2 -127
- package/dist/e2b-desktop-executor.js.map +1 -1
- package/dist/e2b-desktop-session.d.ts +7 -0
- package/dist/e2b-desktop-session.js +29 -0
- package/dist/e2b-desktop-session.js.map +1 -0
- package/dist/e2b-terminal-lab.js +1 -0
- package/dist/e2b-terminal-lab.js.map +1 -1
- package/dist/frame-signature.d.ts +24 -0
- package/dist/frame-signature.js +128 -0
- package/dist/frame-signature.js.map +1 -0
- package/dist/guest-bootstrap.d.ts +43 -0
- package/dist/guest-bootstrap.js +240 -0
- package/dist/guest-bootstrap.js.map +1 -0
- package/dist/guest-browser-tools.d.ts +8 -0
- package/dist/guest-browser-tools.js +66 -0
- package/dist/guest-browser-tools.js.map +1 -0
- package/dist/guest-chromium-text.d.ts +27 -0
- package/dist/guest-chromium-text.js +281 -0
- package/dist/guest-chromium-text.js.map +1 -0
- package/dist/guest-desktop-executor.d.ts +22 -0
- package/dist/guest-desktop-executor.js +177 -0
- package/dist/guest-desktop-executor.js.map +1 -0
- package/dist/guest-desktop-native.d.ts +14 -0
- package/dist/guest-desktop-native.js +131 -0
- package/dist/guest-desktop-native.js.map +1 -0
- package/dist/guest-runtime-desktop.d.ts +35 -0
- package/dist/guest-runtime-desktop.js +231 -0
- package/dist/guest-runtime-desktop.js.map +1 -0
- package/dist/guest-runtime-main.d.ts +1 -0
- package/dist/guest-runtime-main.js +31 -0
- package/dist/guest-runtime-main.js.map +1 -0
- package/dist/guest-runtime-revision.d.ts +1 -0
- package/dist/guest-runtime-revision.js +3 -0
- package/dist/guest-runtime-revision.js.map +1 -0
- package/dist/guest-runtime.d.ts +25 -0
- package/dist/guest-runtime.js +96 -0
- package/dist/guest-runtime.js.map +1 -0
- package/dist/index.d.ts +1 -1
- package/dist/lab-config.js +10 -3
- package/dist/lab-config.js.map +1 -1
- package/dist/lab-engine.js +6 -0
- package/dist/lab-engine.js.map +1 -1
- package/dist/lab-summary.d.ts +5 -1
- package/dist/lab-summary.js +5 -0
- package/dist/lab-summary.js.map +1 -1
- package/dist/local-agent-cli.js +1 -1
- package/dist/local-agent-cli.js.map +1 -1
- package/dist/local-firecracker-desktop.d.ts +13 -0
- package/dist/local-firecracker-desktop.js +150 -0
- package/dist/local-firecracker-desktop.js.map +1 -0
- package/dist/local-firecracker-study.d.ts +9 -0
- package/dist/local-firecracker-study.js +93 -0
- package/dist/local-firecracker-study.js.map +1 -0
- package/dist/local-runtime-config.d.ts +6 -0
- package/dist/local-runtime-config.js +56 -0
- package/dist/local-runtime-config.js.map +1 -0
- package/dist/local-runtime-release.d.ts +3 -0
- package/dist/local-runtime-release.js +8 -0
- package/dist/local-runtime-release.js.map +1 -0
- package/dist/local-runtime.d.ts +25 -0
- package/dist/local-runtime.js +113 -0
- package/dist/local-runtime.js.map +1 -0
- package/dist/observer-app.html +4 -4
- package/dist/pricing.d.ts +22 -1
- package/dist/pricing.js +22 -0
- package/dist/pricing.js.map +1 -1
- package/dist/program.js +50 -9
- package/dist/program.js.map +1 -1
- package/dist/restricted-codex-analysis.d.ts +15 -0
- package/dist/restricted-codex-analysis.js +13 -0
- package/dist/restricted-codex-analysis.js.map +1 -0
- package/dist/restricted-codex-participant-policy.d.ts +39 -0
- package/dist/restricted-codex-participant-policy.js +69 -0
- package/dist/restricted-codex-participant-policy.js.map +1 -0
- package/dist/restricted-codex-participant-run.d.ts +20 -0
- package/dist/restricted-codex-participant-run.js +78 -0
- package/dist/restricted-codex-participant-run.js.map +1 -0
- package/dist/restricted-codex-participant.d.ts +14 -0
- package/dist/restricted-codex-participant.js +178 -0
- package/dist/restricted-codex-participant.js.map +1 -0
- package/dist/restricted-codex-policy.d.ts +56 -0
- package/dist/restricted-codex-policy.js +151 -0
- package/dist/restricted-codex-policy.js.map +1 -0
- package/dist/restricted-codex-session.d.ts +19 -0
- package/dist/restricted-codex-session.js +413 -0
- package/dist/restricted-codex-session.js.map +1 -0
- package/dist/restricted-codex-transport.d.ts +58 -0
- package/dist/restricted-codex-transport.js +233 -0
- package/dist/restricted-codex-transport.js.map +1 -0
- package/dist/run-detail.js +4 -2
- package/dist/run-detail.js.map +1 -1
- package/dist/run.d.ts +12 -5
- package/dist/run.js +17 -1
- package/dist/run.js.map +1 -1
- package/dist/shared-world-lab.js +2 -2
- package/dist/shared-world-lab.js.map +1 -1
- package/dist/study-analysis-codex-config.d.ts +11 -0
- package/dist/study-analysis-codex-config.js +34 -0
- package/dist/study-analysis-codex-config.js.map +1 -0
- package/dist/study-analysis-engine.d.ts +6 -3
- package/dist/study-analysis-engine.js +19 -10
- package/dist/study-analysis-engine.js.map +1 -1
- package/dist/study-analysis-job.d.ts +3 -2
- package/dist/study-analysis-job.js +1 -1
- package/dist/study-analysis-job.js.map +1 -1
- package/dist/study-analysis-provider.d.ts +4 -2
- package/dist/study-analysis-provider.js +1 -1
- package/dist/study-analysis-provider.js.map +1 -1
- package/dist/study-analysis-service.d.ts +3 -0
- package/dist/study-analysis-service.js +24 -6
- package/dist/study-analysis-service.js.map +1 -1
- package/dist/study-analysis-validation.d.ts +43 -19
- package/dist/study-analysis-validation.js +23 -10
- package/dist/study-analysis-validation.js.map +1 -1
- package/dist/study-analysis.d.ts +27 -2
- package/dist/study-costs.js +6 -0
- package/dist/study-costs.js.map +1 -1
- package/dist/tui-app.js +102 -102
- package/docs/architecture/browser-control.md +117 -0
- package/docs/architecture/desktop-sessions.md +80 -0
- package/docs/architecture/guest-desktop.md +87 -0
- package/docs/architecture/local-browser-runtime.md +100 -0
- package/docs/architecture/restricted-codex-analysis.md +91 -0
- package/docs/architecture/runtime-broker-core.md +30 -0
- package/docs/contracts/schemas.md +1 -1
- package/docs/contracts/study-analysis.md +44 -2
- package/docs/goals/current.md +25 -9
- package/docs/product/automatic-analysis.md +24 -3
- package/docs/product/open-source-install-experience.md +7 -0
- package/docs/ramp/README.md +30 -13
- package/docs/release/0.97.0-codex-account-analysis.md +45 -0
- package/docs/release/0.98.0-local-browser-studies.md +31 -0
- package/package.json +4 -2
- package/skills/humanish/SKILL.md +39 -0
package/docs/goals/current.md
CHANGED
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
# Current Goals
|
|
2
2
|
|
|
3
|
-
Status date: 2026-09-
|
|
3
|
+
Status date: 2026-09-24. Release baseline: `0.98.0`.
|
|
4
4
|
|
|
5
5
|
This page guides work on current merged source. Published behavior is described
|
|
6
|
-
in the [release notes](../release/0.
|
|
6
|
+
in the [release notes](../release/0.98.0-local-browser-studies.md).
|
|
7
7
|
The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
|
|
8
8
|
preserves the former status log; its queues do not supersede this page.
|
|
9
9
|
|
|
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
|
|
|
88
88
|
No first-party deletion branch has met that gate. Public demonstrations do not
|
|
89
89
|
substitute for it.
|
|
90
90
|
|
|
91
|
-
## Current Program Truth (source `0.
|
|
91
|
+
## Current Program Truth (source `0.98.0`)
|
|
92
92
|
|
|
93
93
|
| Surface | Available in merged source | Remaining boundary |
|
|
94
94
|
| --- | --- | --- |
|
|
@@ -99,7 +99,7 @@ substitute for it.
|
|
|
99
99
|
| Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
|
|
100
100
|
| Observer | Live/recorded views, shared grid and participant playback, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
|
|
101
101
|
| Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
|
|
102
|
-
| Study findings | Default post-run analysis on supported live routes with a separate disclosed $3 admission estimate limit and opt-out; explicit `analyze`, fairer evidence selection, concern review and versioned findings with exact source links | Model interpretation needs review;
|
|
102
|
+
| Study findings | Default post-run analysis on supported live routes with a separate disclosed $3 admission estimate limit and opt-out; explicit `analyze`, fairer evidence selection, concern review and versioned findings with exact source links; explicit restricted Codex account analysis on the qualified Linux profile | Account dollars/output-token caps are unavailable; Mac/keychain/other CLI profiles are unqualified. Model interpretation needs review; selection limits coverage; opening Observer never dispatches analysis |
|
|
103
103
|
| TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving, run library and AgentMail setup, authentication and lab configuration | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
|
|
104
104
|
| Off-app communication | Recipient-scoped local capture and fresh real AgentMail receiving, supported inline raster images, bounded collection and host-owned recovery | Real mail uses isolated participant surfaces and remains local-only for publication. Hosted mail/model processing, bounded fidelity and interrupted-run recovery are explicit; local-agent, borrowed inboxes and SMS are unsupported |
|
|
105
105
|
| Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed; a synthetic video-only call with separate hosted peers is proven | Audio, TURN, provider-specific rooms, physical-device and touch fidelity remain unproven; unsupported media declarations are rejected |
|
|
@@ -107,13 +107,24 @@ substitute for it.
|
|
|
107
107
|
Use the [task support matrix](../architecture/task-protocol-support.md),
|
|
108
108
|
[actor registry](https://github.com/danielgwilson/humanish/blob/main/src/actor-registry.ts)
|
|
109
109
|
and [CLI reference](https://humanish.dev/docs/cli) when choosing a concrete path.
|
|
110
|
-
Source
|
|
110
|
+
Source and tests establish observed behavior. Resolve conflicts with requirements
|
|
111
|
+
explicitly; neither stale status prose nor a passing test makes a bug correct.
|
|
111
112
|
|
|
112
113
|
The library-assisted `local-app` route now includes a
|
|
113
114
|
[runnable npm example](../architecture/examples/state-driven-local-app/README.md).
|
|
114
115
|
Its deterministic provider demonstrates the integration with a real loopback
|
|
115
116
|
app; it does not establish persona effectiveness or independent adoption.
|
|
116
117
|
|
|
118
|
+
The [local browser runtime](../architecture/local-browser-runtime.md)
|
|
119
|
+
runs isolated Linux browser participants through the same scheduler, recordings
|
|
120
|
+
and automatic analysis as hosted studies. It uses Docker-owned resources and
|
|
121
|
+
ordinary TAP/NAT networking. Continue managed-local work from this complete study
|
|
122
|
+
path; the earlier offline owner/service qualification experiments are historical
|
|
123
|
+
fixtures, not an installation architecture or a prerequisite queue. Explicit
|
|
124
|
+
Linux local labs now use the installed CLI/TUI, with a verified runtime download
|
|
125
|
+
before the first live run. Mac support, inbox integration and optional media
|
|
126
|
+
remain unfinished. Existing hosted labs retain their behavior.
|
|
127
|
+
|
|
117
128
|
## Gates And Deferred Work
|
|
118
129
|
|
|
119
130
|
- Live OSS meta-lab execution remains disabled until repository-derived
|
|
@@ -135,10 +146,12 @@ app; it does not establish persona effectiveness or independent adoption.
|
|
|
135
146
|
Follow [AGENTS.md](../../AGENTS.md), the [invariants](../principles/invariants-and-defaults.md)
|
|
136
147
|
and the [public-readiness standard](../release/public-readiness-standard.md).
|
|
137
148
|
|
|
138
|
-
- Keep `main` clean and work on scoped branches/worktrees.
|
|
139
|
-
|
|
149
|
+
- Keep `main` clean and work on scoped branches/worktrees. Keep the task's scope,
|
|
150
|
+
authority, relevant checks and material failure boundaries in its issue, PR or
|
|
151
|
+
current handoff; do not create a separate packet for routine work.
|
|
140
152
|
- Existing explicit shipping authority governs implementation and merge;
|
|
141
|
-
otherwise issue readiness does not create authority by itself.
|
|
153
|
+
otherwise issue readiness does not create authority by itself. Machine-readiness
|
|
154
|
+
fields gate automated queue pickup, not directly assigned interactive work.
|
|
142
155
|
- Never commit secrets, private transcripts/screenshots, customer data or
|
|
143
156
|
private project context. Keep generated proof in ignored `.humanish/` and
|
|
144
157
|
retain needed evidence before removing a worktree.
|
|
@@ -153,7 +166,10 @@ and the [public-readiness standard](../release/public-readiness-standard.md).
|
|
|
153
166
|
|
|
154
167
|
## Proof Before Shipping
|
|
155
168
|
|
|
156
|
-
|
|
169
|
+
Use the [verification guidance](../../AGENTS.md#verification): check the changed
|
|
170
|
+
behavior and material risks, then stop unless new evidence warrants more work.
|
|
171
|
+
Required CI remains the merge gate. For a release, run the full release gates
|
|
172
|
+
from a clean contributor worktree:
|
|
157
173
|
|
|
158
174
|
```bash
|
|
159
175
|
pnpm install --frozen-lockfile
|
|
@@ -22,7 +22,7 @@ review:
|
|
|
22
22
|
|
|
23
23
|
Omitting `review.analysis` uses these defaults. Set `review.analysis: false` to
|
|
24
24
|
run participants without the additional analysis request. An explicit analysis
|
|
25
|
-
mapping requires `maxCostUsd`. This limits an admission estimate, not the
|
|
25
|
+
mapping using the default OpenAI API provider requires `maxCostUsd`. This limits an admission estimate, not the
|
|
26
26
|
provider's final bill, and is separate from participant spending limits. Analysis
|
|
27
27
|
can decline a large study before dispatch when its conservative estimate exceeds
|
|
28
28
|
that limit. Use `analyze --dry-run --max-cost <usd>` on retained evidence to inspect
|
|
@@ -37,6 +37,27 @@ its existing authentication boundary. Review the separate analysis budget before
|
|
|
37
37
|
zero-dollar cap does not cap post-run analysis. The bundled first-contact
|
|
38
38
|
zero-spend product fixture explicitly disables analysis.
|
|
39
39
|
|
|
40
|
+
To explicitly use your Codex ChatGPT account for the separate analyst:
|
|
41
|
+
|
|
42
|
+
```yaml
|
|
43
|
+
review:
|
|
44
|
+
analysis:
|
|
45
|
+
provider: codex
|
|
46
|
+
model: gpt-6-astra
|
|
47
|
+
timeoutMs: 600000
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
This requires Linux x64, qualified Codex CLI `0.154.0`, and a file-backed ChatGPT account login; the analyst
|
|
51
|
+
uses low reasoning effort and remote inference. Dollar cost and a provider
|
|
52
|
+
enforced output-token ceiling are unknown, so omit `maxCostUsd` and
|
|
53
|
+
`maxOutputTokens`. Numeric values are rejected before participant resources are
|
|
54
|
+
allocated. There is no API fallback. Missing or unsupported account setup leaves
|
|
55
|
+
an explicit failed analysis state and the original recording intact. Use
|
|
56
|
+
`humanish doctor --lab <lab>` for setup checks; account allowance and model access
|
|
57
|
+
remain untested until a request. An omitted provider still means OpenAI, including
|
|
58
|
+
hosted studies whose participant uses a local Codex or Claude login. This setting
|
|
59
|
+
does not enable managed local desktops.
|
|
60
|
+
|
|
40
61
|
The same configuration works through `humanish run <lab>`, `lab run <lab>`,
|
|
41
62
|
`watch <lab>`, and TUI live starts. Direct library calls to the five recording
|
|
42
63
|
producers honor it too. Supported routes are computer-use, scripted-browser,
|
|
@@ -53,7 +74,7 @@ Default analysis also skips recordings containing only setup or failure records
|
|
|
53
74
|
with no retained participant activity. A desktop startup failure does not start
|
|
54
75
|
an analysis request. The original failure remains visible.
|
|
55
76
|
|
|
56
|
-
CLI live starts disclose the separate admission estimate limit before execution.
|
|
77
|
+
CLI live starts disclose the selected analyst and its separate admission estimate limit or unknown account dollars before execution.
|
|
57
78
|
`humanish lab preflight <lab> --json` and the TUI lab screen also expose the
|
|
58
79
|
resolved budget without dispatching analysis. Library callers can inspect
|
|
59
80
|
`resolveAutomaticAnalysis` or `automaticAnalysisBudget` before running.
|
|
@@ -64,7 +85,7 @@ existing job and does not start another request. Concurrent or repeated automati
|
|
|
64
85
|
invocations cannot silently retry a paid attempt. If a process disappears while
|
|
65
86
|
an attempt is in flight, its state can be unknown rather than falsely complete.
|
|
66
87
|
Use manual `humanish analyze --run <exact-run-id> --max-cost 3` for an intentional
|
|
67
|
-
follow-up after inspecting the existing attempt and its accounting.
|
|
88
|
+
follow-up after inspecting the existing attempt and its accounting. For the account branch, use `humanish analyze --run <exact-run-id> --provider codex --rerun` without a dollar limit.
|
|
68
89
|
|
|
69
90
|
Stopping participant execution does not start a fresh automatic analysis. A
|
|
70
91
|
recorded harness cancellation is skipped; ordinary time limits and participant
|
|
@@ -235,3 +235,10 @@ them should install `@e2b/desktop` explicitly instead of receiving that
|
|
|
235
235
|
substrate as part of the default Humanish package install. When a GitHub token is
|
|
236
236
|
present, repo labels are redacted in durable artifacts by default; live stream
|
|
237
237
|
auth URLs are used only by the attached watch server and are not persisted.
|
|
238
|
+
|
|
239
|
+
An explicitly selected [local browser lab](../architecture/local-browser-runtime.md)
|
|
240
|
+
can instead use Linux x64, Docker/KVM and an existing Codex ChatGPT login for
|
|
241
|
+
participants and findings. The installed CLI prepares its pinned image before
|
|
242
|
+
the first live run. It does not install host prerequisites, and Mac, inbox and
|
|
243
|
+
media integration remain separate follow-ups. Existing hosted lab configuration
|
|
244
|
+
is preserved.
|
package/docs/ramp/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Status: public-safe contributor and agent ramp.
|
|
4
4
|
|
|
5
|
-
Package/source version in this tree: `0.
|
|
5
|
+
Package/source version in this tree: `0.98.0` (2026-09-24). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
|
|
6
6
|
`0.15.1` remains in force: managed run and output paths bind to validated
|
|
7
7
|
physical filesystem identities, and stored provider IDs are evidence, not
|
|
8
8
|
cleanup authority. The bundled OSS meta-lab is dry-run only until
|
|
@@ -14,19 +14,26 @@ context.
|
|
|
14
14
|
|
|
15
15
|
## First Read
|
|
16
16
|
|
|
17
|
-
|
|
17
|
+
Start with three things:
|
|
18
18
|
|
|
19
|
-
1. [`AGENTS.md`](../../AGENTS.md) for
|
|
20
|
-
2. [`docs/
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
19
|
+
1. [`AGENTS.md`](../../AGENTS.md) for engineering judgment and public boundaries.
|
|
20
|
+
2. The current task and [`docs/goals/current.md`](../goals/current.md) for current
|
|
21
|
+
product status. Explicit task direction takes precedence over historical queues.
|
|
22
|
+
3. Instructions in the component being changed, then its relevant contracts.
|
|
23
|
+
|
|
24
|
+
Use the references below as needed. Historical plans are context, not a backlog
|
|
25
|
+
to resume automatically. Keep one concise current task handoff with the requested
|
|
26
|
+
outcome, demonstrated behavior, next complete result, constraints and rejected or
|
|
27
|
+
deferred approaches; link evidence rather than repeating its chronology.
|
|
28
|
+
|
|
29
|
+
| When working on | Reference |
|
|
30
|
+
| --- | --- |
|
|
31
|
+
| Install, commands or first-run UX | [`README.md`](../../README.md), [install experience](../product/open-source-install-experience.md) |
|
|
32
|
+
| Security, evidence handling or defaults | [Invariants and defaults](../principles/invariants-and-defaults.md) |
|
|
33
|
+
| Observer | [Observer architecture](../architecture/observer.md) and its component instructions |
|
|
34
|
+
| Bundle formats or policy | [Run bundle](../contracts/run-bundle.md), [policy](../contracts/policy.md) |
|
|
35
|
+
| Public artifacts or packaging | [Public-readiness standard](../release/public-readiness-standard.md), [release procedure](../release/open-source-readiness.md) |
|
|
36
|
+
| Proof architecture or historical decisions | [Proof roadmap](https://github.com/danielgwilson/humanish/blob/main/docs/goals/proof-roadmap/goal.md), [historical delivery roadmap](../roadmap/world-class-open-source-v0.md) |
|
|
30
37
|
|
|
31
38
|
## Mental Model
|
|
32
39
|
|
|
@@ -47,6 +54,16 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
|
|
|
47
54
|
|
|
48
55
|
## Current State
|
|
49
56
|
|
|
57
|
+
The [0.98.0 release note](../release/0.98.0-local-browser-studies.md) describes
|
|
58
|
+
installed Linux browser studies with managed runtime images, Codex account
|
|
59
|
+
participants and automatic analysis, and shared CLI/TUI setup checks.
|
|
60
|
+
|
|
61
|
+
The [0.97.0 release note](../release/0.97.0-codex-account-analysis.md) describes
|
|
62
|
+
explicit Codex account analysis on a qualified Linux CLI/login profile, with
|
|
63
|
+
separate analyst authority, evidence-linked reports and unknown-dollar accounting.
|
|
64
|
+
Existing API defaults remain unchanged; managed local desktops and Mac account
|
|
65
|
+
analysis are not qualified by this release.
|
|
66
|
+
|
|
50
67
|
The [0.96.1 release note](../release/0.96.1-browser-navigation.md) describes
|
|
51
68
|
accurate physical browser measurements and bounded fitting that preserves normal
|
|
52
69
|
browser controls when they fit. Narrow screens retain the fullscreen fallback.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# 0.97.0 — Codex account reports on Linux
|
|
2
|
+
|
|
3
|
+
Saved studies can use an existing Codex ChatGPT login for their separate
|
|
4
|
+
findings report. Select it explicitly:
|
|
5
|
+
|
|
6
|
+
```bash
|
|
7
|
+
humanish analyze --run latest --provider codex --dry-run
|
|
8
|
+
humanish analyze --run latest --provider codex
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
For an automatic report after a live study, set:
|
|
12
|
+
|
|
13
|
+
```yaml
|
|
14
|
+
review:
|
|
15
|
+
analysis:
|
|
16
|
+
provider: codex
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
The initial qualification is Linux x64, Codex CLI 0.154.0, a file-backed
|
|
20
|
+
ChatGPT login, and `gpt-6-astra` with low reasoning effort. Other platforms,
|
|
21
|
+
CLI versions and keychain-only logins are refused. Selected evidence still goes
|
|
22
|
+
to remote inference. Model access and account allowance depend on the account.
|
|
23
|
+
|
|
24
|
+
Account dollars and the generated-token ceiling are unknown; numeric dollar or
|
|
25
|
+
output-token caps are rejected. Time and byte limits still apply. Token counts
|
|
26
|
+
are retained when available, and interrupted usage remains incomplete. Existing
|
|
27
|
+
OpenAI API analysis defaults and participant setup requirements are unchanged.
|
|
28
|
+
|
|
29
|
+
The analyst has its own conversation and restricted tool configuration.
|
|
30
|
+
Evidence admission, exact source references, secret scrubbing, immutable attempt
|
|
31
|
+
accounting and source-change checks apply to both providers. Cancellation
|
|
32
|
+
preserves recordings and earlier valid findings. Unconfirmed process cleanup or
|
|
33
|
+
unexpected replacement credentials produce a failed attempt and private recovery
|
|
34
|
+
guidance. Historical reports remain readable as execution profiles evolve.
|
|
35
|
+
|
|
36
|
+
Doctor checks the selected analyst's setup without a model turn. Dry-run checks
|
|
37
|
+
local evidence and configuration only; it does not test CLI/login/model access.
|
|
38
|
+
The CLI, TUI and Observer disclose account limits consistently. This release
|
|
39
|
+
does not introduce managed local desktops or qualify Mac account analysis.
|
|
40
|
+
|
|
41
|
+
The [qualification receipt](https://github.com/danielgwilson/humanish/blob/46330116726f74080fa18947c36da4fb4b333805/docs/goals/computer-use-actor/receipts/codex-account-analysis-2026-09-23.md)
|
|
42
|
+
records real screenshot analysis, cache reuse, cancellation, automatic reports
|
|
43
|
+
after desktop cleanup, desktop/phone evidence review, failures and proof limits.
|
|
44
|
+
See [configuration](../product/automatic-analysis.md) and the
|
|
45
|
+
[restricted launcher contract](../architecture/restricted-codex-analysis.md).
|
|
@@ -0,0 +1,31 @@
|
|
|
1
|
+
# 0.98.0 — Local browser studies on Linux
|
|
2
|
+
|
|
3
|
+
Explicit local browser labs now run from the installed CLI and TUI. On Linux
|
|
4
|
+
x64 with Docker/KVM and a supported Codex ChatGPT login, isolated Firecracker
|
|
5
|
+
participants use the normal scheduler, recordings, Observer and automatic
|
|
6
|
+
analysis without E2B or OpenAI API keys. Codex inference remains remote and
|
|
7
|
+
consumes account quota; dollar cost is unknown.
|
|
8
|
+
|
|
9
|
+
`humanish runtime setup` prepares a pinned, verified runtime image. A live local
|
|
10
|
+
lab also prepares it automatically before starting participants. `runtime status`
|
|
11
|
+
and `doctor --lab` inspect readiness without downloading or launching a browser.
|
|
12
|
+
The TUI shows that same runtime status. Existing hosted labs retain their
|
|
13
|
+
configuration, and a missing local prerequisite never selects another provider.
|
|
14
|
+
|
|
15
|
+
The [setup guide](../architecture/local-browser-runtime.md) includes a complete
|
|
16
|
+
manifest and the current limits. This release supports loopback apps, fixed
|
|
17
|
+
960×720 Chromium desktops, and optional OpenAI API participants. Mac/Lima,
|
|
18
|
+
local inboxes and local camera/microphone integration remain follow-ups.
|
|
19
|
+
|
|
20
|
+
Private state uses a Docker-owned anonymous volume, removed with its container
|
|
21
|
+
on normal close or controller death. Runtime image downloads are checked for
|
|
22
|
+
exact size and SHA-256 before loading; matching source archives and notices
|
|
23
|
+
are distributed separately.
|
|
24
|
+
|
|
25
|
+
Validation includes a packed installation outside the checkout: two concurrent
|
|
26
|
+
Codex participants saved distinct notes in a local app, the app confirmed both
|
|
27
|
+
saves, the recordings verified, and automatic analysis completed without model
|
|
28
|
+
or desktop API keys. Separate checks cover controller death and state-volume
|
|
29
|
+
removal, download integrity/cancellation, missing prerequisites and preserving
|
|
30
|
+
existing hosted configurations. This is a small Linux integration proof, not
|
|
31
|
+
a claim about Mac readiness, high concurrency or real conferencing apps.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "humanish",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.98.0",
|
|
4
4
|
"description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
|
|
5
5
|
"author": "Daniel G Wilson <daniel@danielgwilson.com>",
|
|
6
6
|
"keywords": [
|
|
@@ -76,13 +76,15 @@
|
|
|
76
76
|
"tui:smoke": "node scripts/tui-smoke.mjs",
|
|
77
77
|
"tui:test": "pnpm --filter humanish-tui test",
|
|
78
78
|
"release:dogfood": "node scripts/release-dogfood.mjs",
|
|
79
|
+
"browser-control:proof": "node scripts/browser-control-proof.mjs",
|
|
79
80
|
"docs:generate": "tsx scripts/generate-cli-docs.ts",
|
|
80
81
|
"docs:check": "tsx scripts/generate-cli-docs.ts --check",
|
|
81
82
|
"observer:browser:proof": "node scripts/observer-browser-proof.mjs",
|
|
82
83
|
"observer:iframe:proof": "node scripts/observer-iframe-proof.mjs",
|
|
83
84
|
"observer:chrome:proof": "node scripts/observer-chrome-proof.mjs",
|
|
84
85
|
"observer:reliability:proof": "node scripts/observer-reliability-proof.mjs",
|
|
85
|
-
"tui:connections:proof": "python3 scripts/tui-connections-proof.py"
|
|
86
|
+
"tui:connections:proof": "python3 scripts/tui-connections-proof.py",
|
|
87
|
+
"guest-desktop:proof": "node scripts/guest-desktop-proof.mjs"
|
|
86
88
|
},
|
|
87
89
|
"repository": {
|
|
88
90
|
"type": "git",
|
package/skills/humanish/SKILL.md
CHANGED
|
@@ -93,6 +93,45 @@ exact returned path, not a basename that could resolve to another manifest.
|
|
|
93
93
|
- keep `.env.example` commit-safe and value-free;
|
|
94
94
|
- never commit generated run bundles.
|
|
95
95
|
|
|
96
|
+
## Choosing a findings analyst
|
|
97
|
+
|
|
98
|
+
Analysis is separate from the participant. Hosted and manual defaults use the
|
|
99
|
+
OpenAI API with its own admission budget. An explicitly local browser study with
|
|
100
|
+
a Codex participant defaults to a separate Codex account analyst. To select that
|
|
101
|
+
restricted Codex ChatGPT account analyst on other supported studies,
|
|
102
|
+
set `review.analysis.provider: codex` or pass `analyze --provider codex` on a
|
|
103
|
+
completed recording. This uses remote inference and the qualified CLI/login,
|
|
104
|
+
not local inference or the participant's existing conversation. See
|
|
105
|
+
[the analysis contract](../../docs/contracts/study-analysis.md) for the current
|
|
106
|
+
CLI/model qualification and setup limits.
|
|
107
|
+
|
|
108
|
+
Do not pass numeric `maxCostUsd`/`maxOutputTokens` or their CLI flags to the
|
|
109
|
+
account branch. It cannot enforce those ceilings and rejects them. Account dollar
|
|
110
|
+
cost remains unknown even when token usage is reported. There is no fallback to
|
|
111
|
+
an API key or another provider. `analyze --dry-run --provider codex` validates
|
|
112
|
+
local evidence/configuration only; `doctor --lab` checks setup without a model
|
|
113
|
+
request. Inspect a failed attempt before explicitly retrying `--provider codex
|
|
114
|
+
--rerun`. Opening Observer never starts analysis.
|
|
115
|
+
|
|
116
|
+
## Local browser setup
|
|
117
|
+
|
|
118
|
+
On Linux x64 with a local rootful Docker Engine, KVM and TUN, an `app-url` lab can
|
|
119
|
+
set `execution.target: local` and `actors[0].type: local-agent` with
|
|
120
|
+
`localAgent: codex`. It uses the supported Codex ChatGPT login, not E2B or an
|
|
121
|
+
OpenAI API key. Inference is remote and consumes account quota. Existing hosted
|
|
122
|
+
labs stay hosted; never silently change their execution or billing provider.
|
|
123
|
+
|
|
124
|
+
Use `humanish runtime status --json` and `humanish doctor --lab <path> --json`
|
|
125
|
+
for read-only setup inspection. `humanish runtime setup` downloads and verifies
|
|
126
|
+
the pinned runtime; a live local run also prepares it automatically. The normal
|
|
127
|
+
`humanish lab run <path>` command and TUI use the same study runner and Observer.
|
|
128
|
+
See [the complete example and limits](../../docs/architecture/local-browser-runtime.md).
|
|
129
|
+
|
|
130
|
+
Local browsers currently require a loopback app URL with an explicit port above
|
|
131
|
+
1023, use a 960×720 Chromium desktop, and reject inbox/media declarations. Mac
|
|
132
|
+
setup is not integrated. Do not claim that installing the CLI also installs
|
|
133
|
+
Docker or makes these prerequisites available.
|
|
134
|
+
|
|
96
135
|
## Format Stack
|
|
97
136
|
|
|
98
137
|
When creating or editing Humanish files:
|