humanish 0.90.0 → 0.91.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +11 -7
- package/dist/automatic-analysis-completion.d.ts +4 -4
- package/dist/automatic-analysis-completion.js +10 -8
- package/dist/automatic-analysis-completion.js.map +1 -1
- package/dist/automatic-analysis-config.d.ts +9 -1
- package/dist/automatic-analysis-config.js +15 -2
- package/dist/automatic-analysis-config.js.map +1 -1
- package/dist/automatic-study-analysis.d.ts +5 -2
- package/dist/automatic-study-analysis.js +34 -15
- package/dist/automatic-study-analysis.js.map +1 -1
- package/dist/chrome-cdp-probe.d.ts +5 -3
- package/dist/chrome-cdp-probe.js.map +1 -1
- package/dist/concurrent-shared-world-lab.js +1 -1
- package/dist/concurrent-shared-world-lab.js.map +1 -1
- package/dist/cua-actor-lab.d.ts +10 -4
- package/dist/cua-actor-lab.js +20 -9
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/device-presets.d.ts +4 -4
- package/dist/device-presets.js +8 -10
- package/dist/device-presets.js.map +1 -1
- package/dist/e2b-terminal-lab.js +6 -1
- package/dist/e2b-terminal-lab.js.map +1 -1
- package/dist/index.d.ts +3 -2
- package/dist/index.js +1 -0
- package/dist/index.js.map +1 -1
- package/dist/init-templates.js +11 -8
- package/dist/init-templates.js.map +1 -1
- package/dist/lab-config.d.ts +4 -3
- package/dist/lab-config.js +2 -2
- package/dist/lab-config.js.map +1 -1
- package/dist/lab-preflight.d.ts +3 -0
- package/dist/lab-preflight.js +3 -0
- package/dist/lab-preflight.js.map +1 -1
- package/dist/lab-summary.js +6 -4
- package/dist/lab-summary.js.map +1 -1
- package/dist/observer-app.html +8 -8
- package/dist/program.js +7 -1
- package/dist/program.js.map +1 -1
- package/dist/run-index.js +2 -2
- package/dist/run-index.js.map +1 -1
- package/dist/scripted-browser-lab.js +5 -5
- package/dist/scripted-browser-lab.js.map +1 -1
- package/dist/shared-world-lab.js +2 -1
- package/dist/shared-world-lab.js.map +1 -1
- package/dist/stats.d.ts +9 -2
- package/dist/stats.js +38 -15
- package/dist/stats.js.map +1 -1
- package/dist/study-analysis-engine.d.ts +1 -1
- package/dist/study-analysis-engine.js +9 -3
- package/dist/study-analysis-engine.js.map +1 -1
- package/dist/study-analysis-evidence.js +233 -47
- package/dist/study-analysis-evidence.js.map +1 -1
- package/dist/study-analysis-job.d.ts +9 -0
- package/dist/study-analysis-job.js +18 -1
- package/dist/study-analysis-job.js.map +1 -1
- package/dist/study-analysis-service.js +10 -3
- package/dist/study-analysis-service.js.map +1 -1
- package/dist/study-analysis-sharing.js +1 -1
- package/dist/study-analysis-sharing.js.map +1 -1
- package/dist/study-analysis-store.d.ts +14 -1
- package/dist/study-analysis-store.js +114 -3
- package/dist/study-analysis-store.js.map +1 -1
- package/dist/study-analysis-validation.d.ts +145 -0
- package/dist/study-analysis-validation.js +55 -22
- package/dist/study-analysis-validation.js.map +1 -1
- package/dist/study-analysis.d.ts +8 -0
- package/dist/study-costs.d.ts +27 -0
- package/dist/study-costs.js +108 -0
- package/dist/study-costs.js.map +1 -0
- package/dist/terminal-participant-activity.d.ts +6 -0
- package/dist/terminal-participant-activity.js +34 -0
- package/dist/terminal-participant-activity.js.map +1 -0
- package/dist/tui-app.js +20 -20
- package/docs/architecture/examples/state-driven-local-app/README.md +5 -0
- package/docs/architecture/examples/state-driven-local-app/runner.mjs +1 -0
- package/docs/contracts/run-bundle.md +4 -0
- package/docs/contracts/schemas.md +2 -1
- package/docs/contracts/study-analysis.md +63 -5
- package/docs/contracts/study-costs.md +69 -0
- package/docs/goals/current.md +4 -4
- package/docs/principles/invariants-and-defaults.md +7 -3
- package/docs/product/automatic-analysis.md +25 -7
- package/docs/ramp/README.md +11 -2
- package/docs/release/0.91.0-analysis-quality-and-defaults.md +33 -0
- package/docs/release/0.91.1-study-review-polish.md +28 -0
- package/package.json +1 -1
|
@@ -73,3 +73,8 @@ npx tsc --allowJs --checkJs --noEmit --strict --skipLibCheck --types node --targ
|
|
|
73
73
|
See the [state-driven executor guide](../../state-driven-executor.md) for progress
|
|
74
74
|
projection limits, runtime-only state, unpinned local-app provenance, and the
|
|
75
75
|
fail-closed guards on this route.
|
|
76
|
+
|
|
77
|
+
The example explicitly sets `review: { analysis: false }` to keep post-run
|
|
78
|
+
analysis free of provider requests too. Supported live recordings otherwise
|
|
79
|
+
use the separate default analysis budget; a deterministic participant provider
|
|
80
|
+
does not replace the analysis provider.
|
|
@@ -68,6 +68,7 @@ try {
|
|
|
68
68
|
// This registry id selects the CUA loop. buildProvider supplies the actual provider.
|
|
69
69
|
actors: [{ type: "openai-computer-use", persona: "pixel-pat", mission: "Greet the app." }],
|
|
70
70
|
scenario: { mode: "live" },
|
|
71
|
+
review: { analysis: false }, // Keep this deterministic example free of provider requests.
|
|
71
72
|
execution: { timeoutMs: 15_000 }
|
|
72
73
|
});
|
|
73
74
|
if (!parsed.ok) throw new Error(parsed.error.message);
|
|
@@ -210,6 +210,10 @@ bundles stay byte-stable. Each lane's own estimate also rides its
|
|
|
210
210
|
from the reserved provider-returned `tokenUsage.costUsd`. See
|
|
211
211
|
[`schemas.md`](schemas.md) → Run Cost Summary And Estimated Actor Cost.
|
|
212
212
|
|
|
213
|
+
This bundle subtotal excludes separate study-analysis requests. Use
|
|
214
|
+
`humanish stats` for the complete retained estimate and explicit accounting
|
|
215
|
+
gaps across run costs and analysis attempts. See [study cost statistics](study-costs.md).
|
|
216
|
+
|
|
213
217
|
`humanish verify` treats cost as ADVISORY on magnitude and FAIL-CLOSED on
|
|
214
218
|
labeling: absence passes, but a claimed dollar figure without its `ratesAsOf`
|
|
215
219
|
date + `source`, or a total that does not match its known lines, fails. Verify
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
Date: 2026-06-02 (current-state note updated 2026-07-14)
|
|
4
4
|
|
|
5
5
|
Status: reference map for the major contracts shipped through source version
|
|
6
|
-
`0.
|
|
6
|
+
`0.91.1`; it is not an exhaustive inventory of command/result envelopes. Exported types,
|
|
7
7
|
schema constants, parsers, and validators in `src/` are authoritative. Rows
|
|
8
8
|
marked "reserved" name layering intent only — no code emits or validates them
|
|
9
9
|
yet. Do not emit a reserved schema.
|
|
@@ -41,6 +41,7 @@ workflow without leaking private upstream truth into core.
|
|
|
41
41
|
| Study analysis | `humanish.study-analysis.v1` | see [study analysis](study-analysis.md) and synthetic analysis fixtures in `tests/` |
|
|
42
42
|
| Study analysis correction | `humanish.study-analysis-correction.v1` | see [study analysis](study-analysis.md#human-review-and-sharing) |
|
|
43
43
|
| Analysis execution receipt | `humanish.analysis-execution.v1` | see [study analysis](study-analysis.md#durable-records) |
|
|
44
|
+
| Analysis execution start | `humanish.analysis-execution-start.v1` | see [study analysis](study-analysis.md#durable-records) |
|
|
44
45
|
| Verification | `humanish.verify-result.v1` | `five-check-verify` |
|
|
45
46
|
| Policy | `humanish.policy.v1` (fixture-only; not engine-validated) | `public-safety-policy` |
|
|
46
47
|
| Feedback | `humanish.feedback.v1` | `public-safe-feedback` |
|
|
@@ -1,8 +1,11 @@
|
|
|
1
1
|
# Study analysis
|
|
2
2
|
|
|
3
|
-
Study analysis is an
|
|
3
|
+
Study analysis is an independent interpretation of retained participant evidence.
|
|
4
4
|
It is separate from the participant's account, recorded outcome, and the run's
|
|
5
5
|
deterministic review verdict. Opening an Observer never starts a provider request.
|
|
6
|
+
Supported live runs request analysis on completion by default, with a separate
|
|
7
|
+
$3 admission estimate limit. Set `review.analysis: false` to disable that request;
|
|
8
|
+
see [automatic analysis](../product/automatic-analysis.md).
|
|
6
9
|
|
|
7
10
|
## Invocation
|
|
8
11
|
|
|
@@ -59,8 +62,15 @@ stopped, an operator can remove the empty lock directory and retry.
|
|
|
59
62
|
|
|
60
63
|
The packet currently admits up to 16 participants, 800 evidence items, 40 PNG
|
|
61
64
|
captures, 160 KiB of text and 20 MiB of images. Individual source files, image
|
|
62
|
-
dimensions and result sizes have separate limits.
|
|
63
|
-
|
|
65
|
+
dimensions and result sizes have separate limits. Count and text budgets are
|
|
66
|
+
distributed across included participants, with unused capacity from short
|
|
67
|
+
sessions available to longer ones. Capture selection prioritizes session endings
|
|
68
|
+
and beginnings, context around recorded failures, and spread across each whole
|
|
69
|
+
session. Failure priority uses structured source status, not application-specific
|
|
70
|
+
keywords or image interpretation. Unflagged visual errors may still be omitted.
|
|
71
|
+
Bounded reads and the total image byte limit can reduce coverage further.
|
|
72
|
+
Selected entries retain their original source order, frame and event identities;
|
|
73
|
+
this is not a statistically representative sample. Selection
|
|
64
74
|
omissions and unreadable or invalid capture files make declared coverage
|
|
65
75
|
incomplete. Coverage records file availability and selection; it does not
|
|
66
76
|
certify visual legibility, correct interpretation, or exhaustive issue discovery.
|
|
@@ -72,7 +82,8 @@ to an earlier capture retain that context without creating another frame.
|
|
|
72
82
|
Scripted lanes use their recorded `ui.intent` goal when no participant assignment
|
|
73
83
|
exists. Missing assignments and declared captures without supported trace
|
|
74
84
|
references are explicit omissions. Artifacts without `captureVersion` continue
|
|
75
|
-
to validate against the original
|
|
85
|
+
to validate against the original capture mapping; previously saved selections
|
|
86
|
+
are not recomputed or rewritten.
|
|
76
87
|
|
|
77
88
|
The standard review covers session summary, apparent intent, observed outcome,
|
|
78
89
|
friction, dead ends, recovery, and participant feedback. Findings are ordered by
|
|
@@ -80,6 +91,21 @@ observed task impact, replication among exposed participants, and recovery.
|
|
|
80
91
|
Impact and confidence remain separate. There is no numeric frustration or
|
|
81
92
|
universal priority score.
|
|
82
93
|
|
|
94
|
+
The standard review also accounts for material participant concerns before
|
|
95
|
+
ranking findings. Reported uncertainty can be useful even when the interface is
|
|
96
|
+
correct or study setup may explain it. Claims distinguish that experience from
|
|
97
|
+
an established product defect, preserve consequential recoveries, and check
|
|
98
|
+
participant accounts against the actual assignment and captured state.
|
|
99
|
+
|
|
100
|
+
New results include `concernReviews`: evidence-linked observations with a
|
|
101
|
+
`finding`, `context`, or `unsupported` disposition and a concise reason. A finding
|
|
102
|
+
disposition references an existing local finding ID; other dispositions use null.
|
|
103
|
+
The Observer's **Concerns considered** disclosure exposes these decisions and
|
|
104
|
+
their original evidence. There is no required finding count or inventory of every
|
|
105
|
+
thought. Older reports may omit this field and remain readable. The disclosure
|
|
106
|
+
is model-generated assessment, not an independent completeness audit or a human
|
|
107
|
+
reviewer annotation.
|
|
108
|
+
|
|
83
109
|
Every observation cites packet-local evidence IDs. The model cannot choose a
|
|
84
110
|
filesystem path or fetch another resource. Validation checks participant
|
|
85
111
|
membership, unique counts, quote fidelity, evidence type and reference
|
|
@@ -92,6 +118,19 @@ offset. Nonvisual events retain event identity without invented frame offsets.
|
|
|
92
118
|
Scripted captures without recorded timestamps keep null analysis times; any
|
|
93
119
|
uniform playback pacing is an estimate, not an observed duration.
|
|
94
120
|
|
|
121
|
+
Finding previews select from the finding's cited entries. They prefer a capture
|
|
122
|
+
directly cited as visual evidence, then the number of distinct visual/action
|
|
123
|
+
observations citing it; equal support keeps citation order. Duplicate claims do
|
|
124
|
+
not increase support. This is a display heuristic, not a confidence score or a
|
|
125
|
+
guarantee that the selected capture is the most relevant. Context-only and
|
|
126
|
+
nonvisual evidence retain their basis and original event. All cited moments stay
|
|
127
|
+
available, with exact recording links.
|
|
128
|
+
|
|
129
|
+
Confidence, recovery and the first full evidence limitation remain visible when
|
|
130
|
+
a finding opens. Exposure and remaining unique limits are one disclosure away;
|
|
131
|
+
observation details retain each original claim and its specific limitation.
|
|
132
|
+
This presentation does not rewrite the saved analysis or reviewer corrections.
|
|
133
|
+
|
|
95
134
|
## Durable records
|
|
96
135
|
|
|
97
136
|
The frozen `humanish.observer-data.v1` schema is unchanged. A companion
|
|
@@ -103,9 +142,10 @@ status, usage and validated findings:
|
|
|
103
142
|
.humanish/runs/<run>/
|
|
104
143
|
analysis/<analysis>/analysis.json
|
|
105
144
|
analysis/<analysis>/corrections/<correction>/correction.json
|
|
145
|
+
analysis-attempts/<analysis>/start.json # before transport; outcome initially unknown
|
|
106
146
|
analysis-attempts/<analysis>/receipt.json
|
|
107
147
|
observer/study-analysis.json
|
|
108
|
-
analysis-automatic/job.json #
|
|
148
|
+
analysis-automatic/job.json # post-run lifecycle; never a retry instruction
|
|
109
149
|
```
|
|
110
150
|
|
|
111
151
|
The optional automatic job is separate from the immutable analysis. Its view
|
|
@@ -120,6 +160,24 @@ model, budget, status and known usage even if source changes prevent report
|
|
|
120
160
|
publication. They contain no question, participant text, images, or findings.
|
|
121
161
|
`analyze list --json` includes these receipts.
|
|
122
162
|
|
|
163
|
+
New requests first claim their execution directory and atomically publish
|
|
164
|
+
`start.json` (`humanish.analysis-execution-start.v1`) before provider transport.
|
|
165
|
+
It contains only the attempt/run IDs, input/config/source digests, prompt version
|
|
166
|
+
and timestamp. The live caller retains a binding to that exact directory and
|
|
167
|
+
publishes the final receipt once. A start without usable final accounting means
|
|
168
|
+
dispatch and spend remain unresolved; it is not proof that a provider charged.
|
|
169
|
+
A cancellation before transport can finalize with `usage.dispatched: false`.
|
|
170
|
+
Admission refusals, missing credentials and rejected dispatch guards do not
|
|
171
|
+
create a potentially paid attempt. Reuse does not create a new start.
|
|
172
|
+
|
|
173
|
+
`humanish stats` reads all retained execution IDs, including failed, cancelled,
|
|
174
|
+
unpriced and unresolved attempts. Report/receipt copies and automatic reuse
|
|
175
|
+
count once per run and analysis ID. Legacy report usage can contribute without
|
|
176
|
+
a fresh source or valid findings, but the missing execution receipt is labeled.
|
|
177
|
+
This accounting read never approves findings or requests a provider. See
|
|
178
|
+
[study cost statistics](study-costs.md) for the additive JSON contract and
|
|
179
|
+
run-date attribution.
|
|
180
|
+
|
|
123
181
|
Analysis and execution-history directories each admit 256 entries, including
|
|
124
182
|
interrupted writes; correction history admits 256 entries per analysis. A new
|
|
125
183
|
attempt requires readable inventories with room for its records before dispatch.
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# Study cost statistics
|
|
2
|
+
|
|
3
|
+
`humanish stats` leads with the known estimated spend across the selected runs:
|
|
4
|
+
participant and desktop estimates plus every distinct retained analysis attempt.
|
|
5
|
+
It separates those components and shows missing usage and incomplete history.
|
|
6
|
+
All amounts are estimates from retained rate-table accounting, not provider bills.
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
humanish stats
|
|
10
|
+
humanish stats --lab sample-study --since 2026-09-01 --json
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
The filters select runs by lab and run start date. All later analysis reruns
|
|
14
|
+
belong to their source run for filtering and daily grouping. This is study-cost
|
|
15
|
+
attribution, not a calendar of provider charges.
|
|
16
|
+
|
|
17
|
+
With no filters, a directory whose source metadata is unreadable can still
|
|
18
|
+
contribute valid analysis receipts under `(no lab)` and `(undated)`. Lab/date
|
|
19
|
+
filters exclude such unattributable directories; they remain named in
|
|
20
|
+
`unreadable`. The command does not guess their date or lab from an analysis.
|
|
21
|
+
|
|
22
|
+
## Additive JSON contract
|
|
23
|
+
|
|
24
|
+
The envelope remains `humanish.stats.v1`. Existing fields keep their meanings:
|
|
25
|
+
`totals.estimatedSpendUsd`, `days[].estimatedSpendUsd`, lab `medianCostUsd`,
|
|
26
|
+
`costSamples` and `unpricedRuns` describe participant/desktop run estimates.
|
|
27
|
+
They do not suddenly include a separate analysis request. The bundle's
|
|
28
|
+
`cost.estimatedTotalUsd` and the cached run-index estimate also remain unchanged.
|
|
29
|
+
Observer and terminal run summaries label this narrower scope.
|
|
30
|
+
|
|
31
|
+
The new `costs` object appears on totals, each lab and each day. `costsByRun`
|
|
32
|
+
contains the same accounting per selected run with stable warning codes.
|
|
33
|
+
|
|
34
|
+
| Field | Meaning |
|
|
35
|
+
| --- | --- |
|
|
36
|
+
| `estimatedTotalUsd` | Sum of the retained run and analysis estimates |
|
|
37
|
+
| `runEstimatedUsd` | Participant/desktop estimate; may be a known subtotal |
|
|
38
|
+
| `analysisEstimatedUsd` | Sum of all distinct retained analysis estimates |
|
|
39
|
+
| `incompleteRunEstimates` | Runs with partial or unknown participant/desktop accounting |
|
|
40
|
+
| `analysisAttempts` | Distinct retained execution IDs, including unresolved claims |
|
|
41
|
+
| `analysisDispatchedAttempts` | Attempts whose final accounting confirms transport |
|
|
42
|
+
| `analysisNotDispatchedAttempts` | Attempts whose final accounting confirms no transport |
|
|
43
|
+
| `analysisUnpricedAttempts` | Dispatched or potentially dispatched attempts without a complete price |
|
|
44
|
+
| `analysisUnresolvedAttempts` | Claims without usable final accounting; a subset of unpriced attempts |
|
|
45
|
+
| `analysisHistoryUncertainRuns` | Runs with absent, legacy report-only, unreadable or conflicting history |
|
|
46
|
+
|
|
47
|
+
The three amount fields are `null` when nothing in that component has an
|
|
48
|
+
estimate. Known zero is retained only when supported: for example, final
|
|
49
|
+
accounting confirms an attempt never dispatched. A known subtotal can coexist
|
|
50
|
+
with unknown costs; inspect the counts alongside it. An absent analysis history
|
|
51
|
+
is not converted into a free analysis.
|
|
52
|
+
|
|
53
|
+
## Accounting rules
|
|
54
|
+
|
|
55
|
+
Execution receipts take part regardless of whether the analysis succeeded,
|
|
56
|
+
its report was published, its findings are current, or the original recording
|
|
57
|
+
later changed. The receipt and report with the same run/analysis ID count once;
|
|
58
|
+
new IDs from explicit reruns count separately. Reusing a prior result adds no
|
|
59
|
+
attempt or expense. Conflicting accounting for the same ID stays unresolved.
|
|
60
|
+
|
|
61
|
+
Older reports without receipts contribute their strictly validated accounting
|
|
62
|
+
metadata with a legacy warning. Provider requests that left no durable record
|
|
63
|
+
in older versions cannot be reconstructed. A missing or corrupt inventory is
|
|
64
|
+
explicitly uncertain. Reads are contained and bounded, and never dispatch,
|
|
65
|
+
repair accounting, update timestamps or write files.
|
|
66
|
+
|
|
67
|
+
The report covers retained study bundles only. It cannot account for separately
|
|
68
|
+
launched preflight desktops, deleted runs, third-party application hosting,
|
|
69
|
+
provider subscriptions, or spending outside the selected project directory.
|
package/docs/goals/current.md
CHANGED
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
# Current Goals
|
|
2
2
|
|
|
3
|
-
Status date: 2026-09-15. Release baseline: `0.
|
|
3
|
+
Status date: 2026-09-15. Release baseline: `0.91.1`.
|
|
4
4
|
|
|
5
5
|
This page guides work on current merged source. Published behavior is described
|
|
6
|
-
in the [release notes](../release/0.
|
|
6
|
+
in the [release notes](../release/0.91.1-study-review-polish.md).
|
|
7
7
|
The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
|
|
8
8
|
preserves the former status log; its queues do not supersede this page.
|
|
9
9
|
|
|
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
|
|
|
88
88
|
No first-party deletion branch has met that gate. Public demonstrations do not
|
|
89
89
|
substitute for it.
|
|
90
90
|
|
|
91
|
-
## Current Program Truth (source `0.
|
|
91
|
+
## Current Program Truth (source `0.91.1`)
|
|
92
92
|
|
|
93
93
|
| Surface | Available in merged source | Remaining boundary |
|
|
94
94
|
| --- | --- | --- |
|
|
@@ -99,7 +99,7 @@ substitute for it.
|
|
|
99
99
|
| Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
|
|
100
100
|
| Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
|
|
101
101
|
| Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
|
|
102
|
-
| Study findings |
|
|
102
|
+
| Study findings | Default post-run analysis on supported live routes with a separate disclosed $3 admission estimate limit and opt-out; explicit `analyze`, fairer evidence selection, concern review and versioned findings with exact source links | Model interpretation needs review; bounded selection and source truncation limit coverage; opening Observer never dispatches analysis |
|
|
103
103
|
| TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
|
|
104
104
|
| Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
|
|
105
105
|
| Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
|
|
@@ -34,9 +34,12 @@ certify (see the conformance suite).
|
|
|
34
34
|
terminal) is only ever pointed at a URL the harness itself issued or validated under a
|
|
35
35
|
declared policy (loopback entry, provisioned subject, declared external target). Never an
|
|
36
36
|
arbitrary URL from unvalidated input.
|
|
37
|
-
3. **Live spend
|
|
38
|
-
|
|
39
|
-
live
|
|
37
|
+
3. **Live spend requires an explicit live invocation.** No omission or fallback may
|
|
38
|
+
turn a dry run, preview, reader or unsupported route into provider or sandbox
|
|
39
|
+
spend. An explicitly live supported study includes the disclosed default
|
|
40
|
+
post-run analysis budget unless `review.analysis: false` disables it. Analysis
|
|
41
|
+
has a separate admission estimate limit, not a provider billing cap or part of
|
|
42
|
+
the actor budget. Spend-bearing tests retain their explicit env gates.
|
|
40
43
|
4. **Evidence verifies fail-closed.** A run bundle that cannot pass verification (schema,
|
|
41
44
|
redaction status, artifact presence, public-safety scan) is a failed run, even when the
|
|
42
45
|
session "worked." The gate applies to the harness's own error reports.
|
|
@@ -82,6 +85,7 @@ silently drifting from one is not.
|
|
|
82
85
|
| Default | Why it is the default | Legitimate override |
|
|
83
86
|
|---|---|---|
|
|
84
87
|
| Dry-run | Spend safety (invariant 3 sets the floor; dry-run keeps the floor far away) | `scenario.mode: live` |
|
|
88
|
+
| Post-run analysis on supported live studies | Findings accompany the recording; a separate $3 admission estimate limit is visible before execution | `review.analysis: false` disables it; an explicit mapping sets another analysis budget. Dry-run and unsupported routes never dispatch; missing default credentials produce a recorded skip |
|
|
85
89
|
| Per-lane worlds | Isolation, attribution, reproducibility | `subject.topology: shared-world` — N seats against ONE provisioned, mutable plane for scenarios that ARE about interaction between roles (#164). `execution.concurrency: 1` (an explicit choice) = SEQUENTIAL turns (one sandbox); higher = CONCURRENT — and since #350 an omitted concurrency fills to the seat count, so every declared seat runs live at once by default (one getHost-exposed subject sandbox + N actor sandboxes driving it at once, synthetic-subject only). The bundle declares the weaker `attributionClass: shared-world` + a verify-enforced `attributionLimits` ceiling (the concurrent set drops `sequential-only` and adds `best-effort-causal-attribution` etc.), so the looser per-role attribution is honest, not hidden. |
|
|
86
90
|
| External key placement | Smallest blast radius: when the keyed process (e.g. a computer-use provider loop) runs outside the sandbox, its key never enters | In-sandbox placement when the keyed process runs inside (an agent harness under test); declared per actor type, with a spend budget |
|
|
87
91
|
| Loopback entry URLs | Public-safety: never drive third-party sites unbidden | `policies.allowPublicTargets` for an owner-declared deployment/preview (a Vercel preview of your own app). Multi-lane public/preview fan-out needs explicit `actors[0].lanes[].target` for every lane, so the adapter-owned topology is declared rather than inferred. Provisioned clone subjects always serve in-sandbox on loopback |
|
|
@@ -1,8 +1,11 @@
|
|
|
1
1
|
# Automatic study analysis
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
Supported live studies automatically request analysis after each recording finishes. Findings remain
|
|
4
4
|
separate from participant feedback and the recorded study verdict.
|
|
5
5
|
|
|
6
|
+
The default is `gpt-6-astra` with high reasoning effort, a separate $3 admission
|
|
7
|
+
estimate limit, a 300-second timeout and 16,384 output tokens. To customize it:
|
|
8
|
+
|
|
6
9
|
```yaml
|
|
7
10
|
review:
|
|
8
11
|
analysis:
|
|
@@ -14,26 +17,37 @@ review:
|
|
|
14
17
|
# question: Where did participants need to recover?
|
|
15
18
|
```
|
|
16
19
|
|
|
17
|
-
|
|
18
|
-
|
|
20
|
+
Omitting `review.analysis` uses these defaults. Set `review.analysis: false` to
|
|
21
|
+
run participants without the additional analysis request. An explicit analysis
|
|
22
|
+
mapping requires `maxCostUsd`. This limits an admission estimate, not the
|
|
19
23
|
provider's final bill, and is separate from participant spending limits. Analysis
|
|
20
24
|
sends selected retained text and captures to OpenAI using `OPENAI_API_KEY`.
|
|
21
25
|
Analysis runs in the Humanish runner using its credentials. This setting adds no
|
|
22
26
|
credential channel to the target application; each participant backend retains
|
|
23
|
-
its existing authentication boundary. Review a manifest
|
|
24
|
-
|
|
27
|
+
its existing authentication boundary. Review the separate analysis budget before running a manifest live; an actor's
|
|
28
|
+
zero-dollar cap does not cap post-run analysis. The bundled first-contact
|
|
29
|
+
zero-spend product fixture explicitly disables analysis.
|
|
25
30
|
|
|
26
31
|
The same configuration works through `humanish run <lab>`, `lab run <lab>`,
|
|
27
32
|
`watch <lab>`, and TUI live starts. Direct library calls to the five recording
|
|
28
33
|
producers honor it too. Supported routes are computer-use, scripted-browser,
|
|
29
34
|
terminal-product, sequential shared-world and concurrent shared-world. Synthetic,
|
|
30
|
-
smoke and meta routes
|
|
35
|
+
smoke and meta routes never enable analysis by default and reject an explicit
|
|
36
|
+
analysis mapping before execution. `false` is accepted on every route. Dry runs show analysis
|
|
31
37
|
as skipped, without reading analysis credentials or making a provider request.
|
|
32
38
|
|
|
33
39
|
Participant execution finishes and its recording is finalized before analysis is
|
|
34
40
|
queued. A participant who was blocked or interrupted can still have useful
|
|
35
41
|
retained evidence; analysis requires a verified live recording, not a successful
|
|
36
42
|
participant outcome. An active, missing or invalid recording is not analyzed.
|
|
43
|
+
Default analysis also skips recordings containing only setup or failure records
|
|
44
|
+
with no retained participant activity. A desktop startup failure does not start
|
|
45
|
+
an analysis request. The original failure remains visible.
|
|
46
|
+
|
|
47
|
+
CLI live starts disclose the separate admission estimate limit before execution.
|
|
48
|
+
`humanish lab preflight <lab> --json` and the TUI lab screen also expose the
|
|
49
|
+
resolved budget without dispatching analysis. Library callers can inspect
|
|
50
|
+
`resolveAutomaticAnalysis` or `automaticAnalysisBudget` before running.
|
|
37
51
|
|
|
38
52
|
The command waits for analysis and reports its separate state. A TUI-launched
|
|
39
53
|
runner continues after the TUI closes; reopening the TUI or Observer reads the
|
|
@@ -54,7 +68,11 @@ accepted. Known usage is retained; missing usage remains unknown.
|
|
|
54
68
|
|
|
55
69
|
The CLI's JSON keeps `runOk` for the original backend result, `automaticAnalysis`
|
|
56
70
|
for post-run analysis, and `ok` for the overall request. Failed, cancelled or
|
|
57
|
-
unknown analysis produces exit code 2 without discarding the recording.
|
|
71
|
+
unknown analysis produces exit code 2 without discarding the recording. A missing
|
|
72
|
+
`OPENAI_API_KEY` skips default analysis and preserves a successful run exit;
|
|
73
|
+
`automaticAnalysisTrigger: "default"` distinguishes that case in JSON. A missing
|
|
74
|
+
key for an explicitly configured analysis remains a failed overall request.
|
|
75
|
+
Either skip is retained for review and does not retry automatically. Partial
|
|
58
76
|
findings remain visibly partial; a valid partial result can succeed, while a
|
|
59
77
|
partial result with an analysis error still fails the command. Recorded task
|
|
60
78
|
outcomes and the deterministic review verdict are never rewritten by analysis.
|
package/docs/ramp/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Status: public-safe contributor and agent ramp.
|
|
4
4
|
|
|
5
|
-
Package/source version in this tree: `0.
|
|
5
|
+
Package/source version in this tree: `0.91.1` (2026-09-15). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
|
|
6
6
|
`0.15.1` remains in force: managed run and output paths bind to validated
|
|
7
7
|
physical filesystem identities, and stored provider IDs are evidence, not
|
|
8
8
|
cleanup authority. The bundled OSS meta-lab is dry-run only until
|
|
@@ -47,6 +47,15 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
|
|
|
47
47
|
|
|
48
48
|
## Current State
|
|
49
49
|
|
|
50
|
+
The [0.91.1 release note](../release/0.91.1-study-review-polish.md) describes
|
|
51
|
+
retained analysis costs in study totals, finding previews and caveats, final
|
|
52
|
+
active-page viewport measurements, and corrected setup guidance.
|
|
53
|
+
|
|
54
|
+
The [0.91.0 release note](../release/0.91.0-analysis-quality-and-defaults.md)
|
|
55
|
+
describes automatic analysis by default on supported live recordings, its
|
|
56
|
+
separate disclosed budget and opt-out, fairer evidence selection, and
|
|
57
|
+
evidence-linked review of material concerns and exclusions.
|
|
58
|
+
|
|
50
59
|
The [0.90.0 release note](../release/0.90.0-automatic-analysis.md) describes
|
|
51
60
|
opt-in analysis after live runs, truthful job states, cancellation, and scripted
|
|
52
61
|
captures and assignments in findings and playback.
|
|
@@ -99,7 +108,7 @@ Implemented:
|
|
|
99
108
|
|
|
100
109
|
- `commander` CLI with stable command help;
|
|
101
110
|
- `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, `analyze`, and `feedback`;
|
|
102
|
-
-
|
|
111
|
+
- study analysis with bounded provider admission, a live-run default and opt-out, immutable findings,
|
|
103
112
|
source-bound corrections and evidence-linked Observer review;
|
|
104
113
|
- synthetic run bundles;
|
|
105
114
|
- public-safety verification with machine-readable `shareSafety.status`
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# Humanish 0.91.0
|
|
2
|
+
|
|
3
|
+
Supported live studies now request analysis when the recording finishes, using
|
|
4
|
+
`gpt-6-astra` with high reasoning and a separate $3 admission estimate limit.
|
|
5
|
+
The CLI, preflight and TUI disclose that budget. Set `review.analysis: false` to
|
|
6
|
+
disable the extra request, or supply an analysis mapping with `maxCostUsd` to
|
|
7
|
+
customize it. The estimate is additional to participant and desktop costs and
|
|
8
|
+
is not a provider billing cap. Missing default credentials record a skip while
|
|
9
|
+
preserving a successful recording; explicit analysis failures remain failures.
|
|
10
|
+
Startup failures with no retained participant activity skip default analysis.
|
|
11
|
+
|
|
12
|
+
Analysis now reviews material participant concerns before ranking findings.
|
|
13
|
+
Useful reported uncertainty and recovered mistakes can remain findings even
|
|
14
|
+
when product fault is unproven. **Concerns considered** shows evidence-linked
|
|
15
|
+
decisions, including why material concerns were left out of the ranked list.
|
|
16
|
+
It is model assessment, separate from original participant feedback and human
|
|
17
|
+
reviewer annotations.
|
|
18
|
+
|
|
19
|
+
Evidence selection shares count and text budgets across participants and samples
|
|
20
|
+
their whole sessions, prioritizing beginnings, endings and context around
|
|
21
|
+
recorded failures. Original frame and event identities remain intact. Missing
|
|
22
|
+
assignments remain explicit; historical assignments are never invented.
|
|
23
|
+
Existing analysis versions remain readable and unchanged.
|
|
24
|
+
|
|
25
|
+
Selection is bounded and can still omit important moments. Concern review does
|
|
26
|
+
not certify exhaustive discovery or accurate causal diagnosis. Opening Observer,
|
|
27
|
+
exporting a recording and reading history never start a model request. Dry-runs
|
|
28
|
+
and unsupported default routes remain request-free, and examples advertised as
|
|
29
|
+
zero-spend explicitly disable automatic analysis.
|
|
30
|
+
|
|
31
|
+
See the [analysis contract](../contracts/study-analysis.md),
|
|
32
|
+
[automatic analysis guide](../product/automatic-analysis.md), and
|
|
33
|
+
[bounded evaluation receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/analysis-quality-and-defaults-2026-09-15.md).
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Humanish 0.91.1
|
|
2
|
+
|
|
3
|
+
Study cost summaries include retained analysis attempts alongside participant
|
|
4
|
+
and desktop estimates. Reusing a saved analysis does not add another charge;
|
|
5
|
+
separate retries do. Missing prices and incomplete histories remain explicit.
|
|
6
|
+
The existing stats JSON fields keep their original participant-and-desktop
|
|
7
|
+
meaning; additive cost fields provide the combined retained estimate.
|
|
8
|
+
|
|
9
|
+
Finding previews prefer directly cited visual evidence over nearby context
|
|
10
|
+
captures. Selection is deterministic, but does not establish that the chosen
|
|
11
|
+
image is the best illustration of an issue. Exact source links and reviewer
|
|
12
|
+
corrections remain available.
|
|
13
|
+
|
|
14
|
+
Report qualifications are easier to scan, with the full limitations available
|
|
15
|
+
through a disclosure. This changes presentation without changing saved analysis
|
|
16
|
+
or participant feedback.
|
|
17
|
+
|
|
18
|
+
Final hosted Chromium geometry follows the active page after tab changes.
|
|
19
|
+
Valid CSS viewport measurements survive missing page-reported outer-window
|
|
20
|
+
bounds. Missing measurements remain missing; physical screen dimensions never
|
|
21
|
+
substitute for the page viewport.
|
|
22
|
+
|
|
23
|
+
The own-app guide now describes the hosted reachability probe and its desktop
|
|
24
|
+
cost. Starter comments document the existing opt-in mobile emulation settings.
|
|
25
|
+
|
|
26
|
+
See the [cost accounting contract](../contracts/study-costs.md) and the
|
|
27
|
+
[verification receipt](https://github.com/danielgwilson/humanish/blob/main/docs/goals/computer-use-actor/receipts/study-review-polish-2026-09-15.md)
|
|
28
|
+
for the methods and limits of these checks.
|
package/package.json
CHANGED