humanish 0.88.1 → 0.89.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +8 -0
- package/dist/actor-contract.d.ts +1 -1
- package/dist/actor-stop-cause.js +1 -0
- package/dist/actor-stop-cause.js.map +1 -1
- package/dist/computer-use-actor.d.ts +2 -0
- package/dist/computer-use-actor.js +6 -3
- package/dist/computer-use-actor.js.map +1 -1
- package/dist/computer-use.d.ts +5 -2
- package/dist/computer-use.js +38 -3
- package/dist/computer-use.js.map +1 -1
- package/dist/cua-actor-lab.js +3 -1
- package/dist/cua-actor-lab.js.map +1 -1
- package/dist/cua-diagnostics.d.ts +1 -1
- package/dist/cua-diagnostics.js +1 -1
- package/dist/cua-diagnostics.js.map +1 -1
- package/dist/export-bundle.js +6 -0
- package/dist/export-bundle.js.map +1 -1
- package/dist/export.js +67 -10
- package/dist/export.js.map +1 -1
- package/dist/feedback.d.ts +10 -0
- package/dist/feedback.js +58 -1
- package/dist/feedback.js.map +1 -1
- package/dist/index.d.ts +5 -1
- package/dist/index.js +2 -0
- package/dist/index.js.map +1 -1
- package/dist/lab-config.d.ts +1 -1
- package/dist/observer-app.html +10 -9
- package/dist/observer.d.ts +2 -0
- package/dist/observer.js +52 -5
- package/dist/observer.js.map +1 -1
- package/dist/openai-responses-cu.d.ts +3 -0
- package/dist/openai-responses-cu.js +4 -1
- package/dist/openai-responses-cu.js.map +1 -1
- package/dist/pricing.js +2 -2
- package/dist/pricing.js.map +1 -1
- package/dist/program.js +146 -3
- package/dist/program.js.map +1 -1
- package/dist/run.d.ts +19 -1
- package/dist/run.js +151 -21
- package/dist/run.js.map +1 -1
- package/dist/shared-world-lab.d.ts +2 -0
- package/dist/shared-world-lab.js +90 -5
- package/dist/shared-world-lab.js.map +1 -1
- package/dist/study-analysis-engine.d.ts +34 -0
- package/dist/study-analysis-engine.js +217 -0
- package/dist/study-analysis-engine.js.map +1 -0
- package/dist/study-analysis-evidence.d.ts +27 -0
- package/dist/study-analysis-evidence.js +420 -0
- package/dist/study-analysis-evidence.js.map +1 -0
- package/dist/study-analysis-provider.d.ts +38 -0
- package/dist/study-analysis-provider.js +140 -0
- package/dist/study-analysis-provider.js.map +1 -0
- package/dist/study-analysis-service.d.ts +45 -0
- package/dist/study-analysis-service.js +207 -0
- package/dist/study-analysis-service.js.map +1 -0
- package/dist/study-analysis-sharing.d.ts +11 -0
- package/dist/study-analysis-sharing.js +27 -0
- package/dist/study-analysis-sharing.js.map +1 -0
- package/dist/study-analysis-store.d.ts +30 -0
- package/dist/study-analysis-store.js +343 -0
- package/dist/study-analysis-store.js.map +1 -0
- package/dist/study-analysis-validation.d.ts +448 -0
- package/dist/study-analysis-validation.js +331 -0
- package/dist/study-analysis-validation.js.map +1 -0
- package/dist/study-analysis.d.ts +167 -0
- package/dist/study-analysis.js +4 -0
- package/dist/study-analysis.js.map +1 -0
- package/docs/architecture/observer.md +10 -1
- package/docs/contracts/feedback.md +9 -3
- package/docs/contracts/schemas.md +52 -3
- package/docs/contracts/study-analysis.md +150 -0
- package/docs/goals/current.md +4 -3
- package/docs/ramp/README.md +13 -2
- package/docs/release/0.88.2-sequential-study-budgets.md +57 -0
- package/docs/release/0.89.0-study-findings.md +79 -0
- package/package.json +3 -2
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
Date: 2026-06-02 (current-state note updated 2026-07-14)
|
|
4
4
|
|
|
5
5
|
Status: reference map for the major contracts shipped through source version
|
|
6
|
-
`0.
|
|
6
|
+
`0.89.0`; it is not an exhaustive inventory of command/result envelopes. Exported types,
|
|
7
7
|
schema constants, parsers, and validators in `src/` are authoritative. Rows
|
|
8
8
|
marked "reserved" name layering intent only — no code emits or validates them
|
|
9
9
|
yet. Do not emit a reserved schema.
|
|
@@ -38,6 +38,9 @@ workflow without leaking private upstream truth into core.
|
|
|
38
38
|
| Substrate | reserved (never shipped) | none |
|
|
39
39
|
| Evidence stream | reserved (streams live inside the run bundle) | see [`run-bundle.md`](run-bundle.md) |
|
|
40
40
|
| Review | `humanish.review.v1` | `contract-proof-review` |
|
|
41
|
+
| Study analysis | `humanish.study-analysis.v1` | see [study analysis](study-analysis.md) and synthetic analysis fixtures in `tests/` |
|
|
42
|
+
| Study analysis correction | `humanish.study-analysis-correction.v1` | see [study analysis](study-analysis.md#human-review-and-sharing) |
|
|
43
|
+
| Analysis execution receipt | `humanish.analysis-execution.v1` | see [study analysis](study-analysis.md#durable-records) |
|
|
41
44
|
| Verification | `humanish.verify-result.v1` | `five-check-verify` |
|
|
42
45
|
| Policy | `humanish.policy.v1` (fixture-only; not engine-validated) | `public-safety-policy` |
|
|
43
46
|
| Feedback | `humanish.feedback.v1` | `public-safe-feedback` |
|
|
@@ -513,6 +516,16 @@ shared-world bundle adds TWO additive, optional fields to `humanish.run-bundle.v
|
|
|
513
516
|
|
|
514
517
|
SEQUENTIAL shape (`topologyMode: sequential`, #164 PR1):
|
|
515
518
|
- `sequence: [roleId, …]` — the role ids that actually took a turn, in declared order.
|
|
519
|
+
- `skippedTail` (optional, live sequential only) — `{ afterRoleId, roles,
|
|
520
|
+
cause, maxTotalUsd?, estimatedTotalUsd? }`. Each ordered `roles` entry names
|
|
521
|
+
`{ roleId, simId, streamId }` for an unstarted participant. Together with the
|
|
522
|
+
executed prefix it must account for the full declared denominator. The
|
|
523
|
+
predecessor must have a matching `harness_error`, explicit `session_error`,
|
|
524
|
+
`usage_unreported`, or measured `study_spend_limit`. Only the last cause
|
|
525
|
+
carries budget figures, using the same per-participant estimates as the
|
|
526
|
+
tracker. Blocked seats need matching simulation, stream and blocked-event
|
|
527
|
+
evidence, with no actor, trace, screenshot or invented timeline turn.
|
|
528
|
+
Historical bundles without this field still require every role in the timeline.
|
|
516
529
|
- `timeline: (checkpoint | turn)[]` — a harness-clocked, strictly alternating
|
|
517
530
|
timeline that starts `cp-baseline`, alternates checkpoint → turn → checkpoint,
|
|
518
531
|
and ends on a checkpoint:
|
|
@@ -1067,8 +1080,9 @@ request's worst-case cost. In-flight requests, retries, concurrent lanes, and
|
|
|
1067
1080
|
unreported usage can exceed or escape these estimates. These thresholds are
|
|
1068
1081
|
not hard provider billing caps and exclude desktop and target-app charges.
|
|
1069
1082
|
|
|
1070
|
-
|
|
1071
|
-
|
|
1083
|
+
Crossing `maxUsd` ends `budget_reached` / `incomplete`, including when no
|
|
1084
|
+
material action has executed. The recorded reason distinguishes prior progress
|
|
1085
|
+
from no material progress; a harness budget stop is not participant abandonment.
|
|
1072
1086
|
Crossing the shared study threshold ends `budget_reached` / `incomplete`, with
|
|
1073
1087
|
sibling lanes stopping when their next post-response check sees it. Reaching a
|
|
1074
1088
|
threshold is not proof of task completion.
|
|
@@ -1081,6 +1095,41 @@ threshold on a model `src/pricing.ts` cannot price is refused at preflight
|
|
|
1081
1095
|
(`HUMANISH_CUA_LAB_UNPRICED_CAP`) before sandbox allocation. This rate-availability
|
|
1082
1096
|
check is separate from the post-response spend check.
|
|
1083
1097
|
|
|
1098
|
+
Sequential shared-world studies (`subject.topology: shared-world` with
|
|
1099
|
+
`execution.concurrency: 1`, using clone or local-tree subjects) enforce these
|
|
1100
|
+
same per-participant and shared model thresholds. Final reported usage,
|
|
1101
|
+
including a closing request, is reconciled before admitting the next participant.
|
|
1102
|
+
After the aggregate threshold is crossed, later participants are `blocked` with
|
|
1103
|
+
a recorded skip reason; they make no model requests and add no executed turn to
|
|
1104
|
+
the checkpoint timeline. An unpriced model with a declared threshold fails
|
|
1105
|
+
before allocation with `HUMANISH_SHARED_WORLD_LAB_INVALID`.
|
|
1106
|
+
|
|
1107
|
+
On this sequential route, an otherwise completed capped interaction that returns
|
|
1108
|
+
missing or partial usage stops with `harness_error`, `stopCause: usage_unreported`, and the label
|
|
1109
|
+
“provider usage unavailable.” It is not recorded as a crossed threshold. A
|
|
1110
|
+
stalled or failed request with unknown spend is not retried by the CUA loop.
|
|
1111
|
+
The default OpenAI adapter also disables HTTP and policy-negotiation retries for
|
|
1112
|
+
these strict capped sessions. The loop cancels its owned request signal when a
|
|
1113
|
+
request ends or its timeout wins; injected providers must honor cancellation
|
|
1114
|
+
and remain responsible for their own internal dispatch. Known usage remains in the
|
|
1115
|
+
trace alongside an explicit unknown; subsequent participants do not start when
|
|
1116
|
+
the shared budget cannot be established. Reported zero input and output counts
|
|
1117
|
+
remain valid zero usage. This stricter unknown-usage policy is specific to
|
|
1118
|
+
sequential capped studies; other routes retain their existing behavior. An
|
|
1119
|
+
explicitly incomplete provider response retains its original interruption cause
|
|
1120
|
+
first, with any missing usage still recorded as unknown.
|
|
1121
|
+
|
|
1122
|
+
Sequential traces persist dated model estimates. Their run cost summary marks
|
|
1123
|
+
desktop compute as unmeasured, so the displayed model subtotal is a lower bound.
|
|
1124
|
+
The sequential route does not provide a running Observer usage stream; its final
|
|
1125
|
+
CLI and Observer projections read these persisted estimates.
|
|
1126
|
+
|
|
1127
|
+
A capped custom session must return the declared model identity on its trace.
|
|
1128
|
+
A mismatch fails orchestration and blocks later participants while preserving
|
|
1129
|
+
the participant's original outcome and the estimate for its returned model.
|
|
1130
|
+
This check does not establish which model an arbitrary custom runner actually
|
|
1131
|
+
called or retrospectively enforce a runner that ignored its cap options.
|
|
1132
|
+
|
|
1084
1133
|
These computer-use rules do not replace the terminal route's separate
|
|
1085
1134
|
`scenario.caps` cost-ledger and product-spend rules described above.
|
|
1086
1135
|
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
# Study analysis
|
|
2
|
+
|
|
3
|
+
Study analysis is an optional interpretation of retained participant evidence.
|
|
4
|
+
It is separate from the participant's account, recorded outcome, and the run's
|
|
5
|
+
deterministic review verdict. Opening an Observer never starts a provider request.
|
|
6
|
+
|
|
7
|
+
## Invocation
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
humanish analyze --run latest --max-cost 3 --dry-run --json
|
|
11
|
+
humanish analyze --run latest --max-cost 3 --json
|
|
12
|
+
humanish observe --run latest
|
|
13
|
+
humanish analyze list --run latest --json
|
|
14
|
+
humanish analyze show --run latest --json
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
The source must be a verified, completed live run. A dry-run contract bundle is
|
|
18
|
+
not a participant study. Here, `analyze --dry-run` means checking an existing
|
|
19
|
+
study's input and admission estimate without credentials, a provider request,
|
|
20
|
+
or a new analysis artifact.
|
|
21
|
+
|
|
22
|
+
The default model is `gpt-6-astra`, with high reasoning effort. A request sends selected retained text and
|
|
23
|
+
captures to OpenAI, without tools, redirects, provider-side response storage, or
|
|
24
|
+
automatic retries. `--question` adds a reviewer question; it never changes the
|
|
25
|
+
participant assignment. `--max-cost` is required, including for dry-run
|
|
26
|
+
admission. It bounds a conservative estimate, not an exact provider bill.
|
|
27
|
+
`--timeout-ms` and `--max-output-tokens` bound the request. An exceeded admission
|
|
28
|
+
estimate retains valid findings and usage but returns a partial result and a
|
|
29
|
+
nonzero command exit, including when that version is reused.
|
|
30
|
+
|
|
31
|
+
The defaults allow five minutes and 16,384 output tokens, including reasoning.
|
|
32
|
+
The analysis checks the assigned requirements against the retained end state;
|
|
33
|
+
an unverified essential result remains unknown even if the participant reported
|
|
34
|
+
success. Findings keep reported concerns and observed recovery distinct across
|
|
35
|
+
participants. Other supported models can be selected explicitly, but evidence
|
|
36
|
+
reference validation does not certify their interpretation of small visual details.
|
|
37
|
+
|
|
38
|
+
Analysis distinguishes participant actions from harness setup and accounting.
|
|
39
|
+
Runtime credentials or model usage do not establish that a participant made an
|
|
40
|
+
external call while performing their task. Observations with an action basis
|
|
41
|
+
must cite an action-bearing source; invalid source bases are rejected with the
|
|
42
|
+
attempt's status and known usage retained.
|
|
43
|
+
|
|
44
|
+
Identical source input, configuration, and prompt version reuse a current valid
|
|
45
|
+
analysis. `--rerun` creates another immutable version. Failed attempts do not
|
|
46
|
+
hide earlier valid findings. Ctrl-C cancels the request; usage remains unknown
|
|
47
|
+
when the provider did not report it. Cancellation cannot undo an already accepted
|
|
48
|
+
provider request.
|
|
49
|
+
|
|
50
|
+
Decoded participant context, evidence text, and the reviewer question are checked
|
|
51
|
+
for known sensitive-text patterns before dispatch. JSON escaping cannot bypass
|
|
52
|
+
that check. Image bytes are not treated as text for pattern matching.
|
|
53
|
+
|
|
54
|
+
Only one analysis command can own a run's `.analysis-lock` directory. Interrupted
|
|
55
|
+
locks are not stolen using stored PIDs. After confirming the owning command has
|
|
56
|
+
stopped, an operator can remove the empty lock directory and retry.
|
|
57
|
+
|
|
58
|
+
## Evidence and findings
|
|
59
|
+
|
|
60
|
+
The packet currently admits up to 16 participants, 800 evidence items, 40 PNG
|
|
61
|
+
captures, 160 KiB of text and 20 MiB of images. Individual source files, image
|
|
62
|
+
dimensions and result sizes have separate limits. Selection follows retained
|
|
63
|
+
source order; it is not a statistically representative sample. Selection
|
|
64
|
+
omissions and unreadable or invalid capture files make declared coverage
|
|
65
|
+
incomplete. Coverage records file availability and selection; it does not
|
|
66
|
+
certify visual legibility, correct interpretation, or exhaustive issue discovery.
|
|
67
|
+
|
|
68
|
+
The standard review covers session summary, apparent intent, observed outcome,
|
|
69
|
+
friction, dead ends, recovery, and participant feedback. Findings are ordered by
|
|
70
|
+
observed task impact, replication among exposed participants, and recovery.
|
|
71
|
+
Impact and confidence remain separate. There is no numeric frustration or
|
|
72
|
+
universal priority score.
|
|
73
|
+
|
|
74
|
+
Every observation cites packet-local evidence IDs. The model cannot choose a
|
|
75
|
+
filesystem path or fetch another resource. Validation checks participant
|
|
76
|
+
membership, unique counts, quote fidelity, evidence type and reference
|
|
77
|
+
integrity. Visual claims require retained captures; screenshot-free evidence
|
|
78
|
+
opens its original event. These checks do not prove that every interpretation
|
|
79
|
+
is correct or every consequential issue was found.
|
|
80
|
+
|
|
81
|
+
Elapsed replay time starts at the first retained capture. It is not a video
|
|
82
|
+
offset. Nonvisual events retain event identity without invented frame offsets.
|
|
83
|
+
|
|
84
|
+
## Durable records
|
|
85
|
+
|
|
86
|
+
The frozen `humanish.observer-data.v1` schema is unchanged. A companion
|
|
87
|
+
`humanish.study-analysis.v1` artifact records source/config/input hashes,
|
|
88
|
+
participant context, evidence manifest, coverage, provider/model/prompt version,
|
|
89
|
+
status, usage and validated findings:
|
|
90
|
+
|
|
91
|
+
```text
|
|
92
|
+
.humanish/runs/<run>/
|
|
93
|
+
analysis/<analysis>/analysis.json
|
|
94
|
+
analysis/<analysis>/corrections/<correction>/correction.json
|
|
95
|
+
analysis-attempts/<analysis>/receipt.json
|
|
96
|
+
observer/study-analysis.json
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Version and correction directories are claimed exclusively; publication is
|
|
100
|
+
atomic. Source evidence is not rewritten. Minimal execution receipts retain
|
|
101
|
+
model, budget, status and known usage even if source changes prevent report
|
|
102
|
+
publication. They contain no question, participant text, images, or findings.
|
|
103
|
+
`analyze list --json` includes these receipts.
|
|
104
|
+
|
|
105
|
+
Analysis and execution-history directories each admit 256 entries, including
|
|
106
|
+
interrupted writes; correction history admits 256 entries per analysis. A new
|
|
107
|
+
attempt requires readable inventories with room for its records before dispatch.
|
|
108
|
+
Valid reuse remains available at capacity. The command does not remove old
|
|
109
|
+
versions automatically. A present correction that cannot be read or validated
|
|
110
|
+
blocks sharing and feedback promotion until the history can be checked.
|
|
111
|
+
|
|
112
|
+
States distinguish no analysis, complete with no findings, complete with
|
|
113
|
+
findings, partial, failed, cancelled, stale and invalid. A copied or modified
|
|
114
|
+
source cannot silently inherit a current analysis. Saved HTML and the HTTP
|
|
115
|
+
companion project validated records; generic HTTP serving of producer-owned raw
|
|
116
|
+
analysis/correction JSON, execution receipts, atomic write temporaries and locks
|
|
117
|
+
is disabled. Existing adapter captures and logs retain their contained routes.
|
|
118
|
+
|
|
119
|
+
## Human review and sharing
|
|
120
|
+
|
|
121
|
+
```bash
|
|
122
|
+
humanish analyze correct --run latest --analysis <id> --finding F1 \
|
|
123
|
+
--status confirmed --reason "The cited captures reproduce the blocker."
|
|
124
|
+
humanish analyze correct --run latest --analysis <id> --finding F1 \
|
|
125
|
+
--status amended --reason "The claim was too broad." --claim "A narrower supported claim."
|
|
126
|
+
humanish feedback issue --run latest --analysis <id> --finding F1 --repo owner/repo
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
Corrections append against exact analysis and finding hashes. Confirming one
|
|
130
|
+
version does not approve a later claim. Dismissed findings cannot become feedback
|
|
131
|
+
drafts; amendments preserve the original and record the replacement. Feedback
|
|
132
|
+
drafts include source/version/evidence references and remain explicitly
|
|
133
|
+
independent of participant-authored candidates. No command above posts to GitHub.
|
|
134
|
+
Each analysis has room for 256 correction inventory entries, including interrupted
|
|
135
|
+
writes. A full or unsafe inventory refuses a new correction before creating its
|
|
136
|
+
entry; prior records remain unchanged. Analysis and correction commands share the
|
|
137
|
+
run lock so concurrent writers cannot overrun that bound.
|
|
138
|
+
|
|
139
|
+
The existing sharing gate scans source and derived text. Export and feedback
|
|
140
|
+
also check the exact in-memory analysis snapshot they include. Sensitive derived
|
|
141
|
+
text blocks sharing and is quarantined from Observer, while independently
|
|
142
|
+
verified source recordings remain viewable. Filesystem containment failures
|
|
143
|
+
still fail closed. Valid failed/cancelled history alone does not downgrade
|
|
144
|
+
sharing.
|
|
145
|
+
|
|
146
|
+
Redacted bundle export omits analysis, corrections and execution receipts and
|
|
147
|
+
records that omission in derivation provenance. Changed source bytes require
|
|
148
|
+
new analysis; old hashes and review approvals cannot survive redaction. Ordinary
|
|
149
|
+
legacy evidence under `analysis/` remains part of the recording and follows the
|
|
150
|
+
normal redaction and sharing checks.
|
package/docs/goals/current.md
CHANGED
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
# Current Goals
|
|
2
2
|
|
|
3
|
-
Status date: 2026-09-
|
|
3
|
+
Status date: 2026-09-14. Release baseline: `0.89.0`.
|
|
4
4
|
|
|
5
5
|
This page guides work on current merged source. Published behavior is described
|
|
6
|
-
in the [release notes](../release/0.
|
|
6
|
+
in the [release notes](../release/0.89.0-study-findings.md).
|
|
7
7
|
The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
|
|
8
8
|
preserves the former status log; its queues do not supersede this page.
|
|
9
9
|
|
|
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
|
|
|
88
88
|
No first-party deletion branch has met that gate. Public demonstrations do not
|
|
89
89
|
substitute for it.
|
|
90
90
|
|
|
91
|
-
## Current Program Truth (source `0.
|
|
91
|
+
## Current Program Truth (source `0.89.0`)
|
|
92
92
|
|
|
93
93
|
| Surface | Available in merged source | Remaining boundary |
|
|
94
94
|
| --- | --- | --- |
|
|
@@ -99,6 +99,7 @@ substitute for it.
|
|
|
99
99
|
| Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
|
|
100
100
|
| Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
|
|
101
101
|
| Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
|
|
102
|
+
| Study findings | Explicit `analyze`, bounded evidence selection, versioned findings, exact source links and append-only corrections within the Observer study shell | Model interpretation needs review; bounded selection and source truncation limit coverage; opening Observer never dispatches analysis |
|
|
102
103
|
| TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
|
|
103
104
|
| Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
|
|
104
105
|
| Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
|
package/docs/ramp/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Status: public-safe contributor and agent ramp.
|
|
4
4
|
|
|
5
|
-
Package/source version in this tree: `0.
|
|
5
|
+
Package/source version in this tree: `0.89.0` (2026-09-14). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
|
|
6
6
|
`0.15.1` remains in force: managed run and output paths bind to validated
|
|
7
7
|
physical filesystem identities, and stored provider IDs are evidence, not
|
|
8
8
|
cleanup authority. The bundled OSS meta-lab is dry-run only until
|
|
@@ -47,6 +47,15 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
|
|
|
47
47
|
|
|
48
48
|
## Current State
|
|
49
49
|
|
|
50
|
+
The [0.89.0 release note](../release/0.89.0-study-findings.md) describes explicit
|
|
51
|
+
analysis of completed studies, versioned findings and review corrections, exact
|
|
52
|
+
evidence links, and Participants / Findings within one Observer shell. Analysis
|
|
53
|
+
remains separate from participant feedback and recorded outcomes.
|
|
54
|
+
|
|
55
|
+
The [0.88.2 release note](../release/0.88.2-sequential-study-budgets.md)
|
|
56
|
+
describes model-spend thresholds on sequential shared-world studies, blocked
|
|
57
|
+
later participants, and explicit unknown-usage accounting.
|
|
58
|
+
|
|
50
59
|
The [0.88.1 release note](../release/0.88.1-completion-evidence-and-local-app.md)
|
|
51
60
|
describes computer-use labels that distinguish participant reports from
|
|
52
61
|
recorded condition matches. It also covers the complete npm local-app example
|
|
@@ -82,7 +91,9 @@ pnpm humanish -- verify --run latest --json
|
|
|
82
91
|
Implemented:
|
|
83
92
|
|
|
84
93
|
- `commander` CLI with stable command help;
|
|
85
|
-
- `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, and `feedback`;
|
|
94
|
+
- `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, `analyze`, and `feedback`;
|
|
95
|
+
- opt-in study analysis with bounded provider admission, immutable findings,
|
|
96
|
+
source-bound corrections and evidence-linked Observer review;
|
|
86
97
|
- synthetic run bundles;
|
|
87
98
|
- public-safety verification with machine-readable `shareSafety.status`
|
|
88
99
|
(`share_ready`, `local_only`, or `blocked`);
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
# Humanish 0.88.2: sequential studies honor model-spend thresholds
|
|
2
|
+
|
|
3
|
+
Sequential shared-world studies now enforce the `execution.caps.maxUsd` and
|
|
4
|
+
`maxTotalUsd` thresholds they previously accepted without applying. This covers
|
|
5
|
+
computer-use participants sharing a clone or local-tree subject with
|
|
6
|
+
`subject.topology: shared-world` and `execution.concurrency: 1`.
|
|
7
|
+
|
|
8
|
+
Each participant's reported usage feeds the per-participant threshold and the
|
|
9
|
+
shared study estimate. Final usage, including a closing report, is reconciled
|
|
10
|
+
before the next participant starts. A participant interrupted by a threshold
|
|
11
|
+
has `budget_reached` / `incomplete`; its reason distinguishes prior activity
|
|
12
|
+
from no material progress. Later participants blocked by the shared threshold
|
|
13
|
+
make no model requests and add no executed turn to the checkpoint timeline.
|
|
14
|
+
A closing report that crosses the threshold preserves the already recorded
|
|
15
|
+
completion condition while blocking subsequent participants.
|
|
16
|
+
The bundle retains all declared participants and an explicit blocked suffix.
|
|
17
|
+
Verification checks the suffix against its preceding interruption, participant
|
|
18
|
+
records and unchanged executed timeline; absent participants cannot masquerade
|
|
19
|
+
as budget-blocked seats.
|
|
20
|
+
|
|
21
|
+
These checks happen after a model response and before its actions or another
|
|
22
|
+
participant turn. The current request can overshoot a threshold. They estimate
|
|
23
|
+
model spend; they do not reserve future requests or cap provider invoices,
|
|
24
|
+
desktop compute, or target-app charges. A zero threshold can still allow the
|
|
25
|
+
first model request. Use an explicit dry-run for a path without provider calls.
|
|
26
|
+
|
|
27
|
+
An unpriced model with a declared threshold is refused before allocation. For
|
|
28
|
+
an otherwise completed response, missing or partial usage ends capped sequential
|
|
29
|
+
execution with `harness_error` and `usage_unreported` (“provider usage unavailable”).
|
|
30
|
+
The CUA loop sends no further participant request, retry or closing request.
|
|
31
|
+
An explicitly incomplete provider response retains its original interruption
|
|
32
|
+
cause first; its missing usage remains unknown. For these strict capped sessions,
|
|
33
|
+
the default OpenAI adapter makes one dispatch per requested turn, including
|
|
34
|
+
HTTP errors and policy negotiation. The loop cancels the request signal when
|
|
35
|
+
its timeout wins, so an outstanding transport cannot retry after the loop ends.
|
|
36
|
+
Injected providers must honor that signal and control their own dispatches;
|
|
37
|
+
cancellation cannot undo an already billed request. Known partial costs stay
|
|
38
|
+
in the trace, and an unknown shared budget blocks later participants.
|
|
39
|
+
|
|
40
|
+
Sequential traces now retain dated model estimates. The run total explicitly
|
|
41
|
+
marks desktop compute as unmeasured, so the model subtotal is a lower bound.
|
|
42
|
+
Observer labels a partial estimate as known cost with the total unknown.
|
|
43
|
+
A custom session returning a different model identity fails orchestration and
|
|
44
|
+
blocks later participants while preserving its original outcome and the
|
|
45
|
+
estimate for its returned model. This does not retrospectively enforce an
|
|
46
|
+
arbitrary custom runner that ignored its cap options.
|
|
47
|
+
|
|
48
|
+
Existing bundles are not rewritten. Uncapped sessions and other execution
|
|
49
|
+
routes retain their existing behavior. The stricter unknown-usage rule applies
|
|
50
|
+
to sequential capped studies; the sequential route still has no running
|
|
51
|
+
Observer usage stream.
|
|
52
|
+
|
|
53
|
+
Verification covers the actual participant loop with captured provider usage,
|
|
54
|
+
individual and shared thresholds, unstarted seats, zero and missing usage,
|
|
55
|
+
closing requests, model mismatches, checkpoint evidence and cleanup. These
|
|
56
|
+
contract checks establish the tested behavior; they do not establish persona
|
|
57
|
+
effectiveness, independent adoption, or exact provider billing.
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# Humanish 0.89.0: evidence-linked study findings
|
|
2
|
+
|
|
3
|
+
Completed studies can now produce ranked findings with links to the participant
|
|
4
|
+
events and captures that support them. Observer keeps Participants and Findings
|
|
5
|
+
inside the same study shell, with the recording grid, original participant
|
|
6
|
+
feedback and playback controls available throughout the review.
|
|
7
|
+
|
|
8
|
+
## Analyze a completed study
|
|
9
|
+
|
|
10
|
+
```bash
|
|
11
|
+
npm install humanish@0.89.0
|
|
12
|
+
npx humanish analyze --run latest --max-cost 3 --dry-run --json
|
|
13
|
+
npx humanish analyze --run latest --max-cost 3 --json
|
|
14
|
+
npx humanish observe --run latest
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Select a verified, completed live run. The first analysis command checks admission
|
|
18
|
+
without credentials or a provider request; the second requires `OPENAI_API_KEY`
|
|
19
|
+
and sends selected text and captures to OpenAI. The required cost ceiling bounds
|
|
20
|
+
a conservative estimate, not the provider's invoice. Opening Observer never
|
|
21
|
+
starts analysis.
|
|
22
|
+
|
|
23
|
+
The standard analysis covers session summary, apparent intent, observed outcome,
|
|
24
|
+
friction, dead ends and recovery. Findings explain task impact, affected and
|
|
25
|
+
exposed participants, evidence and limitations. Participant reports, recorded
|
|
26
|
+
completion conditions and independent analysis remain separate. Missing evidence
|
|
27
|
+
can leave an outcome unknown even when the participant reported success.
|
|
28
|
+
|
|
29
|
+
Identical source input, configuration and prompt version reuse a saved analysis.
|
|
30
|
+
`--rerun` requests a new immutable version; `--question` adds a reviewer question.
|
|
31
|
+
The default is `gpt-6-astra` with high reasoning, a five-minute timeout and 16,384
|
|
32
|
+
output tokens including reasoning. Each attempt makes one stateless request,
|
|
33
|
+
without tools or automatic retries, and retains known usage or explicit
|
|
34
|
+
uncertainty.
|
|
35
|
+
|
|
36
|
+
## Review findings against the recording
|
|
37
|
+
|
|
38
|
+
Participants and Findings share the study library, header and available width.
|
|
39
|
+
Ranked findings expand in place and open their exact source event. A screenshot
|
|
40
|
+
shown only as context is labeled accordingly; nonvisual evidence opens its
|
|
41
|
+
original event without inventing a frame. Playback links, return navigation and
|
|
42
|
+
phone controls retain their source context. The scrubber thumb and track now
|
|
43
|
+
share the same centerline, with component geometry checks at multiple sizes.
|
|
44
|
+
|
|
45
|
+
Use `analyze list` and `analyze show` to inspect history. `analyze correct` appends
|
|
46
|
+
a confirmed, dismissed or amended review against the exact analysis and finding
|
|
47
|
+
hashes. Original claims and evidence remain intact. Full correction history
|
|
48
|
+
refuses another write before existing records become unreadable; analysis and
|
|
49
|
+
correction writers share the run lock.
|
|
50
|
+
|
|
51
|
+
Feedback drafting and export validate the included analysis snapshot and retain
|
|
52
|
+
the existing sharing gates. Redacted bundle derivatives omit analyses tied to
|
|
53
|
+
the original source bytes. A local-only HTML export remains portable and clearly
|
|
54
|
+
labeled. Feedback commands print drafts; they do not post to GitHub.
|
|
55
|
+
|
|
56
|
+
## Compatibility and verification
|
|
57
|
+
|
|
58
|
+
The Observer data v1 contract is unchanged. Older recordings remain readable,
|
|
59
|
+
and absent, failed, stale or invalid analysis does not erase participant evidence.
|
|
60
|
+
The CLI also explains empty participant-feedback candidates and the available
|
|
61
|
+
run-summary drafting path.
|
|
62
|
+
|
|
63
|
+
[The feature implementation and verification](https://github.com/danielgwilson/humanish/pull/773)
|
|
64
|
+
cover analysis validation and storage, sharing boundaries, the existing Observer
|
|
65
|
+
flows, mobile layout, component geometry and fresh-package review/export paths.
|
|
66
|
+
Final source-first model review covered five exposed regression cases and two
|
|
67
|
+
untouched controlled task families, with a separate hostile-evidence challenge.
|
|
68
|
+
That is bounded validation, not a general accuracy benchmark. One conservative
|
|
69
|
+
outcome required an explicit task-scope reconciliation; one accurate positive
|
|
70
|
+
recovery item did not motivate a product change.
|
|
71
|
+
|
|
72
|
+
Evidence selection follows source order and has size/count limits; upstream
|
|
73
|
+
truncation also constrains what can be established. Schema and reference checks
|
|
74
|
+
do not certify interpretation or exhaustive issue discovery. Review the cited
|
|
75
|
+
evidence before acting. Browser proof uses Chromium desktop and phone emulation;
|
|
76
|
+
physical-device and cross-browser behavior remain separate verification work.
|
|
77
|
+
|
|
78
|
+
See the [study-analysis contract](../contracts/study-analysis.md) for limits,
|
|
79
|
+
artifact layout, cancellation, corrections and sharing behavior.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "humanish",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.89.0",
|
|
4
4
|
"description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
|
|
5
5
|
"author": "Daniel G Wilson <daniel@danielgwilson.com>",
|
|
6
6
|
"keywords": [
|
|
@@ -93,7 +93,8 @@
|
|
|
93
93
|
"commander": "^14.0.3",
|
|
94
94
|
"playwright-core": "^1.60.0",
|
|
95
95
|
"pngjs": "^7.0.0",
|
|
96
|
-
"yaml": "^2.9.0"
|
|
96
|
+
"yaml": "^2.9.0",
|
|
97
|
+
"zod": "^4.4.3"
|
|
97
98
|
},
|
|
98
99
|
"peerDependencies": {
|
|
99
100
|
"@anthropic-ai/claude-agent-sdk": "^0.3.0",
|