humanish 0.88.1 → 0.89.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/README.md +8 -0
  2. package/dist/actor-contract.d.ts +1 -1
  3. package/dist/actor-stop-cause.js +1 -0
  4. package/dist/actor-stop-cause.js.map +1 -1
  5. package/dist/computer-use-actor.d.ts +2 -0
  6. package/dist/computer-use-actor.js +6 -3
  7. package/dist/computer-use-actor.js.map +1 -1
  8. package/dist/computer-use.d.ts +5 -2
  9. package/dist/computer-use.js +38 -3
  10. package/dist/computer-use.js.map +1 -1
  11. package/dist/cua-actor-lab.js +3 -1
  12. package/dist/cua-actor-lab.js.map +1 -1
  13. package/dist/cua-diagnostics.d.ts +1 -1
  14. package/dist/cua-diagnostics.js +1 -1
  15. package/dist/cua-diagnostics.js.map +1 -1
  16. package/dist/export-bundle.js +6 -0
  17. package/dist/export-bundle.js.map +1 -1
  18. package/dist/export.js +67 -10
  19. package/dist/export.js.map +1 -1
  20. package/dist/feedback.d.ts +10 -0
  21. package/dist/feedback.js +58 -1
  22. package/dist/feedback.js.map +1 -1
  23. package/dist/index.d.ts +5 -1
  24. package/dist/index.js +2 -0
  25. package/dist/index.js.map +1 -1
  26. package/dist/lab-config.d.ts +1 -1
  27. package/dist/observer-app.html +10 -9
  28. package/dist/observer.d.ts +2 -0
  29. package/dist/observer.js +52 -5
  30. package/dist/observer.js.map +1 -1
  31. package/dist/openai-responses-cu.d.ts +3 -0
  32. package/dist/openai-responses-cu.js +4 -1
  33. package/dist/openai-responses-cu.js.map +1 -1
  34. package/dist/pricing.js +2 -2
  35. package/dist/pricing.js.map +1 -1
  36. package/dist/program.js +146 -3
  37. package/dist/program.js.map +1 -1
  38. package/dist/run.d.ts +19 -1
  39. package/dist/run.js +151 -21
  40. package/dist/run.js.map +1 -1
  41. package/dist/shared-world-lab.d.ts +2 -0
  42. package/dist/shared-world-lab.js +90 -5
  43. package/dist/shared-world-lab.js.map +1 -1
  44. package/dist/study-analysis-engine.d.ts +34 -0
  45. package/dist/study-analysis-engine.js +217 -0
  46. package/dist/study-analysis-engine.js.map +1 -0
  47. package/dist/study-analysis-evidence.d.ts +27 -0
  48. package/dist/study-analysis-evidence.js +420 -0
  49. package/dist/study-analysis-evidence.js.map +1 -0
  50. package/dist/study-analysis-provider.d.ts +38 -0
  51. package/dist/study-analysis-provider.js +140 -0
  52. package/dist/study-analysis-provider.js.map +1 -0
  53. package/dist/study-analysis-service.d.ts +45 -0
  54. package/dist/study-analysis-service.js +207 -0
  55. package/dist/study-analysis-service.js.map +1 -0
  56. package/dist/study-analysis-sharing.d.ts +11 -0
  57. package/dist/study-analysis-sharing.js +27 -0
  58. package/dist/study-analysis-sharing.js.map +1 -0
  59. package/dist/study-analysis-store.d.ts +30 -0
  60. package/dist/study-analysis-store.js +343 -0
  61. package/dist/study-analysis-store.js.map +1 -0
  62. package/dist/study-analysis-validation.d.ts +448 -0
  63. package/dist/study-analysis-validation.js +331 -0
  64. package/dist/study-analysis-validation.js.map +1 -0
  65. package/dist/study-analysis.d.ts +167 -0
  66. package/dist/study-analysis.js +4 -0
  67. package/dist/study-analysis.js.map +1 -0
  68. package/docs/architecture/observer.md +10 -1
  69. package/docs/contracts/feedback.md +9 -3
  70. package/docs/contracts/schemas.md +52 -3
  71. package/docs/contracts/study-analysis.md +150 -0
  72. package/docs/goals/current.md +4 -3
  73. package/docs/ramp/README.md +13 -2
  74. package/docs/release/0.88.2-sequential-study-budgets.md +57 -0
  75. package/docs/release/0.89.0-study-findings.md +79 -0
  76. package/package.json +3 -2
@@ -3,7 +3,7 @@
3
3
  Date: 2026-06-02 (current-state note updated 2026-07-14)
4
4
 
5
5
  Status: reference map for the major contracts shipped through source version
6
- `0.88.1`; it is not an exhaustive inventory of command/result envelopes. Exported types,
6
+ `0.89.0`; it is not an exhaustive inventory of command/result envelopes. Exported types,
7
7
  schema constants, parsers, and validators in `src/` are authoritative. Rows
8
8
  marked "reserved" name layering intent only — no code emits or validates them
9
9
  yet. Do not emit a reserved schema.
@@ -38,6 +38,9 @@ workflow without leaking private upstream truth into core.
38
38
  | Substrate | reserved (never shipped) | none |
39
39
  | Evidence stream | reserved (streams live inside the run bundle) | see [`run-bundle.md`](run-bundle.md) |
40
40
  | Review | `humanish.review.v1` | `contract-proof-review` |
41
+ | Study analysis | `humanish.study-analysis.v1` | see [study analysis](study-analysis.md) and synthetic analysis fixtures in `tests/` |
42
+ | Study analysis correction | `humanish.study-analysis-correction.v1` | see [study analysis](study-analysis.md#human-review-and-sharing) |
43
+ | Analysis execution receipt | `humanish.analysis-execution.v1` | see [study analysis](study-analysis.md#durable-records) |
41
44
  | Verification | `humanish.verify-result.v1` | `five-check-verify` |
42
45
  | Policy | `humanish.policy.v1` (fixture-only; not engine-validated) | `public-safety-policy` |
43
46
  | Feedback | `humanish.feedback.v1` | `public-safe-feedback` |
@@ -513,6 +516,16 @@ shared-world bundle adds TWO additive, optional fields to `humanish.run-bundle.v
513
516
 
514
517
  SEQUENTIAL shape (`topologyMode: sequential`, #164 PR1):
515
518
  - `sequence: [roleId, …]` — the role ids that actually took a turn, in declared order.
519
+ - `skippedTail` (optional, live sequential only) — `{ afterRoleId, roles,
520
+ cause, maxTotalUsd?, estimatedTotalUsd? }`. Each ordered `roles` entry names
521
+ `{ roleId, simId, streamId }` for an unstarted participant. Together with the
522
+ executed prefix it must account for the full declared denominator. The
523
+ predecessor must have a matching `harness_error`, explicit `session_error`,
524
+ `usage_unreported`, or measured `study_spend_limit`. Only the last cause
525
+ carries budget figures, using the same per-participant estimates as the
526
+ tracker. Blocked seats need matching simulation, stream and blocked-event
527
+ evidence, with no actor, trace, screenshot or invented timeline turn.
528
+ Historical bundles without this field still require every role in the timeline.
516
529
  - `timeline: (checkpoint | turn)[]` — a harness-clocked, strictly alternating
517
530
  timeline that starts `cp-baseline`, alternates checkpoint → turn → checkpoint,
518
531
  and ends on a checkpoint:
@@ -1067,8 +1080,9 @@ request's worst-case cost. In-flight requests, retries, concurrent lanes, and
1067
1080
  unreported usage can exceed or escape these estimates. These thresholds are
1068
1081
  not hard provider billing caps and exclude desktop and target-app charges.
1069
1082
 
1070
- A lane with material progress that crosses `maxUsd` ends `budget_reached` /
1071
- `incomplete`; the existing zero-action guard ends `gave_up` / `abandoned`.
1083
+ Crossing `maxUsd` ends `budget_reached` / `incomplete`, including when no
1084
+ material action has executed. The recorded reason distinguishes prior progress
1085
+ from no material progress; a harness budget stop is not participant abandonment.
1072
1086
  Crossing the shared study threshold ends `budget_reached` / `incomplete`, with
1073
1087
  sibling lanes stopping when their next post-response check sees it. Reaching a
1074
1088
  threshold is not proof of task completion.
@@ -1081,6 +1095,41 @@ threshold on a model `src/pricing.ts` cannot price is refused at preflight
1081
1095
  (`HUMANISH_CUA_LAB_UNPRICED_CAP`) before sandbox allocation. This rate-availability
1082
1096
  check is separate from the post-response spend check.
1083
1097
 
1098
+ Sequential shared-world studies (`subject.topology: shared-world` with
1099
+ `execution.concurrency: 1`, using clone or local-tree subjects) enforce these
1100
+ same per-participant and shared model thresholds. Final reported usage,
1101
+ including a closing request, is reconciled before admitting the next participant.
1102
+ After the aggregate threshold is crossed, later participants are `blocked` with
1103
+ a recorded skip reason; they make no model requests and add no executed turn to
1104
+ the checkpoint timeline. An unpriced model with a declared threshold fails
1105
+ before allocation with `HUMANISH_SHARED_WORLD_LAB_INVALID`.
1106
+
1107
+ On this sequential route, an otherwise completed capped interaction that returns
1108
+ missing or partial usage stops with `harness_error`, `stopCause: usage_unreported`, and the label
1109
+ “provider usage unavailable.” It is not recorded as a crossed threshold. A
1110
+ stalled or failed request with unknown spend is not retried by the CUA loop.
1111
+ The default OpenAI adapter also disables HTTP and policy-negotiation retries for
1112
+ these strict capped sessions. The loop cancels its owned request signal when a
1113
+ request ends or its timeout wins; injected providers must honor cancellation
1114
+ and remain responsible for their own internal dispatch. Known usage remains in the
1115
+ trace alongside an explicit unknown; subsequent participants do not start when
1116
+ the shared budget cannot be established. Reported zero input and output counts
1117
+ remain valid zero usage. This stricter unknown-usage policy is specific to
1118
+ sequential capped studies; other routes retain their existing behavior. An
1119
+ explicitly incomplete provider response retains its original interruption cause
1120
+ first, with any missing usage still recorded as unknown.
1121
+
1122
+ Sequential traces persist dated model estimates. Their run cost summary marks
1123
+ desktop compute as unmeasured, so the displayed model subtotal is a lower bound.
1124
+ The sequential route does not provide a running Observer usage stream; its final
1125
+ CLI and Observer projections read these persisted estimates.
1126
+
1127
+ A capped custom session must return the declared model identity on its trace.
1128
+ A mismatch fails orchestration and blocks later participants while preserving
1129
+ the participant's original outcome and the estimate for its returned model.
1130
+ This check does not establish which model an arbitrary custom runner actually
1131
+ called or retrospectively enforce a runner that ignored its cap options.
1132
+
1084
1133
  These computer-use rules do not replace the terminal route's separate
1085
1134
  `scenario.caps` cost-ledger and product-spend rules described above.
1086
1135
 
@@ -0,0 +1,150 @@
1
+ # Study analysis
2
+
3
+ Study analysis is an optional interpretation of retained participant evidence.
4
+ It is separate from the participant's account, recorded outcome, and the run's
5
+ deterministic review verdict. Opening an Observer never starts a provider request.
6
+
7
+ ## Invocation
8
+
9
+ ```bash
10
+ humanish analyze --run latest --max-cost 3 --dry-run --json
11
+ humanish analyze --run latest --max-cost 3 --json
12
+ humanish observe --run latest
13
+ humanish analyze list --run latest --json
14
+ humanish analyze show --run latest --json
15
+ ```
16
+
17
+ The source must be a verified, completed live run. A dry-run contract bundle is
18
+ not a participant study. Here, `analyze --dry-run` means checking an existing
19
+ study's input and admission estimate without credentials, a provider request,
20
+ or a new analysis artifact.
21
+
22
+ The default model is `gpt-6-astra`, with high reasoning effort. A request sends selected retained text and
23
+ captures to OpenAI, without tools, redirects, provider-side response storage, or
24
+ automatic retries. `--question` adds a reviewer question; it never changes the
25
+ participant assignment. `--max-cost` is required, including for dry-run
26
+ admission. It bounds a conservative estimate, not an exact provider bill.
27
+ `--timeout-ms` and `--max-output-tokens` bound the request. An exceeded admission
28
+ estimate retains valid findings and usage but returns a partial result and a
29
+ nonzero command exit, including when that version is reused.
30
+
31
+ The defaults allow five minutes and 16,384 output tokens, including reasoning.
32
+ The analysis checks the assigned requirements against the retained end state;
33
+ an unverified essential result remains unknown even if the participant reported
34
+ success. Findings keep reported concerns and observed recovery distinct across
35
+ participants. Other supported models can be selected explicitly, but evidence
36
+ reference validation does not certify their interpretation of small visual details.
37
+
38
+ Analysis distinguishes participant actions from harness setup and accounting.
39
+ Runtime credentials or model usage do not establish that a participant made an
40
+ external call while performing their task. Observations with an action basis
41
+ must cite an action-bearing source; invalid source bases are rejected with the
42
+ attempt's status and known usage retained.
43
+
44
+ Identical source input, configuration, and prompt version reuse a current valid
45
+ analysis. `--rerun` creates another immutable version. Failed attempts do not
46
+ hide earlier valid findings. Ctrl-C cancels the request; usage remains unknown
47
+ when the provider did not report it. Cancellation cannot undo an already accepted
48
+ provider request.
49
+
50
+ Decoded participant context, evidence text, and the reviewer question are checked
51
+ for known sensitive-text patterns before dispatch. JSON escaping cannot bypass
52
+ that check. Image bytes are not treated as text for pattern matching.
53
+
54
+ Only one analysis command can own a run's `.analysis-lock` directory. Interrupted
55
+ locks are not stolen using stored PIDs. After confirming the owning command has
56
+ stopped, an operator can remove the empty lock directory and retry.
57
+
58
+ ## Evidence and findings
59
+
60
+ The packet currently admits up to 16 participants, 800 evidence items, 40 PNG
61
+ captures, 160 KiB of text and 20 MiB of images. Individual source files, image
62
+ dimensions and result sizes have separate limits. Selection follows retained
63
+ source order; it is not a statistically representative sample. Selection
64
+ omissions and unreadable or invalid capture files make declared coverage
65
+ incomplete. Coverage records file availability and selection; it does not
66
+ certify visual legibility, correct interpretation, or exhaustive issue discovery.
67
+
68
+ The standard review covers session summary, apparent intent, observed outcome,
69
+ friction, dead ends, recovery, and participant feedback. Findings are ordered by
70
+ observed task impact, replication among exposed participants, and recovery.
71
+ Impact and confidence remain separate. There is no numeric frustration or
72
+ universal priority score.
73
+
74
+ Every observation cites packet-local evidence IDs. The model cannot choose a
75
+ filesystem path or fetch another resource. Validation checks participant
76
+ membership, unique counts, quote fidelity, evidence type and reference
77
+ integrity. Visual claims require retained captures; screenshot-free evidence
78
+ opens its original event. These checks do not prove that every interpretation
79
+ is correct or every consequential issue was found.
80
+
81
+ Elapsed replay time starts at the first retained capture. It is not a video
82
+ offset. Nonvisual events retain event identity without invented frame offsets.
83
+
84
+ ## Durable records
85
+
86
+ The frozen `humanish.observer-data.v1` schema is unchanged. A companion
87
+ `humanish.study-analysis.v1` artifact records source/config/input hashes,
88
+ participant context, evidence manifest, coverage, provider/model/prompt version,
89
+ status, usage and validated findings:
90
+
91
+ ```text
92
+ .humanish/runs/<run>/
93
+ analysis/<analysis>/analysis.json
94
+ analysis/<analysis>/corrections/<correction>/correction.json
95
+ analysis-attempts/<analysis>/receipt.json
96
+ observer/study-analysis.json
97
+ ```
98
+
99
+ Version and correction directories are claimed exclusively; publication is
100
+ atomic. Source evidence is not rewritten. Minimal execution receipts retain
101
+ model, budget, status and known usage even if source changes prevent report
102
+ publication. They contain no question, participant text, images, or findings.
103
+ `analyze list --json` includes these receipts.
104
+
105
+ Analysis and execution-history directories each admit 256 entries, including
106
+ interrupted writes; correction history admits 256 entries per analysis. A new
107
+ attempt requires readable inventories with room for its records before dispatch.
108
+ Valid reuse remains available at capacity. The command does not remove old
109
+ versions automatically. A present correction that cannot be read or validated
110
+ blocks sharing and feedback promotion until the history can be checked.
111
+
112
+ States distinguish no analysis, complete with no findings, complete with
113
+ findings, partial, failed, cancelled, stale and invalid. A copied or modified
114
+ source cannot silently inherit a current analysis. Saved HTML and the HTTP
115
+ companion project validated records; generic HTTP serving of producer-owned raw
116
+ analysis/correction JSON, execution receipts, atomic write temporaries and locks
117
+ is disabled. Existing adapter captures and logs retain their contained routes.
118
+
119
+ ## Human review and sharing
120
+
121
+ ```bash
122
+ humanish analyze correct --run latest --analysis <id> --finding F1 \
123
+ --status confirmed --reason "The cited captures reproduce the blocker."
124
+ humanish analyze correct --run latest --analysis <id> --finding F1 \
125
+ --status amended --reason "The claim was too broad." --claim "A narrower supported claim."
126
+ humanish feedback issue --run latest --analysis <id> --finding F1 --repo owner/repo
127
+ ```
128
+
129
+ Corrections append against exact analysis and finding hashes. Confirming one
130
+ version does not approve a later claim. Dismissed findings cannot become feedback
131
+ drafts; amendments preserve the original and record the replacement. Feedback
132
+ drafts include source/version/evidence references and remain explicitly
133
+ independent of participant-authored candidates. No command above posts to GitHub.
134
+ Each analysis has room for 256 correction inventory entries, including interrupted
135
+ writes. A full or unsafe inventory refuses a new correction before creating its
136
+ entry; prior records remain unchanged. Analysis and correction commands share the
137
+ run lock so concurrent writers cannot overrun that bound.
138
+
139
+ The existing sharing gate scans source and derived text. Export and feedback
140
+ also check the exact in-memory analysis snapshot they include. Sensitive derived
141
+ text blocks sharing and is quarantined from Observer, while independently
142
+ verified source recordings remain viewable. Filesystem containment failures
143
+ still fail closed. Valid failed/cancelled history alone does not downgrade
144
+ sharing.
145
+
146
+ Redacted bundle export omits analysis, corrections and execution receipts and
147
+ records that omission in derivation provenance. Changed source bytes require
148
+ new analysis; old hashes and review approvals cannot survive redaction. Ordinary
149
+ legacy evidence under `analysis/` remains part of the recording and follows the
150
+ normal redaction and sharing checks.
@@ -1,9 +1,9 @@
1
1
  # Current Goals
2
2
 
3
- Status date: 2026-09-12. Published baseline: `0.88.1`.
3
+ Status date: 2026-09-14. Release baseline: `0.89.0`.
4
4
 
5
5
  This page guides work on current merged source. Published behavior is described
6
- in the [release notes](../release/0.88.1-completion-evidence-and-local-app.md).
6
+ in the [release notes](../release/0.89.0-study-findings.md).
7
7
  The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
8
8
  preserves the former status log; its queues do not supersede this page.
9
9
 
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
88
88
  No first-party deletion branch has met that gate. Public demonstrations do not
89
89
  substitute for it.
90
90
 
91
- ## Current Program Truth (source `0.88.1`)
91
+ ## Current Program Truth (source `0.89.0`)
92
92
 
93
93
  | Surface | Available in merged source | Remaining boundary |
94
94
  | --- | --- | --- |
@@ -99,6 +99,7 @@ substitute for it.
99
99
  | Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
100
100
  | Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
101
101
  | Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
102
+ | Study findings | Explicit `analyze`, bounded evidence selection, versioned findings, exact source links and append-only corrections within the Observer study shell | Model interpretation needs review; bounded selection and source truncation limit coverage; opening Observer never dispatches analysis |
102
103
  | TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
103
104
  | Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
104
105
  | Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Status: public-safe contributor and agent ramp.
4
4
 
5
- Package/source version in this tree: `0.88.1` (2026-09-12). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
5
+ Package/source version in this tree: `0.89.0` (2026-09-14). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
6
6
  `0.15.1` remains in force: managed run and output paths bind to validated
7
7
  physical filesystem identities, and stored provider IDs are evidence, not
8
8
  cleanup authority. The bundled OSS meta-lab is dry-run only until
@@ -47,6 +47,15 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
47
47
 
48
48
  ## Current State
49
49
 
50
+ The [0.89.0 release note](../release/0.89.0-study-findings.md) describes explicit
51
+ analysis of completed studies, versioned findings and review corrections, exact
52
+ evidence links, and Participants / Findings within one Observer shell. Analysis
53
+ remains separate from participant feedback and recorded outcomes.
54
+
55
+ The [0.88.2 release note](../release/0.88.2-sequential-study-budgets.md)
56
+ describes model-spend thresholds on sequential shared-world studies, blocked
57
+ later participants, and explicit unknown-usage accounting.
58
+
50
59
  The [0.88.1 release note](../release/0.88.1-completion-evidence-and-local-app.md)
51
60
  describes computer-use labels that distinguish participant reports from
52
61
  recorded condition matches. It also covers the complete npm local-app example
@@ -82,7 +91,9 @@ pnpm humanish -- verify --run latest --json
82
91
  Implemented:
83
92
 
84
93
  - `commander` CLI with stable command help;
85
- - `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, and `feedback`;
94
+ - `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, `analyze`, and `feedback`;
95
+ - opt-in study analysis with bounded provider admission, immutable findings,
96
+ source-bound corrections and evidence-linked Observer review;
86
97
  - synthetic run bundles;
87
98
  - public-safety verification with machine-readable `shareSafety.status`
88
99
  (`share_ready`, `local_only`, or `blocked`);
@@ -0,0 +1,57 @@
1
+ # Humanish 0.88.2: sequential studies honor model-spend thresholds
2
+
3
+ Sequential shared-world studies now enforce the `execution.caps.maxUsd` and
4
+ `maxTotalUsd` thresholds they previously accepted without applying. This covers
5
+ computer-use participants sharing a clone or local-tree subject with
6
+ `subject.topology: shared-world` and `execution.concurrency: 1`.
7
+
8
+ Each participant's reported usage feeds the per-participant threshold and the
9
+ shared study estimate. Final usage, including a closing report, is reconciled
10
+ before the next participant starts. A participant interrupted by a threshold
11
+ has `budget_reached` / `incomplete`; its reason distinguishes prior activity
12
+ from no material progress. Later participants blocked by the shared threshold
13
+ make no model requests and add no executed turn to the checkpoint timeline.
14
+ A closing report that crosses the threshold preserves the already recorded
15
+ completion condition while blocking subsequent participants.
16
+ The bundle retains all declared participants and an explicit blocked suffix.
17
+ Verification checks the suffix against its preceding interruption, participant
18
+ records and unchanged executed timeline; absent participants cannot masquerade
19
+ as budget-blocked seats.
20
+
21
+ These checks happen after a model response and before its actions or another
22
+ participant turn. The current request can overshoot a threshold. They estimate
23
+ model spend; they do not reserve future requests or cap provider invoices,
24
+ desktop compute, or target-app charges. A zero threshold can still allow the
25
+ first model request. Use an explicit dry-run for a path without provider calls.
26
+
27
+ An unpriced model with a declared threshold is refused before allocation. For
28
+ an otherwise completed response, missing or partial usage ends capped sequential
29
+ execution with `harness_error` and `usage_unreported` (“provider usage unavailable”).
30
+ The CUA loop sends no further participant request, retry or closing request.
31
+ An explicitly incomplete provider response retains its original interruption
32
+ cause first; its missing usage remains unknown. For these strict capped sessions,
33
+ the default OpenAI adapter makes one dispatch per requested turn, including
34
+ HTTP errors and policy negotiation. The loop cancels the request signal when
35
+ its timeout wins, so an outstanding transport cannot retry after the loop ends.
36
+ Injected providers must honor that signal and control their own dispatches;
37
+ cancellation cannot undo an already billed request. Known partial costs stay
38
+ in the trace, and an unknown shared budget blocks later participants.
39
+
40
+ Sequential traces now retain dated model estimates. The run total explicitly
41
+ marks desktop compute as unmeasured, so the model subtotal is a lower bound.
42
+ Observer labels a partial estimate as known cost with the total unknown.
43
+ A custom session returning a different model identity fails orchestration and
44
+ blocks later participants while preserving its original outcome and the
45
+ estimate for its returned model. This does not retrospectively enforce an
46
+ arbitrary custom runner that ignored its cap options.
47
+
48
+ Existing bundles are not rewritten. Uncapped sessions and other execution
49
+ routes retain their existing behavior. The stricter unknown-usage rule applies
50
+ to sequential capped studies; the sequential route still has no running
51
+ Observer usage stream.
52
+
53
+ Verification covers the actual participant loop with captured provider usage,
54
+ individual and shared thresholds, unstarted seats, zero and missing usage,
55
+ closing requests, model mismatches, checkpoint evidence and cleanup. These
56
+ contract checks establish the tested behavior; they do not establish persona
57
+ effectiveness, independent adoption, or exact provider billing.
@@ -0,0 +1,79 @@
1
+ # Humanish 0.89.0: evidence-linked study findings
2
+
3
+ Completed studies can now produce ranked findings with links to the participant
4
+ events and captures that support them. Observer keeps Participants and Findings
5
+ inside the same study shell, with the recording grid, original participant
6
+ feedback and playback controls available throughout the review.
7
+
8
+ ## Analyze a completed study
9
+
10
+ ```bash
11
+ npm install humanish@0.89.0
12
+ npx humanish analyze --run latest --max-cost 3 --dry-run --json
13
+ npx humanish analyze --run latest --max-cost 3 --json
14
+ npx humanish observe --run latest
15
+ ```
16
+
17
+ Select a verified, completed live run. The first analysis command checks admission
18
+ without credentials or a provider request; the second requires `OPENAI_API_KEY`
19
+ and sends selected text and captures to OpenAI. The required cost ceiling bounds
20
+ a conservative estimate, not the provider's invoice. Opening Observer never
21
+ starts analysis.
22
+
23
+ The standard analysis covers session summary, apparent intent, observed outcome,
24
+ friction, dead ends and recovery. Findings explain task impact, affected and
25
+ exposed participants, evidence and limitations. Participant reports, recorded
26
+ completion conditions and independent analysis remain separate. Missing evidence
27
+ can leave an outcome unknown even when the participant reported success.
28
+
29
+ Identical source input, configuration and prompt version reuse a saved analysis.
30
+ `--rerun` requests a new immutable version; `--question` adds a reviewer question.
31
+ The default is `gpt-6-astra` with high reasoning, a five-minute timeout and 16,384
32
+ output tokens including reasoning. Each attempt makes one stateless request,
33
+ without tools or automatic retries, and retains known usage or explicit
34
+ uncertainty.
35
+
36
+ ## Review findings against the recording
37
+
38
+ Participants and Findings share the study library, header and available width.
39
+ Ranked findings expand in place and open their exact source event. A screenshot
40
+ shown only as context is labeled accordingly; nonvisual evidence opens its
41
+ original event without inventing a frame. Playback links, return navigation and
42
+ phone controls retain their source context. The scrubber thumb and track now
43
+ share the same centerline, with component geometry checks at multiple sizes.
44
+
45
+ Use `analyze list` and `analyze show` to inspect history. `analyze correct` appends
46
+ a confirmed, dismissed or amended review against the exact analysis and finding
47
+ hashes. Original claims and evidence remain intact. Full correction history
48
+ refuses another write before existing records become unreadable; analysis and
49
+ correction writers share the run lock.
50
+
51
+ Feedback drafting and export validate the included analysis snapshot and retain
52
+ the existing sharing gates. Redacted bundle derivatives omit analyses tied to
53
+ the original source bytes. A local-only HTML export remains portable and clearly
54
+ labeled. Feedback commands print drafts; they do not post to GitHub.
55
+
56
+ ## Compatibility and verification
57
+
58
+ The Observer data v1 contract is unchanged. Older recordings remain readable,
59
+ and absent, failed, stale or invalid analysis does not erase participant evidence.
60
+ The CLI also explains empty participant-feedback candidates and the available
61
+ run-summary drafting path.
62
+
63
+ [The feature implementation and verification](https://github.com/danielgwilson/humanish/pull/773)
64
+ cover analysis validation and storage, sharing boundaries, the existing Observer
65
+ flows, mobile layout, component geometry and fresh-package review/export paths.
66
+ Final source-first model review covered five exposed regression cases and two
67
+ untouched controlled task families, with a separate hostile-evidence challenge.
68
+ That is bounded validation, not a general accuracy benchmark. One conservative
69
+ outcome required an explicit task-scope reconciliation; one accurate positive
70
+ recovery item did not motivate a product change.
71
+
72
+ Evidence selection follows source order and has size/count limits; upstream
73
+ truncation also constrains what can be established. Schema and reference checks
74
+ do not certify interpretation or exhaustive issue discovery. Review the cited
75
+ evidence before acting. Browser proof uses Chromium desktop and phone emulation;
76
+ physical-device and cross-browser behavior remain separate verification work.
77
+
78
+ See the [study-analysis contract](../contracts/study-analysis.md) for limits,
79
+ artifact layout, cancellation, corrections and sharing behavior.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.88.1",
3
+ "version": "0.89.0",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [
@@ -93,7 +93,8 @@
93
93
  "commander": "^14.0.3",
94
94
  "playwright-core": "^1.60.0",
95
95
  "pngjs": "^7.0.0",
96
- "yaml": "^2.9.0"
96
+ "yaml": "^2.9.0",
97
+ "zod": "^4.4.3"
97
98
  },
98
99
  "peerDependencies": {
99
100
  "@anthropic-ai/claude-agent-sdk": "^0.3.0",