humanish 0.88.2 → 0.89.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/README.md +8 -0
  2. package/dist/export-bundle.js +6 -0
  3. package/dist/export-bundle.js.map +1 -1
  4. package/dist/export.js +67 -10
  5. package/dist/export.js.map +1 -1
  6. package/dist/feedback.d.ts +10 -0
  7. package/dist/feedback.js +58 -1
  8. package/dist/feedback.js.map +1 -1
  9. package/dist/index.d.ts +4 -0
  10. package/dist/index.js +2 -0
  11. package/dist/index.js.map +1 -1
  12. package/dist/observer-app.html +10 -9
  13. package/dist/observer.d.ts +2 -0
  14. package/dist/observer.js +52 -5
  15. package/dist/observer.js.map +1 -1
  16. package/dist/pricing.js +2 -2
  17. package/dist/pricing.js.map +1 -1
  18. package/dist/program.js +146 -3
  19. package/dist/program.js.map +1 -1
  20. package/dist/run.d.ts +3 -1
  21. package/dist/run.js +46 -18
  22. package/dist/run.js.map +1 -1
  23. package/dist/study-analysis-engine.d.ts +34 -0
  24. package/dist/study-analysis-engine.js +217 -0
  25. package/dist/study-analysis-engine.js.map +1 -0
  26. package/dist/study-analysis-evidence.d.ts +27 -0
  27. package/dist/study-analysis-evidence.js +420 -0
  28. package/dist/study-analysis-evidence.js.map +1 -0
  29. package/dist/study-analysis-provider.d.ts +38 -0
  30. package/dist/study-analysis-provider.js +140 -0
  31. package/dist/study-analysis-provider.js.map +1 -0
  32. package/dist/study-analysis-service.d.ts +45 -0
  33. package/dist/study-analysis-service.js +207 -0
  34. package/dist/study-analysis-service.js.map +1 -0
  35. package/dist/study-analysis-sharing.d.ts +11 -0
  36. package/dist/study-analysis-sharing.js +27 -0
  37. package/dist/study-analysis-sharing.js.map +1 -0
  38. package/dist/study-analysis-store.d.ts +30 -0
  39. package/dist/study-analysis-store.js +343 -0
  40. package/dist/study-analysis-store.js.map +1 -0
  41. package/dist/study-analysis-validation.d.ts +448 -0
  42. package/dist/study-analysis-validation.js +331 -0
  43. package/dist/study-analysis-validation.js.map +1 -0
  44. package/dist/study-analysis.d.ts +167 -0
  45. package/dist/study-analysis.js +4 -0
  46. package/dist/study-analysis.js.map +1 -0
  47. package/docs/architecture/observer.md +10 -1
  48. package/docs/contracts/feedback.md +9 -3
  49. package/docs/contracts/schemas.md +4 -1
  50. package/docs/contracts/study-analysis.md +150 -0
  51. package/docs/goals/current.md +4 -3
  52. package/docs/ramp/README.md +12 -2
  53. package/docs/release/0.89.0-study-findings.md +79 -0
  54. package/docs/release/0.89.1-analysis-finished-notice.md +24 -0
  55. package/package.json +3 -2
@@ -0,0 +1,150 @@
1
+ # Study analysis
2
+
3
+ Study analysis is an optional interpretation of retained participant evidence.
4
+ It is separate from the participant's account, recorded outcome, and the run's
5
+ deterministic review verdict. Opening an Observer never starts a provider request.
6
+
7
+ ## Invocation
8
+
9
+ ```bash
10
+ humanish analyze --run latest --max-cost 3 --dry-run --json
11
+ humanish analyze --run latest --max-cost 3 --json
12
+ humanish observe --run latest
13
+ humanish analyze list --run latest --json
14
+ humanish analyze show --run latest --json
15
+ ```
16
+
17
+ The source must be a verified, completed live run. A dry-run contract bundle is
18
+ not a participant study. Here, `analyze --dry-run` means checking an existing
19
+ study's input and admission estimate without credentials, a provider request,
20
+ or a new analysis artifact.
21
+
22
+ The default model is `gpt-6-astra`, with high reasoning effort. A request sends selected retained text and
23
+ captures to OpenAI, without tools, redirects, provider-side response storage, or
24
+ automatic retries. `--question` adds a reviewer question; it never changes the
25
+ participant assignment. `--max-cost` is required, including for dry-run
26
+ admission. It bounds a conservative estimate, not an exact provider bill.
27
+ `--timeout-ms` and `--max-output-tokens` bound the request. An exceeded admission
28
+ estimate retains valid findings and usage but returns a partial result and a
29
+ nonzero command exit, including when that version is reused.
30
+
31
+ The defaults allow five minutes and 16,384 output tokens, including reasoning.
32
+ The analysis checks the assigned requirements against the retained end state;
33
+ an unverified essential result remains unknown even if the participant reported
34
+ success. Findings keep reported concerns and observed recovery distinct across
35
+ participants. Other supported models can be selected explicitly, but evidence
36
+ reference validation does not certify their interpretation of small visual details.
37
+
38
+ Analysis distinguishes participant actions from harness setup and accounting.
39
+ Runtime credentials or model usage do not establish that a participant made an
40
+ external call while performing their task. Observations with an action basis
41
+ must cite an action-bearing source; invalid source bases are rejected with the
42
+ attempt's status and known usage retained.
43
+
44
+ Identical source input, configuration, and prompt version reuse a current valid
45
+ analysis. `--rerun` creates another immutable version. Failed attempts do not
46
+ hide earlier valid findings. Ctrl-C cancels the request; usage remains unknown
47
+ when the provider did not report it. Cancellation cannot undo an already accepted
48
+ provider request.
49
+
50
+ Decoded participant context, evidence text, and the reviewer question are checked
51
+ for known sensitive-text patterns before dispatch. JSON escaping cannot bypass
52
+ that check. Image bytes are not treated as text for pattern matching.
53
+
54
+ Only one analysis command can own a run's `.analysis-lock` directory. Interrupted
55
+ locks are not stolen using stored PIDs. After confirming the owning command has
56
+ stopped, an operator can remove the empty lock directory and retry.
57
+
58
+ ## Evidence and findings
59
+
60
+ The packet currently admits up to 16 participants, 800 evidence items, 40 PNG
61
+ captures, 160 KiB of text and 20 MiB of images. Individual source files, image
62
+ dimensions and result sizes have separate limits. Selection follows retained
63
+ source order; it is not a statistically representative sample. Selection
64
+ omissions and unreadable or invalid capture files make declared coverage
65
+ incomplete. Coverage records file availability and selection; it does not
66
+ certify visual legibility, correct interpretation, or exhaustive issue discovery.
67
+
68
+ The standard review covers session summary, apparent intent, observed outcome,
69
+ friction, dead ends, recovery, and participant feedback. Findings are ordered by
70
+ observed task impact, replication among exposed participants, and recovery.
71
+ Impact and confidence remain separate. There is no numeric frustration or
72
+ universal priority score.
73
+
74
+ Every observation cites packet-local evidence IDs. The model cannot choose a
75
+ filesystem path or fetch another resource. Validation checks participant
76
+ membership, unique counts, quote fidelity, evidence type and reference
77
+ integrity. Visual claims require retained captures; screenshot-free evidence
78
+ opens its original event. These checks do not prove that every interpretation
79
+ is correct or every consequential issue was found.
80
+
81
+ Elapsed replay time starts at the first retained capture. It is not a video
82
+ offset. Nonvisual events retain event identity without invented frame offsets.
83
+
84
+ ## Durable records
85
+
86
+ The frozen `humanish.observer-data.v1` schema is unchanged. A companion
87
+ `humanish.study-analysis.v1` artifact records source/config/input hashes,
88
+ participant context, evidence manifest, coverage, provider/model/prompt version,
89
+ status, usage and validated findings:
90
+
91
+ ```text
92
+ .humanish/runs/<run>/
93
+ analysis/<analysis>/analysis.json
94
+ analysis/<analysis>/corrections/<correction>/correction.json
95
+ analysis-attempts/<analysis>/receipt.json
96
+ observer/study-analysis.json
97
+ ```
98
+
99
+ Version and correction directories are claimed exclusively; publication is
100
+ atomic. Source evidence is not rewritten. Minimal execution receipts retain
101
+ model, budget, status and known usage even if source changes prevent report
102
+ publication. They contain no question, participant text, images, or findings.
103
+ `analyze list --json` includes these receipts.
104
+
105
+ Analysis and execution-history directories each admit 256 entries, including
106
+ interrupted writes; correction history admits 256 entries per analysis. A new
107
+ attempt requires readable inventories with room for its records before dispatch.
108
+ Valid reuse remains available at capacity. The command does not remove old
109
+ versions automatically. A present correction that cannot be read or validated
110
+ blocks sharing and feedback promotion until the history can be checked.
111
+
112
+ States distinguish no analysis, complete with no findings, complete with
113
+ findings, partial, failed, cancelled, stale and invalid. A copied or modified
114
+ source cannot silently inherit a current analysis. Saved HTML and the HTTP
115
+ companion project validated records; generic HTTP serving of producer-owned raw
116
+ analysis/correction JSON, execution receipts, atomic write temporaries and locks
117
+ is disabled. Existing adapter captures and logs retain their contained routes.
118
+
119
+ ## Human review and sharing
120
+
121
+ ```bash
122
+ humanish analyze correct --run latest --analysis <id> --finding F1 \
123
+ --status confirmed --reason "The cited captures reproduce the blocker."
124
+ humanish analyze correct --run latest --analysis <id> --finding F1 \
125
+ --status amended --reason "The claim was too broad." --claim "A narrower supported claim."
126
+ humanish feedback issue --run latest --analysis <id> --finding F1 --repo owner/repo
127
+ ```
128
+
129
+ Corrections append against exact analysis and finding hashes. Confirming one
130
+ version does not approve a later claim. Dismissed findings cannot become feedback
131
+ drafts; amendments preserve the original and record the replacement. Feedback
132
+ drafts include source/version/evidence references and remain explicitly
133
+ independent of participant-authored candidates. No command above posts to GitHub.
134
+ Each analysis has room for 256 correction inventory entries, including interrupted
135
+ writes. A full or unsafe inventory refuses a new correction before creating its
136
+ entry; prior records remain unchanged. Analysis and correction commands share the
137
+ run lock so concurrent writers cannot overrun that bound.
138
+
139
+ The existing sharing gate scans source and derived text. Export and feedback
140
+ also check the exact in-memory analysis snapshot they include. Sensitive derived
141
+ text blocks sharing and is quarantined from Observer, while independently
142
+ verified source recordings remain viewable. Filesystem containment failures
143
+ still fail closed. Valid failed/cancelled history alone does not downgrade
144
+ sharing.
145
+
146
+ Redacted bundle export omits analysis, corrections and execution receipts and
147
+ records that omission in derivation provenance. Changed source bytes require
148
+ new analysis; old hashes and review approvals cannot survive redaction. Ordinary
149
+ legacy evidence under `analysis/` remains part of the recording and follows the
150
+ normal redaction and sharing checks.
@@ -1,9 +1,9 @@
1
1
  # Current Goals
2
2
 
3
- Status date: 2026-09-13. Published baseline: `0.88.2`.
3
+ Status date: 2026-09-15. Release baseline: `0.89.1`.
4
4
 
5
5
  This page guides work on current merged source. Published behavior is described
6
- in the [release notes](../release/0.88.2-sequential-study-budgets.md).
6
+ in the [release notes](../release/0.89.1-analysis-finished-notice.md).
7
7
  The [September 9 history](https://github.com/danielgwilson/humanish/blob/main/docs/goals/current-history-2026-09-09.md)
8
8
  preserves the former status log; its queues do not supersede this page.
9
9
 
@@ -88,7 +88,7 @@ requires decision-equivalent retained evidence and a real deletion branch.
88
88
  No first-party deletion branch has met that gate. Public demonstrations do not
89
89
  substitute for it.
90
90
 
91
- ## Current Program Truth (source `0.88.2`)
91
+ ## Current Program Truth (source `0.89.1`)
92
92
 
93
93
  | Surface | Available in merged source | Remaining boundary |
94
94
  | --- | --- | --- |
@@ -99,6 +99,7 @@ substitute for it.
99
99
  | Shared state | Sequential and concurrent single-origin shared-world studies with retained evidence | Multi-origin implementation remains gated; concurrent state change does not establish per-action causation |
100
100
  | Observer | Live/recorded views, participant assignments, action-specific links, saved moments, zoom, comparison and phone-width review | Sparse captures cannot prove every action's effect; visual comparison alone is not a controlled experiment |
101
101
  | Review and feedback | Verification grades, feedback drafts, portable HTML, redacted bundle derivatives and computer-use completion-source labels | Sharing requires the appropriate grade; participant reports and condition matches still need task adjudication |
102
+ | Study findings | Explicit `analyze`, bounded evidence selection, versioned findings, exact source links and append-only corrections within the Observer study shell | Model interpretation needs review; bounded selection and source truncation limit coverage; opening Observer never dispatches analysis |
102
103
  | TUI and serving | Detached starts, run stopping, reclamation, Observer attachment, loopback serving and run library | Stopping a process does not itself prove sandbox cleanup; TUI views over CLI `stats`/`export` remain follow-ups |
103
104
  | Off-app communication | In-sandbox email/SMS catch and digest-only thread evidence | This does not establish real-provider delivery |
104
105
  | Mobile and media | Hosted viewport/emulation, desktop geometry checks, bounded dwell and declared camera feed | Physical-device and touch fidelity remain unproven; unsupported microphone declarations are rejected |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Status: public-safe contributor and agent ramp.
4
4
 
5
- Package/source version in this tree: `0.88.2` (2026-09-13). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
5
+ Package/source version in this tree: `0.89.1` (2026-09-15). The Observer is phone-usable as a stated requirement (observer/AGENTS.md); interactive primitives start from Base UI. The Observer renderer is the observer/ workspace artifact only; the legacy string-concat renderer was deleted at cutover (#426), and rollback is a version pin to 0.42.0. The containment boundary introduced in
6
6
  `0.15.1` remains in force: managed run and output paths bind to validated
7
7
  physical filesystem identities, and stored provider IDs are evidence, not
8
8
  cleanup authority. The bundled OSS meta-lab is dry-run only until
@@ -47,6 +47,14 @@ If a change does not improve one of those loops, it probably belongs elsewhere.
47
47
 
48
48
  ## Current State
49
49
 
50
+ The [0.89.1 release note](../release/0.89.1-analysis-finished-notice.md)
51
+ clarifies that an analysis with limitations has finished.
52
+
53
+ The [0.89.0 release note](../release/0.89.0-study-findings.md) describes explicit
54
+ analysis of completed studies, versioned findings and review corrections, exact
55
+ evidence links, and Participants / Findings within one Observer shell. Analysis
56
+ remains separate from participant feedback and recorded outcomes.
57
+
50
58
  The [0.88.2 release note](../release/0.88.2-sequential-study-budgets.md)
51
59
  describes model-spend thresholds on sequential shared-world studies, blocked
52
60
  later participants, and explicit unknown-usage accounting.
@@ -86,7 +94,9 @@ pnpm humanish -- verify --run latest --json
86
94
  Implemented:
87
95
 
88
96
  - `commander` CLI with stable command help;
89
- - `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, and `feedback`;
97
+ - `init`, `doctor`, `run`, `watch`, `verify`, `review`, `runs`, `analyze`, and `feedback`;
98
+ - opt-in study analysis with bounded provider admission, immutable findings,
99
+ source-bound corrections and evidence-linked Observer review;
90
100
  - synthetic run bundles;
91
101
  - public-safety verification with machine-readable `shareSafety.status`
92
102
  (`share_ready`, `local_only`, or `blocked`);
@@ -0,0 +1,79 @@
1
+ # Humanish 0.89.0: evidence-linked study findings
2
+
3
+ Completed studies can now produce ranked findings with links to the participant
4
+ events and captures that support them. Observer keeps Participants and Findings
5
+ inside the same study shell, with the recording grid, original participant
6
+ feedback and playback controls available throughout the review.
7
+
8
+ ## Analyze a completed study
9
+
10
+ ```bash
11
+ npm install humanish@0.89.0
12
+ npx humanish analyze --run latest --max-cost 3 --dry-run --json
13
+ npx humanish analyze --run latest --max-cost 3 --json
14
+ npx humanish observe --run latest
15
+ ```
16
+
17
+ Select a verified, completed live run. The first analysis command checks admission
18
+ without credentials or a provider request; the second requires `OPENAI_API_KEY`
19
+ and sends selected text and captures to OpenAI. The required cost ceiling bounds
20
+ a conservative estimate, not the provider's invoice. Opening Observer never
21
+ starts analysis.
22
+
23
+ The standard analysis covers session summary, apparent intent, observed outcome,
24
+ friction, dead ends and recovery. Findings explain task impact, affected and
25
+ exposed participants, evidence and limitations. Participant reports, recorded
26
+ completion conditions and independent analysis remain separate. Missing evidence
27
+ can leave an outcome unknown even when the participant reported success.
28
+
29
+ Identical source input, configuration and prompt version reuse a saved analysis.
30
+ `--rerun` requests a new immutable version; `--question` adds a reviewer question.
31
+ The default is `gpt-6-astra` with high reasoning, a five-minute timeout and 16,384
32
+ output tokens including reasoning. Each attempt makes one stateless request,
33
+ without tools or automatic retries, and retains known usage or explicit
34
+ uncertainty.
35
+
36
+ ## Review findings against the recording
37
+
38
+ Participants and Findings share the study library, header and available width.
39
+ Ranked findings expand in place and open their exact source event. A screenshot
40
+ shown only as context is labeled accordingly; nonvisual evidence opens its
41
+ original event without inventing a frame. Playback links, return navigation and
42
+ phone controls retain their source context. The scrubber thumb and track now
43
+ share the same centerline, with component geometry checks at multiple sizes.
44
+
45
+ Use `analyze list` and `analyze show` to inspect history. `analyze correct` appends
46
+ a confirmed, dismissed or amended review against the exact analysis and finding
47
+ hashes. Original claims and evidence remain intact. Full correction history
48
+ refuses another write before existing records become unreadable; analysis and
49
+ correction writers share the run lock.
50
+
51
+ Feedback drafting and export validate the included analysis snapshot and retain
52
+ the existing sharing gates. Redacted bundle derivatives omit analyses tied to
53
+ the original source bytes. A local-only HTML export remains portable and clearly
54
+ labeled. Feedback commands print drafts; they do not post to GitHub.
55
+
56
+ ## Compatibility and verification
57
+
58
+ The Observer data v1 contract is unchanged. Older recordings remain readable,
59
+ and absent, failed, stale or invalid analysis does not erase participant evidence.
60
+ The CLI also explains empty participant-feedback candidates and the available
61
+ run-summary drafting path.
62
+
63
+ [The feature implementation and verification](https://github.com/danielgwilson/humanish/pull/773)
64
+ cover analysis validation and storage, sharing boundaries, the existing Observer
65
+ flows, mobile layout, component geometry and fresh-package review/export paths.
66
+ Final source-first model review covered five exposed regression cases and two
67
+ untouched controlled task families, with a separate hostile-evidence challenge.
68
+ That is bounded validation, not a general accuracy benchmark. One conservative
69
+ outcome required an explicit task-scope reconciliation; one accurate positive
70
+ recovery item did not motivate a product change.
71
+
72
+ Evidence selection follows source order and has size/count limits; upstream
73
+ truncation also constrains what can be established. Schema and reference checks
74
+ do not certify interpretation or exhaustive issue discovery. Review the cited
75
+ evidence before acting. Browser proof uses Chromium desktop and phone emulation;
76
+ physical-device and cross-browser behavior remain separate verification work.
77
+
78
+ See the [study-analysis contract](../contracts/study-analysis.md) for limits,
79
+ artifact layout, cancellation, corrections and sharing behavior.
@@ -0,0 +1,24 @@
1
+ # Humanish 0.89.1: finished analysis notice
2
+
3
+ Observer now labels a partial analysis result **“Analysis finished with
4
+ limitations.”** The previous notice described findings as covering evidence
5
+ “included so far,” which could make a finished analysis appear to be running.
6
+
7
+ A partial result can contain validated findings after evidence selection limits
8
+ or an exceeded admission estimate. Review the recorded coverage, limitations and
9
+ usage when interpreting that result. The analysis status and saved evidence keep
10
+ their existing meaning.
11
+
12
+ ## Update an existing recording
13
+
14
+ ```bash
15
+ npm install humanish@0.89.1
16
+ npx humanish observe --run latest
17
+ ```
18
+
19
+ This renders the existing study with the updated Observer. An already exported
20
+ HTML file is a saved snapshot; export it again to get the new wording. Reuse the
21
+ appropriate sharing options for that run. No new analysis request is needed.
22
+
23
+ The [0.89.0 release note](0.89.0-study-findings.md) describes the underlying study
24
+ findings feature and its verification limits.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "humanish",
3
- "version": "0.88.2",
3
+ "version": "0.89.1",
4
4
  "description": "Open-source-safe CLI for persona simulation, observer review, and public-safe feedback drafts.",
5
5
  "author": "Daniel G Wilson <daniel@danielgwilson.com>",
6
6
  "keywords": [
@@ -93,7 +93,8 @@
93
93
  "commander": "^14.0.3",
94
94
  "playwright-core": "^1.60.0",
95
95
  "pngjs": "^7.0.0",
96
- "yaml": "^2.9.0"
96
+ "yaml": "^2.9.0",
97
+ "zod": "^4.4.3"
97
98
  },
98
99
  "peerDependencies": {
99
100
  "@anthropic-ai/claude-agent-sdk": "^0.3.0",