cowork-harness 2.2.0 → 2.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +31 -12
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +15 -13
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -12
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +8 -1
  9. package/CHANGELOG.md +236 -0
  10. package/DESIGN.md +3 -3
  11. package/README.md +41 -11
  12. package/baselines/desktop-1.37937.1.json +816 -0
  13. package/baselines/prompts/cowork-system-prompt-fingerprints.json +27 -0
  14. package/dist/assert.js +77 -0
  15. package/dist/cli.js +15 -4
  16. package/dist/run/cassette.js +255 -77
  17. package/dist/run/execute.js +59 -0
  18. package/dist/run/trace-view.js +39 -1
  19. package/dist/sync/cowork-sync.js +207 -8
  20. package/dist/types.js +39 -9
  21. package/docs/README.md +3 -0
  22. package/docs/cassette.md +36 -11
  23. package/docs/debugging.md +25 -0
  24. package/docs/fidelity-gaps.md +5 -1
  25. package/docs/maintenance.md +15 -1
  26. package/docs/scenario.md +13 -10
  27. package/examples/replays/README.md +1 -1
  28. package/examples/replays/example-multiselect-gate.cassette.json +41 -46
  29. package/examples/replays/example-pdf-skill.cassette.json +90 -90
  30. package/examples/replays/hostloop-computer-links.cassette.json +62 -68
  31. package/llms.txt +3 -1
  32. package/package.json +1 -1
  33. package/schema/scenario.schema.json +28 -9
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.2.0
7
- tracks-harness: cowork-harness 2.2.0 (baseline desktop-1.34493.1)
6
+ version: 2.3.0
7
+ tracks-harness: cowork-harness 2.3.0 (baseline desktop-1.37937.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -16,14 +16,17 @@ in the shell; this skill tells you how to author scenarios, pick a fidelity tier
16
16
  path, place assertions in the right CI lane, and avoid the harness's "✓ passed ≠ actually correct"
17
17
  traps.
18
18
 
19
+ `cowork-harness` is an unofficial, independent project — not affiliated with or endorsed by
20
+ Anthropic. Say so if a user asks what it is.
21
+
19
22
  The single most important idea: **a green run is not automatically a correct run.** The harness has
20
23
  several ways to no-op a check while still producing a green run (skip an assertion on replay — now
21
24
  flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
22
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
26
  the highest-value part. Read it.
24
27
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.2.0` (baseline
26
- > `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.3.0` (baseline
29
+ > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
31
 
29
32
  ## Preflight — make sure the harness can actually run
@@ -39,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
42
 
40
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.2.0"`. **Pin `@^2.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.3.0"`. **Pin `@^2.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
46
 
44
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -309,6 +312,7 @@ them by what you're trying to prove:
309
312
  | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
310
313
  | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
311
314
  | the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
315
+ | the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
312
316
  | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
313
317
  | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
314
318
  | a context compaction happened | `compaction_occurred: true` |
@@ -402,8 +406,15 @@ the discovery/encode/record dance entirely and answer gates **live during the re
402
406
  `run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
403
407
  `cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
404
408
  refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
405
- cassette-portability pre-flight below**, so it cannot green something a paid run would reject. On a directory it reports every offender and the batch
406
- cost estimate. `lint` checks the assertion invariants (both above).
409
+ cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
410
+ guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
411
+ one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
412
+ portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
413
+ the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
414
+ takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
415
+ contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
416
+ real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
417
+ reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
407
418
 
408
419
  **Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
409
420
  Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
@@ -426,7 +437,12 @@ for a cassette that cannot verify staleness from its own location — recoverabl
426
437
 
427
438
  **Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
428
439
  scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
429
- (and `verify-run`) read the gates + offered option labels out of that run's `events.jsonl` for free. Iterate
440
+ (and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
441
+ `events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
442
+ `description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
443
+ only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
444
+ was never delivered — the views are a digest, and the record is wider (`jq` recipes in
445
+ [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
430
446
  your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
431
447
  skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
432
448
  current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
@@ -821,7 +837,7 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
821
837
  channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
822
838
  reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
823
839
  Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
824
- string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` /
840
+ string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
825
841
  `questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
826
842
  old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
827
843
  on unobserved delivery (absence of evidence is failure, not neutral).
@@ -1034,11 +1050,14 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
1034
1050
  ([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
1035
1051
  → "File delivery" has the binary-verified detail; repo-only.)
1036
1052
 
1037
- 25. **Two distinct host-inventory consent flags — a record-time one and a verify-time one.** `record
1038
- --allow-host-inventory-fixture` is the boolean consent to proceed recording a host-inheriting
1053
+ 25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
1054
+ --allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
1039
1055
  (`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
1040
1056
  `record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
1041
- fixture is the risk). That pre-spend check **warns rather than refuses when the cassette already
1057
+ fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
1058
+ scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
1059
+ pass it. Writing a recording the scan DID flag is the separate `record
1060
+ --allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
1042
1061
  exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
1043
1062
  destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
1044
1063
  before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.2.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "2.3.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.237 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.246 # match your scenario's pinned baseline's agentVersion
40
40
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
41
41
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
42
42
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.2.0"
70
+ - run: npm i -g "cowork-harness@^2.3.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -149,7 +149,7 @@ The split is not just about tokens — it decides **where each lane can run**:
149
149
  `max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
150
150
  `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
151
151
  `allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
152
- `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
152
+ `question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
153
153
  (`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
154
154
  manifest. `file_absent` is in neither class — it is live/verify-run only.
155
155
  **That list is illustrative, not the authoritative set** — more keys are replay-checkable than fit a
@@ -237,12 +237,14 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
237
237
  (`hostloop`, `protocol`) has an **empty** redaction policy — that combination commits real host paths
238
238
  the `path` scanner then hard-fails at `verify-cassettes` time. The always-on scanner remains the
239
239
  universal net (container-tier recordings can trip it too).
240
- - **Host-inheriting record refused by default — `--allow-host-inventory-fixture` is the consent.** A
241
- `protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a repo-visible cassette path would
242
- freeze THIS machine's MCP server names, agents, and account metadata into a committed fixture, so
243
- `record` refuses **before the paid spawn**. Pass `--allow-host-inventory-fixture` only when the
244
- recording session genuinely has no personal MCP servers or plugins to leak — it is a per-record
245
- boolean consent, not a pattern.
240
+ - **Host-inheriting record refused by default — `--allow-host-inventory-fixture` bypasses the
241
+ PRE-FLIGHT, not the scan.** A `protocol`/`hostloop`/`cowork`-resolving-to-hostloop record into a
242
+ repo-visible cassette path would freeze THIS machine's MCP server names, agents, and account metadata
243
+ into a committed fixture, so `record` refuses **before the paid spawn**. `--allow-host-inventory-fixture`
244
+ proceeds past that check — and nothing more: the finished recording is still scanned, and a real finding
245
+ still refuses the write and quarantines it. **You do not need to audit the session by hand to pass this
246
+ flag**; that precondition was never decidable by the operator, and the scan is the real gate. Writing a
247
+ recording the scan DID flag is the separate `--allow-host-inventory-findings`.
246
248
 
247
249
  Two details that matter for a re-record loop. The pre-spend check **warns rather than refuses when the
248
250
  cassette already exists**, deliberately: refusing there would fire on every `--rerecord-stale` pass and
@@ -340,7 +342,7 @@ jobs:
340
342
  with: { node-version: '24' }
341
343
  - uses: actions/setup-python@v5
342
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
343
- - run: npm i -g "cowork-harness@^2.2.0"
345
+ - run: npm i -g "cowork-harness@^2.3.0"
344
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
345
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
346
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -369,7 +371,7 @@ jobs:
369
371
  echo "live=true" >> "$GITHUB_OUTPUT"
370
372
  fi
371
373
  - if: steps.guard.outputs.live == 'true'
372
- run: npm i -g "cowork-harness@^2.2.0"
374
+ run: npm i -g "cowork-harness@^2.3.0"
373
375
  - if: steps.guard.outputs.live == 'true'
374
376
  run: cowork-harness run scenarios/ --output-format json
375
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.2.0`
4
- (baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.3.0`
4
+ (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -289,10 +289,10 @@ same set live from the schema.
289
289
  | Assertion | Passes when |
290
290
  |---|---|
291
291
  | `result: success \| error` | the run ended with that status |
292
- | `transcript_contains: <str>` | the assistant transcript includes the literal string |
293
- | `transcript_not_contains: <str>` | it does not |
294
- | `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose |
295
- | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) |
292
+ | `transcript_contains: <str>` | the assistant transcript includes the literal string **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
293
+ | `transcript_not_contains: <str>` | it does not **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
294
+ | `transcript_matches: <regex>` | the transcript matches (case-insensitive) — for stochastic prose **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
295
+ | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
296
296
  | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
297
297
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
298
298
  | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
@@ -344,8 +344,9 @@ same set live from the schema.
344
344
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
345
345
  | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
346
346
  | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
347
- | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
348
- | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
347
+ | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
348
+ | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
349
+ | `question_context: {when_question?, matches}` | a regex over **everything the gate showed the user** — question label + every option label + every option **description**. Use it when the sentence you need to prove reached the founder may land in any of those fields: `question_asked` sees only the question text, `question_options` compares only labels, so a phrase delivered in an option's `description` is invisible to both. `when_question` narrows; omitting it searches every gate (NOT ambiguous here — this key asks whether the text was shown, not which gate offered which set). Ask-time payload only, never a `tool_result` (a producer that also writes the phrase to its own gate-state file would otherwise false-green it). Zero gates FAILS ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
349
350
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
350
351
  | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
351
352
  | `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
@@ -369,8 +370,8 @@ same set live from the schema.
369
370
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
370
371
  | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) |
371
372
  | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
372
- | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) |
373
- | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** |
373
+ | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
374
+ | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
374
375
 
375
376
  `expect_denied: [host, …]` adds one `egress_denied` per host. Run `cowork-harness assertions --list` for this
376
377
  table from the live schema. Example: `artifact_json: { artifact: outputs/cap.json, path: me.run_id, equals: "r1" }`.
@@ -441,7 +442,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
441
442
  modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
442
443
  `allow_stall` are also kept on replay, evaluated as no-op passes.
443
444
 
444
- **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `questions_count_max`,
445
+ **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
445
446
  `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`,
446
447
  `path_denied`, `no_path_denied` (the latter three are also `fidelity: hostloop`-only — see the assertion
447
448
  table). With `controlOut` present they evaluate; on an old
@@ -545,7 +546,8 @@ reference omits). Neither list is a strict superset of the other — reach for t
545
546
  `egress_denied` is dropped. Both now warn loudly. → put egress/live-only checks on a live gate; one
546
547
  concern per item; run the linter. (`LIVE_ONLY_KEYS`/`MANIFEST_KEYS` in `src/run/cassette.ts`.)
547
548
 
548
- 2. **Gate keys need a `controlOut` cassette.** `question_asked`, `questions_count_max`,
549
+ 2. **Gate keys need a `controlOut` cassette.** `question_asked`, `question_options`, `question_context`,
550
+ `questions_count_max`,
549
551
  `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked` only evaluate on
550
552
  replay with `controlOut`; on an old cassette they warn and are excluded (not passed).
551
553
  `gate_answers_delivered` **fails on
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -17,7 +17,7 @@ a paid live re-record? Walk this tree — the answer is usually no:
17
17
  `cowork-harness replay <cassette> --assert-from <scenario.yaml>`. Token-free, no re-record.
18
18
  If the recording genuinely lacks the telemetry a key needs (very old cassettes), the key fails
19
19
  **loud** as `evidence-unavailable` — that is correct behavior, not a bug; only then re-record.
20
- 2. **Gate keys** (`question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`) on a cassette
20
+ 2. **Gate keys** (`question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`) on a cassette
21
21
  **with `controlOut`** (any modern recording) → same token-free `--assert-from` path.
22
22
  3. **Gate keys** on a **pre-`controlOut`** cassette → one re-record unlocks gate asserts for that
23
23
  cassette permanently.
@@ -43,6 +43,7 @@
43
43
  "path_denied",
44
44
  "present_files_called",
45
45
  "question_asked",
46
+ "question_context",
46
47
  "question_options",
47
48
  "questions_count_max",
48
49
  "replay_protocol_fidelity",
@@ -114,6 +114,7 @@ CONTENT_KEYS = {
114
114
  GATE_KEYS = {
115
115
  "question_asked",
116
116
  "question_options",
117
+ "question_context",
117
118
  "questions_count_max",
118
119
  "gate_answers_delivered",
119
120
  "gate_answer_count_min",
@@ -284,6 +285,7 @@ def _load_top_level_keys():
284
285
  # every valid top-level scenario key (generated from the zod ScenarioObject schema; see _load_top_level_keys)
285
286
  TOP_LEVEL_KEYS = _load_top_level_keys()
286
287
  REGEX_KEYS = {
288
+ "matches",
287
289
  "transcript_matches",
288
290
  "transcript_not_matches",
289
291
  "when_question",
@@ -754,7 +756,11 @@ def lint_doc(doc, path, raw_lines):
754
756
  # YAML 1.1 but are rejected by the loader as strings — see _numeric for the same dialect gap.)
755
757
  delivered_values = _assert_values(items, "gate_answers_delivered")
756
758
  if any(v is not False for v in delivered_values):
757
- has_companion = "question_asked" in assert_keys or "question_options" in assert_keys or any(
759
+ has_companion = (
760
+ "question_asked" in assert_keys
761
+ or "question_options" in assert_keys
762
+ or "question_context" in assert_keys
763
+ ) or any(
758
764
  (n := _numeric(v)) is not None and n >= 1 for v in _assert_values(items, "gate_answer_count_min")
759
765
  )
760
766
  if not has_companion:
@@ -837,6 +843,7 @@ def lint_doc(doc, path, raw_lines):
837
843
  ),
838
844
  ("question_asked", "question_asked" in assert_keys),
839
845
  ("question_options", "question_options" in assert_keys),
846
+ ("question_context", "question_context" in assert_keys),
840
847
  ("gate_answers_delivered: false", any(v is False for v in _assert_values(items, "gate_answers_delivered"))),
841
848
  ],
842
849
  "a delivered gate records at least one question, so requiring a gate to be present contradicts requiring zero questions",
package/CHANGELOG.md CHANGED
@@ -4,6 +4,242 @@ All notable changes to this project are documented here. The format is based on
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
5
5
  [Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
6
6
 
7
+ ## [Unreleased]
8
+
9
+ ## [2.3.0] — 2026-08-26
10
+
11
+ ### Parity
12
+
13
+ - **Baseline `desktop-1.37937.1` (agent `2.1.246`).** `sync` had been refusing to write on three
14
+ `spawn.env` deltas; all three are now classified.
15
+
16
+ **Two new pinned keys.** The Cowork spawn sets `CLAUDE_CODE_PROMPT_CACHE_TTL="1h"` and
17
+ `CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL="5m"` unconditionally — no gate, no session or deployment
18
+ branch — so every first-party session receives them. They are **additive**:
19
+ `ENABLE_PROMPT_CACHING_1H="1"` is still set alongside. They read zero times in agent `2.1.241` and six
20
+ times each in `2.1.246`, so the contract went live one agent release after Desktop began setting it.
21
+
22
+ **`MCP_TOOL_TIMEOUT` is now classified per SITE, not per key.** Its first-party construction is
23
+ unchanged (still resolving to `180000`), but the third-party-only branch gained a second, settings-
24
+ conditional construction of the same key whose value expression the const resolver cannot reach.
25
+ Allowlisting the key — the obvious fix — would have been a silent contract loss: the allowlist is
26
+ checked *before* the pin list, so the key would have vanished from the generated env entirely, and it
27
+ is not a `REQUIRED_SPAWN_KEYS` member, so nothing would have hard-failed. Instead the 3p-only branch is
28
+ located by content and its inner keys are classified by name without resolving their values, which is
29
+ what the branch already meant. A brand-new key there still hard-fails.
30
+
31
+ **`CLAUDE_PREVIEW_CLASSIFIER_FLOOR` is now inert on the shipping agent** — recorded, not changed. Agent
32
+ `2.1.246` renamed the flag it reads to `CLAUDE_CHROME_CLASSIFIER_FLOOR` (and its consumer field
33
+ `previewClassifierFloorEnabled` → `chromeClassifierFloorEnabled`) while Desktop still sets only the old
34
+ name, so the classifier floor now falls through to its GrowthBook default. The key stays pinned: the
35
+ baseline records what the spawn constructs, and a Desktop-side rename must surface as a diff line
36
+ rather than as silence. Nothing in the harness reads it behaviourally. Together with the cache-TTL keys
37
+ above this is the same lesson pointing both ways — Desktop and the agent version the spawn env
38
+ independently, so "Desktop sets X" and "the agent reads X" are separately-dated claims.
39
+
40
+ Also found in the same pass and recorded in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md), with no
41
+ harness surface: the deliberately-unmodeled remote device-tool family gained `device_fs` and
42
+ `device_request_delete_permission` plus a folder-access announce mode — the Desktop half of the six
43
+ `cowork_*` risk categories that had appeared in agent `2.1.237`'s auto-mode rubric.
44
+
45
+ Also verified unchanged against the retained previous asar: the Cowork system prompt (the retained raw
46
+ files for 1.32885.1 through 1.37937.1 are byte-identical on disk), both sub-agent appends, `tools[]`,
47
+ the `canUseTool` chain, the mount modes, the egress contract, and the VM rootfs image.
48
+
49
+ - **All three committed example cassettes re-recorded against this baseline** — `example-pdf-skill`
50
+ (`container`), `example-multiselect-gate` (`protocol`) and `hostloop-computer-links` (`hostloop`).
51
+ The container one shows no behavioural change (transcript wording only). `verify-cassettes` is clean,
52
+ and the recorded MCP inventory is the Cowork lane's own servers (`cowork`, `plugins`, `skills`,
53
+ `workspace`) — no host inventory reached the fixtures, checked independently of the built-in scan
54
+ because the 2026-08-04 leak hid from a `mcp__` grep and surfaced only through NAME fields.
55
+
56
+ - **Live end-to-end pass re-run against this baseline.** `npm run test:live` — **4 suites, 19 assertions,
57
+ 19 green, 0 skipped** — across `protocol`, `container` and `hostloop`, on agent `2.1.246`. Every
58
+ `describe` in that lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated
59
+ case reports as skipped rather than passing vacuously; zero skips means the whole population executed.
60
+ `DESIGN.md`'s scope note is re-stamped accordingly and now records that no baseline is unverified for
61
+ want of a live run. Not claimed: the `boundary-check` sandbox proof and the example-scenario suite were
62
+ part of the previous stamp and were not run here.
63
+
64
+ ### Added
65
+
66
+ - **`question_context: {when_question?, matches}` — assert what a gate actually put in front of the user.**
67
+ A regex tested against a gate's founder-visible payload FIELD BY FIELD — the question label, every option
68
+ **label**, and every option **description**, each as a separate string. Matching per field rather than over
69
+ one joined blob is deliberate and load-bearing: a pattern cannot straddle two fields, so
70
+ `invoicing[\s\S]*Audit logging` will not match by stitching one option's description to the next option's
71
+ label — a "sentence" nobody was shown. The neighbouring transcript keys' docs teach `[\s\S]` for spanning
72
+ turns, so that is the habit an author brings here. `matches` is required NON-EMPTY: an empty pattern
73
+ compiles to `//i` and would green any run that fired a gate at all. `question_asked` matches question text only and `question_options`
74
+ compares labels only, so a sentence the model delivered inside an option's `description` was invisible to
75
+ every assertion key — a false-negative generator for any skill that puts context there, which the tool's
76
+ own schema invites. Measured on a consumer's paid run: a producer-authored sentence arrived verbatim in
77
+ the question, reworded inside it, and relocated into the proceed option's `description` across three runs
78
+ of one scenario; the third redded a lane on a run where the founder had in fact been told.
79
+
80
+ Evidence is the **ask-time** `AskUserQuestion` payload, never a `tool_result` — a skill's producer
81
+ typically also writes the same sentence into its own gate-state file, so `tool_result_matches` on that
82
+ phrase grades true whether or not the model ever surfaced it. Unlike `question_options`, omitting
83
+ `when_question` on a multi-gate run is **not** ambiguous: this key asks whether the text was shown at all.
84
+ Zero gates recorded fails; unreadable gate evidence fails evidence-unavailable, never vacuously.
85
+
86
+ ### Changed
87
+
88
+ - **`trace --view questions` now renders each gate's offered options — labels AND `description`s** — under
89
+ an `offered:` block, sub-question-labelled on a bundled gate. It previously printed the question label
90
+ alone, so the option `description` a skill routinely puts the deciding sentence in was reachable only by
91
+ hand-reading `events.jsonl`; a reader who found nothing in the view could reasonably conclude the text was
92
+ never delivered. The payload was always recorded — this was purely a rendering gap, and the same run dir
93
+ answers the question either way. The row (`--output-format json`) gains `subQuestions[]` carrying the
94
+ untruncated ask-time options; the text view caps each description at 240 chars, so nothing is lost, only
95
+ wrapped. This pairs with `question_context`: the view is how you *find* the text, that key is how you
96
+ *gate* on it.
97
+
98
+ - **The `tool_use` blindness of `transcript_contains`/`_not_contains`/`_matches`/`_not_matches` and
99
+ `computer_links_resolve`/`_if_present` is now documented and guarded.** `semantic_matches` has carried a
100
+ ⚠️ spelling out that its corpus excludes every `tool_use` — "a rubric claim about whether a tool was
101
+ called is unassertable" — while the six keys with the identical blindness said only "the assistant
102
+ transcript". A consumer wrote a `transcript_matches` against text living in a gate question; it could not
103
+ have matched at any phrasing, and the recording fail-closed after the spend. The caveat is now a property
104
+ of an enumerable set (`TOOL_USE_BLIND_KEYS`) enforced across every surface that documents a key — the
105
+ docs tables, the zod `.describe()` behind `assertions --list` and the generated JSON schema, and the
106
+ packaged skill reference — so a newly-added blind key cannot ship without it.
107
+ - **`question_asked`/`question_options`/`question_context` now warn that they match model-authored text.**
108
+ Gate question text and option labels are composed by the model and reworded run to run. The `choose:`
109
+ side already documented this (stable leading anchor, 1-based index); the assert side documented it
110
+ nowhere. Guarded by `MODEL_AUTHORED_TEXT_KEYS`.
111
+ - **A bad regex in a NESTED assertion field is now caught at load, not after the paid spawn.** The pre-compile
112
+ pass reached only top-level string keys, so every regex one level down — `artifact_text.matches`,
113
+ `artifact_text.not_matches`, `path_denied.path_matches`, `skill_tool_used.skill`/`.tool`,
114
+ `subagent_dispatch_healthy.type`, `subagent_output_contains.match`, `task_status.match`,
115
+ `question_options.when_question`, and the new `question_context.*` — was first compiled inside the
116
+ evaluator. All eleven are now validated at load, and `test/nested-regex-leaves.test.ts` reads `assert.ts`
117
+ and fails if the evaluator compiles a nested leaf the load-time table does not carry, so the gap cannot
118
+ reopen silently. `lint`'s double-quoted-regex warning also now covers nested `matches:` leaves.
119
+ - **`diff` with a single positional now names the missing operand** instead of printing bare usage. There is
120
+ deliberately no one-argument form: `diff` is polymorphic over baselines, run dirs and cassettes, so it
121
+ would need type dispatch plus a defined source for "the committed version".
122
+ - **The two cost keys are cross-referenced.** `RunResult.cost.usd` is one invocation's SDK
123
+ `total_cost_usd`; the critique report's `costUsd.totalUsd` aggregates the task turn, the reflection turn
124
+ and both evaluator passes. Reading the wrong one returns `undefined` rather than erroring, which reads as
125
+ "no cost recorded". Kept as two shapes on purpose — collapsing them would destroy the per-phase split.
126
+
127
+ ### Fixed
128
+ - **Two spawn-contract sentinels were weaker than their names implied; both now bite.**
129
+
130
+ `checkSpawnContractFacts` pinned the `allowedTools[]` built-in head and the built-in→`mcp__` boundary
131
+ but nothing between the boundary and the closing bracket — so `mcp__plugins__search_connectors` was
132
+ added to the array and both checks stayed green. The `mcp__` membership is now pinned as a set, and the
133
+ flag names the added and removed entries instead of reporting that something "moved". (That tool is
134
+ declared only on the third-party deployment, so the first-party inventory the harness serves is
135
+ unaffected — but the addition should not have been invisible.)
136
+
137
+ `checkMountModeFacts` asserted each read-only mount with a single `regex.test`, while `uploads`,
138
+ `.claude/skills` and `.projects/<uuid>` are each built at **two** sites — the VM-loop mount-set builder
139
+ and host-loop `computeBashMounts`. Either site satisfied the check, so a one-lane `ro`→`rw` flip — a
140
+ containment change on exactly one execution tier — passed green. It now compares the site count to the
141
+ count carrying `mode:"ro"` and names the lane count in the flag. Its fifth fact — the delete-deny
142
+ resolver `?"rwd":"rw"` — had the same shape and the same two-lane reality (it went 1 site to 2 in this
143
+ release) and now guards a floor on that count: a lane losing the resolver flags, a lane gaining one
144
+ does not.
145
+
146
+ - **All eight asar sentinels now carry a committed mutation case.** They are green on the previous
147
+ release's asar too, so a green proves nothing on its own unless the checker is known to bite; five of
148
+ them (`checkCodeTripwires`, `checkWebFetchFacts`, `checkEgressContractFacts`, `checkSyspromptMapFacts`,
149
+ `checkNormalizationSanity`) had no such case. Each new one changes an **inner** character of its
150
+ anchor, because a suffix rename still satisfies a substring regex — a mutation that cannot fail is not
151
+ evidence — and asserts that the mutation actually applied.
152
+
153
+ - **`record --dry-run <dir/>` no longer announces a WARN as a refusal.** The batch arm labelled every
154
+ advisory note `⚠ would-refuse (advisory)`, but `cassettePortabilityPreflight` can only ever return
155
+ `ok`/`warn` — it has no refuse path at all — and `hostInventoryPreflight` returns `warn` whenever the
156
+ target cassette already exists, which is every re-record corpus sitting at the default path. So the
157
+ preview told operators the real `record` would refuse runs it would in fact accept, on the one arm whose
158
+ design principle is that a guess must not gate. The label now follows the verdict kind
159
+ (`⚠ would-warn (advisory)`), the "ADVISORY, not this run's verdict" footer is unchanged, and exit codes
160
+ were never affected either way.
161
+
162
+ - **The 3p-branch rule refuses to blank W1.** W1 is the window every modeled first-party key is derived
163
+ from, so a branch marker appearing there hard-fails instead of blanking; deleting real pinned keys from
164
+ the derived env with nothing failing is the worse of the two outcomes.
165
+
166
+ - **`record --dry-run` now runs every pre-spend refusal the real record runs, and one refusal moved from
167
+ after the paid run to before it.** The rehearsal re-implemented the checks by hand, so it drifted:
168
+ `hostInventoryPreflight` shipped 2026-08-04 and the commit three days later titled *"make `--dry-run`
169
+ refuse what the real record refuses"* swept in the two checks returning `string | undefined` and missed
170
+ the one returning a `{kind}` verdict — for 19 days an operator could not discover that refusal without
171
+ spending. Separately, the slug-collision refusal (*"refusing to overwrite … it belongs to scenario X"*)
172
+ sat **after** `executeScenario`: you paid for the run and were then refused, though it is a pure function
173
+ of path + `--force` + the existing cassette's name. Both now run in one shared pre-spend block.
174
+
175
+ Refusals are also uniform now. `promptPolicyRejection` threw while `hostInventoryPreflight` called
176
+ `fail()`, which `process.exit`s — and the dir-batch loop catches a throw per item but cannot catch an
177
+ exit, so a host-inventory refusal mid-batch abandoned concurrent runs already paid for.
178
+
179
+ Under `--output-format json` a directory batch carries those advisory verdicts in a `notes[]` array, kept
180
+ separate from `refusals[]` so automation cannot read a guess as a binding verdict.
181
+
182
+ **On a directory, the path-dependent verdicts are ADVISORY (`⚠ would-refuse`) and do not affect the exit
183
+ code**, because a directory target takes no `--out`: the preview would have to guess the destination, and
184
+ a guess must not gate. Measured on a real consumer, gating on that guess would have refused 26 of 27
185
+ scenarios the real record accepts. Dry-run a single scenario file with the real flags for a binding
186
+ answer. `--quiet` suppresses the notes; it never suppresses a refusal.
187
+
188
+ Also new in the batch preview: the duplicate-cassette-path refusal the real batch already had.
189
+
190
+ ### Upgrade impact
191
+
192
+ - **`--allow-host-inventory-fixture` no longer waives a MEASURED host-inventory finding.** It was one flag
193
+ doing two jobs: bypassing the pre-flight refusal (a precondition the operator cannot check — "use only
194
+ when the session has no personal MCP servers or plugins") *and* downgrading the write-time scan's refusal
195
+ to a warning. So an operator who passed it to get past the undecidable precondition also switched off the
196
+ scan that would have caught a real leak. It is now the pre-flight bypass only; the finished recording is
197
+ still scanned, and a finding still refuses the write and quarantines the recording. Writing a flagged
198
+ recording needs the new, narrower `--allow-host-inventory-findings`. **A batch recorder that passes the
199
+ old flag on a new-fixture path will now abort on a genuine finding where it previously warned and wrote.**
200
+
201
+ A `--dry-run` that reports what *would* be captured was considered and declined: the inventory does not
202
+ exist until the agent has run, so a preview would either re-implement the scanner against a hypothetical
203
+ (a second oracle free to disagree with the real one) or require the spend it was meant to avoid. Reasoning
204
+ is in [docs/cassette.md](./docs/cassette.md).
205
+
206
+ This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
207
+ no command or flag is removed and no exit-code meaning changes — one flag's consent narrows, and the
208
+ capability it shed is reachable through the new one. So it ships in a minor.
209
+
210
+ ### Documentation
211
+
212
+ Five gaps found by asking a consumer which harness properties actually changed an outcome on a real
213
+ working day, then checking whether the docs said so. Four of the five were documented only as features,
214
+ never as the failure they prevent — the sentence a reader needs to recognise their own situation.
215
+
216
+ - **The blocking gate is now stated as a blocker.** `AskUserQuestion` *blocks*: it is a question to a
217
+ human and `claude -p` has no human, so a gated skill under a plain CLI run stalls or never reaches the
218
+ code behind the gate. That made half the skill untestable, not merely awkward to test — the README had
219
+ only "untestable headless unless something answers it", buried as the last of two afterthought bullets.
220
+ - **"Assert on the run, not the output"** — a new README section naming the class of claim the harness
221
+ exists for (`subagent_tool_absent`, `dispatch_count_max`, `no_delete_in_outputs`, `subagent_file_write`)
222
+ and why no output diff can reach it: a correct run and one that quietly handed a restricted sub-agent
223
+ shell access produce byte-identical files. `subagent_tool_absent` did not appear in the README at all.
224
+ - **The raw-`events.jsonl` escape hatch is documented**, with verified `jq` recipes, in
225
+ [docs/debugging.md](./docs/debugging.md). Every documented route into the event log went through
226
+ `trace`, whose views are a digest — so the run dir's *evidence* surface is wider than any view's
227
+ *observation* surface, and wider still than the assertion catalog. Concluding "it never happened" from
228
+ a view that doesn't render the field is a false negative, and the doc now says so and shows the read.
229
+ - **Sub-agent delivery has a route.** The tier-qualified outputs contract — the reason a hand-off path
230
+ that works on one loop lands in sandbox scratch on the other — was correctly documented in
231
+ [docs/subagents.md](./docs/subagents.md) but filed under "read on demand", reachable only by someone
232
+ who already knew the answer. It now has a "Common tasks" row keyed to the symptom.
233
+ - **Replay's cost claim carries a number.** "Zero spend" is the price; the wall-clock is what makes an
234
+ always-on per-PR gate obviously affordable, and no doc stated it (well under a second per cassette).
235
+
236
+ - **The project's provenance is stated on every surface the framing travels to.** `README.md` (which
237
+ renders on the npm page), `llms.txt`, both `.claude-plugin/marketplace.json` descriptions and `SKILL.md`
238
+ now say the same thing: an independent project, not affiliated with, endorsed by, or supported by
239
+ Anthropic; it bundles no Anthropic code; it is not Cowork. `SKILL.md` additionally tells the agent to say
240
+ so when a user asks what it is. Nothing enforces the four staying in step, so an edit can still drop it
241
+ from one of them unnoticed.
242
+
7
243
  ## [2.2.0] — 2026-08-25
8
244
 
9
245
  ### Upgrade impact