cowork-harness 1.2.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +68 -13
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +8 -5
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +7 -6
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +27 -0
  6. package/CHANGELOG.md +131 -0
  7. package/README.md +19 -17
  8. package/SPEC.md +33 -5
  9. package/baselines/desktop-1.22209.3.json +392 -0
  10. package/dist/agent/session.js +77 -8
  11. package/dist/assert.js +125 -37
  12. package/dist/cli.js +95 -37
  13. package/dist/decide/decider.js +52 -6
  14. package/dist/decide/external-channel.js +19 -6
  15. package/dist/hostloop/webfetch-dedup.js +50 -0
  16. package/dist/hostloop/workspace-handler.js +42 -4
  17. package/dist/loop-decision.js +30 -0
  18. package/dist/run/analyze-artifact-runtime.js +309 -65
  19. package/dist/run/analyze-artifact.js +581 -206
  20. package/dist/run/analyze-skill.js +182 -37
  21. package/dist/run/artifacts.js +95 -19
  22. package/dist/run/cassette.js +37 -1
  23. package/dist/run/chat-result.js +3 -0
  24. package/dist/run/chat.js +10 -1
  25. package/dist/run/diff.js +2 -0
  26. package/dist/run/execute.js +31 -8
  27. package/dist/run/inspect-view.js +19 -0
  28. package/dist/run/pre-run-manifest.js +39 -8
  29. package/dist/run/renderer.js +11 -0
  30. package/dist/run/run-index.js +16 -5
  31. package/dist/run/run-status.js +1 -0
  32. package/dist/run/run.js +49 -7
  33. package/dist/run/trace-view.js +27 -7
  34. package/dist/run/verdict.js +19 -0
  35. package/dist/runtime/hostloop.js +1 -0
  36. package/dist/types.js +2 -3
  37. package/docs/README.md +5 -1
  38. package/docs/cassette.md +9 -2
  39. package/docs/debugging.md +34 -0
  40. package/docs/decisions/README.md +7 -0
  41. package/docs/gotchas.md +9 -0
  42. package/docs/maintenance.md +1 -1
  43. package/docs/scenario.md +55 -4
  44. package/docs/stats.md +5 -1
  45. package/docs/subagents.md +7 -3
  46. package/examples/replays/README.md +1 -1
  47. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  48. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  49. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  50. package/llms.txt +1 -0
  51. package/package.json +1 -1
  52. package/python/README.md +13 -4
  53. package/schema/run-result.json +19 -1
  54. package/schema/scenario.schema.json +2 -2
  55. package/scripts/check-baseline-staleness.ts +147 -0
@@ -1,10 +1,10 @@
1
1
  ---
2
2
  name: cowork-harness
3
- description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
3
+ description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.2.0
7
- tracks-harness: cowork-harness 1.2.0 (baseline desktop-1.21459.0)
6
+ version: 1.4.0
7
+ tracks-harness: cowork-harness 1.4.0 (baseline desktop-1.22209.3)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.2.0` (baseline
26
- > `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.4.0` (baseline
26
+ > `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,9 +39,9 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.2.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.2.0"`. **Pin `@>=1.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.4.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.4.0"`. **Pin `@>=1.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.2.0 floor gates, by release:
44
+ What the ≥ 1.4.0 floor gates, by release:
45
45
 
46
46
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
47
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
@@ -54,6 +54,8 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
54
54
  - **1.0.0:** first stable release — the SPEC §12 compatibility contract takes effect (covered CLI/schema/env/Action surfaces are now stable; breaking changes need a major bump). No new author-facing command; the floor simply tracks the 1.0 release.
55
55
  - **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
56
56
  - **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
57
+ - **1.3.0:** run-identity for the iterate-across-fixes loop — `skill`/`run` take `--label <tag>` (a generation tag surfaced in `result.json` `runLabel`, the run-index row, `inspect`, and `status.json`) alongside an auto-recorded `skillCommit`, layered on the **authoritative** content-exact `fingerprint.skillHash` a harvest step should pair critiques by (`inspect`/index surface a short `skillHash` prefix). `trace --full-results` captures the full input+result of **every** tool call — successful ones too, not just errors — so an external grader can ground a self-critique finding against the call it cites. `verify-run` now **warns** on skillHash drift for answer-less scenarios (was silent unless the scenario declared scripted `answers`). Plus `skill --allow-missing-capability` (the open-ended-run opt-out for a capability FALSE-NEGATIVE on the lean `core` image), a new **warn-severity `ended_with_question`** verdict signal (the agent's final answer contains a question and the run wrote no `outputs/` deliverable — the lenient sibling of the strict `stalled`), and LLM-decider `OTHER:` free-text answers now marked `[via Other free-text]` in gate provenance.
58
+ - **1.4.0:** platform baseline synced to Desktop **1.22209.3** (agent 2.1.215; no prompt/spawn/egress drift), and **`coworkWebFetchDedup` is now enacted** on the hostloop `web_fetch` path — a repeat fetch of the same normalized URL within the TTL (15 min, cap 100) makes **no network request** and returns a marker to re-use the earlier result, with no egress event. Baseline-gated (Desktop ≥ 1.22209.3), keyed under the request + terminal `destination_url`, never caching errors/empty; a hit is assertable via `tool_result_contains: "Already fetched"`.
57
59
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
58
60
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
59
61
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -311,8 +313,10 @@ your `answers:` against that kept run, then record once. **But the kept run is a
311
313
  skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
312
314
  current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
313
315
  deliberately. (Same fail-closed family: corrupt gate evidence — unparseable `events.jsonl` lines, or fewer
314
- gates than `trace.json` recorded questions — and a structurally invalid `result.json` also refuse rather
315
- than certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
316
+ gates than `trace.json` recorded questions — a structurally invalid `result.json`, a `command:"replay"`
317
+ result (a replay is a re-check of a recorded cassette, not run evidence — verify the original live run dir),
318
+ and a `mode:"chat"` result (chat carries no assertions or verdict by contract) also refuse rather than
319
+ certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
316
320
 
317
321
  Run artifacts are written to `~/.cowork-harness/runs/…` by default — **outside any working tree**, so a run
318
322
  launched from a repo root never drops sensitive skill inputs/outputs into it. Pass `--run-dir <path>` (or set
@@ -354,6 +358,15 @@ cassette — has its own recipe:
354
358
  unless the scenario asserts `allow_missing_capability: true`, which downgrades it to a notice and
355
359
  proceeds. Rebuild with `--build-arg COWORK_FULL_PARITY=1` and point `COWORK_AGENT_IMAGE` at it for those
356
360
  skills.
361
+ 6. **Iterate across fixes — verify before you trust, and don't cross-pair generations.** A green run is
362
+ not a correct run, and a skill's self-reported finding is not real until its cited evidence is found in
363
+ the run's own output. Ground each finding against `result.json` (`finalMessage` = the skill's own
364
+ answer/critique; `toolResults` = tool outputs) and the tool-call stream via
365
+ `cowork-harness trace <run-dir> --output-format json` — add `--full-results` so a successful call's full
366
+ input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
367
+ pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
368
+ produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
369
+ kept run predates the current skill). See `docs/debugging.md` (repo-only) for the full loop.
357
370
 
358
371
  #### Interpreting verdict signals
359
372
 
@@ -364,6 +377,42 @@ The run verdict may include `WARN`-severity signals in addition to pass/fail. On
364
377
  the run can still green. If you see it, fix the asset path — a green with a missing asset is
365
378
  not a valid pass.
366
379
 
380
+ **False negatives — signals that are tier/image artifacts, not skill defects.** Some fail-severity
381
+ signals read like a skill gap but are really a property of the reduced test image or the fidelity tier.
382
+ Recognize these before "fixing" a non-bug:
383
+
384
+ - **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
385
+ rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
386
+ `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
387
+ (`magick`) can trip this even though real Cowork **ships** those. The message says so ("likely a FALSE
388
+ NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
389
+ `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
390
+ `allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
391
+ or a declared `requires_capabilities` the tier can't provide, both lanes — an unknown family name
392
+ hard-fails rather than silently passing.) **On an open-ended `skill` run** (no `assert:` block to carry
393
+ the modifier), pass **`--allow-missing-capability`** — the CLI equivalent of the assertion.
394
+ - **`ended_with_question`** (`WARN`, live lane) — a heuristic: the agent's final answer contains a
395
+ question and the run wrote **no deliverable to `outputs/`** — it may have ended on a request for input
396
+ instead of finishing. Warn-only; the fix is scripting/steering the answer (`answer:` / `--answer` / a
397
+ decider, or `--decider-llm --intent`), not editing the skill's prose. The strict, fail-severity sibling
398
+ `stalled` already catches a *trailing*-`?` final turn that did no tool work after the last gate; this
399
+ covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
400
+ the final message before acting — a legitimate question-posing answer that wrote a file never fires.
401
+ Assert `allow_stall: true` if ending on a question is the intended terminal state.
402
+ - **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
403
+ paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
404
+ `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
405
+ `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
406
+ run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
407
+ it's valid.
408
+ - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
409
+ `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
410
+ pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
411
+
412
+ The full 14-code signal table (severity + per-signal opt-out) is in
413
+ [`references/scenario-schema.md`](./references/scenario-schema.md); `docs/scenario.md` (repo-only) carries
414
+ the fuller narrative.
415
+
367
416
  ### Checking whether a background run is alive
368
417
 
369
418
  Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
@@ -472,7 +521,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
472
521
  - **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
473
522
  `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input, with a
474
523
  `trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
475
- **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field.
524
+ **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
525
+ **absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
526
+ collection) is evidence **UNAVAILABLE**, not an empty run — `trace --view files` reports a loud
527
+ UNAVAILABLE marker (`workspaceFilesRecorded: false` in JSON, and no phantom "removed" diff rows) and
528
+ `inspect` prints `artifacts: UNAVAILABLE` (`artifactsRecorded: false`) instead of `artifacts (0):`.
476
529
 
477
530
  ### Debugging with `chat`
478
531
 
@@ -615,9 +668,11 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
615
668
  - `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
616
669
  stalling *after* one is answered — two different failure modes.
617
670
  - **Free-text aside:** a "type-it-in-notes" option has **no scripted deterministic answer** today
618
- (the `OTHER:` directive works only on the LLM-decider path, not scripted `choose:`; on an
619
- options-bearing gate a bare out-of-set LLM answer fails loud (exit 2) — see the LLM-decider
620
- free-text note in `references/fidelity-and-answers.md`).
671
+ (the `OTHER:` directive works only on the LLM-decider path, not scripted `choose:`, and only on
672
+ **single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
673
+ options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
674
+ LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
675
+ `OTHER:` is marked `[via Other free-text]` in its `gateProvenance` rationale.
621
676
  14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
622
677
  but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
623
678
  index lands on a different option (a silent re-record flake). Prefer an exact label when order is
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.2.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.4.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
34
34
  run: |
35
- V=2.1.209 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.215 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.2.0"
60
+ - run: npm i -g "cowork-harness@>=1.4.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.2.0"
200
+ - run: npm i -g "cowork-harness@>=1.4.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.2.0"
229
+ run: npm i -g "cowork-harness@>=1.4.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.2.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.4.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -164,10 +164,13 @@ Caveat: `decide` only builds a **single-select** sample (set choices with `--opt
164
164
  multiSelect flag), so its printed request shows `options[].label` but never `multiSelect:true` — to
165
165
  exercise the array reply path, run a real multiSelect gate or unit-test the helper directly.
166
166
 
167
- LLM-decider free-text goes via `OTHER: <value>` on an **options-bearing** gate; a bare out-of-set
168
- answer (no matching label, no `OTHER:`) fails loud (`UnansweredError` → exit 2) — it never stalls or
169
- guesses an option. Open-ended (no-option) gates need no `OTHER:` prefix: free text is delivered
170
- verbatim. (Scripted scenarios use the separate `answer:` escape hatch.)
167
+ LLM-decider free-text goes via `OTHER: <value>` on an **options-bearing single-select** gate; a bare
168
+ out-of-set answer (no matching label, no `OTHER:`) fails loud (`UnansweredError` → exit 2) — it never
169
+ stalls or guesses an option. A **multi-select** gate is **index-only**: it accepts comma-separated option
170
+ numbers, and `OTHER:` fails loud there (no free-text escape on that path). Open-ended (no-option) gates
171
+ need no `OTHER:` prefix: free text is delivered verbatim. A decision answered via any free-text path is
172
+ marked `[via Other free-text]` in its `gateProvenance` rationale, so a `result.json` consumer can tell it
173
+ from an offered-option pick. (Scripted scenarios use the separate `answer:` escape hatch.)
171
174
 
172
175
  A gate that fails loud (`on_unanswered: fail`, the default) still **salvages a PARTIAL run**: the harness
173
176
  writes a `result.json` (marked `partial: true`) with the artifacts the agent produced before the whiff, so
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.2.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.4.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -252,7 +252,7 @@ same set live from the schema.
252
252
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
253
253
  | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
254
254
  | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
255
- | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
255
+ | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
256
256
  | `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
257
257
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
258
258
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
@@ -334,7 +334,7 @@ dotted path.
334
334
 
335
335
  **VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; most
336
336
  are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
337
- `"success"`) and only three are **warn**-severity (informational, never flip pass/fail). Current signal
337
+ `"success"`) and only four are **warn**-severity (informational, never flip pass/fail). Current signal
338
338
  codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
339
339
 
340
340
  | Code | Severity | Meaning |
@@ -347,16 +347,17 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
347
347
  | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
348
348
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
349
349
  | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
350
- | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`) |
350
+ | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
351
351
  | `infra_error` | fail | A VM/egress sidecar crashed mid-run — not author-suppressible |
352
- | `stalled` | fail | The run ended on an unanswered question (opt out: `allow_stall`) |
352
+ | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
353
353
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
354
354
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
355
355
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
356
+ | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
356
357
 
357
358
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
358
359
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
359
- `result.verdict.signals[].severity` or the run's exit code. Only the three **warn** codes are truly benign.
360
+ `result.verdict.signals[].severity` or the run's exit code. Only the four **warn** codes are truly benign.
360
361
 
361
362
  ## Replay class
362
363
 
@@ -163,3 +163,30 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
163
163
  **Lane note:** `semantic_matches` is **live-only** (the judge is a live model call), so these scenarios
164
164
  run on the `run` lane, never token-free `replay` — the linter's "all assertions live-only" warning is
165
165
  expected and correct here.
166
+
167
+ ## Recipe 6 — Iterate a skill across fixes (ground findings, don't cross-pair generations)
168
+
169
+ Hardening a skill is a loop: run → read what it did → fix → run again. Two disciplines keep it honest.
170
+
171
+ 1. **Verify before you trust.** A green run is not a correct run, and a skill's self-reported finding (a
172
+ self-critique appendix, "I extracted X") is not real until its cited evidence is found in the run's own
173
+ output. The harness emits the substrate; the grader is yours (it lives outside the harness):
174
+ - `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
175
+ - `cowork-harness trace <run-dir> --output-format json` → the tool-call stream. Add `--full-results` so
176
+ a **successful** call's full input + result are captured (the default view slices them to ~100/120
177
+ chars) — this is what lets your grader confirm "the skill claims it read X and derived Y" against the
178
+ actual call.
179
+ - `cowork-harness inspect <run-dir>` → what the run produced, plus the run's `label` and `skillHash`.
180
+ - In-run alternative: dispatch a checker **sub-agent** (maker/checker) whose result folds into the
181
+ verdict.
182
+ 2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
183
+ `result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
184
+ content-exact, on every live run, changes on any tracked edit. **Group/pair on it** (`inspect` and the
185
+ run-index row surface a short prefix). Add `--label <tag>` for a human-readable generation name
186
+ (skillHash is the correctness key; the label is ergonomics). `cowork-harness verify-run <run-dir>
187
+ <scenario.yaml>` is the native staleness guard: it **warns** when a kept run predates the current
188
+ skill, and with scripted `answers` **hard-fails** rather than vouch for a stale gate snapshot.
189
+
190
+ **Lane note:** the exploratory driver is `skill <dir> --decider-llm --intent "<what this run tests>"`,
191
+ which is flagged non-deterministic (a green here is exploration, not a scripted pass) — pin the
192
+ load-bearing gates with `--answer` once you know which fire.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,137 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.4.0] — 2026-07-19
10
+
11
+ ### Added
12
+
13
+ - **`coworkWebFetchDedup` enacted (hostloop `web_fetch`).** Real Cowork keeps a per-session negative-work
14
+ cache: a repeat `web_fetch` of the same normalized URL within a TTL (default 15 min; cap 100; FIFO
15
+ eviction; a hit does not refresh recency) makes **no network request** and returns a marker telling the
16
+ model to re-use the earlier result. The harness now reproduces this on the host-API (`coworkWebFetchViaApi`)
17
+ path — **baseline-gated** (only when the resolved baseline's `coworkWebFetchDedup` gate is on, i.e. Desktop
18
+ ≥ 1.22209.3), keyed under both the request URL and the terminal `destination_url`, never caching errors /
19
+ empty / non-2xx responses, and emitting **no egress event** on a hit (matching production's zero-network
20
+ dedup). A hit is observable via the marker text (`tool_result_contains: "Already fetched"`).
21
+
22
+ ### Changed
23
+
24
+ - **Platform baseline synced to Desktop 1.22209.3** (agent `2.1.215`). No prompt / spawn-env / egress-allowlist
25
+ drift vs 1.21459.0; the sync captured the new `coworkWebFetchDedup` runtime config (enacted above) plus a
26
+ few new (off) GrowthBook gates. The skill/README/reference version floors and agent-binary pins track the
27
+ new baseline.
28
+
29
+ ## [1.3.0] — 2026-07-19
30
+
31
+ ### Added
32
+
33
+ - **`skill --allow-missing-capability`** — the open-ended-run equivalent of a scenario asserting
34
+ `allow_missing_capability: true`. An open-ended `skill` run has no `assert:` block to carry the opt-out,
35
+ so a self-flagged capability FALSE-NEGATIVE on the lean `core` image would hard-fail with no escape
36
+ hatch; the flag merges the modifier onto the synthesized success assertion, suppressing both the
37
+ post-run `missing_capability` fail and the pre-run capability abort.
38
+ - **`ended_with_question` verdict signal (WARN)** — a heuristic run-level classifier: the agent's final
39
+ answer contains a question and the run wrote no deliverable to `outputs/` — a likely conversational
40
+ dead-end that still exited `result:"success"`. The lenient warn-severity sibling of the strict, fail
41
+ `stalled` (which already catches a trailing-`?` final turn with no post-gate tool work); this covers the
42
+ residual — a mid-message `?`, or tool work after the last gate that still ended asking. Never flips a
43
+ verdict; `verify-run` over historical results may surface it on matching old runs (benign).
44
+ - **`--decider-llm` "Other" answers are marked in gate provenance** — a decision answered via the
45
+ `OTHER:` free-text path (or a no-option free-text gate) now carries `[via Other free-text]` in its
46
+ `rationale`, so a `result.json` consumer can distinguish it from an offered-option pick. (Multi-select
47
+ gates remain index-only — `OTHER:` is rejected there, now pinned by a test.)
48
+
49
+ - **Run-identity metadata for the iterate-across-fixes loop.** `skill`/`run` accept `--label <tag>`, a
50
+ human-readable generation tag surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and
51
+ `status.json`. Each live run also records `skillCommit` — best-effort git `HEAD` of the session's skill
52
+ source dirs (commit provenance; `null` when the dirs span >1 repo or aren't a git work tree). These are
53
+ ergonomics on top of the **authoritative** content-exact version key `fingerprint.skillHash` (already
54
+ recorded on every run): a harvest step should group/pair a critique against a matching `skillHash`.
55
+ `inspect` and the run-index row now surface a short `skillHash` prefix so a pairing check needs no
56
+ `result.json` open. (Chat runs carry no `skillHash` and take no `--label`.)
57
+ - **`trace --full-results`** captures the FULL input + result of every tool call — successful ones too,
58
+ not just errors (`resultTextFull`/`detailFull`, 4 KB cap) — so an external grader can ground a
59
+ self-critique finding against the call it cites. The default view keeps its 100/120-char slices, so
60
+ existing JSON consumers are unaffected.
61
+ - **`verify-run` now warns on skill drift for answer-less scenarios.** Previously the skillHash-drift
62
+ check ran only when a scenario declared scripted `answers` (a hard fail). An answer-less `verify-run`
63
+ now emits a `::warning::` ("the kept run predates the current skill … findings describe an older skill
64
+ version") instead of staying silent — a WARN, not a fail, since re-asserting a new `assert:` block
65
+ against a frozen run dir is legitimate.
66
+
67
+ ### Fixed
68
+
69
+ - **Filesystem-evidence assertions no longer pass on incomplete evidence.** `input_unmodified` now fails
70
+ loud when its glob matches **no** pre-run path (a typo or renamed mount was a silent vacuous pass), and
71
+ fails evidence-unavailable when a manifest path escapes the workspace root or a matched file exceeds the
72
+ hash cap. The pre-run baseline records its provenance, so `no_unexpected_files` / `input_unmodified` fail
73
+ evidence-unavailable on an unreadable connected-folder baseline instead of diffing a partial tree
74
+ (`RunResult` gains `preRunOrigin`). `no_lost_write_back` no longer silently misses a modified-but-unreadable
75
+ or over-cap file, an authored file under an unreadable subtree, or a scratchpad deliverable behind a
76
+ symlink/hardlink — each now surfaces as could-not-verify.
77
+ - **Run-dir consumers no longer read absent evidence as empty, or a replay re-check as run evidence.**
78
+ `verify-run` now refuses (exit `2`, "can't verify ⇒ not green") a run dir whose `result.json` was produced
79
+ by `replay` (`command:"replay"` — a re-check of a recorded cassette, not run evidence) or by `chat`
80
+ (`mode:"chat"` — no assertions or verdict by contract); the refusal keys on `command`/`mode`, never on
81
+ `workspaceFiles`, so a live run merely lacking optional evidence fields still verifies. `stats --reindex`
82
+ skips a stray `command:"replay"` `result.json` instead of stripping the label and relabeling it `"run"`,
83
+ and its report line separates `skipped — replay re-check, not evidence` from `skipped — missing/corrupt
84
+ result.json`. `trace --view files` reports workspace-file evidence **UNAVAILABLE**
85
+ (`workspaceFilesRecorded: false` in JSON) when `workspaceFiles` is absent — distinct from a run that
86
+ genuinely wrote nothing — and no longer emits phantom "removed" diff rows against a persisted
87
+ `preRunHashes` in that case. `inspect` likewise prints `artifacts: UNAVAILABLE`
88
+ (`artifactsRecorded: false`) instead of `artifacts (0):` when `result.artifacts` is absent.
89
+ - **Static artifact analysis (`analyze-skill` / `no_lost_write_back` Tier A)** closes several false-green and
90
+ false-positive holes: recognizes `axios`/jQuery and library write-backs, computed member calls
91
+ (`xhr["open"]`), unquoted and submitter-overridden `<form>` actions, and ES-module sources; classifies URLs
92
+ with the WHATWG parser so `localhost.evil.com`, protocol-relative `//host`, and `mailto:`/`data:` are no
93
+ longer misread as local; scope-aware constant folding and member-mutation tracking stop live code being
94
+ proven dead; an unresolved URL or method (including spread request options) yields could-not-verify instead
95
+ of a silent clean; and every write-back in a file is reported, not just the first. The advertised
96
+ `.ts/.tsx/.jsx` source extensions are dropped (no compatible parser) — treated as out-of-scope rather than
97
+ parse-noise.
98
+ - **Runtime artifact confirmation (Tier B)** now confirms edit-fired autosaves and load-time write-backs (not
99
+ only explicit commits), handles `fetch(Request)`, models browser-faithful XHR response/listener semantics,
100
+ skips disabled/hidden controls and honors a submitter's `formaction`, analyzes observed writes before
101
+ downgrading on an external script, matches loopback hosts exactly, and no longer swallows unrelated harness
102
+ exceptions.
103
+ - **`analyze-skill` orchestration** fails could-not-verify (exit `3`) when a `references`/`agents`/`commands`/
104
+ `skills` subtree is unreadable (was a silent clean); an ignore marker inside a fenced or block-quoted
105
+ example no longer suppresses real findings; a mixed invocation fails on a positional that resolves to no
106
+ scannable source; JSON coverage lists clean artifact sources; and runtime mode honors the 3 MB read cap.
107
+ - **Protocol / decision handling fails closed on drift.** Unknown control-request subtypes and duplicate
108
+ outstanding request IDs are rejected as protocol errors; malformed user/tool-result blocks increment a new
109
+ `evidenceErrors.protocolMalformed` counter; the ABSTAIN permission/dialog fallback reconciles answer
110
+ delivery like a normal decision; `present_files` leak classification normalizes `..` paths; gate answers
111
+ require a request id; and an `allow_if` expression no longer fails to compile when a permission input key is
112
+ a reserved word.
113
+ - **CI / release gates.** The baseline-staleness check rejects non-finite, future-dated, and corrupt
114
+ timestamps (extracted to a unit-tested `scripts/check-baseline-staleness.ts`); manual container-image
115
+ publishes require a green CI run unless an explicit break-glass input is set; and the composite Action
116
+ passes inputs via the environment (closing a shell-injection surface) and accepts a JSON-array `extra-args`
117
+ that preserves quoting and spaces.
118
+ - **`analyze-skill` top-level `--help`** now documents the `--runtime` flag and exit code `3`
119
+ (could-not-verify); previously these appeared only in the per-command `analyze-skill --help`.
120
+ - Corrected a phantom assertion key in the `record --margins` documentation
121
+ (`max_tool_calls` → the real `tool_calls_max`).
122
+
123
+ ### Documentation
124
+
125
+ - Documented previously-undocumented CLI flags: `record --force` (narrowly overrides the
126
+ different-scenario slug-collision overwrite refusal) and `record --decider-model`; `probe-dispatch`'s
127
+ inherited `--decider-cmd` / `--decider-dir` / `--on-unanswered` / `--ablate-skill`; and
128
+ `skill --timeout` / `--answer-policy`.
129
+ - Clarified the scenario-schema descriptions for `user_visible_artifact` (the assertion value is
130
+ workRoot-relative — e.g. `outputs/x.md`, not `mnt/`-prefixed) and `gate_answers_delivered`
131
+ (documented the `: false` confirmed-non-delivery inverse). Regenerated `schema/scenario.schema.json`.
132
+ - Qualified the stable JSON-envelope contract: SPEC §11 now maps commands to envelope families by
133
+ mechanism (`jsonEnvelope` / `jsonPayloadEnvelope` / dedicated), and README defers to it; clarified
134
+ that `replay` exit `2` is a whole-cassette operational failure, distinct from an in-cassette
135
+ malformation (which fails as an exit-`1` assertion).
136
+ - Added a `docs/decisions/` ADR index, `python/README.md` cross-links into the main doc spine, and
137
+ `RELEASING.md` to `llms.txt`; signposted the specialized `docs/*.md` guides and added `lint-skill` /
138
+ `analyze-skill` rows to the docs index.
139
+
9
140
  ## [1.2.0] — 2026-07-18
10
141
 
11
142
  ### Added