cowork-harness 2.0.1 → 2.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +5 -5
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +29 -16
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
  7. package/CHANGELOG.md +366 -0
  8. package/README.md +25 -8
  9. package/RELEASING.md +7 -2
  10. package/SPEC.md +24 -4
  11. package/dist/assert.js +18 -3
  12. package/dist/baseline.js +37 -15
  13. package/dist/cli.js +5 -1
  14. package/dist/prompt.js +5 -2
  15. package/dist/redact.js +44 -0
  16. package/dist/run/budget.js +21 -6
  17. package/dist/run/cassette.js +76 -14
  18. package/dist/run/chat-result.js +1 -0
  19. package/dist/run/chat.js +2 -0
  20. package/dist/run/execute.js +21 -6
  21. package/dist/run/run.js +19 -0
  22. package/dist/runtime/argv.js +20 -1
  23. package/dist/runtime/container.js +11 -3
  24. package/dist/runtime/hostloop.js +5 -1
  25. package/dist/runtime/lima.js +7 -1
  26. package/dist/runtime/microvm.js +22 -4
  27. package/dist/session.js +4 -0
  28. package/dist/types.js +1 -1
  29. package/docs/cassette.md +18 -6
  30. package/docs/invariants.md +5 -2
  31. package/docs/maintenance.md +5 -3
  32. package/docs/protocol.md +23 -5
  33. package/docs/scenario.md +9 -6
  34. package/examples/replays/README.md +1 -1
  35. package/fixtures/protocol/v1/dialog-response.json +10 -0
  36. package/fixtures/protocol/v1/elicit-response.json +10 -0
  37. package/fixtures/protocol/v1/elicitation-request.json +17 -0
  38. package/fixtures/protocol/v1/error-response.json +8 -0
  39. package/fixtures/protocol/v1/user-dialog-request.json +11 -0
  40. package/package.json +1 -1
  41. package/schema/cassette.v12.json +2 -2
  42. package/schema/protocol.v1.json +447 -79
  43. package/schema/run-result.json +4 -0
  44. package/schema/scenario.schema.json +1 -1
  45. package/scripts/check-versions.ts +134 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.0.1
7
- tracks-harness: cowork-harness 2.0.1 (baseline desktop-1.34493.1)
6
+ version: 2.2.0
7
+ tracks-harness: cowork-harness 2.2.0 (baseline desktop-1.34493.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.0.1` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.2.0` (baseline
26
26
  > `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.0.1"`. **Pin `@^2.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.2.0"`. **Pin `@^2.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -656,7 +656,7 @@ than a stuck `"running"`.)
656
656
  ### Place assertions in the right CI lane
657
657
 
658
658
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
659
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@main` (a packaged GitHub Action with a
659
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
660
660
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
661
661
  the four-stage pipeline.
662
662
 
@@ -1,19 +1,23 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.0.1` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@main
9
+ - uses: yaniv-golan/cowork-harness@v2
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
+ version: "^2" # hold the major; see below
13
14
  ```
14
15
 
15
- The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "2.0.1"`) for reproducible CI.
16
+ **These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
17
+ CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
+ copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
+ patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
+ (e.g. `version: "2.2.0"`) instead when you want byte-reproducible CI.
17
21
 
18
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -43,10 +47,11 @@ jobs:
43
47
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
44
48
  # Background on the provenance chain: the "Agent-binary provenance" section of
45
49
  # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
46
- - uses: yaniv-golan/cowork-harness@main
50
+ - uses: yaniv-golan/cowork-harness@v2
47
51
  with:
48
52
  command: run
49
53
  path: scenarios/
54
+ version: "^2"
50
55
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
51
56
  ```
52
57
 
@@ -62,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
62
67
  GitHub-hosted runners, no token/Docker/agent:
63
68
 
64
69
  ```yaml
65
- - run: npm i -g "cowork-harness@^2.0.1"
70
+ - run: npm i -g "cowork-harness@^2.2.0"
66
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
67
72
  # no silent false-greens. WITHOUT --strict this
68
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -110,20 +115,28 @@ Action has no input for, and it creates a coupling nothing checks:
110
115
 
111
116
  > **If a flag in `extra-args` was added in release X, floor that step's `version` to `>=X`.**
112
117
 
113
- `version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. So:
118
+ `version` defaults to `latest`, and accepts **any npm range** — not just an exact pin. Leave it off unless
119
+ you have a reason:
114
120
 
115
121
  ```yaml
116
- - uses: yaniv-golan/cowork-harness@main
122
+ - uses: yaniv-golan/cowork-harness@v2
117
123
  with:
118
124
  command: lint
119
125
  path: scenarios/
120
- version: ">=1.11.0" # --min-severity landed in 1.11.0
121
- extra-args: --min-severity WARN
126
+ version: "^2" # holds the major
127
+ extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
122
128
  ```
123
129
 
124
- Without the floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN` (exit 2,
125
- wrapped in an `ok:false` envelope) — it does **not** degrade gracefully. An exact pin is fine for
126
- reproducibility, but it rots the moment a recipe adopts a newer flag; a floor expresses the real dependency.
130
+ **If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
131
+ floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
132
+ recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
133
+ major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
134
+ stopping at the major boundary. An exact
135
+ pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
136
+ moment a recipe adopts a newer flag.
137
+
138
+ Without a satisfied floor, an older CLI fails the step with `unrecognized arguments: --min-severity WARN`
139
+ (exit 2, wrapped in an `ok:false` envelope) — it does **not** degrade gracefully.
127
140
 
128
141
 
129
142
  The harness has two execution lanes with different cost, coverage, AND infrastructure requirements.
@@ -307,7 +320,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
307
320
 
308
321
  ## GitHub Actions sketch
309
322
 
310
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@main` does
323
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
311
324
  in one step (see the top of this doc) — reach for this form when you need independent per-command
312
325
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
313
326
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -327,7 +340,7 @@ jobs:
327
340
  with: { node-version: '24' }
328
341
  - uses: actions/setup-python@v5
329
342
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
330
- - run: npm i -g "cowork-harness@^2.0.1"
343
+ - run: npm i -g "cowork-harness@^2.2.0"
331
344
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
332
345
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
333
346
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -356,7 +369,7 @@ jobs:
356
369
  echo "live=true" >> "$GITHUB_OUTPUT"
357
370
  fi
358
371
  - if: steps.guard.outputs.live == 'true'
359
- run: npm i -g "cowork-harness@^2.0.1"
372
+ run: npm i -g "cowork-harness@^2.2.0"
360
373
  - if: steps.guard.outputs.live == 'true'
361
374
  run: cowork-harness run scenarios/ --output-format json
362
375
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.0.1` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.0.1` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.0.1`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.2.0`
4
4
  (baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -342,8 +342,8 @@ same set live from the schema.
342
342
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
343
343
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
344
344
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
345
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
346
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
345
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
346
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
347
347
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
348
348
  | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
349
349
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.0.1` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -66,7 +66,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
66
66
  | `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
67
67
  | `scenarioSource` | Relative path to the authored YAML this was recorded from |
68
68
  | `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
69
- | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
69
+ | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
70
70
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
71
71
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
72
72
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
package/CHANGELOG.md CHANGED
@@ -4,6 +4,372 @@ All notable changes to this project are documented here. The format is based on
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
5
5
  [Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
6
6
 
7
+ ## [2.2.0] — 2026-08-25
8
+
9
+ ### Upgrade impact
10
+
11
+ Two behaviour changes can turn a previously-green run red. Neither breaks a covered surface
12
+ ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) — both make an assertion report
13
+ what it always documented — so they ship in a minor:
14
+
15
+ - **`no_scratchpad_leak` at `container` can now FAIL.** It was measuring containment against a root in a
16
+ different path space, so every presented file classified `leaked: false` and the check passed
17
+ vacuously. A scenario whose skill genuinely leaves a presented file in the scratchpad will now red —
18
+ that is the leak the key exists to catch.
19
+ - **`baseline: desktop-1.11847.5` is now refused at `container`/`hostloop`/`microvm`.** It carries no
20
+ `spawn` block, so those tiers cannot reproduce Cowork's toolset — a run on it launched an agent with no
21
+ file or bash tools and still reported a verdict. Use a `sync`-recorded baseline, or `fidelity: protocol`.
22
+
23
+ ### Fixed
24
+
25
+ - **`present_files_called` no longer reports "the tool was never called" about a run that called it, and a
26
+ hostloop delivery scenario can be recorded with a host-path redaction policy at last.** The assertion
27
+ read presence off `RunResult.presentedFiles`, which is a *classification* of each presented path
28
+ (scratchpad? promoted? leaked?) and requires an absolute path to compute. At `hostloop` a presented path
29
+ is a real host path, so the shipped redaction policy rewrites it to
30
+ `[REDACTED:local-path:<hash>]/mnt/outputs/report.html` — correct, documented, and deliberately ordered so
31
+ the mount tail survives — and the classifier then drops every entry as un-normalizable. The list came
32
+ back empty and the assertion stated, as fact, that the tool had never been called.
33
+
34
+ Because `record` replays the base and redacted cassettes and refuses to write when the verdict differs,
35
+ this was not merely a wrong message: **no cassette asserting `present_files_called` could be recorded at
36
+ that tier at all**, so the assertion had never executed on the replay lane. The refusal was right — the
37
+ redacted cassette genuinely could not support the assert — but the defect it was reporting was in the
38
+ harness, not in the recording.
39
+
40
+ Presence now comes from `RunResult.presentFilesCalls`, a new count of the `present_files` invocations
41
+ that carried a well-formed `file_path`, taken from the tool_use input's shape and never from a path's
42
+ content, so redaction cannot alter it. Classification is untouched: `presentedFiles` still drops what it
43
+ cannot resolve, and `no_scratchpad_leak` still reads it. A run recorded before the field falls back to
44
+ the old `presentedFiles`-non-empty test, so no existing green changes.
45
+
46
+ A run whose every `present_files` call carried an unusable path now reports **cannot verify** rather than
47
+ "never called" — the tool *was* invoked, and the harness already knew it. (The malformed count is
48
+ deliberately kept out of that message: the record self-check normalizes `[REDACTED…]` tokens out of
49
+ failing messages but not digits, so an interpolated count that differed between the two replays would
50
+ refuse a cassette that is otherwise fine to write.)
51
+
52
+ Pinned as an invariant ([docs/invariants.md](./docs/invariants.md)) with the end-to-end record self-check
53
+ as its test anchor — the case that could not be recorded, and so had never run in CI.
54
+
55
+ - **Guest paths are built from the tree the harness stages, not from a baseline's recorded mount layout.**
56
+ `resolveMounts` returned `mountLayout.mntRoot` verbatim — or, when that field was absent and the recorded
57
+ `sessionRoot` already ended in `/mnt`, the session root itself. But the staged tree is always
58
+ `<sessionRoot>/mnt`: `stageWorkspace` creates it there and `dockerRunArgv` nests its read-only binds
59
+ there. On a baseline recording anything else, `--plugin-dir` therefore pointed one directory **above**
60
+ the staged plugin tree, so the plugin under test never loaded. `mntRoot` is now derived from the session
61
+ root, and a recorded layout the harness cannot stage is reported as a fidelity divergence at spawn
62
+ instead of silently composing a path no stager creates.
63
+
64
+ Guest paths now anchor on `sessionRoot` — the bind target — rather than `cwd`, which is only where the
65
+ agent's process starts (production's own working directory is a folder mount or `outputs`, so the two are
66
+ not interchangeable even though every synced baseline records them equal). `dockerRunArgv` takes an
67
+ explicit `agentCwd` for `-w`.
68
+
69
+ Two more derivations of the same rule are gone: `prompt.ts` had a private `/sessions/<id>` +`/mnt`
70
+ literal (so the prompt could describe a tree the runtimes had not staged), and **microvm** read the
71
+ agent's root from the baseline when lima structurally mounts it at `/sessions/<sessionId>` — which put
72
+ `CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages. A baseline recording a cwd that tier
73
+ cannot honour is now **refused**, not warned about: the guest `cd` would otherwise succeed at the wrong
74
+ directory.
75
+
76
+ - **A baseline with no `spawn` block is refused at the sandbox tiers.** That block carries the tool set,
77
+ pre-approvals, effort default and config-dir location, and the `?? []` fallbacks meant a run would launch
78
+ an agent with **no Read/Write/Bash/Skill/Task at all** and still report a verdict. `fidelity: protocol`
79
+ builds its own argv and is unaffected.
80
+
81
+ - **`no_scratchpad_leak` can see a container leak again — the session root was in the wrong path space.** The
82
+ root that `presentedFiles`' promoted/leaked classification is measured from was derived by the caller as
83
+ `<run-dir>/work/session`, a HOST path, and handed to every non-`protocol` tier. But the path space the
84
+ agent reports in is per-tier: at `container` (the default) it runs inside the sandbox and reports
85
+ `/sessions/<id>/…`, and only at `hostloop` does it run natively and report host paths. Measured against a
86
+ host root, no VM path is ever inside the root, so every presented file classified `leaked: false` —
87
+ including the handler's copy-failure branch, which returns the source path unchanged when a file is
88
+ blocked-extension, a directory, or absent. `no_scratchpad_leak` evaluates at `container` and nowhere else,
89
+ so the assertion that exists to catch that leak could not catch it, and `verdict.ts`'s delivery check read
90
+ the same `leaked: false` as a successful delivery.
91
+
92
+ Each runtime now reports the session root it actually launched the agent with, and the run consumes that
93
+ instead of deriving a second one — the two can no longer drift into different spaces. `chat` sets it on
94
+ both serving tiers as well; it never did, so a hostloop chat's `presentedFiles` was inverted in its result
95
+ file. `protocol` and `microvm` serve no `present_files` and keep the cwd fallback.
96
+
97
+ Independently, the classifier now fails CLOSED on a space mismatch: if the agent's cwd is not at or inside
98
+ the session root, the presented batch counts as malformed (`no_scratchpad_leak` → cannot verify) rather
99
+ than being graded against a root it cannot be compared to. `leaked: true` is not derivable in that state
100
+ either, and a `leaked: false` verdict there is exactly the vacuous pass the key exists to prevent.
101
+
102
+ ### Added
103
+
104
+ - **First live coverage for `present_files_called` / `no_scratchpad_leak`.** No scenario in the repo
105
+ asserted either key, which is how a session root in the wrong path space could record `leaked: false`
106
+ for every presented file without anything noticing. `e2e/scenarios/smoke-present-files.yaml` writes a
107
+ file OUTSIDE `mnt/` and delivers it, so the promotion is real and the pair is non-vacuous; it runs in
108
+ CI's live e2e loop. Measured on a live container run: `presentFilesCalls: 1`, promoted `true`, leaked
109
+ `false`.
110
+
111
+ - **`RunResult.presentFilesCalls`** — the count of `present_files` invocations that carried a well-formed
112
+ `file_path`, in `result.json` and [schema/run-result.json](./schema/run-result.json). Content-class, so a
113
+ replay re-drive reproduces it; read by `run`, `replay` and `verify-run` alike, and absent (not `0`) on a
114
+ result written before this release, which is what the assertion's fallback distinguishes. Use it to
115
+ answer "did the agent deliver anything?" from a result file without interpreting `presentedFiles`'
116
+ promoted/leaked classification.
117
+
118
+ ## [2.1.0] — 2026-08-24
119
+
120
+ ### Changed
121
+
122
+ - **`rehash` leads with the split, and PARTIAL success has its own exit code (`4`).** The summary used to
123
+ print last, after every per-file line had scrolled past — and those two counts *are* the decision:
124
+ commit what migrated and budget a re-record for the rest, versus nothing here is salvageable. It now
125
+ prints first, with the per-file lines as the detail behind it.
126
+
127
+ **`4 migrated, 18 failed` and `0 migrated, 22 failed` both exited `1`**, which made a shell consumer
128
+ unable to tell apart two situations demanding opposite responses. The JSON envelope always carried the
129
+ split as `migrated`/`skipped`/`errors`; a bare terminal run did not. `rehash` now exits `0` (all
130
+ migrated, or nothing needed migrating) / **`4`** (partial) / `1` (nothing migrated and at least one
131
+ could not) / `2` (usage).
132
+
133
+ **This is a behaviour change for anyone branching on `rehash`'s exit code** — a script testing
134
+ `rc == 1` for "something failed" will miss the partial case, though `rc != 0` is unaffected. It ships in
135
+ a minor deliberately: `rehash`'s codes were documented **nowhere** — zero mentions in its `--help`,
136
+ absent from SPEC §11 — so there was no published contract to break, and a consumer on `rc == 1` was
137
+ relying on observed behaviour. They are now documented in both places. `4` rather than `3` because the
138
+ code space is per-command and `3`'s "could not verify" meaning is load-bearing on `verify-cassettes`.
139
+
140
+ - **The batch cost estimate now reports its basis instead of claiming authority.** `estimatedCostUsd` is
141
+ `sum(max(local run history))` — a max over whatever *this machine* has run. The line already qualified
142
+ the partially-priced case with `— LOWER BOUND`, but the fully-priced case said
143
+ `(all N scenario(s) priced from prior runs)`, which is an active claim of completeness and was the one
144
+ case that said nothing qualifying. At a new baseline or a new agent binary the history describes a
145
+ materially different configuration, so the estimate can **under**-predict exactly when it is most
146
+ consulted — a consumer wrote "that is the ceiling, not the scope" into a plan off this line and had to
147
+ retract it.
148
+
149
+ The line now carries `basis: N prior run(s) on THIS machine, thinnest scenario has M; a max over that
150
+ history, NOT a bound`, and the JSON payload gains `estimateBasis`
151
+ (`{source, pricedRuns, thinnestScenarioRuns}`). `thinnest` is the useful half: a scenario with one prior
152
+ run contributes a single sample, not a worst case. Wired through `pricedRunCount`, which had been
153
+ exported and doc-commented *"for messages that report their own basis"* with zero callers.
154
+
155
+ - **Pre-epoch cassettes: we could report ordinary content drift, and chose not to.** When a cassette
156
+ recorded before the 2.0.0 hash-format epoch is read, `rehash` recomputes the **legacy** digest over the
157
+ current tree and compares it to the legacy digest in the cassette. On a mismatch that is a positive
158
+ determination of ordinary content drift — strictly more informative than `unverifiable-skill`, and the
159
+ same determination 1.25.0 reported as warn-only `skill` drift.
160
+
161
+ Replay does not report it, and that is a decision rather than a limit. Reporting it would require
162
+ keeping a fold of the retired hash algorithm alive in the replay path — the most-run lane — permanently,
163
+ to soften one release's migration; the population it would help shrinks with every re-record. The
164
+ runtime cost would be nil (both digests fold from one tree walk); the cost is code that could never be
165
+ deleted. **"We can compute this and chose not to report it, so the legacy fold can die" is a different
166
+ claim from "this cannot be known", and the earlier phrasing implied the latter.** The remedy for a
167
+ pre-epoch cassette is `rehash` where it can prove content unchanged, and a re-record where it cannot.
168
+
169
+ ### Fixed
170
+
171
+ - **The fast test lane had no global timeout, so 93 subprocess-spawning test files inherited vitest's 5s
172
+ default.** They measure 167-888ms locally — a fine margin until you remember the lane runs 344 files in
173
+ parallel across every core, and a CI runner is ~3x slower again. That is how an 888ms test crosses 5s;
174
+ it cost two red CI runs on unrelated PRs before anyone looked at the failure rather than re-running it.
175
+ `vitest.config.ts` now sets `testTimeout: 30_000` — ~34x the slowest measured non-e2e case, while a
176
+ genuine hang still fails ~6x faster than in the live lane (which sets 180s). Per-test values still win.
177
+
178
+ - **Redaction pattern ORDER is load-bearing, and now says so — plus a warning when it is wrong.** Patterns
179
+ apply in sequence over the accumulating output, so a bare catch-all placed ahead of a lookahead-anchored
180
+ rule for the same prefix matches first and eats the `/mnt/` tail the lookahead exists to preserve. The
181
+ shipped policy depends on that order: with it, a run-dir path redacts to
182
+ `[REDACTED:local-path:…]/mnt/outputs/report.md` and still resolves; without it the whole path is consumed
183
+ and `normalizeHostShapedForReplay` returns `null` — **every `computer://` structural-marker resolution
184
+ silently stops working, with no error and no finding.**
185
+
186
+ `loadRedactionPolicy` now warns, naming both pattern indices, when a policy is in the hazardous order.
187
+ Detection is deliberately conservative — it fires only when a later pattern's source is exactly an
188
+ earlier one's plus a trailing lookahead (modulo lazy quantifiers) — because regex subsumption is
189
+ undecidable in general and a false positive would train authors to ignore the warning.
190
+
191
+ The remainder is matched by **shape**, never by parsing the lookahead's body. A first cut used
192
+ `\(\?=[^()]*\)`, whose `[^()]*` silently skipped every lookahead containing a group — so
193
+ `(?=/mnt(?:/|$|[\s"'\\)\]]))`, the natural way to write "slash, end, or delimiter" and arguably more
194
+ correct than a bare `(?=/mnt/)`, went unflagged while being just as dangerous. Caught by a consumer
195
+ running it against their own policy, which is now a regression fixture. Not looking inside also sidesteps
196
+ escape- and char-class-awareness, since that policy carries an escaped `\)` inside a character class.
197
+
198
+ `docs/cassette.md` states the requirement next to the existing "stop before `/mnt/`" guidance, which had
199
+ the shape of the rule but not the ordering half. The new test pins the runtime consequence, not just the
200
+ detector: a reorder must make the link fail to normalize AND be flagged, so the syntactic check cannot
201
+ drift away from what it is standing in for.
202
+
203
+ - **The copy-pasteable Action steps now pin `version: "^2"`, and a guard requires it.** Bounding every
204
+ published npm floor last release fixed the *form* of a floor (`>=1.11.0` reads as a bound and silently
205
+ means "and every future major too") but removed the input from the recipes rather than correcting it —
206
+ so the shipped snippets carried no `version:` at all and fell back to the input's `latest` default. That
207
+ reproduced the exact footgun `action.yml`'s own description warns about two lines earlier: a CLI major
208
+ reaches a workflow the moment it is promoted, even though the `uses:` ref never changed.
209
+
210
+ `^2` was not among the alternatives weighed at the time, and it is the form `action.yml` itself
211
+ recommends: it holds the major, needs no patch number to remember, and only wants a human decision at
212
+ the next major bump. **Five** steps were unpinned, not the three in the CI recipe — `README.md` carries
213
+ two more.
214
+
215
+ `action-docs-sync` now requires every copy-pasteable step (a `uses:` line inside a fenced block with a
216
+ `with:`) to pin `^<package major>`; inline prose mentions are excluded, since there is nothing to pin.
217
+ Verified by mutation: dropping one `version:`, regressing a pin to `^1`, and bumping the package major
218
+ each fail. The guard also asserts it found the steps at all, because a parser that matches nothing
219
+ passes every assertion after it. Guarding a floor's FORM does not guarantee a floor is PRESENT — this
220
+ pins the behaviour instead.
221
+
222
+ - **The `sessionFingerprint` field set is now stated completely, and a guard discovers the sites that
223
+ state it.** The hash covers eight session fields; every place that enumerated them named six or fewer.
224
+ `web_fetch` was missing everywhere, `agent_env` was missing everywhere, and `docs/invariants.md`
225
+ also omitted `skills`. Fourteen sites carried the claim while the working assumption was three: **twelve
226
+ enumerated the set** — four in prose, two in the current cassette schema, six in the retained v9-v11
227
+ schemas — and **two denied it existed at all**. `SPEC.md` and `docs/scenario.md` both said the session
228
+ is "not drift-checked or fingerprinted". `model` genuinely is not hashed; connected folders and plugin/skill/MCP discovery are,
229
+ which is the half a reader would have trusted.
230
+
231
+ The `verify-cassettes` staleness message — the only enumeration a user ever sees — omitted `projects`,
232
+ and `docs/cassette.md` quoted it with `projects` present. Both are fixed and now pinned to each other
233
+ by test, so the doc cannot drift from the string again.
234
+
235
+ New **invariant 14** in `check:versions` derives the field set from `buildSessionFingerprint`'s shape
236
+ and **discovers** the enumeration sites rather than reading a list, so a new one is covered the day it
237
+ lands. Two deliberate limitations are recorded as tests rather than left to look covered: it cannot see
238
+ a flat denial (there is no enumeration to check — the coverage floor is what notices), and it cannot
239
+ see a deleted "only when set" qualifier. `schema/cassette.v{9,10,11}.json` are allowlisted as frozen
240
+ history: a retained schema documents the format as it shipped, and rewriting its description would make
241
+ it describe a shape its own consumers never saw.
242
+
243
+ Verified by mutation: run against the previous revision the guard flags all six then-existing sites
244
+ with the correct missing fields; renaming the shape literal makes it error rather than silently pass;
245
+ a ninth field invalidates a previously-complete site; and a whole-file token check — which would have
246
+ passed today and then never failed again — is rejected in favour of span-scoped matching.
247
+
248
+ - **The documented Action ref is now `@v2`, not `@main`** — 7 references across `README.md`,
249
+ `SKILL.md` and `ci-recipe.md`. `@main` was right when it was written: no alias tag had ever been
250
+ published, so naming one would have sent a copy-pasting reader to a `uses:` that 404s, and the guard's
251
+ own note said to revisit "once 1.0.0 ships". Two things had to be true first, and now are — `v2`/`v2.0`
252
+ point at a real release, and `release.yml` moves them on every stable release rather than leaving it to a
253
+ checklist. Recommending a floating tag nobody remembers to move is worse than recommending `@main`; that
254
+ was the actual situation while `v1` sat at 1.24.0.
255
+
256
+ `action-docs-sync` now derives the expected ref from `package.json`'s major instead of hardcoding it, so
257
+ the next major forces these docs to move with it rather than silently pointing a reader at the previous
258
+ line. `@main` is deliberately no longer accepted there: permitting both would let the recommendation
259
+ drift back with nothing noticing. Verified by mutation — regressing one reference to `@main` fails, and
260
+ setting the package version to 3.0.0 fails all three files.
261
+
262
+ **This changes nothing about which CLI you get.** The ref selects the Action; the CLI still comes from the
263
+ `version:` input, which still defaults to `latest`. `@v2` looks more like a version pin than `@main` did,
264
+ so that distinction matters more now, not less — it is spelled out in `README.md`'s Action section and in
265
+ `action.yml`'s own input description.
266
+
267
+ - **The CI recipe no longer teaches a bare version floor.** It
268
+ carried `version: ">=1.11.0"` — which reads as "at least 1.11.0" and silently means "and every future
269
+ major too", so a copy-paster gets the next major with no say in it. It was **not** broken today
270
+ (`--min-severity` still exists in 2.x, and `lint` reads no cassette, so 2.0.0's hash-format epoch never
271
+ applied to that step) — the defect was latent and in the FORM.
272
+
273
+ The bare floor was first dropped rather than corrected — `^1.11.0` would have frozen every new
274
+ copy-paster on the previous major, and `^2.0.1` needs remembering at each release — so the guidance moved
275
+ to prose. **That went one step too far, and the same release corrects it** (see the `version: "^2"` entry
276
+ above): dropping the input entirely falls back to `action.yml`'s `latest` default, which is the one
277
+ remaining unbounded form and the exact footgun the input's own description warns about two lines earlier.
278
+ `^2` was never among the alternatives weighed at the time, and it is what the recipes now carry. Reach for
279
+ an exact pin only when you want byte-reproducible CI. [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml)'s own description stops offering `>=1.11.0` and
280
+ `^1.11.0` as interchangeable — they are not, and it had been recommending the unbounded one.
281
+
282
+ `check:versions` invariant 13 now covers the Action's `version:` input, which it could not see before:
283
+ it keys on `@>=`, and `version: ">=1.11.0"` has no `@` — the same defect in different syntax, with no
284
+ coverage. Only the **unbounded** floor is rejected; verified against each form, `>=1.11.0 <3`, `^2`,
285
+ `2.0.1` and `latest` all pass.
286
+
287
+ - **`projects[].from` was missing from two places, and the second was a false green.** A connected project
288
+ is a host path exactly like a connected folder, and it was:
289
+ - **not resolved against the session file.** [`docs/session.md`](./docs/session.md) promises, without
290
+ qualification, that *"relative paths resolve from the session file's own directory"* — and it was true
291
+ of every path field except this one, which resolved against the **process CWD**. So the same session
292
+ file mounted different content depending on which directory you invoked from. The doc was right; the
293
+ resolver had simply skipped the field.
294
+ - **not part of the session fingerprint.** Swapping which directory is mounted at `.projects/<uuid>`
295
+ changed the run's inputs and `verify-cassettes` reported nothing — a false green in the gate whose job
296
+ is to notice that inputs moved. Folded in on the same **non-empty-only** terms as `agent_env`, so a
297
+ session with no `projects:` (and one with an explicit `projects: []`) hashes byte-identically to
298
+ before; only sessions that use the feature move. No committed cassette does.
299
+
300
+ **A cassette recorded before this is reported `unverifiable`, not clean.** Its hash contains nothing
301
+ about `projects[]`, so it cannot distinguish "the field was never covered" from "the mount changed since
302
+ record time" — and reporting the mismatch as a benign migration would put the same false green back in
303
+ the remedy. When everything else matches exactly, `verify-cassettes` says so and asks for a re-record to
304
+ gain the coverage. `sessionFingerprintDrift` remains `verify-cassettes`-only: none of this can change a
305
+ `replay` verdict, even under `--strict`.
306
+
307
+ Three enumerations of the covered fields gained `projects` — [`docs/cassette.md`](./docs/cassette.md),
308
+ [`docs/invariants.md`](./docs/invariants.md) and the shipped skill's `task-recipes.md`. `invariants.md`
309
+ also described this as a hash of the **resolved** session; it is the **authored, pre-resolution** shape,
310
+ deliberately, so the digest survives a different checkout — the function's own comment says so, and a
311
+ resolved hash could never match on another clone.
312
+
313
+ - **The published control-protocol schema rejected five kinds of frame the harness sends and answers.**
314
+ `schema/protocol.v1.json` described four request subtypes; the harness has always answered **six** —
315
+ adding `request_user_dialog` and `elicitation`/`side_question` — and it also sends a fail-closed
316
+ `subtype:"error"` response envelope, whose payload is a *string* under `error` rather than an object
317
+ under `response`, so a validator that knew only the success envelope rejected every one. Measured
318
+ against the previous schema: all five representative frames **REJECT**; against the new one, all five
319
+ accept. Anyone validating real traffic was seeing failures on frames the harness handles correctly.
320
+
321
+ Added as new `oneOf`/`anyOf` branches plus one new top-level response shape — **111 insertions, zero
322
+ deletions** in the surface baseline, so nothing existing was narrowed and no frame the schema already
323
+ accepted is affected. Both spellings the parser accepts are admitted (`dialogKind`/`dialog_kind`,
324
+ `mcp_server_name`/`server`, `message`/`prompt`) rather than guessing which the agent sends.
325
+
326
+ The golden vector pack grew with it, generated from the **real** envelope builders rather than
327
+ hand-authored lookalikes. The existing lockstep test — every schema definition must be exercised by a
328
+ vector — caught all five additions immediately, which is what forced those vectors to exist. One of them
329
+ asserts the error envelope does **not** validate as a success envelope, so the two shapes cannot be
330
+ quietly conflated later.
331
+
332
+ [`SPEC.md`](./SPEC.md) §12 now states the additive latitude for this surface explicitly. Its silence had
333
+ read as a prohibition, which is plausibly why three subtypes went undescribed rather than added — while
334
+ [`docs/protocol.md`](./docs/protocol.md)'s own versioning policy had said all along that an additive
335
+ variant is a v1 minor note, which is where the dated entry now lives.
336
+
337
+ - **The Marketplace alias tags are moved by the release workflow instead of by remembering.** `v1` sat at
338
+ 1.24.0 through two releases because moving it was a checklist line. `release.yml`'s last step now points
339
+ `vX` and `vX.Y` at the release it just published, with two guards a hand-run `git tag -f` skips: a
340
+ **prerelease** tag moves nothing (the trigger accepts `v1.0.4-rc.1`, and pointing `v1` at an rc would
341
+ hand every `@v1` consumer a prerelease), and an alias **never moves backwards** — re-releasing an older
342
+ patch on a line moves `vX.Y` and leaves `vX` alone. Verified by executing the logic against a synthetic
343
+ tag set rather than by reading it: releasing `v1.20.5` while `v1.25.0` exists correctly skips `v1` and
344
+ still moves `v1.20`. It runs last, after publish and the GitHub Release, so a failure there cannot
345
+ half-publish anything.
346
+
347
+ Alongside it, the tags are now correct: **`v2` and `v2.0` created** (they did not exist, so nothing
348
+ pointed at the 2.x Action), and **`v1` moved 1.24.0 → 1.25.0**. Worth recording what that move did and
349
+ did not fix: the Action's whole surface — `action.yml` plus the `render.js` it loads — is **byte-identical
350
+ from 1.24.0 through 2.0.1** apart from three lines of input *description*. So a stale `v1` was a promise
351
+ the repo had stopped keeping, not a functional gap, and the one public `@v1` consumer pins `version:` on
352
+ every step and was never exposed to the `latest` default at all.
353
+
354
+ ### Documentation
355
+
356
+ - **The invariants index said `check-versions.ts` has "no dedicated vitest file — it's a standalone script,
357
+ not a unit-testable module boundary".** Three exist, for the invariants whose logic is an exported pure
358
+ function: `check-cassette-version-claims`, `check-fingerprint-field-claims` and
359
+ `check-design-scope-note`. The claim had already been stale before this release.
360
+
361
+ - **The `uses:` ref pins the Action; the `version:` input pins the CLI — and only the second one holds a
362
+ major.** Both are documented as if pinning `@v1` bounded what you install. It does not: they move
363
+ independently, and `version:` defaults to `latest`, so promoting a CLI major reaches a workflow whose
364
+ `uses:` ref has not changed in months. Measured at the time of writing — `v1` points at **1.24.0** and
365
+ has never been moved, yet an `@v1` workflow with no `version:` input installs **2.x**. `README.md`,
366
+ `action.yml` (the text GitHub Marketplace renders) and `RELEASING.md`'s alias-tag step now say so, and
367
+ name the fix: pin the **input** (`version: ^2`), not the ref. Crossing 1.x → 2.x this way means the
368
+ hash-format epoch, so pre-v12 cassettes need `cowork-harness rehash <dir/>`.
369
+
370
+ `RELEASING.md`'s "move the major/minor tags" step additionally records what moving `vX` does *not* do,
371
+ since that step reads as the thing that controls consumer upgrades and is not.
372
+
7
373
  ## [2.0.1] — 2026-08-23
8
374
 
9
375
  ### Added