cowork-harness 3.0.0 → 3.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +10 -4
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +57 -2
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +4 -2
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +14 -6
  7. package/AGENTS.md +1 -1
  8. package/CHANGELOG.md +159 -0
  9. package/CONTRIBUTING.md +35 -0
  10. package/README.md +86 -703
  11. package/SPEC.md +2 -1
  12. package/dist/cli.js +61 -5
  13. package/dist/run/cassette.js +113 -11
  14. package/dist/run/chat-result.js +2 -0
  15. package/dist/run/chat.js +6 -0
  16. package/dist/run/execute.js +23 -2
  17. package/dist/run/hook-events.js +38 -5
  18. package/dist/run/model-provenance.js +100 -0
  19. package/dist/run/run.js +12 -0
  20. package/dist/run/verdict.js +26 -0
  21. package/docs/README.md +9 -6
  22. package/docs/boundary.md +7 -0
  23. package/docs/cassette.md +10 -3
  24. package/docs/ci.md +114 -0
  25. package/docs/cli.md +533 -0
  26. package/docs/companion-skill.md +58 -0
  27. package/docs/debugging.md +10 -7
  28. package/docs/fidelity-gaps.md +65 -1
  29. package/docs/gotchas.md +27 -2
  30. package/docs/invariants.md +1 -1
  31. package/docs/maintenance.md +28 -0
  32. package/docs/scenario.md +7 -1
  33. package/docs/session.md +5 -3
  34. package/docs/stats.md +1 -1
  35. package/examples/README.md +3 -3
  36. package/examples/replays/README.md +1 -1
  37. package/examples/sessions/default.yaml +1 -1
  38. package/llms.txt +3 -0
  39. package/package.json +1 -1
  40. package/schema/cassette.v12.json +15 -3
  41. package/schema/run-result.json +30 -1
  42. package/scripts/bump-version.ts +12 -1
  43. package/scripts/check-versions.ts +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.0.0
7
- tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
6
+ version: 3.1.0
7
+ tracks-harness: cowork-harness 3.1.0 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.1.0` (baseline
29
29
  > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.1.0"`. **Pin `@^3.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -568,6 +568,12 @@ Recognize these before "fixing" a non-bug:
568
568
  would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
569
569
  with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
570
570
  a harness coverage gap (see the *File delivery* section of fidelity-gaps).
571
+ - **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
572
+ `model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
573
+ this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
574
+ transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
575
+ by a different model than the scenario names, so treat a green as evidence about the fallback model.
576
+
571
577
  - **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
572
578
  connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
573
579
  approval, not just outputs — a connected folder shows the identical default — so this run diverged from
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.1.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^3.0.0"
70
+ - run: npm i -g "cowork-harness@^3.1.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^3.0.0"
345
+ - run: npm i -g "cowork-harness@^3.1.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^3.0.0"
374
+ run: npm i -g "cowork-harness@^3.1.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -20,7 +20,8 @@ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.406
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
23
- - A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
23
+ - A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
24
+ or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
24
25
  (`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
25
26
  hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
26
27
  Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
@@ -155,6 +156,60 @@ hands `fn` exactly this dict.
155
156
  - `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
156
157
  `["fail", "prompt", "llm", "first"]`.
157
158
 
159
+ ### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
160
+
161
+ The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
162
+ it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
163
+ carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
164
+ off-registry**.
165
+
166
+ Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
167
+ ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
168
+ both the command and the tier. All rows below are measured, same scenario shape, same session:
169
+
170
+ | Tier | Bash command | tool used | ask | verdict |
171
+ |---|---|---|---|---|
172
+ | `protocol` | `echo hello` | `Bash` | none | ✓ green |
173
+ | `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
174
+ | `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
175
+ | `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
176
+
177
+ So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
178
+ CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
179
+ `container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
180
+ even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
181
+ `protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
182
+ hello-world probe is green everywhere and teaches the wrong expectation.
183
+
184
+ Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
185
+ warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
186
+ allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
187
+ permissive behaviour is deliberately what the scenario is about.
188
+
189
+ > **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
190
+ > traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
191
+
192
+ ### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
193
+
194
+ A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
195
+ a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
196
+ current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
197
+
198
+ Prefer **repinning `baseline:` to an installed version** for anything
199
+ reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
200
+ but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
201
+ failure modes.
202
+
203
+ To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
204
+ definitions shipped with the harness, which are present regardless of what Desktop pruned from this
205
+ machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
206
+ the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
207
+ directory behind, so a directory test passes on exactly the case that fails.
208
+
209
+ `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
210
+ binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
211
+ prevent, so use it to unblock once, never in CI.
212
+
158
213
  ### Determinism contract
159
214
 
160
215
  - `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.1.0`
4
4
  (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -100,7 +100,8 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
100
100
  # Read-only folders and folder-less runs need no opt-in.
101
101
 
102
102
  allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
103
- # declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
103
+ # declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
104
+ # the manifest's `hooks` key: L0 passes
104
105
  # --plugin-dir, so the CLI executes those hooks as NATIVE HOST
105
106
  # processes under your account, with no container sandbox. A plugin
106
107
  # that declares no hooks needs no opt-in, and a misplaced root-level
@@ -420,6 +421,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
420
421
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
421
422
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
422
423
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
424
+ | `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
423
425
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
424
426
  | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
425
427
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -46,6 +46,11 @@ what production would do. Two layers of defense:
46
46
  statically knowable and only produce a non-failing informational note.
47
47
  - **Manual (one-liner):** `grep -h '"effectiveFidelity"' cassettes/*.cassette.json | sort | uniq -c`
48
48
  shows the tier distribution of the fleet at a glance.
49
+ - **Do NOT re-record for a `[note]` alone.** A `·`/`[note]` row is informational and the run still exits
50
+ **0** — `session-fingerprint: predates \`model\` coverage` and `prompt-assets: cassette predates …` both
51
+ mean "this cassette was recorded before that field existed, and everything the hash *does* cover still
52
+ matches". Re-recording buys the new coverage and nothing else, so on a fleet of heavy cassettes it is a
53
+ real bill for no verdict change. Re-record when a `✗` says to (baseline moved, skill drift, tier moved).
49
54
 
50
55
  ### Cassette anatomy (what you're looking at when you open one)
51
56
 
@@ -66,7 +71,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
66
71
  | `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
67
72
  | `scenarioSource` | Relative path to the authored YAML this was recorded from |
68
73
  | `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
69
- | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
74
+ | `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (model/folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
70
75
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
71
76
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
72
77
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
@@ -241,10 +246,13 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
241
246
  unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
242
247
  dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
243
248
  whose source was never committed names a generation that is unrecoverable, which makes the
244
- comparison uninterpretable rather than merely noisy. And with no `model:` in the session (or
245
- `--model` on the `skill` lane) each run uses whatever the staged agent binary defaults to, so a
246
- before/after can silently straddle two models; read `result.json`'s `models` back to confirm —
247
- ignoring any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent fabricated locally, not a model).
249
+ comparison uninterpretable rather than merely noisy. And with no `model:` in the session and no
250
+ `--model` on the command (every lane takes it), each run uses whatever the staged agent binary
251
+ defaults to, so a before/after can silently straddle two models — the run warns when nothing pinned
252
+ one. Read `result.json` back to confirm: `modelSource` says whether anything pinned the model at all,
253
+ and `modelPinHonored` whether the pin survived (**absent means unverifiable, not "yes"**). `models`
254
+ lists what served the run — ignore any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent
255
+ fabricated locally, not a model).
248
256
  3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
249
257
  `result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
250
258
  content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
package/AGENTS.md CHANGED
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
22
22
  Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
23
23
  `npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
24
24
  - CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
25
- see README's [Reproducibility knobs](./README.md#reproducibility-knobs) for the full env-var list.
25
+ see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
26
26
  Node ≥ 22.
27
27
  - `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
28
28
  `baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
package/CHANGELOG.md CHANGED
@@ -6,6 +6,165 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.1.0] — 2026-08-31
10
+
11
+ ### Fixed
12
+
13
+ - **The pre-coverage cassette note was a paragraph repeated per file.** It fires on every cassette
14
+ recorded before `model` joined the session shape — for a consumer with a dozen, a dozen identical
15
+ paragraphs burying the findings that are genuinely per-file. Cut to one line (370 chars to 157), matching
16
+ the prompt-assets note beside it; the reasoning stays in `docs/fidelity-gaps.md`. Notes are deliberately
17
+ **not** aggregated in `verify-cassettes` — it is a per-file audit, and collapsing them would destroy the
18
+ attribution — so length was the part to fix. The note has always been non-failing: a pre-coverage
19
+ cassette stays clean, and `replay --strict` exits 0.
20
+ - **The unpinned-model warning told `run` users the flag lived on other lanes.** It named the `skill`,
21
+ `probe-dispatch` and `chat` lanes for a flag `run` and `record` now accept — caught by running it, not
22
+ by reading it.
23
+
24
+ - **`run --model ""` started a spending run instead of failing.** SPEC §CB-2 requires an empty or
25
+ whitespace model value to be a usage error, never a silently-propagated empty string. `record` inherits
26
+ that from the shared flag parser; `run` hand-rolls its own loop (the session-less lanes take `--model`
27
+ out of the leftovers themselves, so the shared parser cannot consume it) and did not enforce it.
28
+
29
+ ### Added
30
+
31
+ - **`--model <id>` on `run` and `record`.** The `skill`, `probe-dispatch` and `chat` lanes already had it;
32
+ the two lanes that produce cassettes and CI verdicts did not, so the only way to pin them was editing the
33
+ session file. Precedence is explicit: an `--model` flag or a `--matrix` `models:` axis (a per-invocation
34
+ act) outranks the session's `model:`, while `COWORK_HARNESS_MODEL` only fills a gap — a machine-scoped
35
+ variable must never silently outrank a model the scenario declares, or the run's model becomes a
36
+ property of the shell it was launched from.
37
+ - **Model provenance on every run, and a warning when nothing pins the model.** `--model` reaches the
38
+ agent only when a session declares `model:`, so an unset session inherits whatever the local CLI would
39
+ pick — and because the agent selects part of its system prompt by model capability, that moves the
40
+ *instructions* it is given, not just answer quality. `--effort` is not symmetric: it is always emitted
41
+ from the baseline's `spawn.effortDefault`, so an unset session pinned the effort and left the model it
42
+ applies to floating. Runs now warn when no model resolves (`chat` in its own words), naming the key and
43
+ the flag. Omitting it is deprecated and becomes an error in the next major — the same path
44
+ `Scenario.fidelity` took, and deliberately not a hard fail today: the repo chose deprecation for a
45
+ default with a larger blast radius, and a harder gate for a lesser field would be incoherent with that.
46
+ - **A family pin is checked as family membership, not equality.** `--model opus` is resolved by the agent
47
+ to a concrete id (`claude-opus-5`), never echoed back, and which member an account resolves it to is not
48
+ something the harness can know — so comparing the two as strings would report a pinned-and-honored run as
49
+ `modelPinHonored: false`. `opus`/`sonnet`/`haiku`/`fable` now compare by family and still fail on a
50
+ wrong-family model; `best` and `opusplan` name no comparable model and stay unverifiable.
51
+ - **`RunResult.modelFallbacks`, `modelPinHonored` and `modelSource`.** Fallbacks come from the agent's own
52
+ `system`/`model_fallback` event rather than a diff of `models[]` — the event names the trigger, which is
53
+ what separates a retired or blocked pin (every run falls back the same way until the id changes) from a
54
+ transient overload. A new `model_fallback` warn signal surfaces it in the verdict; it cannot fire on a
55
+ healthy run, so it adds no volume to a currently-green one. `modelPinHonored` is three-state on purpose:
56
+ `true`, `false`, and **absent for unverifiable** — nothing pinned, or no live model evidence. The
57
+ natural two-state shape reports "we could not tell" as "the pin held", which is a false green, so the
58
+ no-evidence cases are asserted explicitly rather than left to a truthiness check.
59
+ - **A cassette records the model it was recorded with**, in its `environment` block: the id the agent
60
+ reported (so it survives a mid-run fallback) plus the source. A cassette recorded with nothing pinned
61
+ says `unresolved`, which is the difference between "recorded on opus-5 by choice" and "on whatever that
62
+ laptop happened to pick".
63
+ - **The session fingerprint now covers the pinned model.** Editing `model:` after recording previously
64
+ moved nothing and `verify-cassettes` said nothing, even though the edit changes what the agent was told.
65
+ A cassette recorded before this coverage reports **unverifiable** rather than clean — its hash cannot
66
+ tell "the field was never covered" from "the pin changed since", and claiming otherwise in the remedy
67
+ would reintroduce the false green the coverage exists to close.
68
+
69
+ ## [3.0.1] — 2026-08-30
70
+
71
+ ### Changed
72
+
73
+ - **The README now shows the product, not just the argument for it.** Three independent reviews found
74
+ the same gap: for a test harness, the file a user authors is the most persuasive thing the project
75
+ owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
76
+ answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
77
+ a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
78
+ actual run where a named skill was offered, declined, and the green answer looked fine anyway
79
+ (`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
80
+ it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
81
+ the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
82
+ diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
83
+ can fail before trusting it. The requirements block moves below the `claude -p` argument — three
84
+ reviewers independently reported reading install prerequisites before any reason to want them.
85
+
86
+ - **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
87
+ audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
88
+ prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
89
+ + orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
90
+ Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
91
+ to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
92
+ README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
93
+ the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
94
+
95
+ Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
96
+ `docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
97
+ than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
98
+ moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
99
+
100
+ ### Fixed
101
+
102
+ - **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
103
+ `docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
104
+ but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
105
+ failed its own post-write lockstep check. Caught at release time by that self-check rather than
106
+ shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
107
+ carries the reason.
108
+
109
+
110
+ - **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
111
+ Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
112
+ README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
113
+ `docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
114
+ scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
115
+ suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
116
+ headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
117
+ with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
118
+ regression that shipped. The two nav lines are deleted rather than repointed — they were a
119
+ table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
120
+
121
+ Found by two independent reviewers; the second caught the three non-README instances a README-scoped
122
+ fix would have missed.
123
+
124
+ Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
125
+ **below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
126
+ page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
127
+ This half is the nastier one — the link works, so every link checker stays green forever.
128
+
129
+ - **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
130
+ (3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
131
+ processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
132
+ hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
133
+ the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
134
+ under the operator's own account while the run went green, no consent was asked and no disclosure
135
+ printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
136
+ unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
137
+ the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
138
+ deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
139
+ inert, and a manifest-declared hook is live wherever the manifest sits.
140
+
141
+ Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
142
+ 10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
143
+ declaration, including one recorded at `container`.
144
+
145
+ ### Documentation
146
+
147
+ - **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
148
+ update deletes the prior version's staged ELF while often leaving an empty version directory, so a
149
+ scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
150
+ this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
151
+ ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
152
+ `docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
153
+ repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
154
+ silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
155
+ never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
156
+
157
+ - **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
158
+ `Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
159
+ ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
160
+ hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
161
+ identical command at `container` produces no ask (green), and `hostloop` routes shell through
162
+ `mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
163
+ and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
164
+ expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
165
+ the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
166
+ unmeasured.
167
+
9
168
  ## [3.0.0] — 2026-08-29
10
169
 
11
170
  ### Breaking
package/CONTRIBUTING.md CHANGED
@@ -108,3 +108,38 @@ When a Desktop release moves something `sync` doesn't read, it reports an `unkno
108
108
  ## Reporting issues
109
109
 
110
110
  Use the issue templates. For anything security/sandbox related, see [SECURITY.md](./SECURITY.md).
111
+
112
+ ## This repo's own CI pipeline
113
+
114
+ Contributor-facing. For *consuming* the harness in your own CI — the token-free gate and the packaged
115
+ Action — see [docs/ci.md](./docs/ci.md).
116
+
117
+ The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
118
+
119
+ | Stage | Runs | Needs | Gates |
120
+ |---|---|---|---|
121
+ | **build** | format check · version-lockstep guard · typecheck · source guards · build · CLI smoke · token-free `replay` · `verify-cassettes` · `lint` | nothing | every push/PR |
122
+ | **test** | the unit suite (vitest), sharded 4-way | nothing | every push/PR |
123
+ | **floor** | the unit suite once, unsharded, on Node 22 — the version `engines.node` declares — so the floor is exercised rather than asserted (other jobs run Node 24, the Active LTS line) | nothing | every push/PR; gates the merge context |
124
+ | **action-self-test** | packs this commit and runs the packaged `uses: ./` Action across its full case set — pass (committed example cassette), usage-error fail (nonexistent path), assertion fail (checks the reporter renders a ❌ row), `lint`, and `analyze-skill`; `ci.yml` currently carries 11 `command:` invocations, so re-count here rather than trusting this sentence | nothing | every push/PR |
125
+ | **python** | `pytest` helper self-checks (`python/`, run with `-m 'not cowork'` — the token-free subset; the Docker/token `@pytest.mark.cowork` tests are excluded) | nothing (token-free assertions only) | every push/PR |
126
+ | **boundary** | builds the pinned agent image, brings up the default-deny network, runs `boundary-check`, then `npm run test:live` (live contract tests that guard the binary-resolution assumptions, no token needed) | Docker, arm64 runner | proves the sandbox enforces Cowork's limits — **no API key** |
127
+ | **image-recipe** | compiles `docker/Dockerfile.agent` in **both** variants (lean/core and `COWORK_FULL_PARITY=1`) on an arm64 runner, so a recipe change cannot reach the merge gate uncompiled | Docker, arm64 runner | every push/PR; gates the merge context |
128
+ | **scenarios** | the live scenario suite (mixed `protocol` + `container` fidelity across `examples/scenarios/`), plus the `e2e/scenarios/*.yaml` smoke set (this repo's own L0/L1/hostloop self-tests, `microvm` excluded — needs a real VM); uploads transcripts/egress logs as artifacts; relies on a runner-local staged agent binary (no in-workflow download step). | `ANTHROPIC_API_KEY` | fork PRs: the whole job is skipped (`if:` guard); same-repo without a key: warns and exits 0 |
129
+ | **parity-drift** | reminder to re-`sync` when Desktop updates | nothing | **goes red** if the newest committed baseline is &gt; 90 days old, but sits outside `ci-green`'s `needs:` list, so a red run does not block a merge |
130
+
131
+ This ordering means cheap checks fail fast, the **boundary parity gate runs without secrets** (so forks get it too), and expensive live runs only happen when a key is present.
132
+
133
+
134
+ ## The harness's own suite
135
+
136
+ ```bash
137
+ npm run ci # typecheck + build + test (run format:check separately; NOT the same set as CI's `build` job — see CONTRIBUTING.md)
138
+ npm test # vitest: decider, egress allowlist, launch plan, example validation
139
+ cowork-harness boundary-check # self-verify the sandbox (needs Docker; not part of `npm run ci`)
140
+ ```
141
+
142
+ Unit tests cover the scripted-answer logic, the egress allowlist matcher, the session→launch-plan materialization (mounts + discovery settings + env-strip), and a **schema guard** that fails if any shipped baseline/session/scenario stops validating. Add a test alongside any new schema field or `Decider` rule — see [CONTRIBUTING.md](./CONTRIBUTING.md).
143
+
144
+ > Copy your starting scenarios/sessions from **`examples/`**. The **`e2e/`** directory is the harness's *own* fidelity self-tests (smoke scenarios per tier) — not a template to copy.
145
+