cowork-harness 2.4.0 → 2.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +5 -5
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/critique.md +72 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +5 -1
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +2 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +64 -0
  9. package/CHANGELOG.md +203 -0
  10. package/README.md +6 -6
  11. package/SPEC.md +12 -0
  12. package/dist/assert.js +42 -1
  13. package/dist/cli.js +8 -0
  14. package/dist/critique/command.js +202 -22
  15. package/dist/critique/evaluator.js +9 -3
  16. package/dist/critique/package-evidence.js +39 -5
  17. package/dist/run/cassette.js +23 -1
  18. package/dist/run/chat-result.js +1 -0
  19. package/dist/run/execute.js +31 -1
  20. package/dist/run/probe-dispatch.js +8 -2
  21. package/dist/run/provenance.js +6 -9
  22. package/dist/run/run.js +197 -8
  23. package/dist/run/tier-vacuous-tools.js +69 -0
  24. package/dist/run/tool-name-canonicalization.js +70 -0
  25. package/dist/types.js +36 -3
  26. package/docs/cassette.md +2 -0
  27. package/docs/critique.md +52 -0
  28. package/docs/debugging.md +3 -1
  29. package/docs/scenario.md +8 -5
  30. package/docs/subagents.md +5 -3
  31. package/examples/replays/README.md +1 -1
  32. package/package.json +1 -1
  33. package/schema/critique-report.json +41 -18
  34. package/schema/run-result.json +47 -1
  35. package/schema/scenario.schema.json +13 -3
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.4.0
7
- tracks-harness: cowork-harness 2.4.0 (baseline desktop-1.37937.1)
6
+ version: 2.5.0
7
+ tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.4.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
29
29
  > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.4.0"`. **Pin `@^2.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -147,7 +147,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
147
147
  | Tier | What it gives you | Use when |
148
148
  |---|---|---|
149
149
  | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
150
- | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm`/`protocol` never offer it, so the same `tool_not_called` is vacuous there — moving a scenario between tiers can silently void a web-fetch assertion | Most functional + boundary tests. |
150
+ | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
151
151
  | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
152
152
  | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
153
153
 
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.4.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.4.0"
70
+ - run: npm i -g "cowork-harness@^2.5.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.4.0"
345
+ - run: npm i -g "cowork-harness@^2.5.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.4.0"
374
+ run: npm i -g "cowork-harness@^2.5.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
34
34
  **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
35
35
  UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
36
36
 
37
+ ## An exit-2 report — which turn failed, and whether it was really infrastructure
38
+
39
+ Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
40
+ carry that; read all three before touching anything.
41
+
42
+ | Field | Meaning |
43
+ |---|---|
44
+ | `infraFailure` | the reason |
45
+ | `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
46
+ | `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
47
+ | `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
48
+
49
+ **Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
50
+ unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
51
+ agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
52
+ and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
53
+ **`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
54
+ turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
55
+ unrecognized kind renders as infrastructure):
56
+
57
+ - `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
58
+ - `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
59
+
60
+ **A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
61
+ full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
62
+ a broken instrument.
63
+
64
+ **Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
65
+ verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
66
+ "the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
67
+ `--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
68
+ reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
69
+ is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
70
+ *conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
71
+
72
+ For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
73
+ gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
74
+ (or the scenario's `on_unanswered: llm`).
75
+
76
+ ## Which model was graded — `gradedModels`
77
+
78
+ **The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
79
+ session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
80
+ which may not be the model you are otherwise working under, and nothing about the run announces it.
81
+
82
+ `gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
83
+ is the only record of which model produced the behaviour being graded — distinct from the evaluator's
84
+ resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
85
+ separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
86
+ `--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
87
+ read `gradedModels` back to confirm it took.
88
+
89
+ It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
90
+ messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
91
+ (crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
92
+ can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
93
+
37
94
  ## The report's item shape — no `title`, no `summary`
38
95
 
39
96
  Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
80
137
  specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
81
138
  `not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
82
139
 
140
+ ## Read `referencesAccessed`, not `referencesRead`
141
+
142
+ `referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
143
+ `Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
144
+ `referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
145
+ (`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
146
+
147
+ Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
148
+ entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
149
+ then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
150
+ weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
151
+ nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
152
+ "none".
153
+
83
154
  ## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
84
155
 
85
156
  `result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.4.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
4
4
  (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -305,6 +305,10 @@ same set live from the schema.
305
305
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
306
306
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
307
307
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
308
+ | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
309
+ | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
310
+ | `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
311
+ | `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
308
312
  | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
309
313
  | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
310
314
  | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -35,6 +35,7 @@
35
35
  "no_hook_blocked",
36
36
  "no_lost_write_back",
37
37
  "no_mcp_error",
38
+ "no_observed_reference_access",
38
39
  "no_path_denied",
39
40
  "no_scratchpad_leak",
40
41
  "no_skill_triggered",
@@ -46,6 +47,7 @@
46
47
  "question_context",
47
48
  "question_options",
48
49
  "questions_count_max",
50
+ "reference_read",
49
51
  "replay_protocol_fidelity",
50
52
  "result",
51
53
  "self_heal_ran",
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
80
80
  "tool_result_not_matches",
81
81
  "tool_called",
82
82
  "tool_not_called",
83
+ # Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
84
+ # tool inputs), so they are content keys, not live-only.
85
+ "reference_read",
86
+ "no_observed_reference_access",
83
87
  "subagent_tool_used",
84
88
  "subagent_tool_absent",
85
89
  "subagent_dispatched",
@@ -294,6 +298,8 @@ REGEX_KEYS = {
294
298
  "hook_blocked",
295
299
  "tool_result_matches",
296
300
  "tool_result_not_matches",
301
+ "reference_read",
302
+ "no_observed_reference_access",
297
303
  }
298
304
  VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
299
305
  VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
@@ -583,6 +589,44 @@ def lint_doc(doc, path, raw_lines):
583
589
  # warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
584
590
  # declare the combination incompatible, so authoring it is a bug even if a tool-free run could
585
591
  # accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
592
+ # `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
593
+ # vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
594
+ # (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
595
+ # and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
596
+ # passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
597
+ # must never be flagged here. The harness refuses these at load; this catches them before any run.
598
+ _TIER_VACUOUS = {
599
+ "hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
600
+ "container": {"mcp__workspace__bash": "Bash"},
601
+ "microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
602
+ # `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
603
+ # host CLI registry — machine-dependent and about a different product.
604
+ }
605
+ # BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
606
+ # USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
607
+ for _key in ("tool_not_called", "subagent_tool_absent"):
608
+ for _v in _assert_values(items, _key):
609
+ if not isinstance(_v, str) or "*" in _v or "?" in _v:
610
+ continue # a glob is not a literal claim about one tool
611
+ _repl = _TIER_VACUOUS.get(fidelity, {})
612
+ if _v not in _repl:
613
+ continue
614
+ _instead = _repl[_v]
615
+ findings.append(
616
+ Finding(
617
+ "WARN",
618
+ "tool-not-called-tier-vacuous",
619
+ f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
620
+ f"`{_v}` at all, so this can never be violated and verifies nothing.",
621
+ (
622
+ f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
623
+ if _instead
624
+ else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
625
+ ),
626
+ path,
627
+ )
628
+ )
629
+
586
630
  # linter stays offline — the message carries the gate fact instead of reading a baseline).
587
631
  if "transcript_no_host_path" in assert_keys:
588
632
  if fidelity in ("hostloop", "protocol"):
@@ -1016,6 +1060,26 @@ def lint_doc(doc, path, raw_lines):
1016
1060
  )
1017
1061
  )
1018
1062
 
1063
+ # Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
1064
+ # with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
1065
+ # that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
1066
+ # this only fires on the case that is unambiguously self-defeating.
1067
+ read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
1068
+ unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
1069
+ both_refs = sorted(read_pats & unread_pats)
1070
+ if both_refs:
1071
+ findings.append(
1072
+ Finding(
1073
+ "ERROR",
1074
+ "reference-access-contradiction",
1075
+ f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
1076
+ "(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
1077
+ "Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
1078
+ "another is not, give the two keys different patterns.",
1079
+ path,
1080
+ )
1081
+ )
1082
+
1019
1083
  # W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
1020
1084
  findings.extend(_lint_regex_quoting(path, raw_lines))
1021
1085
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,209 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [2.5.0] — 2026-08-28
10
+
11
+ ### Added
12
+
13
+ - **`cowork-harness assertions --list` gains a "Skill references (progressive disclosure)" family** for the
14
+ two new keys, so they are not appended to a flat dump nobody reads.
15
+
16
+ - **A guard that a committed cassette's `tool_not_called` is actually violable by its own recording.**
17
+ The tool surface at a tier is **gate-conditional**: at `container`, `mcp__workspace__web_fetch` is
18
+ offered only when `coworkWebFetchViaApi` is on (17 of 138 measured container runs), and with the gate
19
+ off `WebFetch` is offered instead. `examples/replays/example-pdf-skill.cassette.json` asserts
20
+ `tool_not_called: "mcp__workspace__web_fetch"` and passes today only because it was recorded gate-ON —
21
+ a re-record with the gate off would leave it naming a tool the run could never have called, so it would
22
+ keep passing while verifying nothing, and nothing in the suite would notice. `verify-cassettes`'s
23
+ `replaced-builtin` note keys on the recorded *inventory*, never on the assertions, and only covers
24
+ built-ins being replaced, never the inverse.
25
+
26
+ The guard checks every committed cassette's `tool_called`/`tool_not_called` against that cassette's own
27
+ frozen init inventory, through the same glob engine the evaluator uses. It deliberately does **not**
28
+ cover `subagent_tool_absent` (judged against the per-dispatch `declaredTools`, a different inventory)
29
+ and cannot see alias-class vacuity (`tool_not_called: "Task"` names a tool that is in every inventory
30
+ yet never emitted — the agent binary canonicalizes `Task` to `Agent`); both are recorded as separate
31
+ work rather than left implied.
32
+
33
+ - **`referencesAccessed` — reference access through EVERY tool channel, not just the `Read` tool.**
34
+ `referencesRead` counts one channel, and `critique`'s headline invited a reader to conclude the agent
35
+ had opened no reference — a claim about *reading* that a one-channel count cannot support. An agent
36
+ that `cat`s, `grep`s or globs a reference has reached it just as much. The new `RunResult` field (and
37
+ its `subagents[]` twin) records each file with the channel(s) it was reached through: `read`
38
+ (`Read.file_path`), `grep` (its `path` input), and `bash` (a `Bash`/`mcp__workspace__bash`
39
+ command naming the path).
40
+
41
+ All four channels apply the **same** `skillReferenceReadPath()` predicate, so a token only counts when
42
+ it is rooted in the mounted plugin — the agent's own `node scripts/build.js` is not a skill-script
43
+ access, and a non-skill filename never reaches `result.json` under a field claiming it is skill
44
+ content. Redirection targets and every argument of a write verb (`rm`/`mv`/`mkdir`/`touch`/`chmod`/`tee`) is excluded, as is a verb that only inspects metadata (`ls`/`test`/`stat`/`echo`): those are files the command
45
+ wrote or destroyed.
46
+
47
+ It **deliberately under-approximates**. A `cd` into the skill dir followed by a bare
48
+ `cat references/x.md`, a heredoc body and a `$VAR`-built path are all invisible, and there are tests
49
+ pinning them as misses so a later widening is a visible change rather than a silent one.
50
+
51
+ **Presence is the cannot-verify channel**, which is the one way it differs from `referencesRead`:
52
+ `[]` means the drive ran and observed nothing (a real negative); **absent** means there was no
53
+ observable drive, and must never be read as "none". Present on live *and* replay — cassettes freeze
54
+ whole tool inputs, so the replay re-drive reconstructs every channel identically.
55
+
56
+ `referencesRead` is unchanged in meaning and is now documented as this field's `read`-channel
57
+ projection, produced by the same capture so the two cannot disagree.
58
+
59
+ Scope is **main agent ∪ sub-agents**, via a single `unionReferenceAccesses()` derivation shared by the
60
+ assertion keys, the critique report and `--probe dispatch` — a dispatcher-shaped skill does all its
61
+ reading a level down, so judging the top-level list alone would report "never reached" on a run where a
62
+ sub-agent read the file cover to cover. A **truncated cassette** (one that could never be driven)
63
+ reports cannot-verify rather than an empty list, for the same reason.
64
+
65
+ - **`reference_read` / `no_observed_reference_access` assertion keys.** Gate on whether a skill's
66
+ progressive disclosure actually works: a well-partitioned skill and one whose second half is dead look
67
+ identical from outside. Regex (unanchored, case-insensitive — the same shared helper every regex key
68
+ uses), main agent ∪ sub-agents, evaluated on replay as well as live.
69
+
70
+ The negative key is named `no_observed_reference_access`, not `no_reference_read`, because the
71
+ detector under-approximates by design: it proves nothing was *seen*, not that the file went unread.
72
+ Both keys **fail evidence-unavailable** when the run recorded no observable access list — including
73
+ the negative one, which is the direction that would otherwise pass vacuously off a missing field. The
74
+ scenario linter rejects asserting both with the same pattern.
75
+
76
+ - **`critique` names why the GRADED turn errored** — `gradedErrorReason` in the JSON report and inline in
77
+ the text NOTE. `taskResult: "error"` is a legitimate **gradeable** outcome and the critique still runs,
78
+ but the report said only "the task run ended in error", so an exhausted quota or a dropped connection
79
+ read as a defect in the skill under review.
80
+
81
+ - **`critique` reports which model produced the GRADED run** — `gradedModels` in the JSON report and
82
+ `graded model(s):` in the text header, read back from the graded turn's own `result.json` and filtered
83
+ of the agent's locally-fabricated `<synthetic>` entries. The report named the *evaluator's* resolved
84
+ model and no other, so the model that produced the behaviour being graded appeared nowhere. It cannot
85
+ be inferred from context either: the turns are a subprocess that inherits no model from whatever
86
+ invoked `critique`, so an omitted `--model` silently grades whatever the spawned agent defaults to. A
87
+ run with nothing recorded now says `graded model(s): unknown` rather than staying silent. The ids are
88
+ **observed, not requested** — read from the model stamped on the graded turn's assistant messages, not
89
+ from the flag.
90
+ - `isLiveModelId` (`src/types.ts`) — the agent-marker filter every consumer of `RunResult.models` must
91
+ apply, now declared once beside the field it governs. It replaces **three** divergent copies, two of
92
+ which disagreed: `src/run/provenance.ts` (which renders `provenance.model` on every JSON envelope)
93
+ matched the angle-bracket **shape**, `scripts/eval-gate.ts` the `<` **prefix** — so a malformed or
94
+ truncated marker such as `"<synthetic"` was dropped by one and rendered as if it were a model id by the
95
+ other. The shared rule is the shape.
96
+
97
+ ### Changed
98
+
99
+ - **Skill and docs updated for the tier refusal.** `SKILL.md` previously told authors a tier-mismatched
100
+ `tool_not_called` "can silently void a web-fetch assertion" when moving a scenario between tiers — the
101
+ harness now refuses it at load instead, so that guidance described behaviour that no longer exists. The
102
+ skill's critique reference gains the `referencesAccessed` field, its channels, and the
103
+ `[]`-vs-absent cannot-verify rule.
104
+
105
+ - **`subagent_declared_but_unused` documents that it fires only on a dispatch that declares a tool list.**
106
+ It reads `subagents[].declaredTools`, populated from a `tools`/`allowedTools` key in the dispatch input.
107
+ The `Agent` tool carries neither, so that list is empty and the key passes on every such dispatch —
108
+ **0 of 1091 real dispatches** carry a non-empty list. The key is not wrong, but a green means "not
109
+ applicable here", never "no fabrication", and neither its description nor its docs said so.
110
+
111
+ - **BEHAVIOUR CHANGE: `tool_not_called` and `subagent_tool_absent` naming a tool the tier does not serve
112
+ are now REFUSED at scenario load.** `tool_not_called: "Bash"` at `hostloop` passed vacuously, always — that tier disallows the
113
+ built-in shell and aliases it to `mcp__workspace__bash`, so the run could never have called `Bash`
114
+ whatever the agent did. The assertion read as a guarantee and verified nothing. The inverse was equally
115
+ broken: `mcp__workspace__bash` at `container`, where the built-in is served instead.
116
+
117
+ The refusal is a `UsageError` at the point the tier resolves — after `fidelity: cowork` becomes a real
118
+ tier, and before any staging, image pull or spawn, so no model spend is wasted. The message names the
119
+ tool to assert instead, and names the sibling keys that behave differently (`tool_called` still fails
120
+ normally; `subagent_tool_absent` is judged against a per-dispatch inventory this check cannot
121
+ determine). `scenario.py lint` WARNs on the same set, before any run at all.
122
+
123
+ **The table is closed to tools the harness itself removes or registers**, and never derived from the
124
+ launch plan. `--tools` gates the built-in set alone — the agent binary's own help says so — while every
125
+ sandbox tier separately passes `--mcp-config`. A launch-set-derived check would therefore have rejected
126
+ `tool_not_called: "mcp__example-fs__write_file"` for a session using the `mcp.config` this repo ships an
127
+ example of, while that tool was registered and callable. Globs are never refused, and the table
128
+ under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the correct
129
+ side to err on when the verdict is a hard refusal.
130
+
131
+ There is **no opt-out**, deliberately. The repo's `allow_*` modifiers all cover cases where the harness
132
+ might be wrong about a real signal; a fired reject here cannot be a false positive, so there is no
133
+ legitimate scenario to rescue. If one is ever found, the table is wrong and the table should change.
134
+
135
+ **`subagent_tool_absent` is covered for the same reason `tool_not_called` is.** It reads the tools
136
+ sub-agents actually USED, not a per-dispatch declared list, so a tool the tier never serves makes it
137
+ vacuous in exactly the same way — corroborated by the run population, where sub-agent `Bash` calls
138
+ appear 20 times, all at `container` and never at `hostloop`. Covering one key and not the other would
139
+ refuse an assertion at hostloop while silently greening the sub-agent form of the identical claim.
140
+
141
+ `e2e/scenarios/canary-hostloop.yaml` carried exactly this defect and is fixed in the same change: its
142
+ `tool_not_called: Bash` never proved anything, and its `tool_called: mcp__workspace__*` already carries
143
+ the canary's stated purpose.
144
+
145
+ - **`critique`'s "no references were Read" headline now says what it observed.** It reads the wide
146
+ signal, names the channels it looked through, and states its own under-approximation in one short
147
+ clause instead of the load-bearing caveat that did all the work. Where the run recorded no observable
148
+ tool stream it makes **no claim** rather than rendering a clean negative. The evaluator's evidence
149
+ section moves with it (previously main-agent `Read`s only — a different population from the headline's,
150
+ so widening one without the other would have handed the evaluator a prompt contradicting the report),
151
+ and the grading prompt now tells the evaluator not to issue a finding whose only support is a path
152
+ missing from that list. `--probe dispatch` prints the wide list too.
153
+
154
+ ### Fixed
155
+
156
+ - **`tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did.** The agent binary
157
+ canonicalizes a set of legacy tool names — `Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, and nine
158
+ more — while the spawn tool list still declares the **legacy** spelling. So the init inventory echoes
159
+ back `Task`, every actual dispatch is emitted as `Agent`, and a literal matcher could never connect the
160
+ two. Measured across 506 kept runs: `Task` offered **506** times and called **0**; `Agent` called **188**
161
+ times and offered **0**. The negative form was a permanent vacuous pass in the most common dispatch
162
+ assertion there is, at every tier.
163
+
164
+ All four tool keys (`tool_called`, `tool_not_called`, `subagent_tool_used`, `subagent_tool_absent`) now
165
+ match either spelling, through the one shared matcher. Globs see both too — the **recorded name** is
166
+ expanded rather than the author's pattern, so `tool_not_called: "Ta*"` and `"*"` are violated by a
167
+ recorded `Agent`, which rewriting the pattern would not have fixed.
168
+
169
+ Recorded data is unchanged: `toolCounts`, `context.tools` and every cassette keep exactly what the agent
170
+ reported, the same verbatim posture `RunResult.models` documents. Only matching is alias-aware, so
171
+ `tool_available: "Task"` still matches the inventory's literal `Task` and no committed cassette changes
172
+ meaning.
173
+
174
+ **Still vacuous, and now documented as such:** `tool_not_called` on a tool the *tier* never offers —
175
+ `Bash` at `hostloop`, or `mcp__workspace__bash` at `container`. That is a separate class with a separate
176
+ fix; the key's docs and the skill reference now warn about it rather than leaving it implied.
177
+
178
+ - **`critique` diagnosed an unanswered gate as an infrastructure failure, and named the wrong turn while
179
+ doing it.** A graded run that stopped at an `AskUserQuestion` gate with no scripted answer exits 2 with
180
+ a fully-formed `{ok:false, error:{category:"unanswered", …}}` envelope, but the report read only
181
+ `results[0].outDir` and answered *"task turn exited nonzero without a parseable result envelope — it
182
+ crashed before completing a gradeable task"* — under a header hardcoded to
183
+ `INFRASTRUCTURE/PROTOCOL FAILURE (reflection turn)`, though it was the **task** turn that stopped and
184
+ no reflection turn had been attempted. Both halves pointed at the wrong subsystem: a reader following
185
+ the report would audit Docker and the staged agent over what is one scenario flag.
186
+
187
+ The report now prefers the failed turn's own diagnosis, on both turns, carrying its message **and hint
188
+ verbatim** (and only once — `UnansweredError`'s message already contains its hint, so appending it
189
+ unconditionally printed the whole question, its options and its remedy tip twice). No remedy text is
190
+ synthesized on top: there are 36 `UnansweredError` sites and only one is "the skill asked an unscripted
191
+ question", so advice keyed on the category is wrong for nearly all of them.
192
+
193
+ Two new fields carry the classification — `infraFailurePhase` (`task turn` | `reflection turn`) and
194
+ `infraFailureKind` (the harness `ErrCategory`). The header keys on **which** category, not on merely
195
+ having one: `unanswered` / `usage` / `boundary` render as `RUN FAILED (<turn>, <kind>)`; everything
196
+ else renders as `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>)` and it **fails closed** — `internal` is the
197
+ CLI's catch-all for any unexpected throw (Docker down, container start failure, missing staged agent, a
198
+ harness bug), so an unrecognized category is treated as an instrument failure rather than assumed
199
+ ordinary.
200
+
201
+ - **A turn that exited 1 was reported as a bare exit code, so an exhausted quota looked like a crash.**
202
+ A turn that RAN and errored exits `1` with a full result envelope whose top-level `error` is `null` —
203
+ its cause lives in `results[0]`, not in an error object. Reading only the error object left the report
204
+ saying `reflection turn exited with code 1 (expected 0)` and nothing more, for a run whose actual cause
205
+ was an exhausted seven-day quota; it was findable only by opening `events.jsonl` by hand. The failed
206
+ turn's `resultErrorKind` / `errorSource` / `resultSubtype` are now read and rendered — `usage_limit`
207
+ as "the account's quota is exhausted; retry after the reset. This is NOT a harness or skill defect",
208
+ `transport` as a retryable tail-end drop — matching how the run renderer has always shown them.
209
+ `usage_limit` and `transport` join the ordinary/actionable set; `agent` does not, because for
210
+ critique's own protocol turn an agent-level failure *is* the instrument breaking.
211
+
9
212
  ## [2.4.0] — 2026-08-27
10
213
 
11
214
  ### Fidelity
package/README.md CHANGED
@@ -120,7 +120,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
120
120
 
121
121
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
122
122
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
123
- > From a global install (`npm i -g "cowork-harness@^2.4.0"`), point at the package root instead:
123
+ > From a global install (`npm i -g "cowork-harness@^2.5.0"`), point at the package root instead:
124
124
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
125
125
  > (or copy the cassette into your own project and pass that path).
126
126
 
@@ -130,7 +130,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
130
130
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
131
131
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
132
132
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
133
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.4.0"`.
133
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.5.0"`.
134
134
 
135
135
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
136
136
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -155,7 +155,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
155
155
  claude plugin install cowork-harness@cowork-harness
156
156
  ```
157
157
 
158
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.4.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
158
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.5.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
159
159
 
160
160
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
161
161
 
@@ -180,7 +180,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
180
180
  ones that still need a source checkout. (The marketplace
181
181
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
182
182
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
183
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.4.0"` — see
183
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.5.0"` — see
184
184
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
185
185
 
186
186
  ### Prerequisites for anything above `protocol` fidelity
@@ -297,7 +297,7 @@ cowork-harness lint examples/scenarios/*.yaml --strict --min-severity WARN
297
297
  >
298
298
  > | When | Assertions |
299
299
  > |---|---|
300
- > | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
300
+ > | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `reference_read`/`no_observed_reference_access`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
301
301
  > | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
302
302
  > | Only if the cassette carries an `artifacts` manifest | `file_exists`, `artifact_text`, `user_visible_artifact`, `artifact_json`, `computer_links_resolve`, `computer_links_resolve_if_present`, `no_unexpected_files`, `input_unmodified` |
303
303
  > | Always skipped (live-only) | `file_absent`, `egress_*`, `expect_denied`, `no_delete_in_outputs`, `no_delete_in_mounts`, `self_heal_ran`, `transcript_no_host_path`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back` — keep these in a periodic live `run` |
@@ -802,7 +802,7 @@ jobs:
802
802
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
803
803
  ```
804
804
 
805
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.4.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
805
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.5.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
806
806
 
807
807
  The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
808
808