cowork-harness 2.3.0 → 2.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +16 -9
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/critique.md +72 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +5 -1
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +2 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +88 -0
  9. package/CHANGELOG.md +333 -0
  10. package/README.md +6 -6
  11. package/SPEC.md +12 -0
  12. package/dist/assert.js +70 -4
  13. package/dist/cli.js +8 -0
  14. package/dist/critique/command.js +202 -22
  15. package/dist/critique/evaluator.js +9 -3
  16. package/dist/critique/package-evidence.js +39 -5
  17. package/dist/hostloop/workspace-handler.js +12 -2
  18. package/dist/run/cassette.js +72 -2
  19. package/dist/run/chat-result.js +1 -0
  20. package/dist/run/execute.js +107 -9
  21. package/dist/run/probe-dispatch.js +8 -2
  22. package/dist/run/provenance.js +6 -9
  23. package/dist/run/run.js +197 -8
  24. package/dist/run/tier-vacuous-tools.js +69 -0
  25. package/dist/run/tool-name-canonicalization.js +70 -0
  26. package/dist/runtime/container.js +53 -2
  27. package/dist/runtime/hostloop.js +52 -8
  28. package/dist/sync/cowork-sync.js +26 -8
  29. package/dist/types.js +37 -4
  30. package/docs/cassette.md +2 -0
  31. package/docs/critique.md +52 -0
  32. package/docs/debugging.md +3 -1
  33. package/docs/fidelity-gaps.md +128 -5
  34. package/docs/scenario.md +61 -12
  35. package/docs/subagents.md +5 -3
  36. package/examples/replays/README.md +1 -1
  37. package/examples/replays/example-pdf-skill.cassette.json +161 -112
  38. package/examples/scenarios/example-pdf-skill.yaml +8 -1
  39. package/package.json +1 -1
  40. package/python/test_scenario_lint.py +39 -0
  41. package/schema/critique-report.json +41 -18
  42. package/schema/run-result.json +47 -1
  43. package/schema/scenario.schema.json +14 -4
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.3.0
7
- tracks-harness: cowork-harness 2.3.0 (baseline desktop-1.37937.1)
6
+ version: 2.5.0
7
+ tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.3.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
29
29
  > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.3.0"`. **Pin `@^2.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -147,9 +147,9 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
147
147
  | Tier | What it gives you | Use when |
148
148
  |---|---|---|
149
149
  | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
150
- | `container` | Real sandbox + real default-deny egress (**default**) | Most functional + boundary tests. |
151
- | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity | Testing untrusted code escape, not network behavior. |
152
- | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with native Bash/WebFetch disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
150
+ | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
151
+ | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
152
+ | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
153
153
 
154
154
  Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
155
155
  (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
@@ -550,8 +550,15 @@ Recognize these before "fixing" a non-bug:
550
550
  only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
551
551
  **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
552
552
  scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
553
- **The fix is lane-dependent.** On `lane: local`, write deliverables under `outputs/` or a connected
554
- folder, or deliver them explicitly. **On `lane: remote`, moving a file under `outputs/` does NOT help** —
553
+ **The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
554
+ can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
555
+ (what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
556
+ `outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
557
+ `fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
558
+ bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
559
+ a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
560
+ decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
561
+ [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
555
562
  nothing is delivered by location there, so only an explicit delivery counts. Assert
556
563
  **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
557
564
  downloaded inputs) rather than a delivery gap.
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.3.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.3.0"
70
+ - run: npm i -g "cowork-harness@^2.5.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.3.0"
345
+ - run: npm i -g "cowork-harness@^2.5.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.3.0"
374
+ run: npm i -g "cowork-harness@^2.5.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
34
34
  **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
35
35
  UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
36
36
 
37
+ ## An exit-2 report — which turn failed, and whether it was really infrastructure
38
+
39
+ Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
40
+ carry that; read all three before touching anything.
41
+
42
+ | Field | Meaning |
43
+ |---|---|
44
+ | `infraFailure` | the reason |
45
+ | `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
46
+ | `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
47
+ | `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
48
+
49
+ **Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
50
+ unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
51
+ agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
52
+ and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
53
+ **`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
54
+ turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
55
+ unrecognized kind renders as infrastructure):
56
+
57
+ - `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
58
+ - `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
59
+
60
+ **A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
61
+ full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
62
+ a broken instrument.
63
+
64
+ **Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
65
+ verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
66
+ "the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
67
+ `--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
68
+ reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
69
+ is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
70
+ *conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
71
+
72
+ For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
73
+ gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
74
+ (or the scenario's `on_unanswered: llm`).
75
+
76
+ ## Which model was graded — `gradedModels`
77
+
78
+ **The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
79
+ session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
80
+ which may not be the model you are otherwise working under, and nothing about the run announces it.
81
+
82
+ `gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
83
+ is the only record of which model produced the behaviour being graded — distinct from the evaluator's
84
+ resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
85
+ separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
86
+ `--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
87
+ read `gradedModels` back to confirm it took.
88
+
89
+ It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
90
+ messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
91
+ (crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
92
+ can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
93
+
37
94
  ## The report's item shape — no `title`, no `summary`
38
95
 
39
96
  Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
80
137
  specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
81
138
  `not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
82
139
 
140
+ ## Read `referencesAccessed`, not `referencesRead`
141
+
142
+ `referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
143
+ `Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
144
+ `referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
145
+ (`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
146
+
147
+ Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
148
+ entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
149
+ then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
150
+ weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
151
+ nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
152
+ "none".
153
+
83
154
  ## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
84
155
 
85
156
  `result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.3.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
4
4
  (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -305,6 +305,10 @@ same set live from the schema.
305
305
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
306
306
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
307
307
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
308
+ | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
309
+ | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
310
+ | `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
311
+ | `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
308
312
  | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
309
313
  | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
310
314
  | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -35,6 +35,7 @@
35
35
  "no_hook_blocked",
36
36
  "no_lost_write_back",
37
37
  "no_mcp_error",
38
+ "no_observed_reference_access",
38
39
  "no_path_denied",
39
40
  "no_scratchpad_leak",
40
41
  "no_skill_triggered",
@@ -46,6 +47,7 @@
46
47
  "question_context",
47
48
  "question_options",
48
49
  "questions_count_max",
50
+ "reference_read",
49
51
  "replay_protocol_fidelity",
50
52
  "result",
51
53
  "self_heal_ran",
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
80
80
  "tool_result_not_matches",
81
81
  "tool_called",
82
82
  "tool_not_called",
83
+ # Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
84
+ # tool inputs), so they are content keys, not live-only.
85
+ "reference_read",
86
+ "no_observed_reference_access",
83
87
  "subagent_tool_used",
84
88
  "subagent_tool_absent",
85
89
  "subagent_dispatched",
@@ -294,6 +298,8 @@ REGEX_KEYS = {
294
298
  "hook_blocked",
295
299
  "tool_result_matches",
296
300
  "tool_result_not_matches",
301
+ "reference_read",
302
+ "no_observed_reference_access",
297
303
  }
298
304
  VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
299
305
  VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
@@ -488,6 +494,30 @@ def lint_doc(doc, path, raw_lines):
488
494
 
489
495
  fidelity = (doc.get("fidelity") or "container")
490
496
  lane = (doc.get("lane") or "local")
497
+
498
+ # W: no `fidelity:` — the default models the WRONG LANE.
499
+ # `container` (the schema default) models VM-loop; production runs HOST-LOOP, gate 1143815894 is
500
+ # force-ON in every shipped baseline. So an omitted key measures the scenario against a lane real
501
+ # users are not on: the file tools resolve a bare relative path differently, the shell starts
502
+ # somewhere else, and the offered tool set differs (measured 2026-08-27).
503
+ # Read the KEY, not the resolved value: `fidelity: container` is a deliberate choice and must not warn.
504
+ # DEPRECATION — `fidelity:` becomes REQUIRED in the next major; this is the warning window.
505
+ if "fidelity" not in doc:
506
+ findings.append(
507
+ Finding(
508
+ "WARN",
509
+ "fidelity-defaulted",
510
+ "no `fidelity:` — defaulting to `container`, which models the VM-LOOP lane. Production "
511
+ "runs HOST-LOOP by default (gate 1143815894), so this scenario is likely measured "
512
+ "against a lane your users are not on.",
513
+ "Name a tier: `fidelity: hostloop` to match production, `fidelity: cowork` to auto-pick "
514
+ "the way Cowork does, or `fidelity: container` to keep today's behaviour deliberately. "
515
+ "Switching tiers can COST you assertions: `no_scratchpad_leak` is container-only (an "
516
+ "error elsewhere) and `transcript_no_host_path` fails by design at hostloop/protocol. "
517
+ "The default is being removed — `fidelity:` becomes REQUIRED in the next major.",
518
+ path,
519
+ )
520
+ )
491
521
  items = _assert_items(doc)
492
522
  assert_keys = _all_assert_keys(items)
493
523
  has_expect_denied = bool(doc.get("expect_denied"))
@@ -559,6 +589,44 @@ def lint_doc(doc, path, raw_lines):
559
589
  # warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
560
590
  # declare the combination incompatible, so authoring it is a bug even if a tool-free run could
561
591
  # accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
592
+ # `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
593
+ # vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
594
+ # (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
595
+ # and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
596
+ # passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
597
+ # must never be flagged here. The harness refuses these at load; this catches them before any run.
598
+ _TIER_VACUOUS = {
599
+ "hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
600
+ "container": {"mcp__workspace__bash": "Bash"},
601
+ "microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
602
+ # `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
603
+ # host CLI registry — machine-dependent and about a different product.
604
+ }
605
+ # BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
606
+ # USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
607
+ for _key in ("tool_not_called", "subagent_tool_absent"):
608
+ for _v in _assert_values(items, _key):
609
+ if not isinstance(_v, str) or "*" in _v or "?" in _v:
610
+ continue # a glob is not a literal claim about one tool
611
+ _repl = _TIER_VACUOUS.get(fidelity, {})
612
+ if _v not in _repl:
613
+ continue
614
+ _instead = _repl[_v]
615
+ findings.append(
616
+ Finding(
617
+ "WARN",
618
+ "tool-not-called-tier-vacuous",
619
+ f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
620
+ f"`{_v}` at all, so this can never be violated and verifies nothing.",
621
+ (
622
+ f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
623
+ if _instead
624
+ else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
625
+ ),
626
+ path,
627
+ )
628
+ )
629
+
562
630
  # linter stays offline — the message carries the gate fact instead of reading a baseline).
563
631
  if "transcript_no_host_path" in assert_keys:
564
632
  if fidelity in ("hostloop", "protocol"):
@@ -992,6 +1060,26 @@ def lint_doc(doc, path, raw_lines):
992
1060
  )
993
1061
  )
994
1062
 
1063
+ # Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
1064
+ # with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
1065
+ # that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
1066
+ # this only fires on the case that is unambiguously self-defeating.
1067
+ read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
1068
+ unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
1069
+ both_refs = sorted(read_pats & unread_pats)
1070
+ if both_refs:
1071
+ findings.append(
1072
+ Finding(
1073
+ "ERROR",
1074
+ "reference-access-contradiction",
1075
+ f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
1076
+ "(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
1077
+ "Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
1078
+ "another is not, give the two keys different patterns.",
1079
+ path,
1080
+ )
1081
+ )
1082
+
995
1083
  # W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
996
1084
  findings.extend(_lint_regex_quoting(path, raw_lines))
997
1085