cowork-harness 2.4.0 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +8 -8
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
  3. package/.claude/skills/cowork-harness/references/critique.md +72 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +6 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +20 -5
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +5 -2
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +67 -2
  9. package/CHANGELOG.md +322 -0
  10. package/DESIGN.md +2 -2
  11. package/README.md +17 -17
  12. package/SPEC.md +14 -2
  13. package/baselines/desktop-1.40609.0.json +878 -0
  14. package/baselines/provisioning/rootfs-provisioning.json +32 -40
  15. package/dist/assert.js +42 -1
  16. package/dist/cli.js +16 -1
  17. package/dist/critique/command.js +202 -22
  18. package/dist/critique/evaluator.js +9 -3
  19. package/dist/critique/package-evidence.js +39 -5
  20. package/dist/run/cassette.js +27 -4
  21. package/dist/run/chat-result.js +2 -1
  22. package/dist/run/chat.js +15 -1
  23. package/dist/run/execute.js +46 -7
  24. package/dist/run/hook-events.js +51 -0
  25. package/dist/run/probe-dispatch.js +8 -2
  26. package/dist/run/provenance.js +6 -9
  27. package/dist/run/run.js +197 -8
  28. package/dist/run/skill-flag-surface.js +11 -0
  29. package/dist/run/tier-vacuous-tools.js +69 -0
  30. package/dist/run/tool-name-canonicalization.js +70 -0
  31. package/dist/run/verdict.js +17 -8
  32. package/dist/runtime/argv.js +16 -1
  33. package/dist/runtime/lima.js +37 -3
  34. package/dist/runtime/protocol.js +85 -13
  35. package/dist/scan.js +1 -0
  36. package/dist/sync/cowork-sync.js +39 -2
  37. package/dist/types.js +47 -6
  38. package/docs/cassette.md +4 -2
  39. package/docs/critique.md +52 -0
  40. package/docs/debugging.md +3 -1
  41. package/docs/discovery.md +1 -1
  42. package/docs/fidelity-gaps.md +78 -6
  43. package/docs/maintenance.md +1 -1
  44. package/docs/scenario.md +14 -8
  45. package/docs/subagents.md +5 -3
  46. package/examples/replays/README.md +1 -1
  47. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  48. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  49. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  50. package/examples/sessions/l0-plugin-delivery.yaml +6 -0
  51. package/package.json +1 -1
  52. package/python/test_scenario_lint.py +1 -1
  53. package/schema/critique-report.json +41 -18
  54. package/schema/run-result.json +49 -3
  55. package/schema/scenario.schema.json +19 -5
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.4.0
7
- tracks-harness: cowork-harness 2.4.0 (baseline desktop-1.37937.1)
6
+ version: 3.0.0
7
+ tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.4.0` (baseline
29
- > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
29
+ > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.4.0"`. **Pin `@^2.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
49
49
  — upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
50
50
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
51
- - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
51
+ - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
52
52
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
53
53
  - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
54
54
 
@@ -147,7 +147,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
147
147
  | Tier | What it gives you | Use when |
148
148
  |---|---|---|
149
149
  | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
150
- | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm`/`protocol` never offer it, so the same `tool_not_called` is vacuous there — moving a scenario between tiers can silently void a web-fetch assertion | Most functional + boundary tests. |
150
+ | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
151
151
  | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
152
152
  | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
153
153
 
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
679
679
  ### Place assertions in the right CI lane
680
680
 
681
681
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
682
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
682
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
683
683
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
684
684
  the four-stage pipeline.
685
685
 
@@ -1,23 +1,23 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@v2
9
+ - uses: yaniv-golan/cowork-harness@v3
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
- version: "^2" # hold the major; see below
13
+ version: "^3" # hold the major; see below
14
14
  ```
15
15
 
16
- **These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
16
+ **These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.4.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.246 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.247 # match your scenario's pinned baseline's agentVersion
40
40
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
41
41
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
42
42
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -47,11 +47,11 @@ jobs:
47
47
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
48
48
  # Background on the provenance chain: the "Agent-binary provenance" section of
49
49
  # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
50
- - uses: yaniv-golan/cowork-harness@v2
50
+ - uses: yaniv-golan/cowork-harness@v3
51
51
  with:
52
52
  command: run
53
53
  path: scenarios/
54
- version: "^2"
54
+ version: "^3"
55
55
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
56
56
  ```
57
57
 
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.4.0"
70
+ - run: npm i -g "cowork-harness@^3.0.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
119
119
  you have a reason:
120
120
 
121
121
  ```yaml
122
- - uses: yaniv-golan/cowork-harness@v2
122
+ - uses: yaniv-golan/cowork-harness@v3
123
123
  with:
124
124
  command: lint
125
125
  path: scenarios/
126
- version: "^2" # holds the major
127
- extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
126
+ version: "^3" # holds the major
127
+ extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
128
128
  ```
129
129
 
130
130
  **If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
131
131
  floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
132
132
  recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
133
- major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
133
+ major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
134
134
  stopping at the major boundary. An exact
135
135
  pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
136
136
  moment a recipe adopts a newer flag.
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
147
147
  Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
148
148
  `transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
149
149
  `max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
150
- `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
150
+ `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
151
151
  `allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
152
152
  `question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
153
153
  (`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
322
322
 
323
323
  ## GitHub Actions sketch
324
324
 
325
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
325
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
326
326
  in one step (see the top of this doc) — reach for this form when you need independent per-command
327
327
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
328
328
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.4.0"
345
+ - run: npm i -g "cowork-harness@^3.0.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.4.0"
374
+ run: npm i -g "cowork-harness@^3.0.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
34
34
  **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
35
35
  UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
36
36
 
37
+ ## An exit-2 report — which turn failed, and whether it was really infrastructure
38
+
39
+ Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
40
+ carry that; read all three before touching anything.
41
+
42
+ | Field | Meaning |
43
+ |---|---|
44
+ | `infraFailure` | the reason |
45
+ | `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
46
+ | `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
47
+ | `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
48
+
49
+ **Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
50
+ unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
51
+ agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
52
+ and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
53
+ **`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
54
+ turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
55
+ unrecognized kind renders as infrastructure):
56
+
57
+ - `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
58
+ - `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
59
+
60
+ **A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
61
+ full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
62
+ a broken instrument.
63
+
64
+ **Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
65
+ verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
66
+ "the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
67
+ `--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
68
+ reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
69
+ is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
70
+ *conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
71
+
72
+ For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
73
+ gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
74
+ (or the scenario's `on_unanswered: llm`).
75
+
76
+ ## Which model was graded — `gradedModels`
77
+
78
+ **The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
79
+ session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
80
+ which may not be the model you are otherwise working under, and nothing about the run announces it.
81
+
82
+ `gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
83
+ is the only record of which model produced the behaviour being graded — distinct from the evaluator's
84
+ resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
85
+ separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
86
+ `--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
87
+ read `gradedModels` back to confirm it took.
88
+
89
+ It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
90
+ messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
91
+ (crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
92
+ can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
93
+
37
94
  ## The report's item shape — no `title`, no `summary`
38
95
 
39
96
  Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
80
137
  specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
81
138
  `not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
82
139
 
140
+ ## Read `referencesAccessed`, not `referencesRead`
141
+
142
+ `referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
143
+ `Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
144
+ `referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
145
+ (`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
146
+
147
+ Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
148
+ entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
149
+ then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
150
+ weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
151
+ nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
152
+ "none".
153
+
83
154
  ## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
84
155
 
85
156
  `result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -20,6 +20,11 @@ Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.379
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
23
+ - A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
24
+ (`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
25
+ hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
26
+ Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
27
+ does not fall back to the default.
23
28
  - **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
24
29
  `skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
25
30
  positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.4.0`
4
- (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
4
+ (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -98,6 +98,17 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
98
98
  # around hostloop's native file tools, that combination gives the
99
99
  # agent genuine, software-checked-only host filesystem access.
100
100
  # Read-only folders and folder-less runs need no opt-in.
101
+
102
+ allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
103
+ # declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
104
+ # --plugin-dir, so the CLI executes those hooks as NATIVE HOST
105
+ # processes under your account, with no container sandbox. A plugin
106
+ # that declares no hooks needs no opt-in, and a misplaced root-level
107
+ # `hooks.json` cannot execute so it does not trigger the gate.
108
+ # Use `--fidelity container` to run them sandboxed instead.
109
+ # NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
110
+ # OLDER CLI does not default it — it hard-errors
111
+ # `Unrecognized key: "allow_host_hooks"` and exits 2.
101
112
  ```
102
113
 
103
114
  Relative paths resolve from the file's own directory, so a scenario + session + referenced files
@@ -305,6 +316,10 @@ same set live from the schema.
305
316
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
306
317
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
307
318
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
319
+ | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
320
+ | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
321
+ | `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
322
+ | `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
308
323
  | `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
309
324
  | `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
310
325
  | `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
@@ -358,7 +373,7 @@ same set live from the schema.
358
373
  | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
359
374
  | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
360
375
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
361
- | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
376
+ | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
362
377
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
363
378
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
364
379
  | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
@@ -400,7 +415,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
400
415
  | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
401
416
  | `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
402
417
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
403
- | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
418
+ | `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
404
419
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
405
420
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
406
421
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
@@ -439,7 +454,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
439
454
  `skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
440
455
  `max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
441
456
  (`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
442
- modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
457
+ modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
443
458
  `allow_stall` are also kept on replay, evaluated as no-op passes.
444
459
 
445
460
  **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -3,7 +3,7 @@
3
3
  "keys": [
4
4
  "all_tasks_completed",
5
5
  "allow_delete_in",
6
- "allow_l0_plugin_divergence",
6
+ "allow_l0_host_config_contamination",
7
7
  "allow_missing_capability",
8
8
  "allow_outputs_delete",
9
9
  "allow_permissive_auto_allow",
@@ -35,6 +35,7 @@
35
35
  "no_hook_blocked",
36
36
  "no_lost_write_back",
37
37
  "no_mcp_error",
38
+ "no_observed_reference_access",
38
39
  "no_path_denied",
39
40
  "no_scratchpad_leak",
40
41
  "no_skill_triggered",
@@ -46,6 +47,7 @@
46
47
  "question_context",
47
48
  "question_options",
48
49
  "questions_count_max",
50
+ "reference_read",
49
51
  "replay_protocol_fidelity",
50
52
  "result",
51
53
  "self_heal_ran",
@@ -81,6 +83,7 @@
81
83
  "vm_path_denied"
82
84
  ],
83
85
  "topLevelKeys": [
86
+ "allow_host_hooks",
84
87
  "allow_host_writes",
85
88
  "answers",
86
89
  "assert",
@@ -99,7 +102,7 @@
99
102
  ],
100
103
  "verdictModifierKeys": [
101
104
  "allow_delete_in",
102
- "allow_l0_plugin_divergence",
105
+ "allow_l0_host_config_contamination",
103
106
  "allow_missing_capability",
104
107
  "allow_outputs_delete",
105
108
  "allow_permissive_auto_allow",
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
80
80
  "tool_result_not_matches",
81
81
  "tool_called",
82
82
  "tool_not_called",
83
+ # Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
84
+ # tool inputs), so they are content keys, not live-only.
85
+ "reference_read",
86
+ "no_observed_reference_access",
83
87
  "subagent_tool_used",
84
88
  "subagent_tool_absent",
85
89
  "subagent_dispatched",
@@ -169,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
169
173
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
170
174
  VERDICT_MODIFIER_KEYS = {
171
175
  "allow_permissive_auto_allow",
172
- "allow_l0_plugin_divergence",
176
+ "allow_l0_host_config_contamination",
173
177
  "allow_missing_capability",
174
178
  "allow_stall",
175
179
  "allow_undelivered_deliverables",
@@ -258,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
258
262
  "assert",
259
263
  "skills", # opt-in skill-staleness hash scope
260
264
  "requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
261
- "allow_host_writes", # hostloop native-split: consent for a writable connected folder (pre-run gate)
265
+ "allow_host_writes",
266
+ "allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
262
267
  }
263
268
 
264
269
 
@@ -294,6 +299,8 @@ REGEX_KEYS = {
294
299
  "hook_blocked",
295
300
  "tool_result_matches",
296
301
  "tool_result_not_matches",
302
+ "reference_read",
303
+ "no_observed_reference_access",
297
304
  }
298
305
  VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
299
306
  VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
@@ -583,6 +590,44 @@ def lint_doc(doc, path, raw_lines):
583
590
  # warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
584
591
  # declare the combination incompatible, so authoring it is a bug even if a tool-free run could
585
592
  # accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
593
+ # `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
594
+ # vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
595
+ # (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
596
+ # and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
597
+ # passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
598
+ # must never be flagged here. The harness refuses these at load; this catches them before any run.
599
+ _TIER_VACUOUS = {
600
+ "hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
601
+ "container": {"mcp__workspace__bash": "Bash"},
602
+ "microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
603
+ # `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
604
+ # host CLI registry — machine-dependent and about a different product.
605
+ }
606
+ # BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
607
+ # USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
608
+ for _key in ("tool_not_called", "subagent_tool_absent"):
609
+ for _v in _assert_values(items, _key):
610
+ if not isinstance(_v, str) or "*" in _v or "?" in _v:
611
+ continue # a glob is not a literal claim about one tool
612
+ _repl = _TIER_VACUOUS.get(fidelity, {})
613
+ if _v not in _repl:
614
+ continue
615
+ _instead = _repl[_v]
616
+ findings.append(
617
+ Finding(
618
+ "WARN",
619
+ "tool-not-called-tier-vacuous",
620
+ f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
621
+ f"`{_v}` at all, so this can never be violated and verifies nothing.",
622
+ (
623
+ f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
624
+ if _instead
625
+ else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
626
+ ),
627
+ path,
628
+ )
629
+ )
630
+
586
631
  # linter stays offline — the message carries the gate fact instead of reading a baseline).
587
632
  if "transcript_no_host_path" in assert_keys:
588
633
  if fidelity in ("hostloop", "protocol"):
@@ -1016,6 +1061,26 @@ def lint_doc(doc, path, raw_lines):
1016
1061
  )
1017
1062
  )
1018
1063
 
1064
+ # Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
1065
+ # with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
1066
+ # that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
1067
+ # this only fires on the case that is unambiguously self-defeating.
1068
+ read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
1069
+ unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
1070
+ both_refs = sorted(read_pats & unread_pats)
1071
+ if both_refs:
1072
+ findings.append(
1073
+ Finding(
1074
+ "ERROR",
1075
+ "reference-access-contradiction",
1076
+ f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
1077
+ "(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
1078
+ "Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
1079
+ "another is not, give the two keys different patterns.",
1080
+ path,
1081
+ )
1082
+ )
1083
+
1019
1084
  # W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
1020
1085
  findings.extend(_lint_regex_quoting(path, raw_lines))
1021
1086