cowork-harness 1.3.0 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +30 -11
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +16 -3
  6. package/CHANGELOG.md +93 -0
  7. package/README.md +9 -8
  8. package/SPEC.md +2 -2
  9. package/baselines/desktop-1.22209.3.json +392 -0
  10. package/dist/cli.js +122 -75
  11. package/dist/critique/armor.js +44 -0
  12. package/dist/critique/command.js +676 -0
  13. package/dist/critique/evaluator.js +388 -0
  14. package/dist/critique/evidence.js +175 -0
  15. package/dist/critique/package-evidence.js +220 -0
  16. package/dist/hostloop/webfetch-dedup.js +50 -0
  17. package/dist/hostloop/workspace-handler.js +42 -4
  18. package/dist/loop-decision.js +30 -0
  19. package/dist/run/cassette.js +30 -14
  20. package/dist/run/chat-result.js +1 -0
  21. package/dist/run/chat.js +10 -1
  22. package/dist/run/diff.js +13 -1
  23. package/dist/run/envelope.js +5 -1
  24. package/dist/run/execute.js +20 -5
  25. package/dist/run/outcome.js +15 -0
  26. package/dist/run/repeat-flags.js +70 -0
  27. package/dist/runtime/hostloop.js +1 -0
  28. package/docs/README.md +1 -0
  29. package/docs/critique.md +117 -0
  30. package/docs/debugging.md +62 -4
  31. package/docs/gotchas.md +3 -2
  32. package/docs/maintenance.md +1 -1
  33. package/docs/scenario.md +3 -2
  34. package/docs/stats.md +43 -1
  35. package/examples/replays/README.md +1 -1
  36. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  37. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  38. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  39. package/llms.txt +2 -1
  40. package/package.json +3 -4
  41. package/schema/run-result.json +160 -54
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.3.0
7
- tracks-harness: cowork-harness 1.3.0 (baseline desktop-1.21459.0)
6
+ version: 1.5.0
7
+ tracks-harness: cowork-harness 1.5.0 (baseline desktop-1.22209.3)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.3.0` (baseline
26
- > `desktop-1.21459.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.5.0` (baseline
26
+ > `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,10 +39,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.3.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.3.0"`. **Pin `@>=1.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
-
44
- What the ≥ 1.3.0 floor gates, by release:
45
-
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.5.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.5.0"`. **Pin `@>=1.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
+
44
+ What the ≥ 1.5.0 floor gates, by release:
45
+
46
+ - **1.5.0:** `critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL) — runs a skill, asks the agent
47
+ what confused it, then grades that self-report against a frozen record of the run: a blinded evaluator
48
+ plus mechanical citation checking. Findings NEVER gate (exit 0); exit 2 means no critique was produced.
49
+ `skill --repeat N` (2-100) with `--min-pass-rate`/`--stop-on-diverge`/`--max-budget-usd`/
50
+ `--allow-budget-stop` brings the variance rollup to the exploratory lane. `result.json` gains
51
+ `outcome` (`errored`/`no_deliverable`/`delivered_with_verdict_fail`/`delivered_clean`). The
52
+ `skill`/`probe-dispatch` lanes now emit `fingerprint.skillHash`/`skillCommit` at all — before 1.5.0
53
+ they emitted NEITHER, so a generation-pairing step over those lanes silently grouped on an absent key.
54
+ `diff`'s exit code now honours its own documented gateable signal (tools/artifacts/meta; a
55
+ transcript-only difference no longer fails it).
46
56
  - **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
47
57
  - **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
48
58
  - **0.22.0:** `computer_links_resolve`.
@@ -55,6 +65,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
55
65
  - **1.1.0:** `analyze-skill` now also flags **interactive-artifact write-backs lost under Cowork** — a relative `fetch`/XHR/`sendBeacon`/`<form method=post>` in an emitted `.html` (or its `.py`/`.js` generator) that silently fails when the artifact is served from Cowork's own origin. `artifact-write-back-lost` gates under `--strict`; `artifact-write-back-suspect` is advisory; an unanalyzable candidate is a could-not-verify exit 3. An optional **`analyze-skill --runtime`** drives the artifact in a headless DOM (needs `jsdom`) to *observe* the lost write-back — enrichment only, never changes the exit code. Plus a `lint` check for a container-only assertion key (`no_scratchpad_leak`/`present_files_called`) used off the `container` tier, and the `doctor --output-format json` envelope frozen as a covered SPEC §12 surface (`schema/doctor.json`).
56
66
  - **1.2.0:** three new assertion keys — `no_lost_write_back: true` (the write-back detector above wired as a per-scenario gate over the run's authored files, live-only) and the regex siblings `tool_result_matches`/`tool_result_not_matches` (case-insensitive per-result, for an error-signature *family* a literal substring can't express). **`microvm` outputs are now observable** — its session tree is snapshotted from the VM into the run dir, so `file_exists`/`artifact_json`/`user_visible_artifact`/`no_unexpected_files`/`input_unmodified`/`no_lost_write_back`/`semantic_matches` all work there, no longer `container`/`hostloop`-only. `status <dir>` also resolves the newest session under a `--run-dir` root; the completion footer prints a `→ result: …/result.json` pointer; the `on_unanswered=fail` error also points to `on_unanswered: llm`. `analyze-skill` hardening: a phantom `<script>` prose block no longer sinks a real verdict, a delete/remove flow claiming success classifies as lost (error) not just suspect, and write-back detection widened (optional-call `?.` spellings, member-spelled/aliased `fetch`/`sendBeacon`, axios instance/config forms).
57
67
  - **1.3.0:** run-identity for the iterate-across-fixes loop — `skill`/`run` take `--label <tag>` (a generation tag surfaced in `result.json` `runLabel`, the run-index row, `inspect`, and `status.json`) alongside an auto-recorded `skillCommit`, layered on the **authoritative** content-exact `fingerprint.skillHash` a harvest step should pair critiques by (`inspect`/index surface a short `skillHash` prefix). `trace --full-results` captures the full input+result of **every** tool call — successful ones too, not just errors — so an external grader can ground a self-critique finding against the call it cites. `verify-run` now **warns** on skillHash drift for answer-less scenarios (was silent unless the scenario declared scripted `answers`). Plus `skill --allow-missing-capability` (the open-ended-run opt-out for a capability FALSE-NEGATIVE on the lean `core` image), a new **warn-severity `ended_with_question`** verdict signal (the agent's final answer contains a question and the run wrote no `outputs/` deliverable — the lenient sibling of the strict `stalled`), and LLM-decider `OTHER:` free-text answers now marked `[via Other free-text]` in gate provenance.
68
+ - **1.4.0:** platform baseline synced to Desktop **1.22209.3** (agent 2.1.215; no prompt/spawn/egress drift), and **`coworkWebFetchDedup` is now enacted** on the hostloop `web_fetch` path — a repeat fetch of the same normalized URL within the TTL (15 min, cap 100) makes **no network request** and returns a marker to re-use the earlier result, with no egress event. Baseline-gated (Desktop ≥ 1.22209.3), keyed under the request + terminal `destination_url`, never caching errors/empty; a hit is assertable via `tool_result_contains: "Already fetched"`.
58
69
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
59
70
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
60
71
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -85,9 +96,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
85
96
  - **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
86
97
  TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
87
98
 
99
+ > **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
100
+ > **npm** install ships `docs/`, `README.md` and `SPEC.md` in the tarball, so try
101
+ > `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
102
+ > install loads a trimmed source-only cache where those pointers genuinely do dangle.
103
+
88
104
  Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
89
105
  lint-skill · analyze-skill · probe-dispatch ·
90
- verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
106
+ verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
91
107
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
92
108
 
93
109
  **Two different `scaffold` tools — don't confuse them.** The native `cowork-harness scaffold <run-id>`
@@ -365,7 +381,10 @@ cassette — has its own recipe:
365
381
  input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
366
382
  pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
367
383
  produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
368
- kept run predates the current skill). See `docs/debugging.md` (repo-only) for the full loop.
384
+ kept run predates the current skill). **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
385
+ pairing step there silently groups on an absent key instead of erroring — check the field is present, or
386
+ require ≥ 1.5.0. See `docs/debugging.md`
387
+ (repo-only) for the full loop.
369
388
 
370
389
  #### Interpreting verdict signals
371
390
 
@@ -461,7 +480,7 @@ Docker, no re-record.
461
480
  | Situation | Symptom | Reach for (in order) |
462
481
  |---|---|---|
463
482
  | **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
464
- | **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
483
+ | **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
465
484
 
466
485
  A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
467
486
  agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.3.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.5.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.3.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.5.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
34
34
  run: |
35
- V=2.1.209 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.215 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.3.0"
60
+ - run: npm i -g "cowork-harness@>=1.5.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.3.0"
200
+ - run: npm i -g "cowork-harness@>=1.5.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.3.0"
229
+ run: npm i -g "cowork-harness@>=1.5.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.3.0` (baseline `desktop-1.20186.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.5.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.3.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.5.0`
4
4
  (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
5
- front-matter (currently 0.32.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ front-matter (currently 1.5.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -170,7 +170,17 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
170
170
 
171
171
  1. **Verify before you trust.** A green run is not a correct run, and a skill's self-reported finding (a
172
172
  self-critique appendix, "I extracted X") is not real until its cited evidence is found in the run's own
173
- output. The harness emits the substrate; the grader is yours (it lives outside the harness):
173
+ output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
174
+ runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
175
+ `--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
176
+ Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`. It rejects `--session-id`/
177
+ `--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
178
+ measurement).
179
+ The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
180
+ docs/debugging.md under "The whole loop, end to end". The harness now SHIPS a grader — `cowork-harness critique <skill-folder> --prompt "<probe>"` runs the
181
+ skill, asks the agent what confused it, and grades that self-report against a frozen record of the run
182
+ (blinded evaluator + mechanical citation checking). See docs/critique.md for cost and limits. If you
183
+ prefer to build your own grader, the substrate is still here:
174
184
  - `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
175
185
  - `cowork-harness trace <run-dir> --output-format json` → the tool-call stream. Add `--full-results` so
176
186
  a **successful** call's full input + result are captured (the default view slices them to ~100/120
@@ -181,7 +191,10 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
181
191
  verdict.
182
192
  2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
183
193
  `result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
184
- content-exact, on every live run, changes on any tracked edit. **Group/pair on it** (`inspect` and the
194
+ content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
195
+ no `skillHash` on the `skill` lane at all, so verify the field is present before pairing on it** (a run that mounts nothing
196
+ records none; the `chat` lane records no fingerprint), changes on any tracked edit. **Group/pair on it**
197
+ (`inspect` and the
185
198
  run-index row surface a short prefix). Add `--label <tag>` for a human-readable generation name
186
199
  (skillHash is the correctness key; the label is ergonomics). `cowork-harness verify-run <run-dir>
187
200
  <scenario.yaml>` is the native staleness guard: it **warns** when a kept run predates the current
package/CHANGELOG.md CHANGED
@@ -6,6 +6,99 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.5.0] — 2026-07-20
10
+
11
+ ### Added
12
+
13
+ - **`cowork-harness critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL).** Runs a skill, asks the
14
+ agent what confused it, then grades that self-report against a frozen record of what actually happened —
15
+ a byte-boundary evidence snapshot taken *before* the reflection turn, a first evaluator pass that is
16
+ structurally blind to the self-report (the text is never put in its prompt, not merely ignored), and
17
+ mechanical citation checking that drops any claim not quoting the evidence verbatim. **A discovery
18
+ instrument, never a gate:** findings of any classification exit 0 (including a task run that itself
19
+ errored — that is a finding about the skill); exit 2 is reserved for a usage error or an instrument
20
+ failure, where no critique was produced. Previously a maintainer-only script that could not run from an installed package at all.
21
+ Costs four model workloads per critique — see [docs/critique.md](./docs/critique.md).
22
+ - **Evidence-package armoring.** The self-report was already fenced, but the evidence package — which
23
+ carries a third-party SKILL.md verbatim into both evaluator prompts — was not, so hostile skill content
24
+ could steer the grader directly. Untrusted content now sits inside per-run nonce markers and only
25
+ nonce-tagged headings outside them count as instructions; a skill cannot pre-author the nonce. Verified
26
+ by a red-team probe across three models: the structural-forgery payload that steered all three now
27
+ matches control. Content that merely *argues* is a documented residual — fencing separates planes, it
28
+ cannot stop persuasion.
29
+ - **`skill --repeat <N>`** (2–100), with `--min-pass-rate` / `--stop-on-diverge` / `--max-budget-usd` /
30
+ `--allow-budget-stop`. The variance rollup already existed but was `run`-only, because the flags were
31
+ parsed inline in the `run` command — the exploratory lane, where an iterate-across-fixes loop actually
32
+ lives, rejected them as unknown. Both lanes now share one parse (`run/repeat-flags.ts`) and one batch
33
+ engine, so rollup shape, JSON envelope, and batch verdict match. `skill --repeat` additionally rejects
34
+ `--session-id`/`--resume` (both pin a single run dir, so iterations would overwrite each other rather
35
+ than produce N independent samples) and `--decider-dir`/`--decider-cmd` (the reproducibility invariant
36
+ `run` already enforced).
37
+ - **`result.json` `outcome`** — a one-field rollup of the `result` × `verdict.pass` × exit-code matrix:
38
+ `errored` / `no_deliverable` / `delivered_with_verdict_fail` / `delivered_clean`. Those three signals
39
+ legitimately disagree (a fail-severity signal flips the verdict while `result` stays `"success"`), and a
40
+ consumer driving a loop had to reconstruct "did this iteration deliver something usable?" from all
41
+ three. A pure function of fields the run already carries, so it cannot disagree with them; the granular
42
+ fields stay authoritative. Absent whenever `verdict` is absent. Note `delivered_*` means "no
43
+ stall/question signal fired", not positive evidence a deliverable exists — check `artifacts`.
44
+ - **Generation-pairing `jq` recipes** in [docs/stats.md](./docs/stats.md) — pass-rate/cost per generation,
45
+ a per-generation verdict-signal histogram, and the before/after of a single fix, grouped on
46
+ `skillHash`/`runLabel` over `index.jsonl`.
47
+
48
+ ### Changed
49
+
50
+ - **`diff`'s exit code now honors its own documented contract.** `--help` has always said "transcript is
51
+ advisory … tools/artifacts/meta are the gateable signal", but `identical` conjoined all four views, so
52
+ two live runs of the SAME skill exited 1 and the signal could not separate "behaviour changed" from
53
+ model stochasticity. `identical` now means the gateable views agree; transcript drift is reported
54
+ separately (`transcriptDiffers` in JSON, rendered in text when you ask for that view) so it stays
55
+ visible. `diff` also carries `skillHash` in the meta view now — a diff across a fix could not previously
56
+ name which two generations it compared.
57
+
58
+ ### Fixed
59
+
60
+ - **`fingerprint.skillHash` and `skillCommit` are now recorded on the `skill` and `probe-dispatch` lanes.**
61
+ Both resolved their skill dirs by re-reading the session *file*, so the lanes that mount via an in-memory
62
+ session — and pass the `"(inline)"` sentinel as the session path — emitted no `skillHash` and a null
63
+ `skillCommit`, even though the mounts were already in scope at the call site. The resolved session object
64
+ is now threaded through instead. Consequences, all previously broken on those lanes: `result.json` carries
65
+ the content-exact generation key, the run-index `skillHash` column populates (so a harvest step can group
66
+ runs by generation), and the run's own "pair critiques by `fingerprint.skillHash`" tip — which is printed
67
+ *only* on the `skill` lane — is no longer advertising a field that lane could not emit. Scoped to the
68
+ sentinel branch: the file-based path is untouched, so recorded cassettes and the
69
+ staleness / `verify-run` recomputes are unaffected. A session that mounts nothing still yields no hash —
70
+ there is nothing to hash.
71
+
72
+ ### Changed
73
+
74
+ - **Reflective skill-critique prompt v2** (`REFLECTION_PROMPT_VERSION` 1 → 2; maintainer instrument, not a
75
+ shipped surface). Adds a sub-agent question (were any dispatched, and was the skill clear about when to
76
+ dispatch, what context to hand them, and what to expect back); replaces the "change ONE thing" cap with
77
+ exhaustive solicitation, since a separate evaluator already triages and drops ungrounded findings, so
78
+ capping at the source loses signal for no quality gain; drops a "fidelity tier" example that is
79
+ cowork-harness vocabulary a third-party skill's agent never encountered; and bounds the pass-2 self-report
80
+ now that the prompt invites longer replies.
81
+
82
+ ## [1.4.0] — 2026-07-19
83
+
84
+ ### Added
85
+
86
+ - **`coworkWebFetchDedup` enacted (hostloop `web_fetch`).** Real Cowork keeps a per-session negative-work
87
+ cache: a repeat `web_fetch` of the same normalized URL within a TTL (default 15 min; cap 100; FIFO
88
+ eviction; a hit does not refresh recency) makes **no network request** and returns a marker telling the
89
+ model to re-use the earlier result. The harness now reproduces this on the host-API (`coworkWebFetchViaApi`)
90
+ path — **baseline-gated** (only when the resolved baseline's `coworkWebFetchDedup` gate is on, i.e. Desktop
91
+ ≥ 1.22209.3), keyed under both the request URL and the terminal `destination_url`, never caching errors /
92
+ empty / non-2xx responses, and emitting **no egress event** on a hit (matching production's zero-network
93
+ dedup). A hit is observable via the marker text (`tool_result_contains: "Already fetched"`).
94
+
95
+ ### Changed
96
+
97
+ - **Platform baseline synced to Desktop 1.22209.3** (agent `2.1.215`). No prompt / spawn-env / egress-allowlist
98
+ drift vs 1.21459.0; the sync captured the new `coworkWebFetchDedup` runtime config (enacted above) plus a
99
+ few new (off) GrowthBook gates. The skill/README/reference version floors and agent-binary pins track the
100
+ new baseline.
101
+
9
102
  ## [1.3.0] — 2026-07-19
10
103
 
11
104
  ### Added
package/README.md CHANGED
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
91
91
 
92
92
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
93
93
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
94
- > From a global install (`npm i -g "cowork-harness@>=1.3.0"`), point at the package root instead:
94
+ > From a global install (`npm i -g "cowork-harness@>=1.5.0"`), point at the package root instead:
95
95
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
96
96
  > (or copy the cassette into your own project and pass that path).
97
97
 
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
101
101
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
102
102
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
103
103
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
104
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.3.0"`.
104
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.5.0"`.
105
105
 
106
106
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
107
107
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
126
126
  claude plugin install cowork-harness@cowork-harness
127
127
  ```
128
128
 
129
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.3.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
129
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.5.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
130
 
131
131
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
132
132
 
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
147
147
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
148
148
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
149
149
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
150
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.3.0"` — see
150
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.5.0"` — see
151
151
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
152
152
 
153
153
  ### Prerequisites for anything above `protocol` fidelity
@@ -392,7 +392,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
392
392
 
393
393
  | Command | What it does | Reach for it when… |
394
394
  |---|---|---|
395
- | `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
395
+ | `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--repeat N` (2-100) runs the same skill+prompt N times and aggregates a variance rollup — "did this finding reproduce, or did it pass once?"; `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
396
396
  | `run <scenario.yaml \| dir/>` | Run authored scenarios with `assert:` + a CI-ready exit code; a decider can answer unscripted gates; `--repeat`/`--matrix` add variance runs / a compatibility matrix (detail below); `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift) | you want a repeatable, **asserted regression test** — to **measure flakiness** instead of trusting one green — or to **test a compatibility matrix** (multiple baselines/models/skill variants) in one run |
397
397
  | `chat <folder> [prompt]` | Interactive multi-turn REPL against a skill (TTY); optional seed prompt is sent as the first turn. `--upload <file>` / `--folder <dir>` (repeatable) attach files/project folders; `--verbose` shows thinking blocks + tool inputs; `--fidelity protocol\|container\|hostloop` (no `microvm`/`cowork` in the REPL); `--allow-host-writes` consents to a writable `hostloop`-fidelity connected folder (same consent as clicking "connect folder" in Desktop); `--raw` skips the control protocol for native `docker run -it` (rejects `--upload`/`--folder`/`--plugin`/`--fidelity`) | debugging a multi-turn flow by hand |
398
398
  | `record` / `replay` | **Record a live run once → replay it token-free, Docker-free thereafter** (key flags below; `replay --explain` prints the evidence behind every passing assert) | **token-free, Docker-free CI** from a once-recorded run |
@@ -415,6 +415,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
415
415
  | `boundary-check [baseline] [--session <file>]` | Prove the **L1 Docker** sandbox enforces Cowork's limitations (sealed FS + default-deny egress; `container`/`hostloop` share this sandbox — `microvm`'s guest firewall is not probed here); `--session` folds a session's `egress.extra_allow` into the probe allowlist | verifying the harness's own fidelity |
416
416
  | `sync` / `list` | Derive/refresh (`sync [--diff] [--allow-empty\|--force]`) & list platform baselines from the Desktop install | after Claude Desktop updates (baselines ship, so it's optional otherwise) |
417
417
  | `diff <a> <b>` | Compare two baselines, two runs, two cassettes, or a run+cassette — kind auto-detected by content (view/normalization flags below). Token-free, no live Desktop/Docker needed | "what changed between two runs/cassettes/baselines?" |
418
+ | `critique <skill-folder>` | **EXPERIMENTAL.** Run a skill, ask the agent what confused it, then grade that self-report against a frozen record of the run — blinded evaluator, mechanical citation checking. Discovery instrument, never a gate (findings always exit 0). Costs four model workloads; see [docs/critique.md](./docs/critique.md) | "what confused the agent about my skill?" |
418
419
  | `doctor [--tier <t>]` | Read-only prerequisite check, per tier (Docker + agent image for `container`/`hostloop`/`cowork`; **Lima** for `microvm`; plus staged agent, token, baseline); prints the exact `docker build` line if the agent image is missing | "can I run the live tiers — what's missing?" before a first live run |
419
420
  | `prune [<runs-dir>] [--keep-last <n>] [--pinned-older-than <N>d\|h\|m]` | Prune accumulated run dirs (keeps the N most recent per scenario; pinned `--session-id` runs are never pruned unless `--pinned-older-than` opts in to reclaiming stale ones by last-activity age); the optional positional overrides the runs root; `--dry-run` | the machine-global runs root has grown and you want space back |
420
421
  | `rehash <dir/>` | Migrate cassette fingerprints to the current format version when the content is provably unchanged (`--dry-run`); no re-record needed | a cassette-format bump flagged committed fixtures as stale |
@@ -659,7 +660,7 @@ jobs:
659
660
  - uses: actions/checkout@v4
660
661
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
661
662
  run: |
662
- V=2.1.209 # match your scenario's pinned baseline's agentVersion
663
+ V=2.1.215 # match your scenario's pinned baseline's agentVersion
663
664
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
664
665
  chmod +x "$RUNNER_TEMP/claude-$V"
665
666
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
@@ -670,7 +671,7 @@ jobs:
670
671
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
671
672
  ```
672
673
 
673
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.3.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
674
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.5.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
674
675
 
675
676
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
676
677
 
@@ -831,6 +832,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
831
832
  ## Status
832
833
 
833
834
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
834
- **`desktop-1.21459.0`**. Release-by-release verification notes (what was re-verified against
835
+ **`desktop-1.22209.3`**. Release-by-release verification notes (what was re-verified against
835
836
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
836
837
  this section used to duplicate lives in the sections above.
package/SPEC.md CHANGED
@@ -309,7 +309,7 @@ Payload sits under an **inner** `response`. Missing the nesting ⇒ `ZodError: e
309
309
  > Channels 2-3 = CLI `--mcp-config` / `.mcp.json` servers (honored in plain cowork mode; dropped only in
310
310
  > hermetic mode). Details per channel below.
311
311
 
312
- 1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
312
+ 1. **SDK servers over the control protocol** — declared via `sdkMcpServers` in `initialize`; tool calls tunnel as `mcp_message`. This is how the **desktop host** bridges its own servers (incl. `claude_desktop_config.json` `mcpServers`, spawned host-side with full host env) and how the harness delivers the workspace shell. The workspace handler (`src/hostloop/workspace-handler.ts`) implements `initialize`/`tools/list`/`tools/call`; `bash`→`docker exec -w <mntRoot> <container> sh -c <cmd>` (container-egress-gated). **CB-8:** `makeWorkspaceHandler` accepts an `onInfraError?: (message: string) => void` callback at parameter position 6 (after `onEgress`); on infrastructure errors (ETIMEDOUT / killed / no code+stdout+stderr) the handler calls `onInfraError?.(e.message)`, returns a textResult with `"[infrastructure error: …]"`, and `spawnHostLoop` wires `onInfraError` to append `{type:"infra_error", ts, message}` to `events.jsonl`. **`web_fetch` is NOT container-egress-gated:** real Cowork routes it through the host API (gate `1978029737` `coworkWebFetchViaApi:true` → `POST /api/organizations/<org>/cowork/web_fetch`), gated by a **separate web-fetch hostname allowlist** (`getWebFetchAllowedUrls`, `*`=unrestricted) + a **URL-provenance** rule (URL must have appeared in a prior message/result). The harness mirrors this with the **two-path model** (`src/hostloop/workspace-handler.ts`, binary-verified `G1t`/`U1t`): **Path A** (provenance engaged — `coworkWebFetchViaApi` on) gates on the **exact-URL provenance set** ONLY (seeded from user-turn + tool-result URLs; `src/hostloop/provenance.ts`), with **no** hostname allowlist — but still an `http(s)`-scheme + private-address **SSRF backstop re-checked on every redirect hop** (a manual redirect loop, not `curl -L`) — a miss raises a per-domain approval (`webfetch:<domain>` permission with options `Allow once | Allow all for website | Deny`) routed through the Decider; "Allow all for website" approves the host for the rest of the run (`Run.approvedDomains`, per-run/ephemeral). On Path A, **AFTER** the provenance/approval gate, **`coworkWebFetchDedup`** (gate `1978029737`, on for baselines ≥ `1.22209.3`) applies a per-session **negative-work cache** (`src/hostloop/webfetch-dedup.ts`, binary-verified): a repeat `web_fetch` of the same **normalized** URL (`normalizeUrl`) within the TTL (baseline-sourced, default `900000` ms; cap `100`, FIFO eviction; a hit does NOT refresh recency) returns a marker telling the model to re-use the earlier result, with **no network request and no egress event** — matching production's zero-network dedup. Only successful (HTTP-2xx, trimmed-nonempty) fetches are cached, keyed under both the request URL and the terminal `destination_url`; errors/empty/non-2xx are never cached. Dedup is baseline-gated (an older baseline lacking the flag never dedups) and Path-A-only. **Path B** (gate off) is a direct host fetch gated by the egress domain list via the same `wen()`/`compile()` matcher container egress uses, with `redirect:"manual"` re-checking `U1t` (scheme + private-address SSRF + allowlist) on **every** redirect hop. web_fetch is thus **decoupled from `plan.egressAllow`** on Path A (egress applies to `bash`/Path B only). An unanswered cold miss is fail-closed; scenarios answer via `--answer "webfetch:<domain>=allow"` (with `grant`), `web_fetch.approved_domains`, or an LLM/external terminal. `bash` stays container-egress-sandboxed.
313
313
  2. **CLI-spawned `--mcp-config` / `.mcp.json` servers — HONORED in plain cowork mode** (NOT ignored). **Verified:** a valid `--mcp-config` populates `mcp_servers` (`[{name,status:"pending"|"connected"}]`); these run in-sandbox with the env-allowlist `CLAUDE_CODE_MCP_ALLOWLIST_ENV` (`RW8`/`oG8`/`LU5` = {HOME, LOGNAME, PATH, SHELL, TERM, USER}). The harness MAY use this as a convenience injection path.
314
314
  3. **The drop is SAFE/HERMETIC-mode-gated, not cowork-gated.** `--mcp-config` is filtered to SDK-only (`ap5()`) **only when** safe mode (`I5()`) or `xB8()` is true, and `xB8()` requires **both** `CLAUDE_CODE_REMOTE` **and** `CLAUDE_CODE_REMOTE_HERMETIC_MODE`. **Verified:** with both set, `mcp_servers:[]`; without them (plain `SESSION_KIND=bg`), the config is honored. The earlier "cowork ignores `--mcp-config`" was a hermetic-session observation over-generalized.
315
315
 
@@ -603,7 +603,7 @@ unused — reserving it now keeps a later addition additive rather than a renumb
603
603
  `0`/`1`/`2`/`3` space. Exit-code space is **per-command**, not global (`status` uses `0`/`1`/`2`/`3`
604
604
  with its own meanings); this reservation applies only to the `run`/`skill` family.
605
605
 
606
- **Per-command exceptions:** `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
606
+ **Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
607
607
  `2` on a **whole-cassette operational failure** — anything `readCassette` rejects (unreadable, invalid
608
608
  shape, unsupported version, unrecognized assertion key) or any per-file throw, plus the batch loop's
609
609
  own source-resolution failures (`--assert-from`/`--reassert` drift, scenario-parse errors, `--write`