cowork-harness 2.5.0 → 3.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +7 -7
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +61 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +17 -5
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -2
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +3 -2
  9. package/AGENTS.md +1 -1
  10. package/CHANGELOG.md +218 -0
  11. package/CONTRIBUTING.md +35 -0
  12. package/DESIGN.md +2 -2
  13. package/README.md +87 -704
  14. package/SPEC.md +2 -2
  15. package/baselines/desktop-1.40609.0.json +878 -0
  16. package/baselines/provisioning/rootfs-provisioning.json +32 -40
  17. package/dist/cli.js +8 -1
  18. package/dist/run/cassette.js +4 -3
  19. package/dist/run/chat-result.js +1 -1
  20. package/dist/run/chat.js +15 -1
  21. package/dist/run/execute.js +15 -6
  22. package/dist/run/hook-events.js +84 -0
  23. package/dist/run/skill-flag-surface.js +11 -0
  24. package/dist/run/verdict.js +17 -8
  25. package/dist/runtime/argv.js +16 -1
  26. package/dist/runtime/lima.js +37 -3
  27. package/dist/runtime/protocol.js +85 -13
  28. package/dist/scan.js +1 -0
  29. package/dist/sync/cowork-sync.js +39 -2
  30. package/dist/types.js +11 -3
  31. package/docs/README.md +9 -6
  32. package/docs/boundary.md +7 -0
  33. package/docs/cassette.md +3 -3
  34. package/docs/ci.md +114 -0
  35. package/docs/cli.md +532 -0
  36. package/docs/companion-skill.md +58 -0
  37. package/docs/debugging.md +3 -3
  38. package/docs/discovery.md +1 -1
  39. package/docs/fidelity-gaps.md +88 -7
  40. package/docs/gotchas.md +27 -2
  41. package/docs/maintenance.md +29 -1
  42. package/docs/scenario.md +6 -3
  43. package/docs/session.md +1 -1
  44. package/docs/stats.md +1 -1
  45. package/examples/README.md +3 -3
  46. package/examples/replays/README.md +1 -1
  47. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  48. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  49. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  50. package/examples/sessions/l0-plugin-delivery.yaml +6 -0
  51. package/llms.txt +3 -0
  52. package/package.json +1 -1
  53. package/python/test_scenario_lint.py +1 -1
  54. package/schema/run-result.json +2 -2
  55. package/schema/scenario.schema.json +6 -2
  56. package/scripts/bump-version.ts +12 -1
  57. package/scripts/check-versions.ts +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.5.0
7
- tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
6
+ version: 3.0.1
7
+ tracks-harness: cowork-harness 3.0.1 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
29
- > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.1` (baseline
29
+ > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.1"`. **Pin `@^3.0.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
49
49
  — upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
50
50
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
51
- - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
51
+ - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
52
52
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
53
53
  - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
54
54
 
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
679
679
  ### Place assertions in the right CI lane
680
680
 
681
681
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
682
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
682
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
683
683
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
684
684
  the four-stage pipeline.
685
685
 
@@ -1,23 +1,23 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@v2
9
+ - uses: yaniv-golan/cowork-harness@v3
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
- version: "^2" # hold the major; see below
13
+ version: "^3" # hold the major; see below
14
14
  ```
15
15
 
16
- **These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
16
+ **These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.0.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.246 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.247 # match your scenario's pinned baseline's agentVersion
40
40
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
41
41
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
42
42
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -47,11 +47,11 @@ jobs:
47
47
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
48
48
  # Background on the provenance chain: the "Agent-binary provenance" section of
49
49
  # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
50
- - uses: yaniv-golan/cowork-harness@v2
50
+ - uses: yaniv-golan/cowork-harness@v3
51
51
  with:
52
52
  command: run
53
53
  path: scenarios/
54
- version: "^2"
54
+ version: "^3"
55
55
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
56
56
  ```
57
57
 
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.5.0"
70
+ - run: npm i -g "cowork-harness@^3.0.1"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
119
119
  you have a reason:
120
120
 
121
121
  ```yaml
122
- - uses: yaniv-golan/cowork-harness@v2
122
+ - uses: yaniv-golan/cowork-harness@v3
123
123
  with:
124
124
  command: lint
125
125
  path: scenarios/
126
- version: "^2" # holds the major
127
- extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
126
+ version: "^3" # holds the major
127
+ extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
128
128
  ```
129
129
 
130
130
  **If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
131
131
  floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
132
132
  recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
133
- major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
133
+ major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
134
134
  stopping at the major boundary. An exact
135
135
  pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
136
136
  moment a recipe adopts a newer flag.
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
147
147
  Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
148
148
  `transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
149
149
  `max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
150
- `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
150
+ `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
151
151
  `allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
152
152
  `question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
153
153
  (`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
322
322
 
323
323
  ## GitHub Actions sketch
324
324
 
325
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
325
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
326
326
  in one step (see the top of this doc) — reach for this form when you need independent per-command
327
327
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
328
328
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.5.0"
345
+ - run: npm i -g "cowork-harness@^3.0.1"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.5.0"
374
+ run: npm i -g "cowork-harness@^3.0.1"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -20,6 +20,12 @@ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.379
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
23
+ - A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
24
+ or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
25
+ (`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
26
+ hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
27
+ Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
28
+ does not fall back to the default.
23
29
  - **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
24
30
  `skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
25
31
  positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
@@ -150,6 +156,60 @@ hands `fn` exactly this dict.
150
156
  - `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
151
157
  `["fail", "prompt", "llm", "first"]`.
152
158
 
159
+ ### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
160
+
161
+ The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
162
+ it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
163
+ carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
164
+ off-registry**.
165
+
166
+ Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
167
+ ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
168
+ both the command and the tier. All rows below are measured, same scenario shape, same session:
169
+
170
+ | Tier | Bash command | tool used | ask | verdict |
171
+ |---|---|---|---|---|
172
+ | `protocol` | `echo hello` | `Bash` | none | ✓ green |
173
+ | `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
174
+ | `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
175
+ | `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
176
+
177
+ So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
178
+ CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
179
+ `container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
180
+ even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
181
+ `protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
182
+ hello-world probe is green everywhere and teaches the wrong expectation.
183
+
184
+ Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
185
+ warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
186
+ allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
187
+ permissive behaviour is deliberately what the scenario is about.
188
+
189
+ > **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
190
+ > traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
191
+
192
+ ### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
193
+
194
+ A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
195
+ a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
196
+ current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
197
+
198
+ Prefer **repinning `baseline:` to an installed version** for anything
199
+ reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
200
+ but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
201
+ failure modes.
202
+
203
+ To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
204
+ definitions shipped with the harness, which are present regardless of what Desktop pruned from this
205
+ machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
206
+ the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
207
+ directory behind, so a directory test passes on exactly the case that fails.
208
+
209
+ `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
210
+ binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
211
+ prevent, so use it to unblock once, never in CI.
212
+
153
213
  ### Determinism contract
154
214
 
155
215
  - `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
4
- (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.1`
4
+ (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -98,6 +98,18 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
98
98
  # around hostloop's native file tools, that combination gives the
99
99
  # agent genuine, software-checked-only host filesystem access.
100
100
  # Read-only folders and folder-less runs need no opt-in.
101
+
102
+ allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
103
+ # declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
104
+ # the manifest's `hooks` key: L0 passes
105
+ # --plugin-dir, so the CLI executes those hooks as NATIVE HOST
106
+ # processes under your account, with no container sandbox. A plugin
107
+ # that declares no hooks needs no opt-in, and a misplaced root-level
108
+ # `hooks.json` cannot execute so it does not trigger the gate.
109
+ # Use `--fidelity container` to run them sandboxed instead.
110
+ # NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
111
+ # OLDER CLI does not default it — it hard-errors
112
+ # `Unrecognized key: "allow_host_hooks"` and exits 2.
101
113
  ```
102
114
 
103
115
  Relative paths resolve from the file's own directory, so a scenario + session + referenced files
@@ -362,7 +374,7 @@ same set live from the schema.
362
374
  | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
363
375
  | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
364
376
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
365
- | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
377
+ | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
366
378
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
367
379
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
368
380
  | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
@@ -404,7 +416,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
404
416
  | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
405
417
  | `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
406
418
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
407
- | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
419
+ | `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
408
420
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
409
421
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
410
422
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
@@ -443,7 +455,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
443
455
  `skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
444
456
  `max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
445
457
  (`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
446
- modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
458
+ modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
447
459
  `allow_stall` are also kept on replay, evaluated as no-op passes.
448
460
 
449
461
  **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.0.1` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -3,7 +3,7 @@
3
3
  "keys": [
4
4
  "all_tasks_completed",
5
5
  "allow_delete_in",
6
- "allow_l0_plugin_divergence",
6
+ "allow_l0_host_config_contamination",
7
7
  "allow_missing_capability",
8
8
  "allow_outputs_delete",
9
9
  "allow_permissive_auto_allow",
@@ -83,6 +83,7 @@
83
83
  "vm_path_denied"
84
84
  ],
85
85
  "topLevelKeys": [
86
+ "allow_host_hooks",
86
87
  "allow_host_writes",
87
88
  "answers",
88
89
  "assert",
@@ -101,7 +102,7 @@
101
102
  ],
102
103
  "verdictModifierKeys": [
103
104
  "allow_delete_in",
104
- "allow_l0_plugin_divergence",
105
+ "allow_l0_host_config_contamination",
105
106
  "allow_missing_capability",
106
107
  "allow_outputs_delete",
107
108
  "allow_permissive_auto_allow",
@@ -173,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
173
173
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
174
174
  VERDICT_MODIFIER_KEYS = {
175
175
  "allow_permissive_auto_allow",
176
- "allow_l0_plugin_divergence",
176
+ "allow_l0_host_config_contamination",
177
177
  "allow_missing_capability",
178
178
  "allow_stall",
179
179
  "allow_undelivered_deliverables",
@@ -262,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
262
262
  "assert",
263
263
  "skills", # opt-in skill-staleness hash scope
264
264
  "requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
265
- "allow_host_writes", # hostloop native-split: consent for a writable connected folder (pre-run gate)
265
+ "allow_host_writes",
266
+ "allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
266
267
  }
267
268
 
268
269
 
package/AGENTS.md CHANGED
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
22
22
  Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
23
23
  `npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
24
24
  - CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
25
- see README's [Reproducibility knobs](./README.md#reproducibility-knobs) for the full env-var list.
25
+ see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
26
26
  Node ≥ 22.
27
27
  - `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
28
28
  `baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
package/CHANGELOG.md CHANGED
@@ -6,6 +6,224 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.0.1] — 2026-08-30
10
+
11
+ ### Changed
12
+
13
+ - **The README now shows the product, not just the argument for it.** Three independent reviews found
14
+ the same gap: for a test harness, the file a user authors is the most persuasive thing the project
15
+ owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
16
+ answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
17
+ a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
18
+ actual run where a named skill was offered, declined, and the green answer looked fine anyway
19
+ (`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
20
+ it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
21
+ the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
22
+ diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
23
+ can fail before trusting it. The requirements block moves below the `claude -p` argument — three
24
+ reviewers independently reported reading install prerequisites before any reason to want them.
25
+
26
+ - **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
27
+ audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
28
+ prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
29
+ + orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
30
+ Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
31
+ to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
32
+ README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
33
+ the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
34
+
35
+ Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
36
+ `docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
37
+ than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
38
+ moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
39
+
40
+ ### Fixed
41
+
42
+ - **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
43
+ `docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
44
+ but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
45
+ failed its own post-write lockstep check. Caught at release time by that self-check rather than
46
+ shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
47
+ carries the reason.
48
+
49
+
50
+ - **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
51
+ Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
52
+ README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
53
+ `docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
54
+ scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
55
+ suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
56
+ headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
57
+ with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
58
+ regression that shipped. The two nav lines are deleted rather than repointed — they were a
59
+ table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
60
+
61
+ Found by two independent reviewers; the second caught the three non-README instances a README-scoped
62
+ fix would have missed.
63
+
64
+ Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
65
+ **below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
66
+ page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
67
+ This half is the nastier one — the link works, so every link checker stays green forever.
68
+
69
+ - **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
70
+ (3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
71
+ processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
72
+ hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
73
+ the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
74
+ under the operator's own account while the run went green, no consent was asked and no disclosure
75
+ printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
76
+ unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
77
+ the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
78
+ deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
79
+ inert, and a manifest-declared hook is live wherever the manifest sits.
80
+
81
+ Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
82
+ 10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
83
+ declaration, including one recorded at `container`.
84
+
85
+ ### Documentation
86
+
87
+ - **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
88
+ update deletes the prior version's staged ELF while often leaving an empty version directory, so a
89
+ scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
90
+ this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
91
+ ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
92
+ `docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
93
+ repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
94
+ silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
95
+ never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
96
+
97
+ - **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
98
+ `Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
99
+ ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
100
+ hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
101
+ identical command at `container` produces no ask (green), and `hostloop` routes shell through
102
+ `mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
103
+ and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
104
+ expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
105
+ the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
106
+ unmeasured.
107
+
108
+ ## [3.0.0] — 2026-08-29
109
+
110
+ ### Breaking
111
+
112
+ - **`l0_plugin_divergence` is renamed `l0_host_config_contamination`**, and its modifier
113
+ `allow_l0_plugin_divergence` is renamed `allow_l0_host_config_contamination`. The `RunResult` field
114
+ `l0PluginDivergence` becomes `l0HostConfigContamination`. Verdict-signal codes and the scenario schema
115
+ are covered surfaces ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)), so this is
116
+ a MAJOR bump. The old name described plugin *delivery* diverging at L0; delivery now works, and what the
117
+ signal actually reports is that the run read the operator's real config dir. A scenario asserting the old
118
+ key must rename it; a consumer keying on the old code must too.
119
+
120
+ - **The signal's firing conditions changed with it.** It fired when a session declared plugin dirs; it now
121
+ fires when `protocol` reads the operator's real config dir — tested on the dir the agent will actually
122
+ read, so a pinned `plugins.config_dir` is caught even with the managed branch nominally active. A
123
+ `skills.local`-only protocol run that previously passed can now fail, and a sealed protocol run with
124
+ plugins that previously failed now passes.
125
+
126
+
127
+ ### Added
128
+
129
+ - **`allow_host_hooks` (scenario key) and `--allow-host-hooks` (`chat` / `skill`).** Consent to
130
+ running a staged plugin's hooks as native host processes at `protocol`. Top-level like
131
+ `allow_host_writes` rather than a verdict modifier — it gates the SPAWN, it does not suppress a signal.
132
+
133
+ **Version floor: `allow_host_hooks` needs cowork-harness ≥ 3.0.0.** The scenario loader is a `z.strictObject`, so an older CLI does NOT fall back to the default — it hard-errors `Unrecognized key: "allow_host_hooks"` and exits 2 (verified). Adopting the key is a floor bump for every consumer of that scenario.
134
+
135
+ - **A committed live probe for L0 plugin delivery** — `examples/probes/l0-plugin-delivery.scenario.yaml`
136
+ plus its fixture and session. It asserts the plugin under test reaches the agent's init inventory at
137
+ `protocol`, which no unit test can see: the argv builder was correct in isolation and the defect was a
138
+ missing call site. It pins the sealed config-dir branch deliberately — off it, the operator's own
139
+ installed plugins are in the inventory and a name collision would satisfy the assertion with or without
140
+ `--plugin-dir`, measuring the machine instead of the argv. The fixture also ships an agent, deliberately
141
+ unasserted: there is no `agent_available` assertion key (tool / skill / connector have one, agents do
142
+ not), and naming that gap is better than inventing a key to hide it.
143
+
144
+ ### Changed
145
+
146
+ - **`protocol` (L0) now passes `--plugin-dir`, so a declared plugin or skill dir is actually delivered.**
147
+ It previously passed no `--plugin-dir` and `local_plugins` never reached the generated `settings.json`,
148
+ so **the positional argument was silently inert**: `cowork-harness chat <dir> --fidelity protocol` (and
149
+ the `skill` / `probe-dispatch` equivalents) measured whatever the operator had installed rather than the
150
+ tree they passed. Live-verified: a bare skill dir and a plugin root both register at L0 now, with their
151
+ skills and their declared agents. Expect scenarios that previously passed *vacuously* — asserting against
152
+ an agent that never had the plugin — to start failing honestly.
153
+
154
+ - **A plugin's hooks and MCP servers now run at `protocol`, and hooks require consent.** Loading a plugin
155
+ means the CLI executes its `<plugin>/hooks/hooks.json` as **native host processes** — the operator's
156
+ account and environment, no container sandbox — and opens its declared MCP servers. `protocol` therefore
157
+ refuses to spawn when a staged plugin declares runnable hooks unless the scenario sets
158
+ `allow_host_hooks: true` (or `--allow-host-hooks` for `chat`/`skill`); a plugin without hooks
159
+ needs no opt-in, and a misplaced root-level `hooks.json` (which cannot execute) does not trigger it. A
160
+ per-run disclosure is printed even when consent was given. Mirrors the existing `allow_host_writes` gate.
161
+
162
+ - **`protocol` takes the managed config dir when any credential is in the environment.**
163
+ `CLAUDE_CODE_OAUTH_TOKEN` and `ANTHROPIC_AUTH_TOKEN` now select it alongside `ANTHROPIC_API_KEY`, and the
164
+ token is injected into the agent's env (a managed config dir with no credential yields "Not logged in").
165
+ This severs host plugin/skill/MCP discovery — verified: a token-only L0 run delivered the plugin under
166
+ test and **zero** host plugins. `COWORK_MANAGED_CONFIG=0` suppresses only the token-derived branch,
167
+ never the `ANTHROPIC_API_KEY` CI path; an unrecognized value is now rejected rather than silently
168
+ selecting the managed branch (`COWORK_MANAGED_CONFIG=false` previously did).
169
+
170
+ - **`l0_plugin_divergence` now reports contamination rather than the `--plugin-dir` layout.** Delivery is
171
+ fixed, so the signal fires when L0 runs against the operator's **real** config dir, and it tests the dir
172
+ the agent will actually read — a pinned `plugins.config_dir` reaches host discovery with the managed
173
+ branch nominally active. It stays a hard fail: nothing else catches this (the host-inventory scan runs
174
+ only at cassette-record time, and `host_path_leak`'s default-fail is skipped at this tier).
175
+
176
+ - **`workflow-authoring` added to the built-in skill roster** (`KNOWN_BUILTIN_SKILLS`), measured against a
177
+ sealed `protocol` run rather than a contaminated one.
178
+
179
+ - **Platform baseline `desktop-1.40609.0` (agent `2.1.247`).** The Cowork system prompt, both sub-agent
180
+ appends, the egress allowlist, `spawn.env`, `tools[]`/`allowedTools` and `mountLayout` are all re-derived
181
+ unchanged. The one `sync` delta was a pure refactor: W2's `CLAUDE_CODE_ENTRYPOINT` deployment ternary is
182
+ now a hoisted helper, and W1 hard-sets `local-agent` after the spread either way, so the derived value
183
+ does not move. The VM rootfs is re-captured — node `v22.22.3` → `v22.23.2`, pip 144 → 136 packages
184
+ (`xlrd` added; the nine removals are Ubuntu system packages, no analysis library dropped).
185
+
186
+ ### Fixed
187
+
188
+ - **`microvm` resolves its agent binary through the shared resolver instead of deriving the path itself.**
189
+ It read `agentBinary.stagedPath` raw and handed it to the guest mount, so a pin that Claude Desktop had
190
+ pruned surfaced as `env: 'claude': No such file or directory` and exit 127 — the least informative
191
+ message possible for a condition the other three tiers name precisely. The quieter half mattered more:
192
+ the raw path also skipped `verifiedElf`, leaving the one tier that actually **executes** the ELF in a VM
193
+ as the only one not verifying it against the baseline pin, while `container` hard-fails on the same
194
+ mismatch. Routing it through `resolveAgentBinary` restores all three safeguards at once — existence
195
+ check, sha verification and the pruned-binary fallback — and removes a fourth derivation of a rule
196
+ `baseline.ts` already documented `microvm` as following. Resolution happens **before** the
197
+ already-Running reuse short-circuit, since a VM created while the binary was present keeps a mount at
198
+ the pruned path — the originally reported state. `vm status` / `vm prune` / `doctor` keep working when
199
+ the binary is missing (that is when an operator reaches for them) and degrade to the pinned path rather
200
+ than throwing.
201
+
202
+ - **`sync` resolves a spawn-env value expression that is a hoisted one-line helper** (`uH(n.type)`), where
203
+ it previously refused the baseline. The branch is narrow by construction: the argument must be a `.type`
204
+ member and the callee body must be exactly the literal deployment ternary over its own parameter, resolved
205
+ in the window's own chunk — a two-character minified name hopped against the joined bundle lands on an
206
+ unrelated helper. Anything else stays unresolvable rather than guessed.
207
+
208
+ - **`provenance.asarGateIds` now includes the gate-defaults map's bare-numeric keys.** The scan matched
209
+ quoted literals only, so any gate that is keyed in that map and never read through a quoted id was
210
+ missing — `provenance.asarGateIds` 252 → 291, and the 1.40609.0 delta corrects from +27/−4 to +30/−4.
211
+ This is not the bare-*number* scan the extractor deliberately rejects: that matches any numeric and adds
212
+ 1687 ids over the same bundle, while the defaults-map entry shape adds 39.
213
+
214
+ ### Documentation
215
+
216
+ - **A plugin's declared MCP servers are a documented fidelity gap.** Production replaces them with
217
+ zero-tool SDK stubs named `plugin:<plugin>:<server>` — remote (`url` + http/sse) unconditionally, local
218
+ and `.mcpb` under an MCP policy — while the harness stages plugins with `--plugin-dir` and lets the CLI
219
+ open the real ones. A plugin under test therefore sees a tool surface production would not give it, in
220
+ both directions and silently. Not modeled: the remote rule is mechanically reproducible, but the
221
+ local/`.mcpb` rule is conditioned on Desktop policy state the harness has no source for.
222
+
223
+ - **The auto-mode permission rubric gap is tier-independent.** `docs/fidelity-gaps.md` scopes it to both
224
+ loops rather than VM-loop only, and records that the rubric's Filesystem section is tier-dependent
225
+ model-visible text — a different kind of divergence from a permission verdict.
226
+
9
227
  ## [2.5.0] — 2026-08-28
10
228
 
11
229
  ### Added