cowork-harness 2.5.0 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +7 -7
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +17 -17
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +6 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +16 -5
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -2
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +3 -2
  9. package/CHANGELOG.md +119 -0
  10. package/DESIGN.md +2 -2
  11. package/README.md +16 -16
  12. package/SPEC.md +2 -2
  13. package/baselines/desktop-1.40609.0.json +878 -0
  14. package/baselines/provisioning/rootfs-provisioning.json +32 -40
  15. package/dist/cli.js +8 -1
  16. package/dist/run/cassette.js +4 -3
  17. package/dist/run/chat-result.js +1 -1
  18. package/dist/run/chat.js +15 -1
  19. package/dist/run/execute.js +15 -6
  20. package/dist/run/hook-events.js +51 -0
  21. package/dist/run/skill-flag-surface.js +11 -0
  22. package/dist/run/verdict.js +17 -8
  23. package/dist/runtime/argv.js +16 -1
  24. package/dist/runtime/lima.js +37 -3
  25. package/dist/runtime/protocol.js +85 -13
  26. package/dist/scan.js +1 -0
  27. package/dist/sync/cowork-sync.js +39 -2
  28. package/dist/types.js +11 -3
  29. package/docs/cassette.md +2 -2
  30. package/docs/discovery.md +1 -1
  31. package/docs/fidelity-gaps.md +78 -6
  32. package/docs/maintenance.md +1 -1
  33. package/docs/scenario.md +6 -3
  34. package/examples/replays/README.md +1 -1
  35. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  36. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  37. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  38. package/examples/sessions/l0-plugin-delivery.yaml +6 -0
  39. package/package.json +1 -1
  40. package/python/test_scenario_lint.py +1 -1
  41. package/schema/run-result.json +2 -2
  42. package/schema/scenario.schema.json +6 -2
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.5.0
7
- tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
6
+ version: 3.0.0
7
+ tracks-harness: cowork-harness 3.0.0 (baseline desktop-1.40609.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
29
- > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.0.0` (baseline
29
+ > `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,13 +42,13 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.0.0"`. **Pin `@^3.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
49
49
  — upgrade rather than work around it, since this skill's `file:line` pointers and flag names track the floor.
50
50
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
51
- - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
51
+ - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
52
52
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
53
53
  - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
54
54
 
@@ -679,7 +679,7 @@ than a stuck `"running"`.)
679
679
  ### Place assertions in the right CI lane
680
680
 
681
681
  CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
682
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v2` (a packaged GitHub Action with a
682
+ (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
683
683
  PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
684
684
  the four-stage pipeline.
685
685
 
@@ -1,23 +1,23 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
7
7
 
8
8
  ```yaml
9
- - uses: yaniv-golan/cowork-harness@v2
9
+ - uses: yaniv-golan/cowork-harness@v3
10
10
  with:
11
11
  command: replay
12
12
  path: cassettes/
13
- version: "^2" # hold the major; see below
13
+ version: "^3" # hold the major; see below
14
14
  ```
15
15
 
16
- **These recipes pin `version: "^2"`.** The Action's `version` input *defaults* to `latest`, which means a
16
+ **These recipes pin `version: "^3"`.** The Action's `version` input *defaults* to `latest`, which means a
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.0.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.246 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.247 # match your scenario's pinned baseline's agentVersion
40
40
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
41
41
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
42
42
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -47,11 +47,11 @@ jobs:
47
47
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
48
48
  # Background on the provenance chain: the "Agent-binary provenance" section of
49
49
  # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md
50
- - uses: yaniv-golan/cowork-harness@v2
50
+ - uses: yaniv-golan/cowork-harness@v3
51
51
  with:
52
52
  command: run
53
53
  path: scenarios/
54
- version: "^2"
54
+ version: "^3"
55
55
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
56
56
  ```
57
57
 
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.5.0"
70
+ - run: npm i -g "cowork-harness@^3.0.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -119,18 +119,18 @@ Action has no input for, and it creates a coupling nothing checks:
119
119
  you have a reason:
120
120
 
121
121
  ```yaml
122
- - uses: yaniv-golan/cowork-harness@v2
122
+ - uses: yaniv-golan/cowork-harness@v3
123
123
  with:
124
124
  command: lint
125
125
  path: scenarios/
126
- version: "^2" # holds the major
127
- extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 2.x satisfies that
126
+ version: "^3" # holds the major
127
+ extra-args: --min-severity WARN # needs a CLI >= 1.11.0; any 3.x satisfies that
128
128
  ```
129
129
 
130
130
  **If a flag you pass in `extra-args` landed in a specific release, bound the range — don't write a bare
131
131
  floor.** `>=1.11.0` reads as "at least 1.11.0" and silently means "and every future major too", so a
132
132
  recipe written that way hands a copy-paster the next major with no say in it. Anchor it at the current
133
- major instead — `version: "^2"`, which is what the steps above use — keeping the floor's intent while
133
+ major instead — `version: "^3"`, which is what the steps above use — keeping the floor's intent while
134
134
  stopping at the major boundary. An exact
135
135
  pin (`version: "2.0.1"`) is the right choice when you want byte-reproducible CI, at the cost of rotting the
136
136
  moment a recipe adopts a newer flag.
@@ -147,7 +147,7 @@ The split is not just about tokens — it decides **where each lane can run**:
147
147
  Docker, no agent binary** — runs on a stock GitHub Actions runner. Evaluates **content** assertions —
148
148
  `transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
149
149
  `max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
150
- `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
150
+ `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
151
151
  `allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
152
152
  `question_context` / `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
153
153
  (`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
@@ -322,7 +322,7 @@ A typical skill repo runs four stages, fastest/cheapest first:
322
322
 
323
323
  ## GitHub Actions sketch
324
324
 
325
- The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v2` does
325
+ The PR gate below is the manual, step-by-step version of what `uses: yaniv-golan/cowork-harness@v3` does
326
326
  in one step (see the top of this doc) — reach for this form when you need independent per-command
327
327
  gating/annotations rather than one action run per command. The nightly live job has no packaged-Action
328
328
  equivalent yet (the Action's `command: run` mode needs a self-hosted runner with Docker + the agent binary
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.5.0"
345
+ - run: npm i -g "cowork-harness@^3.0.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.5.0"
374
+ run: npm i -g "cowork-harness@^3.0.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -20,6 +20,11 @@ Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.379
20
20
  - A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
21
21
  true` — with no container around the native file tools, that combination gives the agent genuine,
22
22
  software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
23
+ - A `protocol` scenario staging a plugin that declares runnable hooks needs `allow_host_hooks: true`
24
+ (`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
25
+ hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
26
+ Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
27
+ does not fall back to the default.
23
28
  - **Set the tier in the scenario's `fidelity:` field — not a flag.** `--fidelity` is accepted only by
24
29
  `skill` (any tier) and `chat` (`protocol`/`container`/`hostloop`; only `microvm`/`cowork` unsupported); `run` rejects an extra `--fidelity`
25
30
  positional ("Fidelity is set by the scenario's `fidelity:` field, not a flag").
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
4
- (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.0.0`
4
+ (baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -98,6 +98,17 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
98
98
  # around hostloop's native file tools, that combination gives the
99
99
  # agent genuine, software-checked-only host filesystem access.
100
100
  # Read-only folders and folder-less runs need no opt-in.
101
+
102
+ allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
103
+ # declares runnable hooks (`<plugin>/hooks/hooks.json`): L0 passes
104
+ # --plugin-dir, so the CLI executes those hooks as NATIVE HOST
105
+ # processes under your account, with no container sandbox. A plugin
106
+ # that declares no hooks needs no opt-in, and a misplaced root-level
107
+ # `hooks.json` cannot execute so it does not trigger the gate.
108
+ # Use `--fidelity container` to run them sandboxed instead.
109
+ # NEEDS cowork-harness >= 3.0.0. The loader is a strict object, so an
110
+ # OLDER CLI does not default it — it hard-errors
111
+ # `Unrecognized key: "allow_host_hooks"` and exits 2.
101
112
  ```
102
113
 
103
114
  Relative paths resolve from the file's own directory, so a scenario + session + referenced files
@@ -362,7 +373,7 @@ same set live from the schema.
362
373
  | `no_path_denied: true` | **`fidelity: hostloop` only** — NO path denial was recorded at all (the channel is already path-scoped, unlike `no_hook_blocked`'s indiscriminate reject). Replay: needs a `controlOut` cassette. Any other tier FAILS "cannot verify". **Only `true` is valid** |
363
374
  | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
364
375
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
365
- | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
376
+ | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
366
377
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
367
378
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
368
379
  | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
@@ -404,7 +415,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
404
415
  | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
405
416
  | `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
406
417
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
407
- | `l0_plugin_divergence` | fail | L0/protocol plugin loading diverged from Cowork (opt out: `allow_l0_plugin_divergence`) |
418
+ | `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
408
419
  | `missing_capability` | fail | A `requires_capabilities` need was unmet, or the skill used a capability the image omits (opt out: `allow_missing_capability`, or `skill --allow-missing-capability` on an open-ended run) |
409
420
  | `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
410
421
  | `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
@@ -443,7 +454,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
443
454
  `skill_tool_used`, `max_cost_usd`, `max_tokens`, `tool_calls_max`, `tool_no_error`,
444
455
  `max_tool_errors`, `max_redundant_tool_calls`, `max_turns`, `compaction_occurred`, `all_tasks_completed`, `task_status`, `task_count_min`, `no_scratchpad_leak`, `present_files_called`, `result`
445
456
  (`max_cost_usd`/`max_tokens` assert the frozen recording's spend on replay, not fresh spend). The verdict
446
- modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
457
+ modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_host_config_contamination` /
447
458
  `allow_stall` are also kept on replay, evaluated as no-op passes.
448
459
 
449
460
  **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `question_context`, `questions_count_max`,
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.0.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -3,7 +3,7 @@
3
3
  "keys": [
4
4
  "all_tasks_completed",
5
5
  "allow_delete_in",
6
- "allow_l0_plugin_divergence",
6
+ "allow_l0_host_config_contamination",
7
7
  "allow_missing_capability",
8
8
  "allow_outputs_delete",
9
9
  "allow_permissive_auto_allow",
@@ -83,6 +83,7 @@
83
83
  "vm_path_denied"
84
84
  ],
85
85
  "topLevelKeys": [
86
+ "allow_host_hooks",
86
87
  "allow_host_writes",
87
88
  "answers",
88
89
  "assert",
@@ -101,7 +102,7 @@
101
102
  ],
102
103
  "verdictModifierKeys": [
103
104
  "allow_delete_in",
104
- "allow_l0_plugin_divergence",
105
+ "allow_l0_host_config_contamination",
105
106
  "allow_missing_capability",
106
107
  "allow_outputs_delete",
107
108
  "allow_permissive_auto_allow",
@@ -173,7 +173,7 @@ LANE_REMOTE_INCOMPATIBLE_KEYS = {"present_files_called", "no_scratchpad_leak", "
173
173
  # verdict modifiers — don't verify anything themselves (e.g. suppress a default-fail)
174
174
  VERDICT_MODIFIER_KEYS = {
175
175
  "allow_permissive_auto_allow",
176
- "allow_l0_plugin_divergence",
176
+ "allow_l0_host_config_contamination",
177
177
  "allow_missing_capability",
178
178
  "allow_stall",
179
179
  "allow_undelivered_deliverables",
@@ -262,7 +262,8 @@ _EMBEDDED_TOP_LEVEL_KEYS = {
262
262
  "assert",
263
263
  "skills", # opt-in skill-staleness hash scope
264
264
  "requires_capabilities", # Fix 4b: scenario-level required-capability declaration (pre-flight gate)
265
- "allow_host_writes", # hostloop native-split: consent for a writable connected folder (pre-run gate)
265
+ "allow_host_writes",
266
+ "allow_host_hooks", # protocol consent: a staged plugin's hooks run as NATIVE HOST processes # hostloop native-split: consent for a writable connected folder (pre-run gate)
266
267
  }
267
268
 
268
269
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,125 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.0.0] — 2026-08-29
10
+
11
+ ### Breaking
12
+
13
+ - **`l0_plugin_divergence` is renamed `l0_host_config_contamination`**, and its modifier
14
+ `allow_l0_plugin_divergence` is renamed `allow_l0_host_config_contamination`. The `RunResult` field
15
+ `l0PluginDivergence` becomes `l0HostConfigContamination`. Verdict-signal codes and the scenario schema
16
+ are covered surfaces ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)), so this is
17
+ a MAJOR bump. The old name described plugin *delivery* diverging at L0; delivery now works, and what the
18
+ signal actually reports is that the run read the operator's real config dir. A scenario asserting the old
19
+ key must rename it; a consumer keying on the old code must too.
20
+
21
+ - **The signal's firing conditions changed with it.** It fired when a session declared plugin dirs; it now
22
+ fires when `protocol` reads the operator's real config dir — tested on the dir the agent will actually
23
+ read, so a pinned `plugins.config_dir` is caught even with the managed branch nominally active. A
24
+ `skills.local`-only protocol run that previously passed can now fail, and a sealed protocol run with
25
+ plugins that previously failed now passes.
26
+
27
+
28
+ ### Added
29
+
30
+ - **`allow_host_hooks` (scenario key) and `--allow-host-hooks` (`chat` / `skill`).** Consent to
31
+ running a staged plugin's hooks as native host processes at `protocol`. Top-level like
32
+ `allow_host_writes` rather than a verdict modifier — it gates the SPAWN, it does not suppress a signal.
33
+
34
+ **Version floor: `allow_host_hooks` needs cowork-harness ≥ 3.0.0.** The scenario loader is a `z.strictObject`, so an older CLI does NOT fall back to the default — it hard-errors `Unrecognized key: "allow_host_hooks"` and exits 2 (verified). Adopting the key is a floor bump for every consumer of that scenario.
35
+
36
+ - **A committed live probe for L0 plugin delivery** — `examples/probes/l0-plugin-delivery.scenario.yaml`
37
+ plus its fixture and session. It asserts the plugin under test reaches the agent's init inventory at
38
+ `protocol`, which no unit test can see: the argv builder was correct in isolation and the defect was a
39
+ missing call site. It pins the sealed config-dir branch deliberately — off it, the operator's own
40
+ installed plugins are in the inventory and a name collision would satisfy the assertion with or without
41
+ `--plugin-dir`, measuring the machine instead of the argv. The fixture also ships an agent, deliberately
42
+ unasserted: there is no `agent_available` assertion key (tool / skill / connector have one, agents do
43
+ not), and naming that gap is better than inventing a key to hide it.
44
+
45
+ ### Changed
46
+
47
+ - **`protocol` (L0) now passes `--plugin-dir`, so a declared plugin or skill dir is actually delivered.**
48
+ It previously passed no `--plugin-dir` and `local_plugins` never reached the generated `settings.json`,
49
+ so **the positional argument was silently inert**: `cowork-harness chat <dir> --fidelity protocol` (and
50
+ the `skill` / `probe-dispatch` equivalents) measured whatever the operator had installed rather than the
51
+ tree they passed. Live-verified: a bare skill dir and a plugin root both register at L0 now, with their
52
+ skills and their declared agents. Expect scenarios that previously passed *vacuously* — asserting against
53
+ an agent that never had the plugin — to start failing honestly.
54
+
55
+ - **A plugin's hooks and MCP servers now run at `protocol`, and hooks require consent.** Loading a plugin
56
+ means the CLI executes its `<plugin>/hooks/hooks.json` as **native host processes** — the operator's
57
+ account and environment, no container sandbox — and opens its declared MCP servers. `protocol` therefore
58
+ refuses to spawn when a staged plugin declares runnable hooks unless the scenario sets
59
+ `allow_host_hooks: true` (or `--allow-host-hooks` for `chat`/`skill`); a plugin without hooks
60
+ needs no opt-in, and a misplaced root-level `hooks.json` (which cannot execute) does not trigger it. A
61
+ per-run disclosure is printed even when consent was given. Mirrors the existing `allow_host_writes` gate.
62
+
63
+ - **`protocol` takes the managed config dir when any credential is in the environment.**
64
+ `CLAUDE_CODE_OAUTH_TOKEN` and `ANTHROPIC_AUTH_TOKEN` now select it alongside `ANTHROPIC_API_KEY`, and the
65
+ token is injected into the agent's env (a managed config dir with no credential yields "Not logged in").
66
+ This severs host plugin/skill/MCP discovery — verified: a token-only L0 run delivered the plugin under
67
+ test and **zero** host plugins. `COWORK_MANAGED_CONFIG=0` suppresses only the token-derived branch,
68
+ never the `ANTHROPIC_API_KEY` CI path; an unrecognized value is now rejected rather than silently
69
+ selecting the managed branch (`COWORK_MANAGED_CONFIG=false` previously did).
70
+
71
+ - **`l0_plugin_divergence` now reports contamination rather than the `--plugin-dir` layout.** Delivery is
72
+ fixed, so the signal fires when L0 runs against the operator's **real** config dir, and it tests the dir
73
+ the agent will actually read — a pinned `plugins.config_dir` reaches host discovery with the managed
74
+ branch nominally active. It stays a hard fail: nothing else catches this (the host-inventory scan runs
75
+ only at cassette-record time, and `host_path_leak`'s default-fail is skipped at this tier).
76
+
77
+ - **`workflow-authoring` added to the built-in skill roster** (`KNOWN_BUILTIN_SKILLS`), measured against a
78
+ sealed `protocol` run rather than a contaminated one.
79
+
80
+ - **Platform baseline `desktop-1.40609.0` (agent `2.1.247`).** The Cowork system prompt, both sub-agent
81
+ appends, the egress allowlist, `spawn.env`, `tools[]`/`allowedTools` and `mountLayout` are all re-derived
82
+ unchanged. The one `sync` delta was a pure refactor: W2's `CLAUDE_CODE_ENTRYPOINT` deployment ternary is
83
+ now a hoisted helper, and W1 hard-sets `local-agent` after the spread either way, so the derived value
84
+ does not move. The VM rootfs is re-captured — node `v22.22.3` → `v22.23.2`, pip 144 → 136 packages
85
+ (`xlrd` added; the nine removals are Ubuntu system packages, no analysis library dropped).
86
+
87
+ ### Fixed
88
+
89
+ - **`microvm` resolves its agent binary through the shared resolver instead of deriving the path itself.**
90
+ It read `agentBinary.stagedPath` raw and handed it to the guest mount, so a pin that Claude Desktop had
91
+ pruned surfaced as `env: 'claude': No such file or directory` and exit 127 — the least informative
92
+ message possible for a condition the other three tiers name precisely. The quieter half mattered more:
93
+ the raw path also skipped `verifiedElf`, leaving the one tier that actually **executes** the ELF in a VM
94
+ as the only one not verifying it against the baseline pin, while `container` hard-fails on the same
95
+ mismatch. Routing it through `resolveAgentBinary` restores all three safeguards at once — existence
96
+ check, sha verification and the pruned-binary fallback — and removes a fourth derivation of a rule
97
+ `baseline.ts` already documented `microvm` as following. Resolution happens **before** the
98
+ already-Running reuse short-circuit, since a VM created while the binary was present keeps a mount at
99
+ the pruned path — the originally reported state. `vm status` / `vm prune` / `doctor` keep working when
100
+ the binary is missing (that is when an operator reaches for them) and degrade to the pinned path rather
101
+ than throwing.
102
+
103
+ - **`sync` resolves a spawn-env value expression that is a hoisted one-line helper** (`uH(n.type)`), where
104
+ it previously refused the baseline. The branch is narrow by construction: the argument must be a `.type`
105
+ member and the callee body must be exactly the literal deployment ternary over its own parameter, resolved
106
+ in the window's own chunk — a two-character minified name hopped against the joined bundle lands on an
107
+ unrelated helper. Anything else stays unresolvable rather than guessed.
108
+
109
+ - **`provenance.asarGateIds` now includes the gate-defaults map's bare-numeric keys.** The scan matched
110
+ quoted literals only, so any gate that is keyed in that map and never read through a quoted id was
111
+ missing — `provenance.asarGateIds` 252 → 291, and the 1.40609.0 delta corrects from +27/−4 to +30/−4.
112
+ This is not the bare-*number* scan the extractor deliberately rejects: that matches any numeric and adds
113
+ 1687 ids over the same bundle, while the defaults-map entry shape adds 39.
114
+
115
+ ### Documentation
116
+
117
+ - **A plugin's declared MCP servers are a documented fidelity gap.** Production replaces them with
118
+ zero-tool SDK stubs named `plugin:<plugin>:<server>` — remote (`url` + http/sse) unconditionally, local
119
+ and `.mcpb` under an MCP policy — while the harness stages plugins with `--plugin-dir` and lets the CLI
120
+ open the real ones. A plugin under test therefore sees a tool surface production would not give it, in
121
+ both directions and silently. Not modeled: the remote rule is mechanically reproducible, but the
122
+ local/`.mcpb` rule is conditioned on Desktop policy state the harness has no source for.
123
+
124
+ - **The auto-mode permission rubric gap is tier-independent.** `docs/fidelity-gaps.md` scopes it to both
125
+ loops rather than VM-loop only, and records that the rubric's Filesystem section is tier-dependent
126
+ model-visible text — a different kind of divergence from a permission verdict.
127
+
9
128
  ## [2.5.0] — 2026-08-28
10
129
 
11
130
  ### Added
package/DESIGN.md CHANGED
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
47
47
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
48
48
 
49
49
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
50
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.246**, per `baselines/desktop-1.37937.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
50
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.247**, per `baselines/desktop-1.40609.0.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
51
51
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
52
52
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
53
53
 
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
176
176
 
177
177
  ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
178
178
 
179
- > **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing in `baselines/` is currently unverified for want of a live run. The pass ran against agent **2.1.246** (the staged VM ELF and the native `.app` are both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
179
+ > **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is no longer the newest committed baseline: **one** baseline has shipped since (`1.40609.0`), **one** of which moved the agent ELF, most recently to **2.1.247** — so the newest baseline is **not** live-verified, and this paragraph describes the 1.37937.1 pass only. The pass ran on agent `2.1.246` (the staged VM ELF and the native `.app` were both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
180
180
 
181
181
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
182
182