cowork-harness 3.8.0 → 3.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +6 -6
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +7 -7
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +4 -2
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/CHANGELOG.md +225 -0
  8. package/DESIGN.md +2 -2
  9. package/README.md +4 -4
  10. package/RELEASING.md +34 -20
  11. package/baselines/desktop-2.7032.0.json +42 -2
  12. package/baselines/desktop-2.9939.2.json +1116 -0
  13. package/baselines/provisioning/rootfs-provisioning.json +2 -2
  14. package/dist/baseline.js +42 -9
  15. package/dist/cli.js +27 -10
  16. package/dist/decide/external-channel.js +84 -15
  17. package/dist/errors.js +14 -0
  18. package/dist/loop-decision.js +22 -1
  19. package/dist/run/cassette.js +5 -1
  20. package/dist/runtime/vm-baseline-arg.js +52 -0
  21. package/dist/session.js +4 -3
  22. package/dist/sync/baseline-diff.js +129 -12
  23. package/dist/sync/cowork-sync.js +138 -29
  24. package/dist/sync/desktop-init-surface.js +206 -0
  25. package/dist/types.js +49 -0
  26. package/docs/ci.md +33 -3
  27. package/docs/cli.md +6 -6
  28. package/docs/companion-skill.md +2 -2
  29. package/docs/fidelity-gaps.md +52 -18
  30. package/docs/invariants.md +1 -0
  31. package/docs/maintenance.md +26 -1
  32. package/docs/scenario.md +8 -1
  33. package/docs/session.md +5 -5
  34. package/docs/subagents.md +4 -2
  35. package/examples/replays/README.md +1 -1
  36. package/examples/replays/example-multiselect-gate.cassette.json +51 -51
  37. package/examples/replays/example-pdf-skill.cassette.json +96 -126
  38. package/examples/replays/hostloop-computer-links.cassette.json +57 -57
  39. package/examples/scenarios/subagent-manifest-probe.yaml +23 -12
  40. package/package.json +3 -3
  41. package/schema/session.schema.json +1 -1
  42. package/scripts/check-versions.ts +4 -2
  43. package/scripts/release-preflight.ts +38 -3
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.8.0
7
- tracks-harness: cowork-harness 3.8.0 (baseline desktop-2.7032.0)
6
+ version: 3.9.0
7
+ tracks-harness: cowork-harness 3.9.0 (baseline desktop-2.9939.2)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.8.0` (baseline
29
- > `desktop-2.7032.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.9.0` (baseline
29
+ > `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.8.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.8.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.8.0"`. **Pin `@^3.8.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.9.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.9.0"`. **Pin `@^3.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -569,7 +569,7 @@ Recognize these before "fixing" a non-bug:
569
569
  rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
570
570
  `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
571
571
  (`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
572
- Desktop `2.7032.0` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
572
+ Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
573
573
  behind that sentence). The message says so ("likely a FALSE
574
574
  NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
575
575
  `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.8.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.9.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.280 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.281 # match your scenario's pinned baseline's agentVersion
40
40
  # The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
41
41
  # served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
42
42
  # (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
@@ -46,7 +46,7 @@ jobs:
46
46
  # from your pinned baseline's agentBinary.releaseBaseUrl. The checksum step fails closed if you
47
47
  # don't, but it cannot tell you why. Baselines written before that field existed were
48
48
  # stable-staged, so their base is the plain https://downloads.claude.ai/claude-code-releases.
49
- B=https://downloads.claude.ai/claude-code-releases/rc/bddba3abd5da53d0c540cfc76a8d18b44633d568
49
+ B=https://downloads.claude.ai/claude-code-releases
50
50
  # The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
51
51
  # read it with jq if you vendor the baseline. An unverified download is an unverified agent:
52
52
  # this step FAILS rather than staging one, which is the whole point of naming it "verified".
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
77
77
  GitHub-hosted runners, no token/Docker/agent:
78
78
 
79
79
  ```yaml
80
- - run: npm i -g "cowork-harness@^3.8.0"
80
+ - run: npm i -g "cowork-harness@^3.9.0"
81
81
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
82
82
  # no silent false-greens. WITHOUT --strict this
83
83
  # step cannot fail on a WARN-class rule (e.g.
@@ -364,7 +364,7 @@ jobs:
364
364
  with: { node-version: '24' }
365
365
  - uses: actions/setup-python@v5
366
366
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
367
- - run: npm i -g "cowork-harness@^3.8.0"
367
+ - run: npm i -g "cowork-harness@^3.9.0"
368
368
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
369
369
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
370
370
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -393,7 +393,7 @@ jobs:
393
393
  echo "live=true" >> "$GITHUB_OUTPUT"
394
394
  fi
395
395
  - if: steps.guard.outputs.live == 'true'
396
- run: npm i -g "cowork-harness@^3.8.0"
396
+ run: npm i -g "cowork-harness@^3.9.0"
397
397
  - if: steps.guard.outputs.live == 'true'
398
398
  run: cowork-harness run scenarios/ --output-format json
399
399
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`).
3
+ Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -68,7 +68,9 @@ cowork-harness vm prune # remove orphaned cowork-vm-* VMs from past configs
68
68
 
69
69
  The instance is `cowork-vm-<config-hash>` — a config or agent-version change yields a new name, so a
70
70
  stale VM is never silently reused (the old one is orphaned until `vm prune`). Pin a fixed name with
71
- `COWORK_LIMA_INSTANCE`.
71
+ `COWORK_LIMA_INSTANCE`. The optional argument to every `vm` subcommand is a **baseline**
72
+ (`desktop-<version>`, default `latest`), never the `cowork-vm-<hash>` VM name — passing a VM name is a
73
+ usage error that names the baseline(s) deriving it.
72
74
 
73
75
  ## Answer paths (resolving gates: AskUserQuestion + tool-permission)
74
76
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.8.0`
4
- (baseline `desktop-2.7032.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.9.0`
4
+ (baseline `desktop-2.9939.2`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -192,7 +192,7 @@ plugins:
192
192
  skills:
193
193
  local: [] # extra host skill dirs
194
194
  suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
195
- proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = synced baseline gate (ON from 1.24012.11)
195
+ proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
196
196
  mcp:
197
197
  config: null # --mcp-config file (standard mcpServers map)
198
198
  enabled: []
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.8.0` (baseline `desktop-2.7032.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,231 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.9.0] — 2026-09-25
10
+
11
+ ### Upgrade notes
12
+
13
+ - **Your own cassettes go stale against `latest`.** `baseline: latest` now resolves to
14
+ `desktop-2.9939.2`, so a cassette recorded against `desktop-2.7032.0` reports `[stale] baseline moved`
15
+ in `verify-cassettes`. From a first-party session's point of view the contract did not move: the
16
+ spawn environment, the Cowork system prompt, the sub-agent append and the egress allowlist are all
17
+ byte-identical. So re-stamping `fingerprint.baseline` is sound. Re-record instead if you want the
18
+ recording itself to come from agent 2.1.281.
19
+ - **The three committed cassettes are re-recorded against `desktop-2.9939.2`**, each at its own tier:
20
+ - `example-multiselect-gate` at `protocol`, on the sealed managed config dir and the same
21
+ `claude-opus-5-5[1m]` model. It records host CLI 2.1.282, since the protocol tier runs the host CLI
22
+ by design.
23
+ - `example-pdf-skill` at `container`, agent 2.1.281.
24
+ - `hostloop-computer-links` at `hostloop`, agent 2.1.281. It was first recorded outside the repository
25
+ and its captured inventory compared: the same 4 product MCP servers, agents and skills. The only
26
+ additions (the `focus` command and the built-in `agents-md` plugin) come from the new agent build and
27
+ appear in the sealed recordings too.
28
+
29
+ `verify-cassettes` is clean and all three replay green.
30
+ - **Live-validated against `desktop-2.9939.2`** (agent 2.1.281) on 2026-09-25, all four tiers:
31
+ `boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
32
+ `run examples/scenarios/` 6/7 on its first run. Both reds were model variance and passed on re-runs:
33
+ - `live-matrix` asked its A-or-B question in plain text instead of calling `AskUserQuestion`. It runs
34
+ at the protocol tier, which uses the host `claude` CLI (2.1.282 here), not the staged agent.
35
+ - `example-pdf-skill` wrote its file one directory too high.
36
+
37
+ Details are in `DESIGN.md`'s scope note. Code that landed after that pass (the decider process-group
38
+ fix, the baseline-name usage errors and a `zod` patch) was re-verified live on the release code
39
+ (`f395d07`), again across all four tiers:
40
+ - e2e self-tests 8/8, and `run examples/scenarios/` 7/7. `subagent-manifest-probe` first went 8/9,
41
+ because the sub-agent's first write used a VM path, which the hook refused. It passed 9/9 on a
42
+ re-run.
43
+ - `smoke-l2-microvm` passed.
44
+ - A VM-name argument to `vm` is a usage error (exit 2, category `usage`).
45
+ - A `--decider-cmd` timeout leaves no surviving process group.
46
+ - Ctrl-C at the container tier exits 130 and reaps the containers, the proxy and the network.
47
+
48
+ CI does not live-validate: without an `ANTHROPIC_API_KEY` secret its live scenario suite is skipped
49
+ (see Fixed).
50
+
51
+ ### Added
52
+
53
+ - **`sync` records the tool surface real Cowork sessions declared** for Desktop's own `cowork`, `plugins`
54
+ and `skills` servers, as `provenance.desktopInitSurface`. A production change such as `save_skill`
55
+ turning on or off then shows up in `sync --diff`.
56
+ - It is read from the synced Desktop's own session logs, limited to that release's agent version and
57
+ install time.
58
+ - No other server name is recorded, and a strict schema over every committed baseline enforces that.
59
+ - When no Cowork session has run since the install, `sync` records `observed: false` with a warning,
60
+ instead of reusing the previous release's surface.
61
+ - `desktop-2.7032.0` carries the first recorded block (`save_skill` declared in every session read).
62
+ - **`npm run preflight` refuses to release an unobserved Desktop init surface.** New check: the newest
63
+ baseline's `desktopInitSurface` must be observed. `--allow-unobserved-init-surface` downgrades it to a
64
+ warning for an emergency release; `--allow-empty` does not.
65
+
66
+ ### Parity — Desktop 2.9939.2 (agent 2.1.281)
67
+
68
+ - **Baseline `desktop-2.9939.2`**, with the agent staged from the **stable** release channel, not an RC
69
+ build. This is what `baseline: latest` now resolves to. From a first-party session's point of view
70
+ nothing moved:
71
+ - The Cowork system prompt and all four sub-agent append fingerprints are byte-identical to 2.7032.0.
72
+ - `spawn.env` (24 keys) and the egress allowlist are unchanged.
73
+ - The VM rootfs is re-captured at `2.9939.2` and its provisioning is identical: node `v22.23.2`, the
74
+ same 136 pip packages, the same apt doc stack and the same npm globals (`tsx` included).
75
+ - The recorded changes: `spawnEnvKeys` gains the 3p-only `CLAUDE_CODE_DISABLE_FAST_MODE`, and
76
+ `asarGateIds` gains 28 ids and loses 1.
77
+ - `network.$comment` now describes the resolver's HIPAA filter (see Fixed).
78
+ - Gate `4202409342` (`builtinToolsApprovableByAutoMode`) is still on. It is now served by default
79
+ rather than forced, so a server rule can flip it. The harness does not read it, because auto mode
80
+ is unreachable here.
81
+ - `desktopInitSurface` is observed from **2 init frames, both from a scheduled task**. The four
82
+ artifact tools now appear in every frame read rather than some. Treat that as a property of those
83
+ two sessions, not as a Desktop change.
84
+
85
+ ### Fixed
86
+
87
+ - **A baseline name that names no baseline is now a usage error at every entry point, not a stack
88
+ trace.** The baseline loader read the file with no existence check, so an unknown name escaped as an
89
+ uncaught `ENOENT`. How that surfaced depended on the entry point:
90
+ - `boundary-check <name>`, `vm init|status|delete|prune <name>`, and `run` on a scenario whose
91
+ `baseline:` names nothing: category `internal` in JSON mode, a raw stack trace in text mode.
92
+ - `record`: the `ENOENT` text, with exit 1.
93
+ - `diff <a> <b>`: exit 2, but the raw `ENOENT` text as the message.
94
+ - `run --matrix` with such a `baselines:` entry: the `ENOENT` text as the cell's error.
95
+
96
+ The loader now throws a usage error for a name or path that names no baseline. Every entry point
97
+ reports it as category `usage`, exit 2, with the committed baselines listed newest first in the
98
+ `hint`. The one exception is `run --matrix`, which reports it on the cell and exits 1, as for any
99
+ failed cell. `vm` adds a case of its own: the natural thing to pass is the `cowork-vm-<hash>` name
100
+ `vm status` prints, so a VM-name argument says so and names the baseline(s) that derive that VM on
101
+ this machine. A VM name is still not accepted as the argument: several baselines can share one VM,
102
+ and the name depends on local install paths. The `vm` usage lines now say `<baseline>` means a
103
+ baseline, defaulting to `latest`.
104
+ - **A `--decider-cmd` helper no longer leaves processes running after the harness kills it.** The helper
105
+ runs under `shell: true`, so the harness held the shell's pid and killed only the shell, on a timeout
106
+ and on close. Anything the shell had started kept running. On Linux, where `/bin/sh` is dash, that
107
+ includes a plain `sleep 30`; on any shell it includes work the helper puts in the background. On POSIX
108
+ the helper now runs in its own process group, and a timeout, `close()`, a normal exit and Ctrl-C/SIGTERM
109
+ all kill the whole group. The helper's own process group means Ctrl-C from the terminal no longer
110
+ reaches it directly, so the harness forwards it. Windows is unchanged. New tests in
111
+ `test/decider-cmd-process-group.test.ts` fail on the old code on both macOS and Linux (dash). Under
112
+ vitest 5, the old code also made the pool print `Timeout terminating forks worker` for
113
+ `test/decider-extra.test.ts`; that warning is gone.
114
+ - **CI's live `scenario suite` no longer reports a pass when it ran nothing.** Without an
115
+ `ANTHROPIC_API_KEY` repository secret it used to run with every real step skipped and show **success**.
116
+ Now a small `live-key` job checks for the key, and without it the whole `scenario suite` job is
117
+ **skipped**, so it shows as skipped. The warning and the run-summary marker move to `live-key`.
118
+ Merges and releases are unaffected: the job is not a required check, and a skipped job leaves the
119
+ `ci.yml` run conclusion `success`. `docs/ci.md` documents the pattern for consumers, including the
120
+ caveat that GitHub counts a skipped job as passing a required-status rule. A structural test in
121
+ `test/workflow-structure.test.ts` fails if the live suite is ever gated at step level again, or made
122
+ reachable from a required check. **Known gap, now documented:** adding the key alone would not make
123
+ the suite run. The job never stages the agent binary, and a run with the key path forced on failed its
124
+ first scenario on the missing binary. `RELEASING.md` now says so, instead of "set the secret to run it".
125
+ - **`sync` admits Desktop 2.9939.2's asar without disarming its guards.** It would otherwise have
126
+ refused the healthy build:
127
+ - **Egress:** the resolver now passes the session allowlist through a HIPAA filter. For a
128
+ HIPAA-restricted org whose list holds `*`, the filter drops `*` and appends four fixed hosts; any
129
+ other list is returned unchanged. The fall-through check accepts that wrapper only after resolving
130
+ it and confirming its first statement returns the list unchanged unless both conditions hold. A
131
+ wrapper that adds hosts unconditionally, or that drops either condition, is still an unknown delta.
132
+ - **Minified names containing `$`:** eight dynamic regexes interpolated a captured name without
133
+ escaping it. In a regex, `$` is an end-of-input anchor, so the lookup could never match.
134
+ - Six sites refused a healthy build. The sub-agent trailing sentence is one: Desktop named it `$D`.
135
+ - Two sites could never fire at all: the un-awaited-async checks on the path hook's pre-pass and
136
+ chain links.
137
+
138
+ All eight now escape the name, and a new test fails on any unescaped interpolation into a dynamic
139
+ RegExp in `src/sync/`.
140
+ - **`CLAUDE_CODE_DISABLE_FAST_MODE`** is new in the 3p-only spawn branch. It is allowlisted, not
141
+ pinned, and a default first-party session never receives it.
142
+ - **`network.$comment`** used to be copied forward from the previous baseline. It is now generated in
143
+ code, and it describes the HIPAA filter instead of saying the OTLP endpoint is the only host the
144
+ bundle adds.
145
+ - **`sync` dated the Desktop install from the wrong clock.** It skips session logs written before the
146
+ synced Desktop was installed, and it took that time from `app.asar`'s mtime. The updater preserves the
147
+ packaged file's timestamps, so that mtime is when the release was built. On Desktop 2.9939.2 it read
148
+ 17:40 on the release day, while the install and first launch were at 00:37 the next day. Any session in
149
+ between would have been attributed to the new release whenever the two releases share an agent
150
+ version, which 22 of 36 committed baselines do. `sync` now uses the file's ctime, which the kernel
151
+ sets when the file is renamed into place.
152
+ - **`subagent-manifest-probe` tests what the sub-agent manifest says now.** Since Desktop 2.7032.0 the
153
+ manifest tells a sub-agent to pass absolute paths to the file tools; it no longer says where a
154
+ relative path resolves. The probe still graded a relative write, and it passed without one. It now
155
+ asserts that a sub-agent write reaches the outputs folder through its host path (`subagent_file_write`
156
+ with a `/`-anchored suffix) and that no file tool was sent a `/sessions/` path (`no_vm_path_file_op`).
157
+ Both assertions go red on a relative, wrong-folder or VM-path write, which the old pair did not. The
158
+ prompt now asks for the outputs folder by name. It passes live against both `desktop-2.7032.0` and
159
+ `desktop-2.9939.2`.
160
+ - **A committed-cassette guard contributed no tests in a fresh checkout.** The tool-assertion
161
+ satisfiability test listed the gitignored `cassettes/` directory by hand. In a fresh checkout it threw
162
+ at module load and ran nothing. In CI it passed only because an earlier test in the same worker had
163
+ created that directory and left it behind, and on a developer machine it scanned private, uncommitted
164
+ recordings. It now takes the committed cassettes from git, and the test that created the directory
165
+ removes it again.
166
+ - **`check:versions` accepts "1 baseline has shipped since"** in the rootfs-manifest lag clause, as its
167
+ `DESIGN.md` check already did. Its error message now suggests wording that passes: the old
168
+ suggestion, "1 baseline(s) have…", ended the citation match at the `)` in "(s)", so following it could
169
+ never pass.
170
+
171
+ ### Changed
172
+
173
+ - **Dependencies.** Runtime: `zod` 4.6.2 → 4.6.5 (patch, via the lockfile). Dev toolchain: `vitest` 4 → 5,
174
+ plus patch bumps to `prettier`, `tsx`, `jsdom` and `@types/node`. Under vitest 5 the old decider code
175
+ printed a pool warning; the process-group fix above removes its cause.
176
+
177
+ ### Documentation
178
+
179
+ - `docs/fidelity-gaps.md` said `save_skill`, being ToolSearch-deferred, does not appear in
180
+ `system/init.tools`. Only its schema is deferred: real init frames list it by name.
181
+ - `docs/maintenance.md` documents reading `provenance.desktopInitSurface`: which sessions count, the
182
+ unobserved remedy, how to read a `toolsAll`/`toolsSome` move, and what it records about the operator.
183
+
184
+ ## [3.8.1] — 2026-09-24
185
+
186
+ ### Upgrade notes
187
+
188
+ - **Cassettes: no re-record needed.** Nothing in this release changes what a run declares or spawns for
189
+ any committed baseline: the proactive-mode fix resolves to the same value on every baseline it touches
190
+ (all from 1.46388.3 on already carry the gate on), and the other edits under `src/runtime`,
191
+ `src/hostloop`, `src/session.ts` and `baselines/` are comments, schema description text and one
192
+ baseline annotation string. `verify-cassettes examples/replays` reports all three committed cassettes
193
+ clean.
194
+ - **Live-validated against `desktop-2.7032.0`**, which 3.8.0 shipped without. On agent 2.1.280, all four
195
+ tiers: `boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
196
+ `run examples/scenarios/` 6/7. The `test:live` red was model variance in `live-matrix` (the model asked
197
+ in plain text instead of calling `AskUserQuestion`); the file passed on two re-runs. The seventh example,
198
+ `subagent-manifest-probe`, stopped when the operator account hit its usage limit and was not re-run.
199
+ Details in `DESIGN.md`'s scope note.
200
+
201
+ ### Fixed
202
+
203
+ - **Proactive `suggest_skills` mode follows the Desktop version, not a gate Desktop no longer reads.**
204
+ From Desktop 1.46388.3 the asar has no reference to gate `1598976391` (`proactiveSkillSuggestEnabled`):
205
+ `suggest_skills` always carries the proactive description and `trigger` param when it is declared. The
206
+ server still serves the gate, and the harness read it for every baseline, so a server-side flip to off
207
+ would have switched proactive mode off here while production kept it on. For a baseline at 1.46388.3 or
208
+ later the row is now ignored and proactive mode is on; older baselines still read the gate, and
209
+ `skills.proactive_suggest_enabled` still overrides both. No change for any committed baseline: all of
210
+ them from 1.46388.3 on carry the gate on.
211
+
212
+ ### Documentation
213
+
214
+ - **`docs/fidelity-gaps.md` — what the plugin-MCP shadow file carries depends on enforcement.** The page
215
+ said the whole rewritten server set goes into `cowork-plugin-mcp-shadow.json`. Read in the 2.7032.0
216
+ asar, that holds only when Desktop enforces remote shadowing (gate `2529235968` on and no enterprise
217
+ managed-configuration override); otherwise the file names only the remote servers a stand-in replaced,
218
+ policy stubs travel in the in-process SDK server map alone, and a failed write is logged rather than
219
+ refusing the session. The section also states that any failure building the stubs refuses the session
220
+ while an MCP policy is active, that a stub never appears as a `LocalMcpServerManager` connection, and
221
+ which `main.log` lines record the remote arm.
222
+ - **Two pinned gate rows are records, not sentinels.** `canSaveSkill` (`3246569822`) has no reference in
223
+ any Desktop asar from 1.44121.1 on, and `proactiveSkillSuggestEnabled` (`1598976391`) none from
224
+ 1.46388.3 on; the server still serves both, so `sync` keeps recording them. The gates `$comment` in the
225
+ 2.7032.0 baseline, which `sync` carries forward, says so, and a flip of either row changes no run
226
+ against a current baseline.
227
+ - **The unmodeled proactive skills-prompt line applies to every current session.** From Desktop 1.46388.3,
228
+ Desktop's generated `<skills_instructions>` block always carries its proactive suggestion guidance when
229
+ `suggest_skills` and `search_plugins` are available; the harness renders no such block, and
230
+ `docs/fidelity-gaps.md` states the gap at that scope. `skills.proactive_suggest_enabled: false` on such
231
+ a baseline builds a non-proactive `suggest_skills` that production does not ship there —
232
+ `docs/session.md`, the session schema's description and the companion skill's schema reference say so.
233
+
9
234
  ## [3.8.0] — 2026-09-22
10
235
 
11
236
  ### Upgrade notes
package/DESIGN.md CHANGED
@@ -52,7 +52,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
52
52
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
53
53
 
54
54
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
55
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.280**, per `baselines/desktop-2.7032.0.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
55
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.281**, per `baselines/desktop-2.9939.2.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
56
56
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
57
57
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
58
58
 
@@ -203,7 +203,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
203
203
  > Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
204
204
  > 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
205
205
 
206
- > **Scope of that claim.** `2026-09-20 / desktop-2.2553.1` is the baseline carrying the latest live pass, and it was the newest committed baseline when that pass ran. **One baseline has shipped since (`2.7032.0`), one of which moved the agent ELF most recently to **2.1.280**.** **No live pass has been run against that newer baseline — it was skipped, not blocked.** The agent it pins (2.1.280) IS the staged one, so nothing prevented it; the decision was to ship the sync without re-running the suites. What the baseline rests on instead: a sync with zero unknown deltas, and a prompt-text change derived from the generator source of both asars rather than from a run. The pass below ran against the then-staged agent 2.1.275 (its VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.7.0**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 tests — 19 passed, 0 failed, 0 skipped**; the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **7/7 success** on its first run (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real VM with its own kernel. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` OAuth token for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. **One fixture was repaired mid-pass, and the repair is the interesting part.** `smoke-semantic-evidence-files` failed twice with the same reasoning — its deliverable asked the agent to write a factual claim about a migration that never happened, and current models refuse to ship that unflagged. Fixture rot, not a regression: no commit had touched the file since v3.6.0 and its pass rate had already decayed to 56% over 9 runs. The payload now reports what the run itself did. Its falsifiability was re-verified rather than assumed — the unscoped twin still refuses with `evidence_incomplete` naming the dropped files, which is what makes the scoped form's green mean anything. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise — and the one red here was re-run, reproduced, and traced to the fixture rather than to the code. (c) The first `examples/scenarios/` run is the one counted; there was no second. Separately and not a live matter: all three committed cassettes in `examples/replays/` remain those re-recorded against `desktop-2.2553.1` on 2026-09-18; `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
206
+ > **Scope of that claim.** `2026-09-25 / desktop-2.9939.2` is the baseline carrying the latest live pass, run against agent **2.1.281**, and it is the newest committed baseline — no baselines have shipped since. The pass ran against the staged agent (the VM ELF `claude-code-vm/2.1.281`, sha256 matching the baseline per `doctor`) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.8.1** (main `0be9dd8`), from a fresh worktree. The protocol tier runs the host `claude` CLI by design, not the staged agent; on this pass that CLI was **2.1.282**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **5 files, 20 tests — 19 passed, 1 failed, 0 skipped** on the first run (the collection was audited beforehand with `vitest list`: live-contract 14, live-matrix 2, live-outputs-delete 2, live-resume-continuity 1, live-stop-hook 1); the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **6/7 success** on its first run (csv-fx-normalize, csv-metrics, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe — the last at 9/9 assertions with one sub-agent). Every run's `result.json` was read from disk: all assertions evaluated, none skipped. **The two non-greens, stated rather than rounded away.** (1) The `test:live` red is `live-matrix`'s second cell (protocol tier, `desktop-1.18286.0`): the model asked its A-or-B question in plain text instead of calling `AskUserQuestion`, so the run ended on an unanswered question. It is the same red as in the previous pass, and it cannot come from this baseline, because the protocol tier runs the host CLI (2.1.282), not the staged agent. The file passed on a re-run (2/2). (2) `example-pdf-skill` (container) failed one of five assertions on the counted run — "user-visible artifact not found: `project/outputs/actions.md`". The agent wrote the right content one directory too high, into the connected folder, having read the prompt's relative `outputs/actions.md` as "the connected folder is the outputs folder". It passed on a separate re-run (5/5). Both are model variance, not a 2.9939.2 regression. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real Apple-VZ VM with its own kernel, created for the run and deleted after (unpinned, so it ran on `claude-sonnet-5`). Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` credentials for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. Spend: **$4.05**, summed from all 35 `result.json` files; a floor, since a live-contract test that spawns the agent without writing a harness run dir is not counted. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (c) The first `examples/scenarios/` run is the one counted; there was no second. (d) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all three committed cassettes in `examples/replays/` were re-recorded against `desktop-2.9939.2` on 2026-09-25, each at its own tier (`protocol`, `container`, `hostloop`); `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). The `protocol` cassette records host CLI 2.1.282 for the reason above. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
207
207
 
208
208
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
209
209
 
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^3.8.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^3.9.0"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.8.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.8.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.9.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.9.0"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Not sure a harness is what you need? The next two sections are the argument.
@@ -407,6 +407,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
407
407
  ## Status
408
408
 
409
409
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
410
- **`desktop-2.7032.0`**. Release-by-release verification notes (what was re-verified against
410
+ **`desktop-2.9939.2`**. Release-by-release verification notes (what was re-verified against
411
411
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
412
412
  this section would otherwise duplicate lives in the sections above.
package/RELEASING.md CHANGED
@@ -9,15 +9,27 @@ requires an OTP and is not how this repo ships.
9
9
  ## The live scenario suite is best-effort (not a publish gate)
10
10
 
11
11
  The live scenario suite (the `scenarios` job in `ci.yml`) runs live inference only when
12
- `ANTHROPIC_API_KEY` is available to the runner. Without the key it **soft-skips (green)** on every
13
- event — pushes to `main` included — emitting a loud `⚠️ NOT live-validated` marker in the run summary.
14
- It does **not** block the run and is **not** a publish gate: `release.yml`'s `require-ci-success` still
15
- requires the `ci.yml` run for the tagged commit to be green, but a green run does not by itself prove
16
- the scenarios were validated against a real model.
17
-
18
- To actually run the live suite in CI, set the `ANTHROPIC_API_KEY` repo secret. There is no
19
- `SKIP_LIVE_SCENARIOS` override — the suite never hard-fails on a missing key, so there is nothing to
20
- override.
12
+ `ANTHROPIC_API_KEY` is available to the runner. Without the key the whole job is **skipped**, on every
13
+ event, pushes to `main` included. It shows as *skipped*, not green. The small `live-key` job before it
14
+ makes that decision and carries the `⚠️ NOT live-validated` warning and run-summary marker.
15
+
16
+ It is **not** a publish gate. `release.yml`'s `require-ci-success` requires the `ci.yml` run for the
17
+ tagged commit to conclude `success`, and a skipped job leaves the run `success`. So a release can still
18
+ ship without live CI validation. The difference is that the check no longer pretends otherwise. (A live
19
+ run that actually executes and FAILS does make the run `failure`, which blocks the release.)
20
+
21
+ **Setting the `ANTHROPIC_API_KEY` repo secret is NOT enough to run the live suite in CI.** The
22
+ `scenarios` job never stages the agent binary on the runner. A run with the key path forced on
23
+ (2026-09-25, no real key) built the image and then failed its first scenario with `Staged agent binary
24
+ not found at …/claude-code-vm/<ver>/claude`, before inference was reached. Set the key only together with
25
+ a step that stages and sha256-verifies the agent ELF (see the self-hosted example in
26
+ [docs/ci.md](./docs/ci.md)). Otherwise the job turns red, the `ci.yml` run concludes `failure`, and
27
+ `require-ci-success` blocks the release. There is no `SKIP_LIVE_SCENARIOS` override, because the suite
28
+ never hard-fails on a missing key.
29
+
30
+ **Do not add `scenario suite` to the branch ruleset's required checks without setting the key first.**
31
+ GitHub counts a job skipped by a conditional as **passing** a required-status rule, so a required but
32
+ key-less live job would satisfy the rule while validating nothing.
21
33
 
22
34
  > **Ran a live pass? Re-stamp `DESIGN.md`'s "Scope of that claim" note — it is the single authority for
23
35
  > the live pin,** naming the baseline, the agent version, which suites and which tiers. Nothing enforces
@@ -37,7 +49,7 @@ branch and opening a PR lets CI prove the exact SHA before anything lands on `ma
37
49
  Phase 1: git checkout -b release/X.Y.Z
38
50
  git push origin release/X.Y.Z
39
51
  gh pr create --base main --head release/X.Y.Z --title "release: X.Y.Z"
40
- # CI runs on the PR. The live scenario stage soft-skips whenever ANTHROPIC_API_KEY is
52
+ # CI runs on the PR. The live scenario job is SKIPPED whenever ANTHROPIC_API_KEY is
41
53
  # unavailable — which is the case today; see "best-effort (not a publish gate)" above.
42
54
  ↓ CI passes
43
55
  Phase 2: gh pr merge <number> --merge (or merge via GitHub UI)
@@ -200,8 +212,10 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
200
212
  could never change its behaviour. The CHANGELOG is the release record.)
201
213
  - [ ] `npm run preflight` — local pre-release gate (`check:versions`, CHANGELOG heading present + non-empty,
202
214
  tag `vX.Y.Z` not already used, clean tree; warns if the `ANTHROPIC_API_KEY` repo secret is missing so
203
- the push-to-main live suite will soft-skip and this release won't be live-validated in CI; warns if a
204
- ruleset **required status check** names no job in `ci.yml`).
215
+ the push-to-main live suite will be skipped and this release won't be live-validated in CI; warns if a
216
+ ruleset **required status check** names no job in `ci.yml`; fails if the newest baseline's
217
+ `provenance.desktopInitSurface` is unobserved — start one Cowork session and re-run `sync`, or pass
218
+ `--allow-unobserved-init-surface` for an emergency release).
205
219
  - [ ] `npm run format:check` — fix any issues (`npm run format:write`).
206
220
  A format failure is the most common first-pass CI red.
207
221
  - [ ] `npx tsc -p tsconfig.test.json --noEmit` — typecheck including tests.
@@ -314,13 +328,13 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
314
328
  - Planning notes belong in a gitignored location excluded from the npm tarball; never commit or publish them.
315
329
  - If the tag was placed on the wrong commit (e.g. a follow-up fix was needed), delete the local tag
316
330
  (`git tag -d vX.Y.Z`), re-create it on the correct commit, and push it.
317
- - The live `scenario suite` CI stage is skipped on **fork** PRs and, independently, soft-skips whenever
318
- `ANTHROPIC_API_KEY` is unset — logging `ANTHROPIC_API_KEY not set — skipping live scenario suite` and
319
- exiting 0. **Observed 2026-08-06 on PR #104/#105: the key was not available and the suite skipped** —
320
- the job log carries `##[warning]ANTHROPIC_API_KEY not set`, which is the authoritative evidence
321
- (`gh secret list` is also empty, but it sees only repo-level Actions secrets, so absence there alone
322
- would not prove it). A green check on that job is therefore NOT evidence of live validation —
323
- `ci.yml` prints that warning in the job summary itself, and `npm run preflight` raises its own
324
- `live-suite key reminder` WARN for the same reason.
331
+ - The live `scenario suite` CI job is skipped on **fork** PRs and, independently, whenever
332
+ `ANTHROPIC_API_KEY` is unset. In both cases the check shows as **skipped**, never green. The
333
+ `live-key` job logs `##[warning]ANTHROPIC_API_KEY not set — the live scenario suite is SKIPPED`,
334
+ which is the authoritative evidence that the key was absent (`gh secret list` sees only repo-level
335
+ Actions secrets, so absence there alone would not prove it). Until 2026-09 the job instead ran with
336
+ every real step skipped and reported **success**; a green `scenario suite` check from before then is
337
+ NOT evidence of live validation. `npm run preflight` raises its own `live-suite key reminder` WARN for
338
+ the same reason.
325
339
  Re-read this bullet if a key is ever added; the `build` + `test` + `image-recipe` + `boundary` stages
326
340
  are what actually gate a release today.