cowork-harness 3.8.0 → 3.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +6 -6
- package/.claude/skills/cowork-harness/references/ci-recipe.md +7 -7
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +4 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +225 -0
- package/DESIGN.md +2 -2
- package/README.md +4 -4
- package/RELEASING.md +34 -20
- package/baselines/desktop-2.7032.0.json +42 -2
- package/baselines/desktop-2.9939.2.json +1116 -0
- package/baselines/provisioning/rootfs-provisioning.json +2 -2
- package/dist/baseline.js +42 -9
- package/dist/cli.js +27 -10
- package/dist/decide/external-channel.js +84 -15
- package/dist/errors.js +14 -0
- package/dist/loop-decision.js +22 -1
- package/dist/run/cassette.js +5 -1
- package/dist/runtime/vm-baseline-arg.js +52 -0
- package/dist/session.js +4 -3
- package/dist/sync/baseline-diff.js +129 -12
- package/dist/sync/cowork-sync.js +138 -29
- package/dist/sync/desktop-init-surface.js +206 -0
- package/dist/types.js +49 -0
- package/docs/ci.md +33 -3
- package/docs/cli.md +6 -6
- package/docs/companion-skill.md +2 -2
- package/docs/fidelity-gaps.md +52 -18
- package/docs/invariants.md +1 -0
- package/docs/maintenance.md +26 -1
- package/docs/scenario.md +8 -1
- package/docs/session.md +5 -5
- package/docs/subagents.md +4 -2
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +51 -51
- package/examples/replays/example-pdf-skill.cassette.json +96 -126
- package/examples/replays/hostloop-computer-links.cassette.json +57 -57
- package/examples/scenarios/subagent-manifest-probe.yaml +23 -12
- package/package.json +3 -3
- package/schema/session.schema.json +1 -1
- package/scripts/check-versions.ts +4 -2
- package/scripts/release-preflight.ts +38 -3
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.9.0
|
|
7
|
+
tracks-harness: cowork-harness 3.9.0 (baseline desktop-2.9939.2)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
29
|
-
> `desktop-2.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.9.0` (baseline
|
|
29
|
+
> `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
32
32
|
## Preflight — make sure the harness can actually run
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.9.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.9.0"`. **Pin `@^3.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -569,7 +569,7 @@ Recognize these before "fixing" a non-bug:
|
|
|
569
569
|
rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
|
|
570
570
|
`markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
|
|
571
571
|
(`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
|
|
572
|
-
Desktop `2.
|
|
572
|
+
Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
|
|
573
573
|
behind that sentence). The message says so ("likely a FALSE
|
|
574
574
|
NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
|
|
575
575
|
`COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.9.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -36,7 +36,7 @@ jobs:
|
|
|
36
36
|
- uses: actions/checkout@v4
|
|
37
37
|
- name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
|
|
38
38
|
run: |
|
|
39
|
-
V=2.1.
|
|
39
|
+
V=2.1.281 # match your scenario's pinned baseline's agentVersion
|
|
40
40
|
# The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
|
|
41
41
|
# served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
|
|
42
42
|
# (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
|
|
@@ -46,7 +46,7 @@ jobs:
|
|
|
46
46
|
# from your pinned baseline's agentBinary.releaseBaseUrl. The checksum step fails closed if you
|
|
47
47
|
# don't, but it cannot tell you why. Baselines written before that field existed were
|
|
48
48
|
# stable-staged, so their base is the plain https://downloads.claude.ai/claude-code-releases.
|
|
49
|
-
B=https://downloads.claude.ai/claude-code-releases
|
|
49
|
+
B=https://downloads.claude.ai/claude-code-releases
|
|
50
50
|
# The expected digest is baselines/desktop-<ver>.json -> agentBinary.sha256. Paste it here, or
|
|
51
51
|
# read it with jq if you vendor the baseline. An unverified download is an unverified agent:
|
|
52
52
|
# this step FAILS rather than staging one, which is the whole point of naming it "verified".
|
|
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
77
77
|
GitHub-hosted runners, no token/Docker/agent:
|
|
78
78
|
|
|
79
79
|
```yaml
|
|
80
|
-
- run: npm i -g "cowork-harness@^3.
|
|
80
|
+
- run: npm i -g "cowork-harness@^3.9.0"
|
|
81
81
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
82
82
|
# no silent false-greens. WITHOUT --strict this
|
|
83
83
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -364,7 +364,7 @@ jobs:
|
|
|
364
364
|
with: { node-version: '24' }
|
|
365
365
|
- uses: actions/setup-python@v5
|
|
366
366
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
367
|
-
- run: npm i -g "cowork-harness@^3.
|
|
367
|
+
- run: npm i -g "cowork-harness@^3.9.0"
|
|
368
368
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
369
369
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
370
370
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -393,7 +393,7 @@ jobs:
|
|
|
393
393
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
394
394
|
fi
|
|
395
395
|
- if: steps.guard.outputs.live == 'true'
|
|
396
|
-
run: npm i -g "cowork-harness@^3.
|
|
396
|
+
run: npm i -g "cowork-harness@^3.9.0"
|
|
397
397
|
- if: steps.guard.outputs.live == 'true'
|
|
398
398
|
run: cowork-harness run scenarios/ --output-format json
|
|
399
399
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -68,7 +68,9 @@ cowork-harness vm prune # remove orphaned cowork-vm-* VMs from past configs
|
|
|
68
68
|
|
|
69
69
|
The instance is `cowork-vm-<config-hash>` — a config or agent-version change yields a new name, so a
|
|
70
70
|
stale VM is never silently reused (the old one is orphaned until `vm prune`). Pin a fixed name with
|
|
71
|
-
`COWORK_LIMA_INSTANCE`.
|
|
71
|
+
`COWORK_LIMA_INSTANCE`. The optional argument to every `vm` subcommand is a **baseline**
|
|
72
|
+
(`desktop-<version>`, default `latest`), never the `cowork-vm-<hash>` VM name — passing a VM name is a
|
|
73
|
+
usage error that names the baseline(s) deriving it.
|
|
72
74
|
|
|
73
75
|
## Answer paths (resolving gates: AskUserQuestion + tool-permission)
|
|
74
76
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
4
|
-
(baseline `desktop-2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.9.0`
|
|
4
|
+
(baseline `desktop-2.9939.2`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -192,7 +192,7 @@ plugins:
|
|
|
192
192
|
skills:
|
|
193
193
|
local: [] # extra host skill dirs
|
|
194
194
|
suggest_enabled: true # gate 245679952 override — `mcp__skills__suggest_skills` on/off (default true)
|
|
195
|
-
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset =
|
|
195
|
+
proactive_suggest_enabled: false # gate 1598976391 override — proactive description + `trigger` param; unset = always on from the 1.46388.3 baseline (where `false` models a surface production does not ship there), synced gate before it
|
|
196
196
|
mcp:
|
|
197
197
|
config: null # --mcp-config file (standard mcpServers map)
|
|
198
198
|
enabled: []
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,231 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.9.0] — 2026-09-25
|
|
10
|
+
|
|
11
|
+
### Upgrade notes
|
|
12
|
+
|
|
13
|
+
- **Your own cassettes go stale against `latest`.** `baseline: latest` now resolves to
|
|
14
|
+
`desktop-2.9939.2`, so a cassette recorded against `desktop-2.7032.0` reports `[stale] baseline moved`
|
|
15
|
+
in `verify-cassettes`. From a first-party session's point of view the contract did not move: the
|
|
16
|
+
spawn environment, the Cowork system prompt, the sub-agent append and the egress allowlist are all
|
|
17
|
+
byte-identical. So re-stamping `fingerprint.baseline` is sound. Re-record instead if you want the
|
|
18
|
+
recording itself to come from agent 2.1.281.
|
|
19
|
+
- **The three committed cassettes are re-recorded against `desktop-2.9939.2`**, each at its own tier:
|
|
20
|
+
- `example-multiselect-gate` at `protocol`, on the sealed managed config dir and the same
|
|
21
|
+
`claude-opus-5-5[1m]` model. It records host CLI 2.1.282, since the protocol tier runs the host CLI
|
|
22
|
+
by design.
|
|
23
|
+
- `example-pdf-skill` at `container`, agent 2.1.281.
|
|
24
|
+
- `hostloop-computer-links` at `hostloop`, agent 2.1.281. It was first recorded outside the repository
|
|
25
|
+
and its captured inventory compared: the same 4 product MCP servers, agents and skills. The only
|
|
26
|
+
additions (the `focus` command and the built-in `agents-md` plugin) come from the new agent build and
|
|
27
|
+
appear in the sealed recordings too.
|
|
28
|
+
|
|
29
|
+
`verify-cassettes` is clean and all three replay green.
|
|
30
|
+
- **Live-validated against `desktop-2.9939.2`** (agent 2.1.281) on 2026-09-25, all four tiers:
|
|
31
|
+
`boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
|
|
32
|
+
`run examples/scenarios/` 6/7 on its first run. Both reds were model variance and passed on re-runs:
|
|
33
|
+
- `live-matrix` asked its A-or-B question in plain text instead of calling `AskUserQuestion`. It runs
|
|
34
|
+
at the protocol tier, which uses the host `claude` CLI (2.1.282 here), not the staged agent.
|
|
35
|
+
- `example-pdf-skill` wrote its file one directory too high.
|
|
36
|
+
|
|
37
|
+
Details are in `DESIGN.md`'s scope note. Code that landed after that pass (the decider process-group
|
|
38
|
+
fix, the baseline-name usage errors and a `zod` patch) was re-verified live on the release code
|
|
39
|
+
(`f395d07`), again across all four tiers:
|
|
40
|
+
- e2e self-tests 8/8, and `run examples/scenarios/` 7/7. `subagent-manifest-probe` first went 8/9,
|
|
41
|
+
because the sub-agent's first write used a VM path, which the hook refused. It passed 9/9 on a
|
|
42
|
+
re-run.
|
|
43
|
+
- `smoke-l2-microvm` passed.
|
|
44
|
+
- A VM-name argument to `vm` is a usage error (exit 2, category `usage`).
|
|
45
|
+
- A `--decider-cmd` timeout leaves no surviving process group.
|
|
46
|
+
- Ctrl-C at the container tier exits 130 and reaps the containers, the proxy and the network.
|
|
47
|
+
|
|
48
|
+
CI does not live-validate: without an `ANTHROPIC_API_KEY` secret its live scenario suite is skipped
|
|
49
|
+
(see Fixed).
|
|
50
|
+
|
|
51
|
+
### Added
|
|
52
|
+
|
|
53
|
+
- **`sync` records the tool surface real Cowork sessions declared** for Desktop's own `cowork`, `plugins`
|
|
54
|
+
and `skills` servers, as `provenance.desktopInitSurface`. A production change such as `save_skill`
|
|
55
|
+
turning on or off then shows up in `sync --diff`.
|
|
56
|
+
- It is read from the synced Desktop's own session logs, limited to that release's agent version and
|
|
57
|
+
install time.
|
|
58
|
+
- No other server name is recorded, and a strict schema over every committed baseline enforces that.
|
|
59
|
+
- When no Cowork session has run since the install, `sync` records `observed: false` with a warning,
|
|
60
|
+
instead of reusing the previous release's surface.
|
|
61
|
+
- `desktop-2.7032.0` carries the first recorded block (`save_skill` declared in every session read).
|
|
62
|
+
- **`npm run preflight` refuses to release an unobserved Desktop init surface.** New check: the newest
|
|
63
|
+
baseline's `desktopInitSurface` must be observed. `--allow-unobserved-init-surface` downgrades it to a
|
|
64
|
+
warning for an emergency release; `--allow-empty` does not.
|
|
65
|
+
|
|
66
|
+
### Parity — Desktop 2.9939.2 (agent 2.1.281)
|
|
67
|
+
|
|
68
|
+
- **Baseline `desktop-2.9939.2`**, with the agent staged from the **stable** release channel, not an RC
|
|
69
|
+
build. This is what `baseline: latest` now resolves to. From a first-party session's point of view
|
|
70
|
+
nothing moved:
|
|
71
|
+
- The Cowork system prompt and all four sub-agent append fingerprints are byte-identical to 2.7032.0.
|
|
72
|
+
- `spawn.env` (24 keys) and the egress allowlist are unchanged.
|
|
73
|
+
- The VM rootfs is re-captured at `2.9939.2` and its provisioning is identical: node `v22.23.2`, the
|
|
74
|
+
same 136 pip packages, the same apt doc stack and the same npm globals (`tsx` included).
|
|
75
|
+
- The recorded changes: `spawnEnvKeys` gains the 3p-only `CLAUDE_CODE_DISABLE_FAST_MODE`, and
|
|
76
|
+
`asarGateIds` gains 28 ids and loses 1.
|
|
77
|
+
- `network.$comment` now describes the resolver's HIPAA filter (see Fixed).
|
|
78
|
+
- Gate `4202409342` (`builtinToolsApprovableByAutoMode`) is still on. It is now served by default
|
|
79
|
+
rather than forced, so a server rule can flip it. The harness does not read it, because auto mode
|
|
80
|
+
is unreachable here.
|
|
81
|
+
- `desktopInitSurface` is observed from **2 init frames, both from a scheduled task**. The four
|
|
82
|
+
artifact tools now appear in every frame read rather than some. Treat that as a property of those
|
|
83
|
+
two sessions, not as a Desktop change.
|
|
84
|
+
|
|
85
|
+
### Fixed
|
|
86
|
+
|
|
87
|
+
- **A baseline name that names no baseline is now a usage error at every entry point, not a stack
|
|
88
|
+
trace.** The baseline loader read the file with no existence check, so an unknown name escaped as an
|
|
89
|
+
uncaught `ENOENT`. How that surfaced depended on the entry point:
|
|
90
|
+
- `boundary-check <name>`, `vm init|status|delete|prune <name>`, and `run` on a scenario whose
|
|
91
|
+
`baseline:` names nothing: category `internal` in JSON mode, a raw stack trace in text mode.
|
|
92
|
+
- `record`: the `ENOENT` text, with exit 1.
|
|
93
|
+
- `diff <a> <b>`: exit 2, but the raw `ENOENT` text as the message.
|
|
94
|
+
- `run --matrix` with such a `baselines:` entry: the `ENOENT` text as the cell's error.
|
|
95
|
+
|
|
96
|
+
The loader now throws a usage error for a name or path that names no baseline. Every entry point
|
|
97
|
+
reports it as category `usage`, exit 2, with the committed baselines listed newest first in the
|
|
98
|
+
`hint`. The one exception is `run --matrix`, which reports it on the cell and exits 1, as for any
|
|
99
|
+
failed cell. `vm` adds a case of its own: the natural thing to pass is the `cowork-vm-<hash>` name
|
|
100
|
+
`vm status` prints, so a VM-name argument says so and names the baseline(s) that derive that VM on
|
|
101
|
+
this machine. A VM name is still not accepted as the argument: several baselines can share one VM,
|
|
102
|
+
and the name depends on local install paths. The `vm` usage lines now say `<baseline>` means a
|
|
103
|
+
baseline, defaulting to `latest`.
|
|
104
|
+
- **A `--decider-cmd` helper no longer leaves processes running after the harness kills it.** The helper
|
|
105
|
+
runs under `shell: true`, so the harness held the shell's pid and killed only the shell, on a timeout
|
|
106
|
+
and on close. Anything the shell had started kept running. On Linux, where `/bin/sh` is dash, that
|
|
107
|
+
includes a plain `sleep 30`; on any shell it includes work the helper puts in the background. On POSIX
|
|
108
|
+
the helper now runs in its own process group, and a timeout, `close()`, a normal exit and Ctrl-C/SIGTERM
|
|
109
|
+
all kill the whole group. The helper's own process group means Ctrl-C from the terminal no longer
|
|
110
|
+
reaches it directly, so the harness forwards it. Windows is unchanged. New tests in
|
|
111
|
+
`test/decider-cmd-process-group.test.ts` fail on the old code on both macOS and Linux (dash). Under
|
|
112
|
+
vitest 5, the old code also made the pool print `Timeout terminating forks worker` for
|
|
113
|
+
`test/decider-extra.test.ts`; that warning is gone.
|
|
114
|
+
- **CI's live `scenario suite` no longer reports a pass when it ran nothing.** Without an
|
|
115
|
+
`ANTHROPIC_API_KEY` repository secret it used to run with every real step skipped and show **success**.
|
|
116
|
+
Now a small `live-key` job checks for the key, and without it the whole `scenario suite` job is
|
|
117
|
+
**skipped**, so it shows as skipped. The warning and the run-summary marker move to `live-key`.
|
|
118
|
+
Merges and releases are unaffected: the job is not a required check, and a skipped job leaves the
|
|
119
|
+
`ci.yml` run conclusion `success`. `docs/ci.md` documents the pattern for consumers, including the
|
|
120
|
+
caveat that GitHub counts a skipped job as passing a required-status rule. A structural test in
|
|
121
|
+
`test/workflow-structure.test.ts` fails if the live suite is ever gated at step level again, or made
|
|
122
|
+
reachable from a required check. **Known gap, now documented:** adding the key alone would not make
|
|
123
|
+
the suite run. The job never stages the agent binary, and a run with the key path forced on failed its
|
|
124
|
+
first scenario on the missing binary. `RELEASING.md` now says so, instead of "set the secret to run it".
|
|
125
|
+
- **`sync` admits Desktop 2.9939.2's asar without disarming its guards.** It would otherwise have
|
|
126
|
+
refused the healthy build:
|
|
127
|
+
- **Egress:** the resolver now passes the session allowlist through a HIPAA filter. For a
|
|
128
|
+
HIPAA-restricted org whose list holds `*`, the filter drops `*` and appends four fixed hosts; any
|
|
129
|
+
other list is returned unchanged. The fall-through check accepts that wrapper only after resolving
|
|
130
|
+
it and confirming its first statement returns the list unchanged unless both conditions hold. A
|
|
131
|
+
wrapper that adds hosts unconditionally, or that drops either condition, is still an unknown delta.
|
|
132
|
+
- **Minified names containing `$`:** eight dynamic regexes interpolated a captured name without
|
|
133
|
+
escaping it. In a regex, `$` is an end-of-input anchor, so the lookup could never match.
|
|
134
|
+
- Six sites refused a healthy build. The sub-agent trailing sentence is one: Desktop named it `$D`.
|
|
135
|
+
- Two sites could never fire at all: the un-awaited-async checks on the path hook's pre-pass and
|
|
136
|
+
chain links.
|
|
137
|
+
|
|
138
|
+
All eight now escape the name, and a new test fails on any unescaped interpolation into a dynamic
|
|
139
|
+
RegExp in `src/sync/`.
|
|
140
|
+
- **`CLAUDE_CODE_DISABLE_FAST_MODE`** is new in the 3p-only spawn branch. It is allowlisted, not
|
|
141
|
+
pinned, and a default first-party session never receives it.
|
|
142
|
+
- **`network.$comment`** used to be copied forward from the previous baseline. It is now generated in
|
|
143
|
+
code, and it describes the HIPAA filter instead of saying the OTLP endpoint is the only host the
|
|
144
|
+
bundle adds.
|
|
145
|
+
- **`sync` dated the Desktop install from the wrong clock.** It skips session logs written before the
|
|
146
|
+
synced Desktop was installed, and it took that time from `app.asar`'s mtime. The updater preserves the
|
|
147
|
+
packaged file's timestamps, so that mtime is when the release was built. On Desktop 2.9939.2 it read
|
|
148
|
+
17:40 on the release day, while the install and first launch were at 00:37 the next day. Any session in
|
|
149
|
+
between would have been attributed to the new release whenever the two releases share an agent
|
|
150
|
+
version, which 22 of 36 committed baselines do. `sync` now uses the file's ctime, which the kernel
|
|
151
|
+
sets when the file is renamed into place.
|
|
152
|
+
- **`subagent-manifest-probe` tests what the sub-agent manifest says now.** Since Desktop 2.7032.0 the
|
|
153
|
+
manifest tells a sub-agent to pass absolute paths to the file tools; it no longer says where a
|
|
154
|
+
relative path resolves. The probe still graded a relative write, and it passed without one. It now
|
|
155
|
+
asserts that a sub-agent write reaches the outputs folder through its host path (`subagent_file_write`
|
|
156
|
+
with a `/`-anchored suffix) and that no file tool was sent a `/sessions/` path (`no_vm_path_file_op`).
|
|
157
|
+
Both assertions go red on a relative, wrong-folder or VM-path write, which the old pair did not. The
|
|
158
|
+
prompt now asks for the outputs folder by name. It passes live against both `desktop-2.7032.0` and
|
|
159
|
+
`desktop-2.9939.2`.
|
|
160
|
+
- **A committed-cassette guard contributed no tests in a fresh checkout.** The tool-assertion
|
|
161
|
+
satisfiability test listed the gitignored `cassettes/` directory by hand. In a fresh checkout it threw
|
|
162
|
+
at module load and ran nothing. In CI it passed only because an earlier test in the same worker had
|
|
163
|
+
created that directory and left it behind, and on a developer machine it scanned private, uncommitted
|
|
164
|
+
recordings. It now takes the committed cassettes from git, and the test that created the directory
|
|
165
|
+
removes it again.
|
|
166
|
+
- **`check:versions` accepts "1 baseline has shipped since"** in the rootfs-manifest lag clause, as its
|
|
167
|
+
`DESIGN.md` check already did. Its error message now suggests wording that passes: the old
|
|
168
|
+
suggestion, "1 baseline(s) have…", ended the citation match at the `)` in "(s)", so following it could
|
|
169
|
+
never pass.
|
|
170
|
+
|
|
171
|
+
### Changed
|
|
172
|
+
|
|
173
|
+
- **Dependencies.** Runtime: `zod` 4.6.2 → 4.6.5 (patch, via the lockfile). Dev toolchain: `vitest` 4 → 5,
|
|
174
|
+
plus patch bumps to `prettier`, `tsx`, `jsdom` and `@types/node`. Under vitest 5 the old decider code
|
|
175
|
+
printed a pool warning; the process-group fix above removes its cause.
|
|
176
|
+
|
|
177
|
+
### Documentation
|
|
178
|
+
|
|
179
|
+
- `docs/fidelity-gaps.md` said `save_skill`, being ToolSearch-deferred, does not appear in
|
|
180
|
+
`system/init.tools`. Only its schema is deferred: real init frames list it by name.
|
|
181
|
+
- `docs/maintenance.md` documents reading `provenance.desktopInitSurface`: which sessions count, the
|
|
182
|
+
unobserved remedy, how to read a `toolsAll`/`toolsSome` move, and what it records about the operator.
|
|
183
|
+
|
|
184
|
+
## [3.8.1] — 2026-09-24
|
|
185
|
+
|
|
186
|
+
### Upgrade notes
|
|
187
|
+
|
|
188
|
+
- **Cassettes: no re-record needed.** Nothing in this release changes what a run declares or spawns for
|
|
189
|
+
any committed baseline: the proactive-mode fix resolves to the same value on every baseline it touches
|
|
190
|
+
(all from 1.46388.3 on already carry the gate on), and the other edits under `src/runtime`,
|
|
191
|
+
`src/hostloop`, `src/session.ts` and `baselines/` are comments, schema description text and one
|
|
192
|
+
baseline annotation string. `verify-cassettes examples/replays` reports all three committed cassettes
|
|
193
|
+
clean.
|
|
194
|
+
- **Live-validated against `desktop-2.7032.0`**, which 3.8.0 shipped without. On agent 2.1.280, all four
|
|
195
|
+
tiers: `boundary-check` 6/6, e2e self-tests 9/9, `test:live` 19/20 on the first run, and
|
|
196
|
+
`run examples/scenarios/` 6/7. The `test:live` red was model variance in `live-matrix` (the model asked
|
|
197
|
+
in plain text instead of calling `AskUserQuestion`); the file passed on two re-runs. The seventh example,
|
|
198
|
+
`subagent-manifest-probe`, stopped when the operator account hit its usage limit and was not re-run.
|
|
199
|
+
Details in `DESIGN.md`'s scope note.
|
|
200
|
+
|
|
201
|
+
### Fixed
|
|
202
|
+
|
|
203
|
+
- **Proactive `suggest_skills` mode follows the Desktop version, not a gate Desktop no longer reads.**
|
|
204
|
+
From Desktop 1.46388.3 the asar has no reference to gate `1598976391` (`proactiveSkillSuggestEnabled`):
|
|
205
|
+
`suggest_skills` always carries the proactive description and `trigger` param when it is declared. The
|
|
206
|
+
server still serves the gate, and the harness read it for every baseline, so a server-side flip to off
|
|
207
|
+
would have switched proactive mode off here while production kept it on. For a baseline at 1.46388.3 or
|
|
208
|
+
later the row is now ignored and proactive mode is on; older baselines still read the gate, and
|
|
209
|
+
`skills.proactive_suggest_enabled` still overrides both. No change for any committed baseline: all of
|
|
210
|
+
them from 1.46388.3 on carry the gate on.
|
|
211
|
+
|
|
212
|
+
### Documentation
|
|
213
|
+
|
|
214
|
+
- **`docs/fidelity-gaps.md` — what the plugin-MCP shadow file carries depends on enforcement.** The page
|
|
215
|
+
said the whole rewritten server set goes into `cowork-plugin-mcp-shadow.json`. Read in the 2.7032.0
|
|
216
|
+
asar, that holds only when Desktop enforces remote shadowing (gate `2529235968` on and no enterprise
|
|
217
|
+
managed-configuration override); otherwise the file names only the remote servers a stand-in replaced,
|
|
218
|
+
policy stubs travel in the in-process SDK server map alone, and a failed write is logged rather than
|
|
219
|
+
refusing the session. The section also states that any failure building the stubs refuses the session
|
|
220
|
+
while an MCP policy is active, that a stub never appears as a `LocalMcpServerManager` connection, and
|
|
221
|
+
which `main.log` lines record the remote arm.
|
|
222
|
+
- **Two pinned gate rows are records, not sentinels.** `canSaveSkill` (`3246569822`) has no reference in
|
|
223
|
+
any Desktop asar from 1.44121.1 on, and `proactiveSkillSuggestEnabled` (`1598976391`) none from
|
|
224
|
+
1.46388.3 on; the server still serves both, so `sync` keeps recording them. The gates `$comment` in the
|
|
225
|
+
2.7032.0 baseline, which `sync` carries forward, says so, and a flip of either row changes no run
|
|
226
|
+
against a current baseline.
|
|
227
|
+
- **The unmodeled proactive skills-prompt line applies to every current session.** From Desktop 1.46388.3,
|
|
228
|
+
Desktop's generated `<skills_instructions>` block always carries its proactive suggestion guidance when
|
|
229
|
+
`suggest_skills` and `search_plugins` are available; the harness renders no such block, and
|
|
230
|
+
`docs/fidelity-gaps.md` states the gap at that scope. `skills.proactive_suggest_enabled: false` on such
|
|
231
|
+
a baseline builds a non-proactive `suggest_skills` that production does not ship there —
|
|
232
|
+
`docs/session.md`, the session schema's description and the companion skill's schema reference say so.
|
|
233
|
+
|
|
9
234
|
## [3.8.0] — 2026-09-22
|
|
10
235
|
|
|
11
236
|
### Upgrade notes
|
package/DESIGN.md
CHANGED
|
@@ -52,7 +52,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
52
52
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
53
53
|
|
|
54
54
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
55
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
55
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.281**, per `baselines/desktop-2.9939.2.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
|
|
56
56
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
57
57
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
58
58
|
|
|
@@ -203,7 +203,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
203
203
|
> Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
|
|
204
204
|
> 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
|
|
205
205
|
|
|
206
|
-
> **Scope of that claim.** `2026-09-
|
|
206
|
+
> **Scope of that claim.** `2026-09-25 / desktop-2.9939.2` is the baseline carrying the latest live pass, run against agent **2.1.281**, and it is the newest committed baseline — no baselines have shipped since. The pass ran against the staged agent (the VM ELF `claude-code-vm/2.1.281`, sha256 matching the baseline per `doctor`) on macOS arm64, agent image `cowork-agent-base:2`, on harness **3.8.1** (main `0be9dd8`), from a fresh worktree. The protocol tier runs the host `claude` CLI by design, not the staged agent; on this pass that CLI was **2.1.282**. It covered **all four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **5 files, 20 tests — 19 passed, 1 failed, 0 skipped** on the first run (the collection was audited beforehand with `vitest list`: live-contract 14, live-matrix 2, live-outputs-delete 2, live-resume-continuity 1, live-stop-hook 1); the e2e self-tests **9/9 success** (canary-hostloop, smoke-askuserquestion, smoke-l1-container, smoke-l1-egress, smoke-l2-microvm, smoke-multiselect-deciderdir, smoke-multiselect, smoke-present-files, smoke-semantic-evidence-files); and `run examples/scenarios/` **6/7 success** on its first run (csv-fx-normalize, csv-metrics, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe — the last at 9/9 assertions with one sub-agent). Every run's `result.json` was read from disk: all assertions evaluated, none skipped. **The two non-greens, stated rather than rounded away.** (1) The `test:live` red is `live-matrix`'s second cell (protocol tier, `desktop-1.18286.0`): the model asked its A-or-B question in plain text instead of calling `AskUserQuestion`, so the run ended on an unanswered question. It is the same red as in the previous pass, and it cannot come from this baseline, because the protocol tier runs the host CLI (2.1.282), not the staged agent. The file passed on a re-run (2/2). (2) `example-pdf-skill` (container) failed one of five assertions on the counted run — "user-visible artifact not found: `project/outputs/actions.md`". The agent wrote the right content one directory too high, into the connected folder, having read the prompt's relative `outputs/actions.md` as "the connected folder is the outputs folder". It passed on a separate re-run (5/5). Both are model variance, not a 2.9939.2 regression. **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously. `smoke-multiselect-deciderdir` runs the `--decider-llm` path, and `smoke-l2-microvm` passed in a real Apple-VZ VM with its own kernel, created for the run and deleted after (unpinned, so it ran on `claude-sonnet-5`). Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`**. Both credential paths were exercised: the `.env` credentials for the container/protocol tiers, and the agent's own macOS-Keychain self-sourcing at `hostloop`. Spend: **$4.05**, summed from all 35 `result.json` files; a floor, since a live-contract test that spawns the agent without writing a harness run dir is not counted. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction. (b) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (c) The first `examples/scenarios/` run is the one counted; there was no second. (d) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all three committed cassettes in `examples/replays/` were re-recorded against `desktop-2.9939.2` on 2026-09-25, each at its own tier (`protocol`, `container`, `hostloop`); `verify-cassettes` exits 0 with all three clean and one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). The `protocol` cassette records host CLI 2.1.282 for the reason above. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
207
207
|
|
|
208
208
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
209
209
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^3.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^3.9.0"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.9.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.9.0"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Not sure a harness is what you need? The next two sections are the argument.
|
|
@@ -407,6 +407,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
407
407
|
## Status
|
|
408
408
|
|
|
409
409
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
410
|
-
**`desktop-2.
|
|
410
|
+
**`desktop-2.9939.2`**. Release-by-release verification notes (what was re-verified against
|
|
411
411
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
412
412
|
this section would otherwise duplicate lives in the sections above.
|
package/RELEASING.md
CHANGED
|
@@ -9,15 +9,27 @@ requires an OTP and is not how this repo ships.
|
|
|
9
9
|
## The live scenario suite is best-effort (not a publish gate)
|
|
10
10
|
|
|
11
11
|
The live scenario suite (the `scenarios` job in `ci.yml`) runs live inference only when
|
|
12
|
-
`ANTHROPIC_API_KEY` is available to the runner. Without the key
|
|
13
|
-
event
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
12
|
+
`ANTHROPIC_API_KEY` is available to the runner. Without the key the whole job is **skipped**, on every
|
|
13
|
+
event, pushes to `main` included. It shows as *skipped*, not green. The small `live-key` job before it
|
|
14
|
+
makes that decision and carries the `⚠️ NOT live-validated` warning and run-summary marker.
|
|
15
|
+
|
|
16
|
+
It is **not** a publish gate. `release.yml`'s `require-ci-success` requires the `ci.yml` run for the
|
|
17
|
+
tagged commit to conclude `success`, and a skipped job leaves the run `success`. So a release can still
|
|
18
|
+
ship without live CI validation. The difference is that the check no longer pretends otherwise. (A live
|
|
19
|
+
run that actually executes and FAILS does make the run `failure`, which blocks the release.)
|
|
20
|
+
|
|
21
|
+
**Setting the `ANTHROPIC_API_KEY` repo secret is NOT enough to run the live suite in CI.** The
|
|
22
|
+
`scenarios` job never stages the agent binary on the runner. A run with the key path forced on
|
|
23
|
+
(2026-09-25, no real key) built the image and then failed its first scenario with `Staged agent binary
|
|
24
|
+
not found at …/claude-code-vm/<ver>/claude`, before inference was reached. Set the key only together with
|
|
25
|
+
a step that stages and sha256-verifies the agent ELF (see the self-hosted example in
|
|
26
|
+
[docs/ci.md](./docs/ci.md)). Otherwise the job turns red, the `ci.yml` run concludes `failure`, and
|
|
27
|
+
`require-ci-success` blocks the release. There is no `SKIP_LIVE_SCENARIOS` override, because the suite
|
|
28
|
+
never hard-fails on a missing key.
|
|
29
|
+
|
|
30
|
+
**Do not add `scenario suite` to the branch ruleset's required checks without setting the key first.**
|
|
31
|
+
GitHub counts a job skipped by a conditional as **passing** a required-status rule, so a required but
|
|
32
|
+
key-less live job would satisfy the rule while validating nothing.
|
|
21
33
|
|
|
22
34
|
> **Ran a live pass? Re-stamp `DESIGN.md`'s "Scope of that claim" note — it is the single authority for
|
|
23
35
|
> the live pin,** naming the baseline, the agent version, which suites and which tiers. Nothing enforces
|
|
@@ -37,7 +49,7 @@ branch and opening a PR lets CI prove the exact SHA before anything lands on `ma
|
|
|
37
49
|
Phase 1: git checkout -b release/X.Y.Z
|
|
38
50
|
git push origin release/X.Y.Z
|
|
39
51
|
gh pr create --base main --head release/X.Y.Z --title "release: X.Y.Z"
|
|
40
|
-
# CI runs on the PR. The live scenario
|
|
52
|
+
# CI runs on the PR. The live scenario job is SKIPPED whenever ANTHROPIC_API_KEY is
|
|
41
53
|
# unavailable — which is the case today; see "best-effort (not a publish gate)" above.
|
|
42
54
|
↓ CI passes
|
|
43
55
|
Phase 2: gh pr merge <number> --merge (or merge via GitHub UI)
|
|
@@ -200,8 +212,10 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
200
212
|
could never change its behaviour. The CHANGELOG is the release record.)
|
|
201
213
|
- [ ] `npm run preflight` — local pre-release gate (`check:versions`, CHANGELOG heading present + non-empty,
|
|
202
214
|
tag `vX.Y.Z` not already used, clean tree; warns if the `ANTHROPIC_API_KEY` repo secret is missing so
|
|
203
|
-
the push-to-main live suite will
|
|
204
|
-
ruleset **required status check** names no job in `ci.yml
|
|
215
|
+
the push-to-main live suite will be skipped and this release won't be live-validated in CI; warns if a
|
|
216
|
+
ruleset **required status check** names no job in `ci.yml`; fails if the newest baseline's
|
|
217
|
+
`provenance.desktopInitSurface` is unobserved — start one Cowork session and re-run `sync`, or pass
|
|
218
|
+
`--allow-unobserved-init-surface` for an emergency release).
|
|
205
219
|
- [ ] `npm run format:check` — fix any issues (`npm run format:write`).
|
|
206
220
|
A format failure is the most common first-pass CI red.
|
|
207
221
|
- [ ] `npx tsc -p tsconfig.test.json --noEmit` — typecheck including tests.
|
|
@@ -314,13 +328,13 @@ tagging `1.0.0`, deliberately review and freeze the surfaces with no machine-rea
|
|
|
314
328
|
- Planning notes belong in a gitignored location excluded from the npm tarball; never commit or publish them.
|
|
315
329
|
- If the tag was placed on the wrong commit (e.g. a follow-up fix was needed), delete the local tag
|
|
316
330
|
(`git tag -d vX.Y.Z`), re-create it on the correct commit, and push it.
|
|
317
|
-
- The live `scenario suite` CI
|
|
318
|
-
`ANTHROPIC_API_KEY` is unset
|
|
319
|
-
|
|
320
|
-
the
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
331
|
+
- The live `scenario suite` CI job is skipped on **fork** PRs and, independently, whenever
|
|
332
|
+
`ANTHROPIC_API_KEY` is unset. In both cases the check shows as **skipped**, never green. The
|
|
333
|
+
`live-key` job logs `##[warning]ANTHROPIC_API_KEY not set — the live scenario suite is SKIPPED`,
|
|
334
|
+
which is the authoritative evidence that the key was absent (`gh secret list` sees only repo-level
|
|
335
|
+
Actions secrets, so absence there alone would not prove it). Until 2026-09 the job instead ran with
|
|
336
|
+
every real step skipped and reported **success**; a green `scenario suite` check from before then is
|
|
337
|
+
NOT evidence of live validation. `npm run preflight` raises its own `live-suite key reminder` WARN for
|
|
338
|
+
the same reason.
|
|
325
339
|
Re-read this bullet if a key is ever added; the `build` + `test` + `image-recipe` + `boundary` stages
|
|
326
340
|
are what actually gate a release today.
|