cowork-harness 1.22.0 → 1.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.22.0
7
- tracks-harness: cowork-harness 1.22.0 (baseline desktop-1.28929.0)
6
+ version: 1.23.0
7
+ tracks-harness: cowork-harness 1.23.0 (baseline desktop-1.30096.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.22.0` (baseline
26
- > `desktop-1.28929.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.23.0` (baseline
26
+ > `desktop-1.30096.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.22.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.22.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.22.0"`. **Pin `@>=1.22.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.23.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.23.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.23.0"`. **Pin `@>=1.23.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`).
3
+ Self-contained reference. Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.22.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.23.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
34
34
  run: |
35
- V=2.1.227 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.229 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -58,7 +58,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
58
58
  GitHub-hosted runners, no token/Docker/agent:
59
59
 
60
60
  ```yaml
61
- - run: npm i -g "cowork-harness@>=1.22.0"
61
+ - run: npm i -g "cowork-harness@>=1.23.0"
62
62
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
63
63
  # no silent false-greens. WITHOUT --strict this
64
64
  # step cannot fail on a WARN-class rule (e.g.
@@ -279,7 +279,7 @@ jobs:
279
279
  with: { node-version: '24' }
280
280
  - uses: actions/setup-python@v5
281
281
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
282
- - run: npm i -g "cowork-harness@>=1.22.0"
282
+ - run: npm i -g "cowork-harness@>=1.23.0"
283
283
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
284
284
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
285
285
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -308,7 +308,7 @@ jobs:
308
308
  echo "live=true" >> "$GITHUB_OUTPUT"
309
309
  fi
310
310
  - if: steps.guard.outputs.live == 'true'
311
- run: npm i -g "cowork-harness@>=1.22.0"
311
+ run: npm i -g "cowork-harness@>=1.23.0"
312
312
  - if: steps.guard.outputs.live == 'true'
313
313
  run: cowork-harness run scenarios/ --output-format json
314
314
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`).
3
+ Self-contained reference. Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.22.0`
4
- (baseline `desktop-1.28929.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.23.0`
4
+ (baseline `desktop-1.30096.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 1.23.0` (baseline `desktop-1.30096.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -69,6 +69,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](htt
69
69
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
70
70
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
71
71
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
72
+ | `environment.agentImage` | The `agentImage` sub-block records the rootfs image the recording ran against, stamped only for the tiers whose capabilities come from it (`container`, `hostloop`; `microvm` probes the Lima guest instead). `ref` is the resolved image ref, `configId` the local config id (built **and** pulled images, NOT comparable across machines), `registryDigest` the registry manifest digest (pulled only, and the only cross-machine-comparable identity). The image decides `missingCapabilityUse`, which `computeVerdict` fails on, so replaying against a different rootfs can move a verdict with nothing in the cassette having changed — `replay` compares `registryDigest` first and prints an advisory note, never a failure. Re-inspected at cassette-WRITE time, so a rebuild between run and record records the later image. **Absent** on cassettes recorded before this field existed, and never backfilled |
72
73
 
73
74
  ## Recipe 3 — Set up redaction BEFORE your first hostloop/protocol record
74
75
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,137 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.23.0] — 2026-08-14
10
+
11
+ ### Added
12
+
13
+ - **Platform baseline for Claude Desktop 1.30096.1 (bundled agent ELF `2.1.229`).** The agent ELF's
14
+ `sha256` was verified against the official release manifest for 2.1.229. The modeled spawn contract is
15
+ unchanged across the bump — `spawn.tools` stays 20 entries, `allowedTools` 19, the egress allowlist 15
16
+ domains, the spawn-env key set 61 with 31 conditional spreads, and the Cowork system prompt is
17
+ byte-identical (its fingerprint is recorded for 1.30096.1 in
18
+ `baselines/prompts/cowork-system-prompt-fingerprints.json`). `sync` refused to write this baseline until
19
+ the two sentinel defects below were corrected.
20
+
21
+ - **Two drift sentinels pinned in the synced baseline's `provenance.gates`**: the artifact-mount gate
22
+ (`coworkArtifacts`) and the CIC `can_use_tool` handler (`cicCanUseToolEnabled`). Both are force-ON in
23
+ production, and the artifact-mount gap documented in `docs/fidelity-gaps.md` rests on that fact — which
24
+ was previously read from a live feature cache the baseline never recorded, so nothing would have
25
+ noticed it changing.
26
+
27
+ - **`check:versions` guards DESIGN.md's live-verification scope note (invariant 11).** That note is the
28
+ repo's disclosure of how much of the *current* baseline has actually been verified live, and every
29
+ figure in it is derivable from `baselines/desktop-*.json` — yet it sat in unguarded prose and had
30
+ drifted twice, the baseline list having been extended without recounting. Understating how much is
31
+ unverified is the doc error least worth shipping, so it is now checked. Two forms, selected by whether
32
+ the note's live-pass baseline is the newest one: with a gap, the listed baselines must run contiguously
33
+ from wherever the list starts through the newest baseline, and both counts must match the list and the
34
+ real `agentVersion` transitions; with no gap, the note must say so explicitly and carry no stale
35
+ enumeration. Either way the named agent must be the newest baseline's. Because shipping a baseline flips
36
+ the no-gap form into the gap form, a new release now forces the note to be rewritten rather than
37
+ silently overstating coverage. The list's *start* is deliberately not derived — the note omits baselines
38
+ covered by the live pass itself, and encoding that rule would only relocate the drift. A missing or
39
+ unrecognisable note is an error, never a skip.
40
+
41
+ - **Cassettes now record the rootfs image they were recorded against**, closing the last gap in the
42
+ agent-image provenance work. The image decides `missingCapabilityUse`, which `computeVerdict` fails
43
+ on, so the rootfs is verdict-affecting — yet no cassette field named it, and a recording silently
44
+ inherited whatever image happened to be on the machine. `environment.agentImage` now carries the
45
+ resolved `ref` plus whichever identities exist: `configId` (the local config id, present for built
46
+ **and** pulled images but not comparable across machines) and `registryDigest` (the registry manifest
47
+ digest, pulled images only, and the only identity comparable across machines).
48
+
49
+ The field is stamped only for the tiers whose capabilities actually come from that image —
50
+ `container` and `hostloop`. `microvm` probes the Lima guest instead, so it records nothing rather
51
+ than naming an image that had no bearing on the run. Additive: no `cassetteVersion` bump, absent on
52
+ cassettes recorded before the field existed, and never backfilled — the absence is meaningful.
53
+
54
+ `ref` is a verbatim `COWORK_AGENT_IMAGE` value, so a private registry ref (`registry.acme.corp/…`)
55
+ would otherwise be committed into a public fixture with `grep`-clean transcript text. It is scanned
56
+ by `verify-cassettes` and rewritten by redaction like every other user-controlled string; the digests
57
+ are content hashes and are deliberately left intact.
58
+
59
+ - **`replay` warns when the rootfs image differs from the recording.** It compares `registryDigest`
60
+ first — the only identity stable across machines, which is the case the field exists to serve — and
61
+ falls back to the local config id only when neither side has a registry digest. A recording made
62
+ against a pulled image and replayed against a local rebuild is reported as drift rather than passing
63
+ silently. Advisory only: a legitimately re-pulled image is the common case, so this names the
64
+ difference instead of failing the replay.
65
+
66
+ The current image is inspected at most once per `replay` invocation, and only when a cassette
67
+ actually recorded one — so replay stays usable with no container runtime present, and the
68
+ `verify-cassettes` privacy scan never shells out.
69
+
70
+ ### Fixed
71
+
72
+ - **The host-loop `canUseTool` chain sentinel could be widened silently.** Desktop 1.30096.1 inserts a
73
+ fourth link into the chain — an allow carve-out that rewrites the tool's input and runs ahead of the
74
+ existing deny. The sentinel was a prefix match with no terminator, so it accepted any chain that merely
75
+ *started* with the expected calls: an inserted **synchronous** link would have passed unnoticed, and
76
+ the block that surfaced this release only happened because the new link is `await`ed.
77
+
78
+ The check now decomposes the chain with a real scanner (the assignment must be brace/paren-balanced,
79
+ and `??` legitimately occurs inside a template interpolation in the chain's own log line) and asserts
80
+ it end to end: the terminal operand must be a bare call to the saved original, every operand whose
81
+ callee is an `async function` must be awaited, no link that can return an allow may precede the
82
+ VM-path deny, and each link must resolve to a definition. The await rule is the load-bearing one — an
83
+ un-awaited async link returns a Promise, which is never nullish, so `??` short-circuits and every later
84
+ link *including the original callback* is skipped. Three and four link chains are both accepted so an
85
+ older Desktop still syncs; a fifth is reported for classification.
86
+
87
+ - **The early-allow ordering check had never fired.** It searched for a containment helper by two
88
+ readable names, neither of which occurs in any shipped asar — production mangles the call — so the
89
+ guard was permanently inert and its test passed only because the fixture hard-coded a token that does
90
+ not exist in the product. The helper is now identified by shape and resolved through the export map.
91
+
92
+ - **S6c no longer hard-blocks on a minifier rename.** It pinned the HIPAA-restriction call by its
93
+ minified *member name*, so the rename `A.r()` → `t.hu()` failed a predicate that is otherwise
94
+ byte-for-byte identical. The callee is now resolved through the chunk's export map and verified two
95
+ hops to the reader that consults the restriction, with resolution failure treated as a miss.
96
+
97
+ - **The `protocol`-tier live suite had been silently skipping itself for many releases.** `live-matrix`
98
+ required a staged agent binary for a baseline pinned years back, but protocol fidelity spawns the host
99
+ `claude` from `PATH` and never resolves a staged binary — the requirement was never real for that tier.
100
+ Because Claude Desktop prunes old staged agents on update, the gate went false as soon as a machine
101
+ moved past that agent version, and the suite dropped out on every developer machine and in CI without
102
+ naming itself. It now gates on what the tier actually uses and emits a skip notice identifying the
103
+ failing precondition, since this is the only protocol-tier live coverage there is.
104
+
105
+ - **The `hostloop` uploads-are-`Read`-able live case no longer fails on the model's choice of exploration
106
+ tool.** It asserted that neither native nor workspace bash ran at all, as a proxy for "the agent needed a
107
+ workaround" — so a run where the agent listed the uploads directory with `ls` went red, while the next
108
+ run, which used `Glob` instead, went green. In both the upload was `Read` directly at the advertised
109
+ path and no outputs-delete fired, so the regression the case guards was absent either way. It now reads
110
+ the recorded bash commands and fails only when one **names the uploaded file**, which is what reading or
111
+ copying it as a workaround requires and what listing its directory cannot do. Verified against both
112
+ recorded runs plus `cat`/`cp` mutations: the false red is gone and the workaround chain still trips it.
113
+
114
+ ### Documentation
115
+
116
+ - **A full live end-to-end pass now covers baseline `desktop-1.30096.1` / agent 2.1.229**, across all
117
+ three tiers (`protocol`, `container`, `hostloop`), superseding the `desktop-1.20186.0` pin. DESIGN.md's
118
+ claim and its scope note are re-stamped accordingly, including the caveats: two `live-outputs-delete`
119
+ cases skipped as model-behaviour misses, the tiers were covered across two invocations because the
120
+ `protocol` suite was repaired mid-pass, and the suites are model-dependent enough that a single red is
121
+ evidence of variance until a re-run says otherwise.
122
+
123
+ - **`docs/fidelity-gaps.md`'s artifacts section records the agent-side consent floor.** The server-delivered
124
+ session flag no longer only decides Desktop's spawned tool list — the agent reads the corresponding
125
+ spawn-env key itself and uses it to select its own artifact publish surface and a read-only mode. The
126
+ agent also refuses artifact publishes, comment replies, comment-thread resolves and artifact database
127
+ writes outright in a session with no answerable approval surface, and fails closed if it cannot confirm
128
+ one. The section now also states why this is recorded rather than modeled: artifact operations are
129
+ server-backed, on hosts outside the sandbox egress allowlist, so supplying the flag would offer a tool
130
+ resolving against a service the sandbox cannot reach.
131
+
132
+ - **The same section's mount-kind list is corrected.** It named a synthetic root that is not a mount kind
133
+ and omitted one that is, which mattered because the artifact-mount gap is stated in terms of what the
134
+ harness does and does not mount.
135
+
136
+ - **`docs/maintenance.md` corrects what moves the floating agent-image `:2` tag.** It is a curated pointer
137
+ moved deliberately by a manual publish with `immutable_only` unchecked — explicitly *not* something a
138
+ release tag push moves.
139
+
9
140
  ## [1.22.0] — 2026-08-12
10
141
 
11
142
  ### Added
package/DESIGN.md CHANGED
@@ -44,7 +44,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
44
44
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
45
45
 
46
46
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
47
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.227**, per `baselines/desktop-1.28929.0.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
47
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.229**, per `baselines/desktop-1.30096.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
48
48
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
49
49
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
50
50
 
@@ -171,9 +171,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
171
171
 
172
172
  > **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
173
173
 
174
- ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS build 2.1.177+. The staged in-VM agent that L1/L2 ran at the time of this pass was **2.1.202** (the `desktop-1.20186.1` baseline has since re-synced the staged VM ELF to **2.1.205** — see the note below), the native host app that `hostloop` runs is **2.1.205**, baseline **`desktop-1.20186.0`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-07-11, superseding the prior `1.19367.0` pin. The intervening `1.19367.0`→`1.20186.0` bump was a minifier re-anchor with a byte/behaviourally-identical value-resolved spawn contract, now confirmed live.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
174
+ ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.229**, the native host app that `hostloop` runs is **2.1.229**, baseline **`desktop-1.30096.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-14, superseding the prior `1.20186.0` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
175
175
 
176
- > **Scope of that claim, stated plainly.** `2026-07-11 / desktop-1.20186.0` is the last baseline carrying a **full live end-to-end pass**; nine baselines have shipped since (`1.20186.9`, `1.21459.0`, `1.22209.3`, `1.24012.0`, `1.24012.1`, `1.24012.9`, `1.24012.11`, `1.25927.0`, `1.26832.0`), four of which moved the agent ELF — most recently to **2.1.222**. Those were verified the cheaper way — `sync` reporting no unknown deltas, asar analysis, and a full local suite — not by re-running the live tiers. So read the sentence above as *"the value-resolved spawn contract is live-verified as of 1.20186.0"*, **not** as *"current-baseline behaviour has been re-passed live"*. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
176
+ > **Scope of that claim, stated plainly.** `2026-08-14 / desktop-1.30096.1` is the last baseline carrying a **full live end-to-end pass**, and it is the newest committed baseline — **no baselines have shipped since**. The pass ran against agent **2.1.229** and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe and resume continuity on the native binary). Three caveats, so this is not read as more than it is: two `live-outputs-delete` cases **skipped** — the agent declined to issue the pinned command, a model-behaviour miss the suite itself classifies as not a guard defect — so this is *every suite exercised*, not *every live assertion green*; the tiers were covered across **two invocations**, not one, because the `protocol` suite was repaired mid-pass (it had been silently skip-gated on a staged binary its tier never uses); and a live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — the `live-outputs-delete` skip set differs run to run for exactly that reason, which is why those cases skip loudly rather than fail. One such flake was root-caused rather than retried: the `hostloop` uploads-are-`Read`-able case counted bash calls, so its verdict turned on which exploration tool the model happened to pick (an `ls` of the uploads directory failed it; a `Glob` in the next run passed), even though the upload was `Read` directly and no outputs-delete fired either time. It now inspects what bash actually did and fails only when a command names the uploaded file — the workaround it exists to catch — so exploration is free and `cat`/`cp` of the upload still trips it. Every assertion in the lane has passed at least once against this baseline. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
177
177
 
178
178
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
179
179
 
package/README.md CHANGED
@@ -92,7 +92,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
92
92
 
93
93
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
94
94
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
95
- > From a global install (`npm i -g "cowork-harness@>=1.22.0"`), point at the package root instead:
95
+ > From a global install (`npm i -g "cowork-harness@>=1.23.0"`), point at the package root instead:
96
96
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
97
97
  > (or copy the cassette into your own project and pass that path).
98
98
 
@@ -102,7 +102,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
102
102
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
103
103
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
104
104
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
105
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.22.0"`.
105
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.23.0"`.
106
106
 
107
107
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
108
108
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -127,7 +127,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
127
127
  claude plugin install cowork-harness@cowork-harness
128
128
  ```
129
129
 
130
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.22.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.23.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
131
131
 
132
132
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
133
133
 
@@ -149,7 +149,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
149
149
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
150
150
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
151
151
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
152
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.22.0"` — see
152
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.23.0"` — see
153
153
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
154
154
 
155
155
  ### Prerequisites for anything above `protocol` fidelity
@@ -461,7 +461,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
461
461
 
462
462
  **Flags worth knowing** (the full list is always `<command> --help`):
463
463
  - `run`: a decider (`--decider-cmd <helper>`/`--decider-dir <dir>`, or a scenario's `on_unanswered: llm` — a scenario-YAML key, not a CLI flag; `run` rejects `--on-unanswered llm`) can answer unscripted gates; `--repeat N` (2-100) runs each scenario N times and aggregates a variance rollup instead of a single pass/fail (`--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd` tune the batch verdict/loop — and `--max-budget-usd` applies without `--repeat` too, as a pre-flight refusal from the scenario's cost history); `--matrix <matrix.yaml>` runs ONE scenario across the cross-product of baseline/model/skill_dir axes (worked example: `examples/matrices/csv-metrics-matrix.yaml`; `--max-cells`/`--concurrency` tune the cap/pool — any cell failing, assertion or infra, fails the run); `--compact`/`--demo` trim `run`/`skill` output for shareable screenshots/GIFs; `--label <tag>` (both `run` and `skill`) stamps a generation name for the iterate-across-fixes loop — surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and `status.json`, alongside the auto-recorded `skillCommit` (git provenance) and the authoritative `fingerprint.skillHash` a harvest step should pair critiques by.
464
- - `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error).
464
+ - `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error). `replay` also prints an **advisory** note — never a failure — when the rootfs agent image differs from the one the cassette recorded (`environment.agentImage`): the image decides which capabilities exist, so replaying against a different one can move a verdict with nothing in the cassette having changed. It compares the registry digest first (the only identity comparable across machines) and falls back to the local image config id only when neither side has one; cassettes recorded before that field existed simply carry no image to compare.
465
465
  - `verify-cassettes`: privacy scan (email/currency/domain/path/machine-inventory) + staleness — exit 1 when verification RAN and found a real problem (a finding, a genuine drift, or scenario-prompt drift), exit 3 when it could NOT complete (an unverifiable-class staleness finding, a too-new cassette format, or a read error); whole-token allows via `--allow <regex>` (a pattern) / class-scoped `--allow-domain` / `--allow-email` / `--allow-path` / `--allow-machine-inventory` / `--allow-patterns-file <path>` (a **file** of patterns, one regex per line); `--skip-privacy` or `--skip-staleness` runs only part of the gate; a diverged scenario `prompt` vs. the cassette's frozen prompt is also a hard fail (its own `scenarioDrift` bucket), opt out with `--skip-scenario-drift`; `--margins` adds a per-cassette recorded-vs-budget report for count-bound assertions (a single-sample estimate — diagnostic only, never changes the gate verdict); `--allow-empty` makes an **existing but cassette-free directory** exit 0 instead of the default loud exit 2 — for a repo that deliberately commits no cassettes (a *missing* path still fails, so the flag can never green a typo'd path).
466
466
  - `stats`: reads `<runsRoot>/index.jsonl`, written automatically at every result; `--since`/`--baseline`/`--branch` filter; `--skill-hash <prefix>`/`--label <tag>` narrow to one generation of the iterate-across-fixes loop and `--group-by scenario|skill-hash|label|fidelity` splits per generation — or per effective fidelity tier — instead of aggregating across them (a window spanning >1 generation or >1 tier warns; under `skill-hash`/`label` grouping, rows lacking the field are excluded from grouping and counted, never bucketed blank — the `fidelity` key is total, so nothing is ever excluded under that grouping); `--runs` lists the individual runs behind each summary with their `skillHash`/`runLabel` (each `runs[]` entry also carries `fidelity` — the tier that run actually executed at, `effectiveFidelity ?? fidelity` — in the `--output-format json` envelope only; the text listing is unchanged); `--last <n>` windows per-group; `--reindex` rebuilds the index from the physical run-dir tree (the migration path for pre-index runs), reconstructing each critique's cost roll-up from its run dir along the way.
467
467
  - `diff`: `--changelog` renders known-field prose for a baseline diff; `--view tools|transcript|artifacts|meta` narrows a run/cassette diff to one section; normalization (default on) masks per-run noise (ids/timestamps/session markers/host paths) so two runs of the same scenario diff as identical — `--no-normalize` compares raw values.
@@ -714,7 +714,7 @@ jobs:
714
714
  - uses: actions/checkout@v4
715
715
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
716
716
  run: |
717
- V=2.1.227 # match your scenario's pinned baseline's agentVersion
717
+ V=2.1.229 # match your scenario's pinned baseline's agentVersion
718
718
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
719
719
  chmod +x "$RUNNER_TEMP/claude-$V"
720
720
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
@@ -725,7 +725,7 @@ jobs:
725
725
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
726
726
  ```
727
727
 
728
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.22.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](./action.yml) for the full input/output reference.
728
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.23.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](./action.yml) for the full input/output reference.
729
729
 
730
730
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
731
731
 
@@ -893,6 +893,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
893
893
  ## Status
894
894
 
895
895
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
896
- **`desktop-1.28929.0`**. Release-by-release verification notes (what was re-verified against
896
+ **`desktop-1.30096.1`**. Release-by-release verification notes (what was re-verified against
897
897
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
898
898
  this section would otherwise duplicate lives in the sections above.
@@ -446,6 +446,11 @@
446
446
  },
447
447
  "note": "coworkWebFetchViaApi=true coworkWebFetchPrompt=true workspaceBashWaitLonger=true sessionsBridgePollBlockMs=30 — web_fetch is host/API-routed (POST /api/organizations/<org>/cowork/web_fetch), NOT container egress; gated by a separate web-fetch hostname allowlist + URL provenance."
448
448
  },
449
+ "cicCanUseToolEnabled:2051942385": {
450
+ "on": true,
451
+ "source": "force",
452
+ "value": true
453
+ },
449
454
  "cliPlugin:2307090146": {
450
455
  "on": false,
451
456
  "source": "defaultValue",
@@ -467,6 +472,11 @@
467
472
  "source": "defaultValue",
468
473
  "value": "## Sensitive personal information\n\nDo not save the following to memory unless the user explicitly asks you to remember it:\n\n- Protected attributes: race, ethnicity, national origin, religion, age, sex, sexual orientation, gender identity, immigration status, disability, serious illness, union membership\n- Government identifiers: Social Security numbers, driver's license numbers, passport numbers, government ID numbers\n- Financial account details: credit card numbers, bank account numbers\n- Health information: medical conditions, diagnoses, lab results, mental health details, therapy or counseling\n- Home or personal mailing addresses (work addresses are fine)\n- Account passwords, secret tokens, or secret keys\n\nIf any of the above appears in conversation context, complete the task but do not persist it to a memory file. If the user explicitly says \"remember my address is X\", saving it is acceptable — they've given consent."
469
474
  },
475
+ "coworkArtifacts:2940196192": {
476
+ "on": true,
477
+ "source": "force",
478
+ "value": true
479
+ },
470
480
  "canSaveSkill:3246569822": {
471
481
  "on": true,
472
482
  "source": "force",
@@ -558,9 +568,9 @@
558
568
  ],
559
569
  "spawnEnvSpreadCount": 31,
560
570
  "fcache": {
561
- "content16": "8ae36e12f8857433",
562
- "embeddedTimestamp": 1786530845855,
563
- "featureCount": 250
571
+ "content16": "341351c3593dc6d8",
572
+ "embeddedTimestamp": 1786562513613,
573
+ "featureCount": 253
564
574
  }
565
575
  },
566
576
  "requireFullVmSandbox": null