cowork-harness 3.3.0 → 3.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +17 -5
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +6 -6
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/CHANGELOG.md +132 -0
  8. package/DESIGN.md +2 -2
  9. package/README.md +31 -4
  10. package/SPEC.md +1 -1
  11. package/baselines/desktop-1.46388.3.json +942 -0
  12. package/baselines/prompts/cowork-system-prompt-fingerprints.json +7 -0
  13. package/baselines/prompts/desktop-1.46388.3/subagent-append-hl.md +37 -0
  14. package/dist/prompt/subagent-manifest.js +86 -0
  15. package/dist/prompt.js +22 -2
  16. package/dist/run/cassette.js +8 -0
  17. package/dist/run/chat.js +14 -0
  18. package/dist/run/execute.js +14 -0
  19. package/dist/session.js +4 -0
  20. package/dist/sync/cowork-sync.js +252 -29
  21. package/docs/cassette.md +5 -1
  22. package/docs/ci.md +2 -2
  23. package/docs/cli.md +4 -4
  24. package/docs/companion-skill.md +3 -3
  25. package/docs/fidelity-gaps.md +38 -5
  26. package/docs/invariants.md +1 -1
  27. package/docs/maintenance.md +42 -13
  28. package/docs/scenario.md +3 -2
  29. package/docs/subagents.md +43 -0
  30. package/examples/data/manifest-probe/ledger.csv +5 -0
  31. package/examples/replays/README.md +1 -1
  32. package/examples/replays/example-multiselect-gate.cassette.json +3 -3
  33. package/examples/replays/example-pdf-skill.cassette.json +97 -114
  34. package/examples/replays/hostloop-computer-links.cassette.json +3 -3
  35. package/examples/scenarios/subagent-manifest-probe.yaml +75 -0
  36. package/examples/sessions/subagent-manifest-probe.yaml +17 -0
  37. package/package.json +1 -1
  38. package/schema/cassette.v10.json +1 -1
  39. package/schema/cassette.v11.json +1 -1
  40. package/schema/cassette.v12.json +1 -1
  41. package/schema/run-result.json +1 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.3.0
7
- tracks-harness: cowork-harness 3.3.0 (baseline desktop-1.44121.1)
6
+ version: 3.4.1
7
+ tracks-harness: cowork-harness 3.4.1 (baseline desktop-1.46388.3)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.3.0` (baseline
29
- > `desktop-1.44121.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.4.1` (baseline
29
+ > `desktop-1.46388.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.3.0"`. **Pin `@^3.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.4.1**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.4.1" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.4.1"`. **Pin `@^3.4.1`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -158,6 +158,18 @@ Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejec
158
158
  (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
159
159
  `references/fidelity-and-answers.md`.
160
160
 
161
+ **Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
162
+ `/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
163
+ **remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
164
+ different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
165
+ reproduces it and none can — that container is not something a local tool can stand up. Which lane a
166
+ real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
167
+ So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
168
+ asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
169
+ `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
170
+ rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
171
+ above).
172
+
161
173
  ### Choose an answer path (gates: AskUserQuestion + tool-permission)
162
174
 
163
175
  Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.3.0` (baseline `desktop-1.44121.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.3.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.4.1"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.258 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.260 # match your scenario's pinned baseline's agentVersion
40
40
  # The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
41
41
  # served only from .../claude-code-releases/rc/<commit>/ — the stable path 404s for those, and
42
42
  # 2.1.255 is one. Take B from your pinned baseline's agentBinary.releaseBaseUrl; baselines
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
73
73
  GitHub-hosted runners, no token/Docker/agent:
74
74
 
75
75
  ```yaml
76
- - run: npm i -g "cowork-harness@^3.3.0"
76
+ - run: npm i -g "cowork-harness@^3.4.1"
77
77
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
78
78
  # no silent false-greens. WITHOUT --strict this
79
79
  # step cannot fail on a WARN-class rule (e.g.
@@ -350,7 +350,7 @@ jobs:
350
350
  with: { node-version: '24' }
351
351
  - uses: actions/setup-python@v5
352
352
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
353
- - run: npm i -g "cowork-harness@^3.3.0"
353
+ - run: npm i -g "cowork-harness@^3.4.1"
354
354
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
355
355
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
356
356
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -379,7 +379,7 @@ jobs:
379
379
  echo "live=true" >> "$GITHUB_OUTPUT"
380
380
  fi
381
381
  - if: steps.guard.outputs.live == 'true'
382
- run: npm i -g "cowork-harness@^3.3.0"
382
+ run: npm i -g "cowork-harness@^3.4.1"
383
383
  - if: steps.guard.outputs.live == 'true'
384
384
  run: cowork-harness run scenarios/ --output-format json
385
385
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.3.0` (baseline `desktop-1.44121.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.3.0` (baseline `desktop-1.44121.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.3.0`
4
- (baseline `desktop-1.44121.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.4.1`
4
+ (baseline `desktop-1.46388.3`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.3.0` (baseline `desktop-1.44121.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.4.1` (baseline `desktop-1.46388.3`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,138 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.4.1] — 2026-09-05
10
+
11
+ ### Documentation
12
+
13
+ - **Which Cowork *lane* the harness models is now stated, in the places a consumer reads.** Every
14
+ fidelity tier reproduces the desktop-local lane; Cowork's remote lane runs server-side in a cloud
15
+ container with a different filesystem, shell tool, delivery mechanism and a server-authored prompt.
16
+ Which lane a real session gets is a Cowork setting ("Only on this computer"), and it was observed
17
+ **off** on a current install. README, the companion skill, `docs/fidelity-gaps.md` and
18
+ `docs/maintenance.md` now say so. The distinction that matters: behaviour conclusions (triggering,
19
+ tool sequencing, gate handling) travel between lanes; anything asserting a **path, mount or delivery
20
+ mechanism** is a local-lane claim only. No remote tier is planned — that container is Anthropic's, so
21
+ emulating it would mean authoring an environment rather than reproducing one; `lane: remote` already
22
+ makes the affected assertions refuse to grade instead of passing.
23
+
24
+ ### Changed
25
+
26
+ - **The sub-agent override sentinel's note now carries a current probe.** Gate `124685897` reads ON,
27
+ which only enables a server-delivered replacement of the `## Cowork environment` section. Re-probed
28
+ against the 1.46388.3 composition in a real host-loop session: all three composed parts arrived
29
+ byte-identical to the shipped fallback, so the gate is on with no payload. The two non-overridable
30
+ parts are the control that makes it conclusive. The note also states the probe's precondition, which
31
+ it previously lacked.
32
+
33
+ ### Verification
34
+
35
+ - **`microvm` exercised after the 3.4.0 tag** — `smoke-l2-microvm` passed 3/3 against the *published*
36
+ 3.4.0 artifact (real Apple-VZ VM, separate kernel; guest egress reached allowlisted
37
+ `api.anthropic.com` and denied an off-list telemetry host). That completes all four tiers for the
38
+ 1.46388.3 baseline; DESIGN.md's scope note is re-stamped accordingly. The 3.4.0 entry below is left
39
+ as the record of what was verified at tag time. This tier is macOS-arm64 only and CI runners are
40
+ Linux, so it is only ever exercised by hand.
41
+
42
+ ## [3.4.0] — 2026-09-05
43
+
44
+ ### Upgrade notes
45
+
46
+ - **`promptAssetsHash` now also covers prompt text the harness GENERATES**, not only the committed
47
+ prompt-asset files — specifically the host-loop sub-agent folder manifest and the trailing skills
48
+ sentence added in this release. If you hold a cassette recorded with 3.3.0 against a baseline whose
49
+ `appVersion` still matches your live one, it will report `prompt-assets` staleness; re-record it. A
50
+ cassette on an older baseline already reports `baseline` staleness and is unaffected by this change.
51
+ - **A host-loop sub-agent is now told which folders exist and how to address them.** If you assert on a
52
+ sub-agent's prose, note it now has per-session path information it did not have before.
53
+ - **Sub-agent prompt assets are no longer the whole append.** If you maintain your own baseline,
54
+ `spawn.subagentAppendHostLoop` now points at the overridable SECTION only; the manifest and trailing
55
+ sentence are generated. Do not restore them into the asset — they would render twice.
56
+
57
+ ### Verification
58
+
59
+ Verified **locally**, against the Desktop install this baseline was synced from — Desktop `1.46388.3`,
60
+ agent `2.1.260`, macOS arm64, agent image `cowork-agent-base:2` — on 2026-09-05:
61
+
62
+ | suite | result |
63
+ |---|---|
64
+ | `boundary-check` | **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress) |
65
+ | `npm run test:live` | 4 files, **19 passed, 0 skipped** |
66
+ | `run examples/scenarios/` | **7/7 success**, 39 assertions, 0 failed |
67
+ | e2e self-tests | **8/8 success** — askuserquestion, multiselect, multiselect-deciderdir, l1-container, l1-egress, present-files, semantic-evidence-files, canary-hostloop |
68
+
69
+ Tiers exercised: **`protocol`, `container`, `hostloop`**. **`microvm` was NOT exercised** (it needs a
70
+ real VM; CI excludes it for the same reason), so `smoke-l2-microvm` did not run.
71
+
72
+ `subagent-manifest-probe` is new here and is the first live coverage the sub-agent append has ever
73
+ had. It proves the composed append is *actionable* — the sub-agent shelled against
74
+ `/sessions/<id>/mnt/<name>/`, summed the same ledger through its file tools to the same total, and a
75
+ bare relative write landed at the host cwd — not that the paraphrase matches Desktop's wording, which
76
+ is what the `manifest`/`suffix` fingerprint axes are for.
77
+
78
+ ### Parity — Desktop 1.46388.3 (agent 2.1.260)
79
+
80
+ New baseline `desktop-1.46388.3`. `sync` refused with 7 unknown deltas; all are classified and it now
81
+ runs clean.
82
+
83
+ #### Fixed — the sub-agent append sentinel could not see part of the append
84
+
85
+ Desktop 1.46388.3 restructured the per-session sub-agent append into three parts: the overridable
86
+ `## Cowork environment` section, a host-loop-only folder manifest, and one trailing sentence appended to
87
+ **both** branches. The sentinel fingerprinted only the two ternary branch texts, so the trailing sentence
88
+ changed the rendered append on both branches while the VM fingerprint did not move — and the pointer
89
+ coupling guard only demands a repoint on an axis whose fingerprint moved. A counterfactual confirms it:
90
+ Desktop 1.44121.1 with only that sentence added raises **no flags and syncs green**.
91
+
92
+ Two fingerprint axes now cover the composed parts (`manifest`, `suffix`). Both become mandatory once
93
+ either is recorded, and only the newest entry is compared, so entries predating the restructure are not
94
+ retro-failed.
95
+
96
+ Related, and separately load-bearing: the host-loop branch extractor was selecting the **wrong template**.
97
+ It anchored on a phrase that left the branch at this release and survives in unrelated prose in the same
98
+ module, so the drift it reported named text that is not the branch. It is now anchored on the ternary's
99
+ structure rather than on wording. The module selector keyed on the same dead phrase and was one chunk
100
+ move away from dropping the entire branch sentinel behind a single generic flag.
101
+
102
+ #### Added — the host-loop folder manifest is modeled
103
+
104
+ A host-loop sub-agent is now told the same thing production tells it: which folders exist on the user's
105
+ computer, each with its host path and its shell path, or a different sentence when there is nothing to
106
+ list. It is generated from the run's real mounts rather than read from a static asset — the same reason
107
+ the main-loop shell section became a generator at Desktop 1.14271.0 — and it uses **canonical** folder
108
+ paths, matching what production renders. Folders dropped by `COWORK_HARNESS_SOFT_MISSING=1` are listed as
109
+ unreachable rather than silently disappearing, which is production's behaviour for a folder that fails to
110
+ mount.
111
+
112
+ The text this generator produces is mixed into `promptAssetsHash`, so editing it stales recorded cassettes
113
+ exactly as editing a committed prompt asset does.
114
+
115
+ #### Added — `CLAUDE_CODE_QUESTION_EXTENDED`
116
+
117
+ A new gate-conditional key in the Cowork spawn window, pinned along with its gate. The gate reads
118
+ served-and-off today, so the key stays out of `spawn.env` and an upstream flip will surface as a diff. It
119
+ drives an AskUserQuestion behaviour the harness models, and Desktop and the agent shipped it in the same
120
+ release.
121
+
122
+ #### Cassettes
123
+
124
+ `example-pdf-skill` was **re-recorded** at `container`: it embeds the VM sub-agent append in its
125
+ recorded frames, and that text changed, so re-stamping would have left a committed fixture
126
+ contradicting what the harness now sends. Tool calls moved 5 → 4 against the cassette it replaced.
127
+
128
+ `example-multiselect-gate` (protocol) and `hostloop-computer-links` (hostloop) were **re-stamped**
129
+ rather than re-recorded. No sub-agent was dispatched in either recording, protocol receives no
130
+ sub-agent append at all by design, and neither cassette embeds the changed text — so the prompt change
131
+ provably never entered either recording, while re-recording at those tiers would bake the recording
132
+ machine's own MCP servers and skills into a public fixture.
133
+
134
+ #### Fixed — a mutation test made ambiguous by the new build
135
+
136
+ A second `=31999` literal appeared in an unrelated chunk, and the two were in different chunks at
137
+ 1.44121.1, so no chunk marker separates them durably. The S4 mutation now changes every occurrence and
138
+ asserts the guard's **resolved value** rather than merely that it flagged, so a mutation that never
139
+ reached the guarded site cannot read as a pass.
140
+
9
141
  ## [3.3.0] — 2026-09-02
10
142
 
11
143
  ### Verification
package/DESIGN.md CHANGED
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
47
47
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
48
48
 
49
49
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
50
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.258**, per `baselines/desktop-1.44121.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
50
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.260**, per `baselines/desktop-1.46388.3.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
51
51
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
52
52
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
53
53
 
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
176
176
 
177
177
  ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
178
178
 
179
- > **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is no longer the newest committed baseline: **three** baselines have shipped since (`1.40609.0`, `1.40609.1`, `1.44121.1`), **three** of which moved the agent ELF, most recently to **2.1.258** — so the newest baseline is **not** live-verified, and this paragraph describes the 1.37937.1 pass only. The pass ran on agent `2.1.246` (the staged VM ELF and the native `.app` were both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
179
+ > **Scope of that claim, stated plainly.** `2026-09-05 / desktop-1.46388.3` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 passed, 0 skipped**; `run examples/scenarios/` **7/7 success, 39 assertions, 0 failed** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **9/9 success** (smoke-askuserquestion, smoke-multiselect, smoke-multiselect-deciderdir, smoke-l1-container, smoke-l1-egress, smoke-present-files, smoke-semantic-evidence-files, canary-hostloop, smoke-l2-microvm). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported none. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`.** The microvm case ran AFTER the 3.4.0 tag, against the published artifact rather than the checkout, and is recorded here rather than in that release's CHANGELOG entry, which states what was verified at tag time. It is the tier CI can never cover — GitHub's runners are Linux and Apple-VZ is macOS-arm64 only — so it is only ever exercised by hand here: `smoke-l2-microvm` passed 3/3 in a real VM with its own kernel, and its guest egress behaved (allowlisted `api.anthropic.com` reached, an off-list telemetry host denied). **Scope-out, so this is not read as more than it is.** (a) `subagent-manifest-probe` is new in this release and is the first live coverage the sub-agent append has ever had; it proves the composed append is ACTIONABLE — a sub-agent reaches the right files through both tool families, and a bare relative write lands at the host cwd — **not** that the harness's paraphrase matches Desktop's wording, which is the `manifest`/`suffix` fingerprint axes' job. (b) A live pass verifies observed behaviour, not the whole spawn contract by construction. (c) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: `example-pdf-skill` was re-recorded against this baseline and the other two committed cassettes were re-stamped (see the CHANGELOG for why each), so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
180
180
 
181
181
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
182
182
 
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^3.3.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^3.4.1"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.3.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.3.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.4.1"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.4.1"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Not sure a harness is what you need? The next two sections are the argument.
@@ -215,6 +215,32 @@ L2 microvm parity Optional. Agent inside a real Linux microVM (Lima/Apple-VZ
215
215
  synced baseline. "Do what real Cowork does."
216
216
  ```
217
217
 
218
+ ### Which Cowork *lane* this models — read this before trusting an environment assertion
219
+
220
+ Every tier above reproduces Cowork's **desktop-local lane**: the agent runs on your machine, shell
221
+ commands land in a Linux sandbox rooted at `/sessions/<id>`, attached folders appear under
222
+ `/sessions/<id>/mnt/<name>`, and finished files reach the user through `present_files`.
223
+
224
+ Cowork also has a **remote lane**, where the session runs server-side in an ephemeral cloud container
225
+ that reaches your machine over a link. There the filesystem, the shell tool, and file delivery are all
226
+ different — folders arrive under `$HOME/mnt/`, deliverables go to `/mnt/user-data/outputs/` and are
227
+ handed over with `SendUserFile`, and the environment prompt is authored by the server rather than by
228
+ Desktop. **Which lane you get is a Cowork setting** ("Only on this computer", Settings → Cowork), and
229
+ it has been observed **off** — i.e. remote — on a current install.
230
+
231
+ The harness **cannot execute the remote lane**: that container is Anthropic's, not something a local
232
+ tool can stand up. What it does instead is refuse to fake it. Declare `lane: remote` on a scenario and
233
+ the assertions that depend on observing a local filesystem degrade honestly — `file_absent` reports
234
+ evidence-unavailable rather than passing, and delivery is reported as unobservable — so a green never
235
+ means more than it should.
236
+
237
+ **What this means for you.** Behaviour-shaped conclusions travel between lanes: whether your skill
238
+ triggers, how it sequences tools, which questions it asks, whether it respects a permission gate.
239
+ Environment-shaped conclusions do not: anything asserting a path, a mount, or a delivery mechanism is a
240
+ statement about the **local** lane specifically. Scope your claims accordingly, and if you are probing
241
+ real Cowork to compare, turn "Only on this computer" **on** first or you will be measuring a lane this
242
+ tool does not model.
243
+
218
244
  **Decision guide** — `fidelity:` takes exactly one of these five values (`protocol`/`container`/`microvm` vary isolation strength; `hostloop`/`cowork` are overlays that instead pick *where the loop runs* — there's no combining the two groups):
219
245
 
220
246
  | Question | Choose |
@@ -281,6 +307,7 @@ See [DESIGN.md](./DESIGN.md) for the full parity matrix, the known deltas vs. re
281
307
 
282
308
  ## Limitations
283
309
 
310
+ - **One lane, deliberately.** Every tier models Cowork's desktop-local lane. The remote (cloud) lane runs server-side with a different filesystem, shell tool, delivery mechanism and a server-authored prompt; no tier reproduces it, and `lane: remote` exists to make the resulting blind spots refuse to grade rather than pass. See [Which Cowork lane this models](#which-cowork-lane-this-models--read-this-before-trusting-an-environment-assertion).
284
311
  - **Not the full Desktop network transport.** L1 is a container, not a VM; L2 *is* a real Apple-VZ microVM but still does not reproduce Cowork's gVisor netstack — its egress is the same allowlist proxy as L1 (with a guest iptables firewall in front). If your skill depends on VM-kernel specifics, validate at L2; if it depends on packet-level gVisor behavior, no tier reproduces it.
285
312
  - **Cowork in-guest context is partial.** Desktop supplies host-loop staging, runtime `mountPath` RPC, and the bridge. We reproduce the *filesystem and cowork mode*, not those host-side services. Skills that call Desktop-only host RPCs won't run here (they wouldn't be portable anyway).
286
313
  - **The agent binary is the staged ELF** (`claude-code-vm/<ver>/claude`), **bind-mounted** from your own Claude Desktop install — nothing Anthropic-owned is bundled or installed. There is **no npm path**; override the path with `COWORK_AGENT_BINARY`. Check licensing/ToS for your use.
@@ -353,6 +380,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
353
380
  ## Status
354
381
 
355
382
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
356
- **`desktop-1.44121.1`**. Release-by-release verification notes (what was re-verified against
383
+ **`desktop-1.46388.3`**. Release-by-release verification notes (what was re-verified against
357
384
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
358
385
  this section would otherwise duplicate lives in the sections above.
package/SPEC.md CHANGED
@@ -630,7 +630,7 @@ and a fail observed) always fails that batch, regardless of the numeric rate.
630
630
  "nonDeterministic?": bool, // true if any decision came from a non-deterministic source → not reproducible
631
631
  "gateProvenance?": { "total": number, "bySource": {…}, "gates": [{ "question","answeredBy","answer","model?" }] }, // how each AskUserQuestion gate was answered; informational (never fails the verdict); live/partial lane only (absent on replay)
632
632
  "permissiveAutoAllow?": ["string"], // tools auto-allowed by cowork parity that real Cowork BLOCKS → green is NOT faithful
633
- "staleness?": [{ "class": "baseline|skill|shared-root|format|unverifiable-baseline|unverifiable-skill|resolved-tier|unverifiable-tier|prompt-assets|unverifiable-prompt-assets", "message" }], // replay only; cassette-staleness findings, surfaced for a JSON gate. Drift classes are non-failing by default (a stale but passing replay stays ok:true); `unverifiable-skill` FAILS by default since 2.0.0. `--strict` fails on every class, `--fail-on-skill-drift` adds skill/shared-root. `resolved-tier` = a `fidelity: cowork` cassette's recorded effectiveFidelity no longer matches the tier the scenario's baseline (pinned `baseline:` or `latest`) resolves to today — the recording exercises the wrong tier; `unverifiable-tier` = the tier check couldn't run for a baseline-dependent (`fidelity: cowork`) cassette (no recorded effectiveFidelity, or its pinned baseline failed to load). Tier resolution is baseline-only (the CLAUDE_FORCE_HOST_LOOP env override is suppressed) so verify results can't differ across machines. `prompt-assets` = the baseline's committed prompt-asset files (spawn.promptTemplate/subagentAppend/subagentAppendHostLoop) changed since record under the SAME appVersion — warn by default, `--strict` fails, re-record; `unverifiable-prompt-assets` = a recorded `fingerprint.promptAssetsHash` exists but the live baseline's prompt assets can't be hashed (a moved/dangling pointer) — can't verify ⇒ not green.
633
+ "staleness?": [{ "class": "baseline|skill|shared-root|format|unverifiable-baseline|unverifiable-skill|resolved-tier|unverifiable-tier|prompt-assets|unverifiable-prompt-assets", "message" }], // replay only; cassette-staleness findings, surfaced for a JSON gate. Drift classes are non-failing by default (a stale but passing replay stays ok:true); `unverifiable-skill` FAILS by default since 2.0.0. `--strict` fails on every class, `--fail-on-skill-drift` adds skill/shared-root. `resolved-tier` = a `fidelity: cowork` cassette's recorded effectiveFidelity no longer matches the tier the scenario's baseline (pinned `baseline:` or `latest`) resolves to today — the recording exercises the wrong tier; `unverifiable-tier` = the tier check couldn't run for a baseline-dependent (`fidelity: cowork`) cassette (no recorded effectiveFidelity, or its pinned baseline failed to load). Tier resolution is baseline-only (the CLAUDE_FORCE_HOST_LOOP env override is suppressed) so verify results can't differ across machines. `prompt-assets` = the baseline's committed prompt-asset files (spawn.promptTemplate/subagentAppend/subagentAppendHostLoop), or the sub-agent prompt text the harness generates rather than reads from an asset (the folder manifest + trailing sentence, Desktop >=1.46388.3), changed since record under the SAME appVersion — warn by default, `--strict` fails, re-record; `unverifiable-prompt-assets` = a recorded `fingerprint.promptAssetsHash` exists but the live baseline's prompt assets can't be hashed (a moved/dangling pointer) — can't verify ⇒ not green.
634
634
  "skippedAssertions?": { "full": number, "partial": number }, // replay only; count of live-only assertions NOT evaluated (full = whole assertion skipped; partial = content half ran, fs/egress half dropped). The skipped ones are absent from `assertions[]`.
635
635
  "toolResults?": [{ "toolUseId?","isError","text","assertText?" }], // tool-result text at assertion-fidelity cap (10 KB); backs tool_result_contains/tool_result_not_contains and their regex siblings tool_result_matches/tool_result_not_matches
636
636
  "skillsInvoked?": ["string"], // Wave 1: skill/plugin ids invoked via the Skill tool_use event, call order, duplicates kept. Backs skill_triggered/no_skill_triggered.