cowork-harness 3.4.0 → 3.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +17 -5
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +29 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +54 -21
  9. package/CHANGELOG.md +173 -0
  10. package/DESIGN.md +2 -2
  11. package/README.md +31 -4
  12. package/baselines/desktop-1.46388.3.json +2 -2
  13. package/baselines/desktop-1.46388.4.json +952 -0
  14. package/dist/agent/session.js +88 -10
  15. package/dist/run/hook-events.js +20 -8
  16. package/dist/runtime/host-env.js +25 -5
  17. package/dist/runtime/hostloop.js +1 -1
  18. package/dist/session.js +34 -4
  19. package/dist/sync/cowork-sync.js +79 -8
  20. package/dist/types.js +23 -5
  21. package/docs/ci.md +1 -1
  22. package/docs/cli.md +3 -3
  23. package/docs/companion-skill.md +3 -3
  24. package/docs/fidelity-gaps.md +169 -12
  25. package/docs/invariants.md +2 -0
  26. package/docs/maintenance.md +31 -0
  27. package/docs/session.md +11 -5
  28. package/docs/subagents.md +13 -6
  29. package/examples/replays/README.md +1 -1
  30. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  31. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  32. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  33. package/package.json +1 -1
  34. package/scripts/gen-schema.ts +6 -1
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.4.0
7
- tracks-harness: cowork-harness 3.4.0 (baseline desktop-1.46388.3)
6
+ version: 3.5.0
7
+ tracks-harness: cowork-harness 3.5.0 (baseline desktop-1.46388.4)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,8 +25,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.4.0` (baseline
29
- > `desktop-1.46388.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.5.0` (baseline
29
+ > `desktop-1.46388.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
32
32
  ## Preflight — make sure the harness can actually run
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.4.0"`. **Pin `@^3.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.5.0"`. **Pin `@^3.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -158,6 +158,18 @@ Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejec
158
158
  (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
159
159
  `references/fidelity-and-answers.md`.
160
160
 
161
+ **Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
162
+ `/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
163
+ **remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
164
+ different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
165
+ reproduces it and none can — that container is not something a local tool can stand up. Which lane a
166
+ real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
167
+ So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
168
+ asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
169
+ `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
170
+ rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
171
+ above).
172
+
161
173
  ### Choose an answer path (gates: AskUserQuestion + tool-permission)
162
174
 
163
175
  Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`).
3
+ Self-contained reference. Tracks `cowork-harness 3.5.0` (baseline `desktop-1.46388.4`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.4.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.5.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
73
73
  GitHub-hosted runners, no token/Docker/agent:
74
74
 
75
75
  ```yaml
76
- - run: npm i -g "cowork-harness@^3.4.0"
76
+ - run: npm i -g "cowork-harness@^3.5.0"
77
77
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
78
78
  # no silent false-greens. WITHOUT --strict this
79
79
  # step cannot fail on a WARN-class rule (e.g.
@@ -350,7 +350,7 @@ jobs:
350
350
  with: { node-version: '24' }
351
351
  - uses: actions/setup-python@v5
352
352
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
353
- - run: npm i -g "cowork-harness@^3.4.0"
353
+ - run: npm i -g "cowork-harness@^3.5.0"
354
354
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
355
355
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
356
356
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -379,7 +379,7 @@ jobs:
379
379
  echo "live=true" >> "$GITHUB_OUTPUT"
380
380
  fi
381
381
  - if: steps.guard.outputs.live == 'true'
382
- run: npm i -g "cowork-harness@^3.4.0"
382
+ run: npm i -g "cowork-harness@^3.5.0"
383
383
  - if: steps.guard.outputs.live == 'true'
384
384
  run: cowork-harness run scenarios/ --output-format json
385
385
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.5.0` (baseline `desktop-1.46388.4`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`).
3
+ Self-contained reference. Tracks `cowork-harness 3.5.0` (baseline `desktop-1.46388.4`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.4.0`
4
- (baseline `desktop-1.46388.3`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.5.0`
4
+ (baseline `desktop-1.46388.4`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.4.0` (baseline `desktop-1.46388.3`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.5.0` (baseline `desktop-1.46388.4`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -113,14 +113,43 @@
113
113
  "PreToolUse"
114
114
  ],
115
115
  "knownHookEvents": [
116
+ "ConfigChange",
117
+ "CwdChanged",
118
+ "DirectoryAdded",
119
+ "Elicitation",
120
+ "ElicitationResult",
121
+ "FileChanged",
122
+ "InstructionsLoaded",
123
+ "MessageDisplay",
116
124
  "Notification",
125
+ "PermissionDenied",
126
+ "PermissionRequest",
127
+ "PostCompact",
128
+ "PostModelSwitch",
129
+ "PostToolBatch",
117
130
  "PostToolUse",
131
+ "PostToolUseFailure",
118
132
  "PreCompact",
133
+ "PreModelSwitch",
119
134
  "PreToolUse",
120
135
  "SessionEnd",
121
136
  "SessionStart",
137
+ "Setup",
122
138
  "Stop",
139
+ "StopFailure",
140
+ "SubagentStart",
123
141
  "SubagentStop",
142
+ "TaskCompleted",
143
+ "TaskCreated",
144
+ "TeammateIdle",
145
+ "UserPromptExpansion",
146
+ "UserPromptSubmit",
147
+ "WorktreeCreate",
148
+ "WorktreeRemove"
149
+ ],
150
+ "liveVerifiedHookEvents": [
151
+ "PostToolUse",
152
+ "SessionStart",
124
153
  "UserPromptSubmit"
125
154
  ],
126
155
  "enums": {
@@ -215,26 +215,49 @@ ASSERT_KEYS = _load_assert_keys()
215
215
  # reason the key lists are: a hand-copied served-set silently stops warning about the event it was later
216
216
  # extended to cover — the exact drift this check exists to prevent.
217
217
  _FALLBACK_SERVED_HOOK_EVENTS = {"PreToolUse"}
218
+ # Re-sourced 2026-09-06 from the agent's OWN hooks-config validator array (ELF 2.1.260), not from a grep
219
+ # for event-name constants. The previous 9-name set reported the other 24 -- PostCompact and
220
+ # MessageDisplay among them -- identically to a misspelling, at ERROR severity below.
218
221
  _FALLBACK_KNOWN_HOOK_EVENTS = {
219
- "PreToolUse", "PostToolUse", "UserPromptSubmit", "SessionStart",
220
- "SessionEnd", "SubagentStop", "PreCompact", "Notification", "Stop",
222
+ "PreToolUse", "PostToolUse", "PostToolUseFailure", "PostToolBatch",
223
+ "Notification", "UserPromptSubmit", "UserPromptExpansion", "SessionStart",
224
+ "SessionEnd", "Stop", "StopFailure", "SubagentStart", "SubagentStop",
225
+ "PreCompact", "PostCompact", "PreModelSwitch", "PostModelSwitch",
226
+ "PermissionRequest", "PermissionDenied", "Setup", "TeammateIdle",
227
+ "TaskCreated", "TaskCompleted", "Elicitation", "ElicitationResult",
228
+ "ConfigChange", "WorktreeCreate", "WorktreeRemove", "InstructionsLoaded",
229
+ "CwdChanged", "FileChanged", "DirectoryAdded", "MessageDisplay",
221
230
  }
231
+ # The subset a plugin hook has been OBSERVED to fire for here (live-verified 2026-08-01, container +
232
+ # hostloop). Kept apart from the known set because the message wording depends on which claim we can
233
+ # make: accepted-by-the-validator is not reached-by-a-run.
234
+ _FALLBACK_LIVE_VERIFIED_HOOK_EVENTS = {"SessionStart", "UserPromptSubmit", "PostToolUse"}
222
235
 
223
236
 
224
237
  def _load_hook_events():
225
- """(served, known) hook-event sets, from the generated assertion-keys.json sidecar."""
238
+ """(served, known, live-verified) hook-event sets, from the generated assertion-keys.json sidecar."""
226
239
  p = Path(__file__).resolve().parent / "assertion-keys.json"
227
240
  try:
228
241
  d = json.loads(p.read_text(encoding="utf-8"))
229
242
  served, known = set(d["servedHookEvents"]), set(d["knownHookEvents"])
243
+ # Tolerate a sidecar generated before liveVerifiedHookEvents existed: fall back rather than
244
+ # dropping to the embedded KNOWN set wholesale, which would undo the 33-name fix. Membership
245
+ # test, NOT truthiness: an explicitly EMPTY list means the firing receipt was withdrawn, and
246
+ # `or` would silently restore the 3-name claim -- widening a receipt, which is the one direction
247
+ # this constant exists to prevent.
248
+ live = set(d["liveVerifiedHookEvents"]) if "liveVerifiedHookEvents" in d else set(_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS)
230
249
  if served and known:
231
- return served, known
250
+ return served, known, live
232
251
  except Exception:
233
252
  pass
234
- return set(_FALLBACK_SERVED_HOOK_EVENTS), set(_FALLBACK_KNOWN_HOOK_EVENTS)
253
+ return (
254
+ set(_FALLBACK_SERVED_HOOK_EVENTS),
255
+ set(_FALLBACK_KNOWN_HOOK_EVENTS),
256
+ set(_FALLBACK_LIVE_VERIFIED_HOOK_EVENTS),
257
+ )
235
258
 
236
259
 
237
- SERVED_HOOK_EVENTS, KNOWN_HOOK_EVENTS = _load_hook_events()
260
+ SERVED_HOOK_EVENTS, KNOWN_HOOK_EVENTS, LIVE_VERIFIED_HOOK_EVENTS = _load_hook_events()
238
261
 
239
262
  # Self-check: every valid assertion key must be classified, else the replay-class lint logic mishandles it.
240
263
  # Surfaced loudly at load AND as a lint ERROR in cmd_lint (so --strict / exit codes flow). Never sys.exit here.
@@ -1489,15 +1512,17 @@ def _lint_hook_events(path):
1489
1512
  """Flag hook events a plugin DECLARES that this harness does not SERVE.
1490
1513
 
1491
1514
  Why this exists: the harness installs `PreToolUse` only, while real Cowork installs three event types
1492
- and the agent binary understands nine. A plugin declaring `UserPromptSubmit` therefore mounted, ran,
1515
+ and the agent binary accepts thirty-three. A plugin declaring `UserPromptSubmit` therefore mounted, ran,
1493
1516
  and produced no comment of any kind — the surface was discoverable only by grepping the harness's own
1494
1517
  compiled output, which is exactly what one consumer had to do.
1495
1518
 
1496
- Deliberately WARN, not ERROR, and deliberately worded as uncertainty: a declared-but-unserved event is
1497
- not a *skill* defect, and the harness cannot currently prove the event never fires. It only knows it
1498
- adds no handling of its own — the agent binary loads a plugin's hooks.json through its own
1499
- `--plugin-dir` channel, which is a separate path the harness neither serves nor blocks and which has
1500
- not been probed. Claiming "this will not fire" would assert more than is known; claiming nothing
1519
+ Severity is three-way and each level means something different. A name the agent's validator ACCEPTS
1520
+ but this harness does not serve is INFO, not a skill defect: the agent loads a plugin's hooks.json
1521
+ through its own `--plugin-dir` channel, which the harness neither serves nor blocks. Of those, only
1522
+ the events in LIVE_VERIFIED_HOOK_EVENTS have been observed firing here, so the confident wording is
1523
+ scoped to them and every other accepted name says so. A name the validator REJECTS is ERROR — that
1524
+ one genuinely never runs on any surface. Claiming "this will not fire" for an accepted event would
1525
+ assert more than is known; claiming nothing
1501
1526
  leaves the consumer to reverse-engineer it. So say precisely what is known.
1502
1527
  """
1503
1528
  findings = []
@@ -1533,17 +1558,25 @@ def _lint_hook_events(path):
1533
1558
  continue
1534
1559
  line_no = next((i for i, ln in enumerate(lines, 1) if f'"{name}"' in ln), 1)
1535
1560
  if name in KNOWN_HOOK_EVENTS:
1561
+ fires = (
1562
+ "fires here — a plugin's own `hooks/hooks.json` is loaded and executed by the agent "
1563
+ "binary (live-verified at both `container` and `hostloop`)"
1564
+ if name in LIVE_VERIFIED_HOOK_EVENTS
1565
+ else "is a hook event the agent accepts, and it loads a plugin's own `hooks/hooks.json` "
1566
+ "itself — though whether a harness run ever reaches this event's trigger has not "
1567
+ "been verified here"
1568
+ )
1536
1569
  findings.append(Finding(
1537
1570
  "INFO", "hook-event-not-served",
1538
- f"`{name}` fires here — a plugin's own `hooks/hooks.json` is loaded and executed by the "
1539
- f"agent binary (live-verified at both `container` and `hostloop`) — but cowork-harness "
1571
+ f"`{name}` {fires} — but cowork-harness "
1540
1572
  f"itself installs only {', '.join(sorted(SERVED_HOOK_EVENTS))} on `initialize`. Two "
1541
1573
  f"consequences: there is no assertion key for this event, so a scenario cannot GATE on it; "
1542
- f"and the harness does not reproduce the additional `{name}` hooks real Cowork installs, so "
1543
- f"anything driven by those is absent here.",
1544
- "Your hook still runs — this is about assertability, not breakage. To gate on its effect, "
1545
- "assert the OBSERVABLE result instead (a file it writes, a tool it blocks), not the hook "
1546
- "itself.",
1574
+ f"and if real Cowork installs a `{name}` hook of its own, the harness does not reproduce it, "
1575
+ f"so anything driven by that is absent here. (Cowork installs hooks of its own for "
1576
+ f"PreToolUse, PostToolUse and UserPromptSubmit only.)",
1577
+ "The harness does not block your hook — this is about assertability, not breakage. To gate "
1578
+ "on its effect, assert the OBSERVABLE result instead (a file it writes, a tool it blocks), "
1579
+ "not the hook itself.",
1547
1580
  path, line_no,
1548
1581
  ))
1549
1582
  elif name.lower() in {e.lower() for e in KNOWN_HOOK_EVENTS}:
@@ -1558,8 +1591,8 @@ def _lint_hook_events(path):
1558
1591
  else:
1559
1592
  findings.append(Finding(
1560
1593
  "ERROR", "hook-event-unknown",
1561
- f"`{name}` is not a recognized hook event — an unrecognized event name is ignored, so this "
1562
- f"hook would never run on any surface.",
1594
+ f"`{name}` is not a hook event the agent recognizes — an unrecognized event name is "
1595
+ f"ignored, so this hook would never run on any surface.",
1563
1596
  f"Check spelling and capitalization. Valid events: {', '.join(sorted(KNOWN_HOOK_EVENTS))}.",
1564
1597
  path, line_no,
1565
1598
  ))
package/CHANGELOG.md CHANGED
@@ -6,6 +6,179 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [3.5.0] — 2026-09-06
10
+
11
+ **Live verification for this release** (macOS arm64, agent **2.1.260**, agent image `cowork-agent-base:2`,
12
+ Desktop 1.46388.4):
13
+
14
+ | Suite | Result |
15
+ |---|---|
16
+ | `boundary-check` | **6/6** — host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress |
17
+ | `npm run test:live` | **4 files, 18 passed, 1 skipped** — the skip is `live-outputs-delete`'s silent-guard case, which reports SKIPPED when the agent issues no Bash call (documented as expected-rare, not a failure) |
18
+ | `run examples/scenarios/` | **7/7 success** across `container`, `hostloop` and `protocol` |
19
+ | e2e self-tests | **8/8 success** — askuserquestion, multiselect, multiselect-deciderdir, l1-container, l1-egress, present-files, semantic-evidence-files, canary-hostloop |
20
+
21
+ **Two things stated rather than glossed.** (a) The first `test:live` run showed one red — a sub-agent
22
+ WebSearch that the model simply did not perform — and the first scenario batch showed one red on
23
+ `example-pdf-skill`. Both were **model variance**: each passed on re-run with nothing changed, which is
24
+ the disposition the tests themselves prescribe for this class. They are recorded because a suite that
25
+ only ever reports its green run is not evidence of anything. (b) `smoke-l2-microvm` was **not run** this
26
+ release. It is the ninth e2e scenario and the tier CI can never cover (Apple-VZ is macOS-arm64 only); the
27
+ most recent microvm pass remains the one recorded under 3.4.0.
28
+
29
+ ### Added
30
+
31
+ - **`liveVerifiedHookEvents` in the generated `assertion-keys.json` sidecar.** A new key alongside
32
+ `servedHookEvents` and `knownHookEvents`, naming the subset of hook events a plugin's own hook has
33
+ actually been **observed** to fire for in a harness run (three, live-verified at `container` and
34
+ `hostloop`). It exists so the linter can distinguish "the agent's validator accepts this name" from "a
35
+ run reaches this trigger" — two claims a single list would have conflated. Additive: a consumer reading
36
+ only the two existing keys is unaffected, and an older `scenario.py` ignores it.
37
+ - **`test/hook-events-elf-parity.test.ts`** — re-extracts the hook-event list from the staged agent binary
38
+ at test time instead of comparing two hand-maintained files. It **skips where no Desktop is staged**, CI
39
+ included, and says so in its own header: a green CI is not evidence for this invariant.
40
+ - **Two rows in [`docs/invariants.md`](./docs/invariants.md)**: the hook-event invariant below, and
41
+ `provenance.spawnEnvKeys` + `spawnEnvSpreadCount` as the spawn-env drift alarm. The latter documents
42
+ machinery that has existed since Desktop 1.24012.1 but appeared **zero times** outside `src/` — which is
43
+ why a later design exercise set about re-inventing it. [`docs/maintenance.md`](./docs/maintenance.md)
44
+ now tells a maintainer to read those two values in a `sync` diff first, and why a key-set delta beats a
45
+ count: it names the key.
46
+
47
+ ### Fixed
48
+
49
+ - **`lint-skill` and the mount-time hook warning reported 24 valid hook events as misspellings.**
50
+ `KNOWN_HOOK_EVENTS` held **9** names, assembled by grepping the agent binary for event-name constants;
51
+ the agent's own hooks-config validator accepts **33**. A plugin declaring `PostCompact` or
52
+ `MessageDisplay` — both accepted by the agent, both of which run — got "not a recognized hook event …
53
+ Check spelling/capitalization", at **ERROR** severity in `lint-skill` and as a `::warning::` on the run
54
+ path. **The message text changed on both paths**, so a CI job grepping for the old strings needs
55
+ updating: the ERROR now reads "is not a hook event the agent recognizes", and an accepted-but-unserved
56
+ event reports at INFO / `::notice::` with the confident "it WILL fire" wording reserved for the three
57
+ events actually observed firing here. The list is now sourced from the validator array itself, and
58
+ `test/hook-events-elf-parity.test.ts` re-reads it from the staged agent binary rather than from a
59
+ committed fixture, because a fixture-vs-const test compares two hand-maintained files and only moves
60
+ when someone re-extracts by hand — the step that had failed. That test **skips where no Desktop is
61
+ staged**, CI included, so a green CI is not evidence for this invariant.
62
+ - **The sub-agent model precedence was documented backwards in three places.** `docs/session.md`,
63
+ `docs/subagents.md` and `src/session.ts` stated `env > dispatch param > frontmatter > inherit`. The
64
+ binary resolves **dispatch param > frontmatter > env > inherit** (verified in agent 2.1.260; promoting
65
+ the env override to the top is what `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` does, and the Cowork spawn sets
66
+ neither that flag nor `CLAUDE_CODE_COORDINATOR_FORCE_WORKER_INHERIT_MODEL`). So
67
+ `agent_env.subagent_model` does **not** outrank a sub-agent's own `model:` frontmatter, contrary to
68
+ what the docs promised.
69
+ - **Two model-forcing env vars leaked asymmetrically into `hostloop` and `protocol`.**
70
+ `SCRUBBED_AGENT_ENV_KEYS` matches exact keys, so `CLAUDE_CODE_SUBAGENT_MODEL` never covered
71
+ `CLAUDE_CODE_SUBAGENT_MODEL_FORCE` or `CLAUDE_CODE_COORDINATOR_FORCE_WORKER_INHERIT_MODEL`. An
72
+ operator with either exported got different sub-agent model resolution on the two inheriting tiers than
73
+ on `container`/`microvm`, which is the asymmetry that constant exists to prevent. Both are now scrubbed.
74
+ **Operator-visible consequence, and the reason it is filed as a fix rather than a cleanup:** unlike the
75
+ three keys already on that list, these two have **no `agent_env` knob**, so exporting one in your shell
76
+ now has it silently deleted with nothing authored to put it back. `docs/session.md` says so. If you
77
+ need either, set it inside the run rather than in the environment the harness inherits. Real Cowork
78
+ sets neither. Companion tests drive the real `hostloop` and `protocol` env builders with each key set —
79
+ not a hand-built object — and pin the exact-key *mechanism*, plus the deliberate decision to leave
80
+ `CLAUDE_CODE_COORDINATOR_MODE` (which enables one of them, and leaks the same way) unscrubbed while no
81
+ coordinator surface is modelled.
82
+
83
+ ### Documentation
84
+
85
+ - **Corrected what this repo says about Cowork's force-ask PreToolUse hook.** It gates **nine** tools,
86
+ not four (the four named ones plus `create`/`update`/`delete_scheduled_task` and
87
+ `start`/`stop_watching`), and its decision is **not** unconditional: two gate-conditioned early returns
88
+ defer to the auto-mode permission classifier, so in a real auto-mode session 7 of the 9 raise no
89
+ prompt. The scheduled-task branch ships in Desktop 1.22209.0 — before 1.24012.9, the first baseline to
90
+ record this hook — so the note in `spawn.hooks` was already wrong for 5 of the 9 tools when it was
91
+ first written; the second branch lands three baselines later at 1.26832.0, making it wrong for all
92
+ nine from there on. Inaccurate in all fourteen baselines carrying it, either way.
93
+ **Nothing about the harness changes**: auto mode is structurally unreachable here, so for every mode a
94
+ scenario can express production still answers `ask`, and serving that hook unconditionally would remain
95
+ faithful. Corrected in `desktop-1.46388.3` forward; the older baselines keep their wording.
96
+ - **`spawn.hooks` is documented as hand-pinned documentation, not as a drift tripwire.** It cannot be
97
+ one — `sync` spreads `spawn` forward from the base baseline, so the field carries through untouched and
98
+ nothing re-derives it from the app bundle. The claim was disproved by its own subject, above.
99
+ - **The desktop-local lane boundary is stated as a missing transport flag rather than an entrypoint
100
+ string.** `--sdk-url` is absent from the entire app bundle and present 41× in the agent, and the
101
+ `ccr-session` host resolver throws without it — so features routed through that host (including
102
+ `cowork_memory_context`) are *structurally* unreachable locally, not merely disabled. An entrypoint
103
+ test can be relaxed in one release; a flag Desktop never passes cannot be worked around agent-side.
104
+ - **New fidelity note: the silent-turn reminder.** Whether the agent narrates between tool calls is
105
+ decided by a server-delivered model capability, and that narration lands in the corpus
106
+ `semantic_matches` grades — so an assertion resting on the presence or wording of inter-tool text rests
107
+ on something an account-level capability can change. Assert the observable result instead.
108
+ - **New fidelity note: auto-memory is delivered through four env keys the harness never sets.** When a
109
+ session has an auto-memory directory, Desktop ships it to the agent as `CLAUDE_COWORK_MEMORY_PATH_OVERRIDE`,
110
+ `CLAUDE_COWORK_MEMORY_INDEX_CONTENT`, `CLAUDE_COWORK_MEMORY_EXTRA_GUIDELINES` and — separately gated —
111
+ `CLAUDE_COWORK_MEMORY_GUIDELINES`. The last two carry **prompt text**: the agent reads them into its
112
+ memory prompt, so this is model-visible content rather than configuration. With no memory directory the
113
+ same branch sends `CLAUDE_CODE_DISABLE_AUTO_MEMORY:"1"` instead. `docs/fidelity-gaps.md` now also
114
+ tabulates the **three distinct gates** involved, which are easy to conflate: the one keyed on the memory
115
+ *directory* is not the one governing the guidelines env key, and neither is the one that gates only the
116
+ resolver's third branch — so a session can get a memory directory with that third gate off.
117
+ - **Corrected an overstated divergence in how the agent's credential reaches it.** A comment in
118
+ `src/runtime/host-env.ts` said real Cowork passes only `CLAUDE_CODE_OAUTH_TOKEN`, without scoping the
119
+ claim. Desktop does swap that env var for a file descriptor — staging the token into a `0600` temp file,
120
+ opening it, and unlinking it — but that wrapper is reached from the **host-loop branch only**. At
121
+ `container` and `microvm`, production passes the plain token exactly as this harness does. The comment
122
+ now records the host-loop-only scope and why the divergence is deliberately not emulated: the env var
123
+ remains a first-class credential source for the agent, Desktop's own helper falls back to it on I/O
124
+ failure, and no boundary is crossed at host-loop that was not already open.
125
+ - **Release-channel facts added to the recovery runbook.** `/stable` is a rollout pointer, not "latest";
126
+ not every published version is served (2.1.255 is 404 on both channels while its neighbours are 200),
127
+ so **a 404 is not evidence of a wrong channel** — use a positive control on a neighbouring version; and
128
+ `manifest.zst.json` is served *beside* `manifest.json`, additive rather than a migration.
129
+
130
+ ### Changed
131
+
132
+ - **Parity sync to Desktop 1.46388.4** (agent **2.1.260**, unmoved from 1.46388.3). Baseline
133
+ `desktop-1.46388.4` written with zero unknown deltas; the whole app-bundle delta is three build chunks.
134
+ Two gates newly pinned and both recorded `force`/ON: `builtinToolsApprovableByAutoMode:4202409342` (the
135
+ unpinned sibling of `scheduledTaskToolsApprovableByAutoMode`, and one of the two gates that release
136
+ tools from the force-ask hook) and `cuCanUseToolEnabled:2486083521`, which had moved `off` → `ON` while
137
+ unpinned. Gates deliberately left unpinned now carry their reasoning in `cowork-sync.ts` rather than
138
+ being silently absent. Committed cassettes re-stamped rather than re-recorded, on a field-level diff of the two baselines:
139
+ the entire delta is `appVersion`, `capturedAt`, the `$comment` date, `provenance.asarFingerprint`,
140
+ `provenance.fcache.embeddedTimestamp` and the two new gate rows. Nothing a replay depends on —
141
+ `spawn.env`, `spawn.hooks`, `permissionMode`, `agentBinary.sha256`, the egress allowlist, the mount
142
+ layout — moved at all.
143
+
144
+ ### Security
145
+
146
+ - **Bumped the transitive `fast-uri` 3.1.5 → 3.1.7**, clearing four Dependabot high-severity advisories.
147
+ Lockfile only — no direct dependency changed and no runtime behaviour is affected.
148
+
149
+ ## [3.4.1] — 2026-09-05
150
+
151
+ ### Documentation
152
+
153
+ - **Which Cowork *lane* the harness models is now stated, in the places a consumer reads.** Every
154
+ fidelity tier reproduces the desktop-local lane; Cowork's remote lane runs server-side in a cloud
155
+ container with a different filesystem, shell tool, delivery mechanism and a server-authored prompt.
156
+ Which lane a real session gets is a Cowork setting ("Only on this computer"), and it was observed
157
+ **off** on a current install. README, the companion skill, `docs/fidelity-gaps.md` and
158
+ `docs/maintenance.md` now say so. The distinction that matters: behaviour conclusions (triggering,
159
+ tool sequencing, gate handling) travel between lanes; anything asserting a **path, mount or delivery
160
+ mechanism** is a local-lane claim only. No remote tier is planned — that container is Anthropic's, so
161
+ emulating it would mean authoring an environment rather than reproducing one; `lane: remote` already
162
+ makes the affected assertions refuse to grade instead of passing.
163
+
164
+ ### Changed
165
+
166
+ - **The sub-agent override sentinel's note now carries a current probe.** Gate `124685897` reads ON,
167
+ which only enables a server-delivered replacement of the `## Cowork environment` section. Re-probed
168
+ against the 1.46388.3 composition in a real host-loop session: all three composed parts arrived
169
+ byte-identical to the shipped fallback, so the gate is on with no payload. The two non-overridable
170
+ parts are the control that makes it conclusive. The note also states the probe's precondition, which
171
+ it previously lacked.
172
+
173
+ ### Verification
174
+
175
+ - **`microvm` exercised after the 3.4.0 tag** — `smoke-l2-microvm` passed 3/3 against the *published*
176
+ 3.4.0 artifact (real Apple-VZ VM, separate kernel; guest egress reached allowlisted
177
+ `api.anthropic.com` and denied an off-list telemetry host). That completes all four tiers for the
178
+ 1.46388.3 baseline; DESIGN.md's scope note is re-stamped accordingly. The 3.4.0 entry below is left
179
+ as the record of what was verified at tag time. This tier is macOS-arm64 only and CI runners are
180
+ Linux, so it is only ever exercised by hand.
181
+
9
182
  ## [3.4.0] — 2026-09-05
10
183
 
11
184
  ### Upgrade notes
package/DESIGN.md CHANGED
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
47
47
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
48
48
 
49
49
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
50
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.260**, per `baselines/desktop-1.46388.3.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
50
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.260**, per `baselines/desktop-1.46388.4.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel — from the channel that baseline's `agentBinary.releaseBaseUrl` names, which is **not always the stable path** (Desktop stages release candidates too); the exact command is in `docs/maintenance.md`.
51
51
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
52
52
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
53
53
 
@@ -176,7 +176,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
176
176
 
177
177
  ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
178
178
 
179
- > **Scope of that claim, stated plainly.** `2026-09-05 / desktop-1.46388.3` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 19 passed, 0 skipped**; `run examples/scenarios/` **7/7 success, 39 assertions, 0 failed** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **8/8 success** (smoke-askuserquestion, smoke-multiselect, smoke-multiselect-deciderdir, smoke-l1-container, smoke-l1-egress, smoke-present-files, smoke-semantic-evidence-files, canary-hostloop). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported none. Tiers exercised: **`protocol`, `container`, `hostloop`**. **`microvm` was NOT exercised** — it needs a real VM and CI excludes it for the same reason, so the ninth e2e scenario (`smoke-l2-microvm`) did not run. **Scope-out, so this is not read as more than it is.** (a) `subagent-manifest-probe` is new in this release and is the first live coverage the sub-agent append has ever had; it proves the composed append is ACTIONABLE — a sub-agent reaches the right files through both tool families, and a bare relative write lands at the host cwd — **not** that the harness's paraphrase matches Desktop's wording, which is the `manifest`/`suffix` fingerprint axes' job. (b) A live pass verifies observed behaviour, not the whole spawn contract by construction. (c) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: `example-pdf-skill` was re-recorded against this baseline and the other two committed cassettes were re-stamped (see the CHANGELOG for why each), so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
179
+ > **Scope of that claim, stated plainly.** `2026-09-06 / desktop-1.46388.4` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing committed here is live-unverified for want of a newer run. The pass ran against agent **2.1.260** (the staged VM ELF and the native `.app` both at that version) on macOS arm64, agent image `cowork-agent-base:2`, and covered **four** suites: `boundary-check` **6/6** (host-fs-sealed, direct-egress-denied, allowlist-enforced, allowlist-permits, loopback-not-proxied, hostloop-bash-egress); `npm run test:live` **4 files, 18 passed, 1 skipped**; `run examples/scenarios/` **7/7 success** (csv-fx-normalize, csv-metrics, example-pdf-skill, hostloop-computer-links, protocol-smoke, skill-loads, subagent-manifest-probe); and the e2e self-tests **9/9 success** (smoke-askuserquestion, smoke-multiselect, smoke-multiselect-deciderdir, smoke-l1-container, smoke-l1-egress, smoke-present-files, smoke-semantic-evidence-files, canary-hostloop, smoke-l2-microvm). **Zero skips** in the live lane is the part worth stating: every `describe` there is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported none. Tiers exercised: **all four — `protocol`, `container`, `hostloop` and `microvm`.** The microvm case ran AFTER the 3.4.0 tag, against the published artifact rather than the checkout, and is recorded here rather than in that release's CHANGELOG entry, which states what was verified at tag time. It is the tier CI can never cover — GitHub's runners are Linux and Apple-VZ is macOS-arm64 only — so it is only ever exercised by hand here: `smoke-l2-microvm` passed 3/3 in a real VM with its own kernel, and its guest egress behaved (allowlisted `api.anthropic.com` reached, an off-list telemetry host denied). **Scope-out, so this is not read as more than it is.** (a) `subagent-manifest-probe` is new in this release and is the first live coverage the sub-agent append has ever had; it proves the composed append is ACTIONABLE — a sub-agent reaches the right files through both tool families, and a bare relative write lands at the host cwd — **not** that the harness's paraphrase matches Desktop's wording, which is the `manifest`/`suffix` fingerprint axes' job. (b) A live pass verifies observed behaviour, not the whole spawn contract by construction. (c) These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: `example-pdf-skill` was re-recorded against this baseline and the other two committed cassettes were re-stamped (see the CHANGELOG for why each), so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
180
180
 
181
181
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
182
182
 
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^3.4.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^3.5.0"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.4.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.4.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^3.5.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^3.5.0"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v3`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Not sure a harness is what you need? The next two sections are the argument.
@@ -215,6 +215,32 @@ L2 microvm parity Optional. Agent inside a real Linux microVM (Lima/Apple-VZ
215
215
  synced baseline. "Do what real Cowork does."
216
216
  ```
217
217
 
218
+ ### Which Cowork *lane* this models — read this before trusting an environment assertion
219
+
220
+ Every tier above reproduces Cowork's **desktop-local lane**: the agent runs on your machine, shell
221
+ commands land in a Linux sandbox rooted at `/sessions/<id>`, attached folders appear under
222
+ `/sessions/<id>/mnt/<name>`, and finished files reach the user through `present_files`.
223
+
224
+ Cowork also has a **remote lane**, where the session runs server-side in an ephemeral cloud container
225
+ that reaches your machine over a link. There the filesystem, the shell tool, and file delivery are all
226
+ different — folders arrive under `$HOME/mnt/`, deliverables go to `/mnt/user-data/outputs/` and are
227
+ handed over with `SendUserFile`, and the environment prompt is authored by the server rather than by
228
+ Desktop. **Which lane you get is a Cowork setting** ("Only on this computer", Settings → Cowork), and
229
+ it has been observed **off** — i.e. remote — on a current install.
230
+
231
+ The harness **cannot execute the remote lane**: that container is Anthropic's, not something a local
232
+ tool can stand up. What it does instead is refuse to fake it. Declare `lane: remote` on a scenario and
233
+ the assertions that depend on observing a local filesystem degrade honestly — `file_absent` reports
234
+ evidence-unavailable rather than passing, and delivery is reported as unobservable — so a green never
235
+ means more than it should.
236
+
237
+ **What this means for you.** Behaviour-shaped conclusions travel between lanes: whether your skill
238
+ triggers, how it sequences tools, which questions it asks, whether it respects a permission gate.
239
+ Environment-shaped conclusions do not: anything asserting a path, a mount, or a delivery mechanism is a
240
+ statement about the **local** lane specifically. Scope your claims accordingly, and if you are probing
241
+ real Cowork to compare, turn "Only on this computer" **on** first or you will be measuring a lane this
242
+ tool does not model.
243
+
218
244
  **Decision guide** — `fidelity:` takes exactly one of these five values (`protocol`/`container`/`microvm` vary isolation strength; `hostloop`/`cowork` are overlays that instead pick *where the loop runs* — there's no combining the two groups):
219
245
 
220
246
  | Question | Choose |
@@ -281,6 +307,7 @@ See [DESIGN.md](./DESIGN.md) for the full parity matrix, the known deltas vs. re
281
307
 
282
308
  ## Limitations
283
309
 
310
+ - **One lane, deliberately.** Every tier models Cowork's desktop-local lane. The remote (cloud) lane runs server-side with a different filesystem, shell tool, delivery mechanism and a server-authored prompt; no tier reproduces it, and `lane: remote` exists to make the resulting blind spots refuse to grade rather than pass. See [Which Cowork lane this models](#which-cowork-lane-this-models--read-this-before-trusting-an-environment-assertion).
284
311
  - **Not the full Desktop network transport.** L1 is a container, not a VM; L2 *is* a real Apple-VZ microVM but still does not reproduce Cowork's gVisor netstack — its egress is the same allowlist proxy as L1 (with a guest iptables firewall in front). If your skill depends on VM-kernel specifics, validate at L2; if it depends on packet-level gVisor behavior, no tier reproduces it.
285
312
  - **Cowork in-guest context is partial.** Desktop supplies host-loop staging, runtime `mountPath` RPC, and the bridge. We reproduce the *filesystem and cowork mode*, not those host-side services. Skills that call Desktop-only host RPCs won't run here (they wouldn't be portable anyway).
286
313
  - **The agent binary is the staged ELF** (`claude-code-vm/<ver>/claude`), **bind-mounted** from your own Claude Desktop install — nothing Anthropic-owned is bundled or installed. There is **no npm path**; override the path with `COWORK_AGENT_BINARY`. Check licensing/ToS for your use.
@@ -353,6 +380,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
353
380
  ## Status
354
381
 
355
382
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
356
- **`desktop-1.46388.3`**. Release-by-release verification notes (what was re-verified against
383
+ **`desktop-1.46388.4`**. Release-by-release verification notes (what was re-verified against
357
384
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
358
385
  this section would otherwise duplicate lives in the sections above.