cowork-harness 2.3.0 → 2.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.3.0
7
- tracks-harness: cowork-harness 2.3.0 (baseline desktop-1.37937.1)
6
+ version: 2.4.0
7
+ tracks-harness: cowork-harness 2.4.0 (baseline desktop-1.37937.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.3.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.4.0` (baseline
29
29
  > `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.3.0"`. **Pin `@^2.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.4.0"`. **Pin `@^2.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -147,9 +147,9 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
147
147
  | Tier | What it gives you | Use when |
148
148
  |---|---|---|
149
149
  | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
150
- | `container` | Real sandbox + real default-deny egress (**default**) | Most functional + boundary tests. |
151
- | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity | Testing untrusted code escape, not network behavior. |
152
- | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with native Bash/WebFetch disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
150
+ | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm`/`protocol` never offer it, so the same `tool_not_called` is vacuous there — moving a scenario between tiers can silently void a web-fetch assertion | Most functional + boundary tests. |
151
+ | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
152
+ | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
153
153
 
154
154
  Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
155
155
  (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
@@ -550,8 +550,15 @@ Recognize these before "fixing" a non-bug:
550
550
  only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
551
551
  **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
552
552
  scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
553
- **The fix is lane-dependent.** On `lane: local`, write deliverables under `outputs/` or a connected
554
- folder, or deliver them explicitly. **On `lane: remote`, moving a file under `outputs/` does NOT help** —
553
+ **The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
554
+ can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
555
+ (what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
556
+ `outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
557
+ `fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
558
+ bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
559
+ a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
560
+ decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
561
+ [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
555
562
  nothing is delivered by location there, so only an explicit delivery counts. Assert
556
563
  **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
557
564
  downloaded inputs) rather than a delivery gap.
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.3.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "2.4.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.3.0"
70
+ - run: npm i -g "cowork-harness@^2.4.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -342,7 +342,7 @@ jobs:
342
342
  with: { node-version: '24' }
343
343
  - uses: actions/setup-python@v5
344
344
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
345
- - run: npm i -g "cowork-harness@^2.3.0"
345
+ - run: npm i -g "cowork-harness@^2.4.0"
346
346
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
347
347
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
348
348
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -371,7 +371,7 @@ jobs:
371
371
  echo "live=true" >> "$GITHUB_OUTPUT"
372
372
  fi
373
373
  - if: steps.guard.outputs.live == 'true'
374
- run: npm i -g "cowork-harness@^2.3.0"
374
+ run: npm i -g "cowork-harness@^2.4.0"
375
375
  - if: steps.guard.outputs.live == 'true'
376
376
  run: cowork-harness run scenarios/ --output-format json
377
377
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.3.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.4.0`
4
4
  (baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.3.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.4.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -488,6 +488,30 @@ def lint_doc(doc, path, raw_lines):
488
488
 
489
489
  fidelity = (doc.get("fidelity") or "container")
490
490
  lane = (doc.get("lane") or "local")
491
+
492
+ # W: no `fidelity:` — the default models the WRONG LANE.
493
+ # `container` (the schema default) models VM-loop; production runs HOST-LOOP, gate 1143815894 is
494
+ # force-ON in every shipped baseline. So an omitted key measures the scenario against a lane real
495
+ # users are not on: the file tools resolve a bare relative path differently, the shell starts
496
+ # somewhere else, and the offered tool set differs (measured 2026-08-27).
497
+ # Read the KEY, not the resolved value: `fidelity: container` is a deliberate choice and must not warn.
498
+ # DEPRECATION — `fidelity:` becomes REQUIRED in the next major; this is the warning window.
499
+ if "fidelity" not in doc:
500
+ findings.append(
501
+ Finding(
502
+ "WARN",
503
+ "fidelity-defaulted",
504
+ "no `fidelity:` — defaulting to `container`, which models the VM-LOOP lane. Production "
505
+ "runs HOST-LOOP by default (gate 1143815894), so this scenario is likely measured "
506
+ "against a lane your users are not on.",
507
+ "Name a tier: `fidelity: hostloop` to match production, `fidelity: cowork` to auto-pick "
508
+ "the way Cowork does, or `fidelity: container` to keep today's behaviour deliberately. "
509
+ "Switching tiers can COST you assertions: `no_scratchpad_leak` is container-only (an "
510
+ "error elsewhere) and `transcript_no_host_path` fails by design at hostloop/protocol. "
511
+ "The default is being removed — `fidelity:` becomes REQUIRED in the next major.",
512
+ path,
513
+ )
514
+ )
491
515
  items = _assert_items(doc)
492
516
  assert_keys = _all_assert_keys(items)
493
517
  has_expect_denied = bool(doc.get("expect_denied"))
package/CHANGELOG.md CHANGED
@@ -6,6 +6,136 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [2.4.0] — 2026-08-27
10
+
11
+ ### Fidelity
12
+
13
+ - **Path resolution: the shell and the file tools use DIFFERENT roots, and the harness now models that.**
14
+ Measured on desktop-local Cowork 2026-08-27: `mcp__workspace__bash` starts every call at the bare
15
+ session root (`/sessions/<id>`), while the agent process sits at the outputs dir, so a bare `Write`
16
+ lands in `mnt/outputs` and is user-visible immediately. The two are different path spaces — a file the
17
+ shell creates with a relative path is *not* where a relative `Write` puts one.
18
+
19
+ `hostloop` ran its workspace bash at `${sessionRoot}/mnt/<firstFolder ?? outputs>`, collapsing the two.
20
+ The replaced derivation came from the asar's `cwd: c.vmCwd` spawn argument, which is **not
21
+ load-bearing on the cowork path** — only the `chat` branch prepends an explicit `cd ${vmCwd}`, which
22
+ would be redundant if the argument worked. It reproduced a prompt claim rather than an observed
23
+ behaviour. Both cwds now come from one function (`hostLoopCwds`) and are pinned together in a single
24
+ test: a single-value assertion cannot express "shell and file tools disagree, on purpose", which is the
25
+ contract every previous version of this bug flattened.
26
+
27
+ **What it changes for you:** a skill that writes deliverables from a shell script using relative paths
28
+ looked correct at `hostloop` and delivered nothing in production. It now fails here too.
29
+
30
+ - **`container` models production's VM-loop `web_fetch` swap.** When `coworkWebFetchViaApi` resolves on
31
+ (true in every shipped baseline), the tier registers a workspace SDK-MCP server exposing **web_fetch
32
+ only**, disallows the built-in `WebFetch`, and aliases the name to `mcp__workspace__web_fetch`
33
+ (`VM_LOOP_TOOL_ALIASES`). Bash is deliberately untouched — "Bash is the only tool that truly diverges
34
+ between loops" — which is why `container` keeps the built-in shell while `hostloop` replaces both.
35
+
36
+ The three parts ship together by necessity: disallowing the built-in without the alias turns a fidelity
37
+ fix into a regression, because the bare name stops resolving instead of landing on the workspace tool.
38
+ The tool is advertised but deliberately **not** pre-approved — production's VM-loop registration passes
39
+ the same approval hook the host loop does, so the call is gated at `can_use_tool`. Pre-approving it
40
+ would make a scripted `webfetch:<domain>` answer, and `decide: deny` on it, silently inert.
41
+
42
+ Its fetch is bound to the session egress allowlist and to URL provenance, exactly as the host loop's is.
43
+ Both matter: this fetch runs in the harness's own process, outside the container network namespace,
44
+ so the sidecar proxy never sees it and only these two gates constrain it. The handler's allowlist now
45
+ defaults to **deny-all** rather than `["*"]`, so a caller that forgets to pass one gets nothing through
46
+ instead of everything.
47
+
48
+ The workspace handler now gates **dispatch** on the same set it advertises, not just `tools/list`. An
49
+ unadvertised tool that still executed when named was a real hole: the VM-loop registration exposes
50
+ web_fetch only, and a `bash` call arriving there must be refused rather than quietly exec'd into the
51
+ container.
52
+
53
+ **`microvm` is unchanged** and still offers the built-in `WebFetch` — `spawnMicroVm` does not receive
54
+ the gate. **`chat` is unchanged too**: the swap applies to `run`/`record`, so `chat --fidelity container`
55
+ still offers the built-in and the two surfaces differ at the same declared tier.
56
+
57
+ ### Upgrade impact
58
+
59
+ - **At `fidelity: container` under `run`/`record`, `WebFetch` is no longer in the offered tool set.** An assertion naming it
60
+ no longer describes a callable tool. `tool_not_called: WebFetch` is the dangerous direction: it now
61
+ passes **vacuously** rather than failing loudly, so a scenario that was genuinely testing "this skill
62
+ does not fetch from the web" silently stops testing anything. Rename to
63
+ `tool_not_called: mcp__workspace__web_fetch`. The shipped `example-pdf-skill` scenario carried exactly
64
+ this defect and is fixed.
65
+
66
+ `verify-cassettes` grew a `replaced-builtin` note for this class: it reads a cassette's recorded init
67
+ inventory and reports built-in names the current build no longer offers at that tier. It is a **note,
68
+ not a finding** — the swap is gate-conditional, so a recording made with the gate off is legitimately
69
+ different, and an init event carrying no `tools` key is *no evidence* rather than a missing surface.
70
+ Neither should be told to re-record.
71
+
72
+ This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
73
+ no command, flag, schema or exit-code meaning changes. The tier's modelled tool inventory is a fidelity
74
+ property, and moving it toward production is the project's purpose — so it ships in a minor.
75
+
76
+ - **DEPRECATION: `fidelity:` becomes REQUIRED in the next major.** A scenario that omits it currently
77
+ defaults to `container`, which models the **VM loop** — but production runs the **host loop** (gate
78
+ `1143815894` is force-ON in every shipped baseline). So an omitted key silently measures the scenario
79
+ against a lane real users are not on, and the two lanes differ in exactly the ways that bite: where a
80
+ bare relative path lands, where the shell starts, and which tools are offered.
81
+
82
+ Both `run` and the skill's `scenario.py` lint now warn (`fidelity-defaulted`). The check reads the RAW
83
+ document, because Zod's `.default()` makes an omitted key indistinguishable from a deliberate
84
+ `fidelity: container` — and those two deserve different treatment. Naming the tier explicitly silences
85
+ it; `fidelity: container` remains a valid, non-warning choice.
86
+
87
+ ### Fixed
88
+
89
+ - **The semantic judge was told which authored files were never delivered.** The authored-file capture
90
+ deliberately includes the scratchpad, but production discards anything outside `mnt/` — "never reaches
91
+ the user or your file tools". Unlabelled, a rubric like "the report was written" graded TRUE on a file
92
+ the user never receives: a false green inside the one evaluator that reads free-form prose and cannot
93
+ infer the convention. Scratch files are now tagged `— SCRATCH, NOT delivered to the user`, with a note
94
+ explaining they are evidence of what the run DID and not that anything was delivered.
95
+
96
+ - **`sync` no longer refuses to write on the `subagentPromptServerOverride` gate.** Gate-ON only enables
97
+ the lookup; the payload that would actually override is delivered **per session by the server** and
98
+ appears in neither the asar, the fcache nor `config.json`. So gate state alone cannot separate
99
+ "override active" from "gate on, no payload, fallback still correct" — the guard blocked forever
100
+ rather than tripping, because it could never clear itself from its own inputs. It is now a
101
+ non-blocking note that says what it can and cannot know.
102
+
103
+ Settled by a live sub-agent probe instead: a real sub-agent's environment section matched the committed
104
+ paraphrase on all four load-bearing claims, so no override was reaching that account. That is evidence,
105
+ not proof — one account, one session, and a server rule can be segment-targeted. The note says so, and
106
+ says to re-probe if the sub-agent append matters to what you are shipping.
107
+
108
+ ### Documentation
109
+
110
+ - **[docs/scenario.md](./docs/scenario.md) now states where a relative path actually lands**, as a
111
+ measured table per tier, replacing a sentence that was simply wrong for the file tools. The guidance
112
+ that follows it is lane-dependent, because the correct answer inverts between lanes: on the desktop
113
+ host loop the file tools are already rooted at `outputs/`, so writing `outputs/x.md` doubles to
114
+ `outputs/outputs/x.md` and the user never sees it — a **bare filename** is right there. At
115
+ `container`/`microvm` a bare name lands in the scratchpad instead. Addressing a connected folder by
116
+ name never reaches it on either lane: it builds a same-named decoy inside `outputs`, reports success,
117
+ and gives no signal.
118
+
119
+ - **[docs/fidelity-gaps.md](./docs/fidelity-gaps.md)** gained the path-resolution split, and its
120
+ "VM tiers have no workspace tool aliases" entry is updated — closed at `container`, still open at
121
+ `microvm`.
122
+
123
+ - The skill's undelivered-deliverables guidance no longer hardcodes the literal prefix `outputs/`, which
124
+ was wrong on the lane production actually runs.
125
+
126
+ - **The product's vocabulary is mapped to this project's**, because one directory had four names and none
127
+ of them was the one Cowork's UI shows. "Working folder" is the user-visible roots (`outputs/` plus each
128
+ connected folder); "Scratchpad" is everything outside `mnt/`; `{{workspaceFolder}}` is a prompt token
129
+ that renders to the *first* user-visible root. Also recorded: Cowork's **Scratchpad panel is an activity
130
+ log, not a location listing** — it lists files as "wrote to" wherever they landed, so a file appearing
131
+ there is not evidence it was undelivered.
132
+
133
+ - **`Write`'s tool result echoes the raw path it was given and never absolutizes** (read from the agent
134
+ binary). Nothing in the harness parses a path out of a `Write` result; this is recorded so nothing
135
+ starts, since such an assertion would be reading something production does not emit. Cowork's own
136
+ chat-surface prompt claims the opposite, so the product's description of its own tool is wrong here.
137
+
138
+
9
139
  ## [2.3.0] — 2026-08-26
10
140
 
11
141
  ### Parity
package/README.md CHANGED
@@ -120,7 +120,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
120
120
 
121
121
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
122
122
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
123
- > From a global install (`npm i -g "cowork-harness@^2.3.0"`), point at the package root instead:
123
+ > From a global install (`npm i -g "cowork-harness@^2.4.0"`), point at the package root instead:
124
124
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
125
125
  > (or copy the cassette into your own project and pass that path).
126
126
 
@@ -130,7 +130,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
130
130
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
131
131
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
132
132
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
133
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.3.0"`.
133
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.4.0"`.
134
134
 
135
135
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
136
136
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -155,7 +155,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
155
155
  claude plugin install cowork-harness@cowork-harness
156
156
  ```
157
157
 
158
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.3.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
158
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.4.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
159
159
 
160
160
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
161
161
 
@@ -180,7 +180,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
180
180
  ones that still need a source checkout. (The marketplace
181
181
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
182
182
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
183
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.3.0"` — see
183
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.4.0"` — see
184
184
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
185
185
 
186
186
  ### Prerequisites for anything above `protocol` fidelity
@@ -802,7 +802,7 @@ jobs:
802
802
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
803
803
  ```
804
804
 
805
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.3.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
805
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.4.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
806
806
 
807
807
  The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
808
808
 
package/dist/assert.js CHANGED
@@ -373,7 +373,7 @@ function capForJudge(text, cap) {
373
373
  return text;
374
374
  return `${text.slice(0, cap)}\n…[${text.length - cap} chars truncated for the judge input budget — evidence beyond this point was NOT shown; do not infer absence from this cut]`;
375
375
  }
376
- function buildJudgedDocument(ctx, includeSubagentText = false) {
376
+ export function buildJudgedDocument(ctx, includeSubagentText = false) {
377
377
  // SCRUB BEFORE CAP: scrub is exact-string replacement, so a secret straddling a cap boundary would
378
378
  // be truncated mid-token and slip past scrub into the doc sent to the (external) judge. Scrub each raw
379
379
  // section FIRST, then cap the already-redacted text — capping redacted content can never re-expose a secret.
@@ -405,8 +405,33 @@ function buildJudgedDocument(ctx, includeSubagentText = false) {
405
405
  parts.push(`## Sub-agent output: ${s(label)}\n${capForJudge(s(text), JUDGE_SUBAGENT_CAP)}`);
406
406
  }
407
407
  }
408
- for (const f of ctx.authoredFiles ?? [])
409
- parts.push(`## Authored file: ${s(f.path)}${f.truncated ? " (truncated)" : ""}\n${s(f.content)}`);
408
+ // AUTHORED is not DELIVERED. The capture deliberately includes the scratchpad (the run DID write those
409
+ // files, and production's sandbox would too), but production DISCARDS anything outside `mnt/` — its own
410
+ // sub-agent prompt: "never reaches the user or your file tools". Unlabelled, a rubric like "the report
411
+ // was written" grades TRUE on a file the user never receives — the harness's own false-green, inside the
412
+ // one evaluator that reads free-form prose and cannot infer the convention.
413
+ //
414
+ // The distinction is already carried in the path: the scratchpad walk emits a synthetic `scratchpad/`
415
+ // prefix (run/artifacts.ts). This only makes it legible to the judge, rather than restructuring the set.
416
+ // `scratchpad` is NOT in RESERVED_MOUNT_NAMES, so a user may connect a folder with that exact name.
417
+ // Its files then arrive workRoot-relative as `scratchpad/…` — indistinguishable by prefix from the
418
+ // synthetic walk. Labelling those would hand the judge a FALSE statement ("NOT delivered" about a file
419
+ // the user does receive) and false-RED a "was the report delivered?" rubric. A wrong claim is worse
420
+ // than a missing one, so the collision disables the label rather than guessing.
421
+ // `?? []` because a partial AssertContext (tests, and any caller building one by hand) may omit this;
422
+ // an absent prefix list means "nothing is user-visible", which keeps the label ON — the safe direction.
423
+ const scratchIsUserVisible = (ctx.userVisiblePrefixes ?? []).some((p) => p === "scratchpad" || p === SCRATCHPAD_PREFIX);
424
+ let sawScratch = false;
425
+ for (const f of ctx.authoredFiles ?? []) {
426
+ const scratch = !scratchIsUserVisible && f.path.startsWith(SCRATCHPAD_PREFIX);
427
+ sawScratch ||= scratch;
428
+ const tag = scratch ? " — SCRATCH, NOT delivered to the user" : "";
429
+ parts.push(`## Authored file: ${s(f.path)}${tag}${f.truncated ? " (truncated)" : ""}\n${s(f.content)}`);
430
+ }
431
+ if (sawScratch)
432
+ parts.push("## Note on scratch files\nFiles marked SCRATCH were written to the session's scratch area, which is " +
433
+ "discarded and never reaches the user. Treat them as working intermediates: they are evidence of what " +
434
+ "the run DID, and are NOT evidence that anything was delivered, saved, shared, or produced for the user.");
410
435
  // Surface authored-file incompleteness to the judge so it never reads an omitted/unreadable file's
411
436
  // ABSENCE as evidence the skill didn't produce it (#14/#16). The verdict is separately forced to
412
437
  // evidence-unavailable in the semantic_matches check; this note keeps a still-produced grade honest.
@@ -228,10 +228,12 @@ const defaultRawFetch = async (url, pinnedAddresses) => {
228
228
  const BASH_DESC = "Run a shell command in the session's isolated Linux workspace. Your connected folders are mounted under {{mnt}}/ — the Shell access section of your system prompt lists the exact path for each folder. Each bash call is independent (no cwd/env carryover). Use absolute paths.";
229
229
  const FETCH_DESC = "Fetch a URL from the session network (subject to the egress allowlist). web_fetch can only retrieve URLs that appeared in a user message or a prior result.";
230
230
  export function makeWorkspaceHandler(opts) {
231
- const { containerName, vmMnt, runner = "docker", webFetchAllow = ["*"], onEgress, onInfraError, provenanceRef, dedup, rawFetch = defaultRawFetch, resolve = defaultResolver, execCwd = vmMnt, } = opts;
231
+ const { containerName, vmMnt, runner = "docker", webFetchAllow = [], // DENY-ALL when unset — see the option doc; an open default was a real hole
232
+ onEgress, onInfraError, provenanceRef, dedup, rawFetch = defaultRawFetch, resolve = defaultResolver, execCwd = vmMnt, } = opts;
232
233
  // Per-handler (per-spawn) latch for the provenance-unenforced warning — was module-level, which
233
234
  // silenced the gap after the first run in a long-lived process. Each fresh handler warns once.
234
235
  const provWarned = { value: false };
236
+ const exposed = opts.tools ?? ["bash", "web_fetch"];
235
237
  const tools = [
236
238
  {
237
239
  name: "bash",
@@ -243,7 +245,7 @@ export function makeWorkspaceHandler(opts) {
243
245
  description: FETCH_DESC,
244
246
  inputSchema: { type: "object", properties: { url: { type: "string" } }, required: ["url"] },
245
247
  },
246
- ];
248
+ ].filter((td) => exposed.includes(td.name));
247
249
  return async (_server, jr) => {
248
250
  const method = jr.method;
249
251
  if (method === "initialize")
@@ -259,6 +261,12 @@ export function makeWorkspaceHandler(opts) {
259
261
  if (method === "tools/call") {
260
262
  const name = jr.params?.name;
261
263
  const a = jr.params?.arguments ?? {};
264
+ // Gate the DISPATCH on the same set as the advertisement, not just the tools/list response. An
265
+ // unadvertised tool that still executes when named is a real hole: the VM-loop registration exposes
266
+ // web_fetch ONLY, and a `bash` call arriving there must be refused, not quietly exec'd into the
267
+ // container.
268
+ if (!exposed.includes(name))
269
+ return { result: textResult(`error: unknown tool "${String(name)}"`, true) };
262
270
  if (name === "bash")
263
271
  return {
264
272
  result: await execInContainer(runner, containerName, execCwd, String(a.command ?? ""), clampTimeout(a.timeout_ms), onInfraError),
@@ -267,6 +275,8 @@ export function makeWorkspaceHandler(opts) {
267
275
  return {
268
276
  result: await fetchViaHost(String(a.url ?? ""), webFetchAllow, onEgress, provenanceRef?.current, provWarned, rawFetch, resolve, dedup),
269
277
  };
278
+ // Unreachable: the dispatch gate above rejects every name outside `exposed`, and `exposed` is a
279
+ // subset of the two handled here. Kept as a total-function backstop if that gate is ever loosened.
270
280
  return { error: { code: -32602, message: `unknown tool: ${name}` } };
271
281
  }
272
282
  return { result: {} }; // ping / notifications
@@ -1484,10 +1484,58 @@ function computeDiscoverySurfaceNote(cassette) {
1484
1484
  `Re-record if this scenario asserts on those tools; harmless otherwise.`,
1485
1485
  ];
1486
1486
  }
1487
+ /** Built-ins this build REPLACES with a workspace MCP tool, per tier. Production swaps different sets per
1488
+ * loop, so the expectation is tier-specific:
1489
+ * host-loop — the patch disallows Bash AND WebFetch and aliases both;
1490
+ * VM-loop — the site registers web_fetch ONLY (gated on coworkWebFetchViaApi) and never touches Bash,
1491
+ * which is why container legitimately keeps the built-in shell. */
1492
+ const REPLACED_BUILTINS_BY_TIER = {
1493
+ // NotebookEdit is in spawnHostLoop's `disallowed` alongside Bash/WebFetch, so a cassette asserting
1494
+ // `tool_not_called: NotebookEdit` at this tier is just as vacuous. It has no workspace replacement —
1495
+ // the tier simply removes it — which the note's wording accommodates.
1496
+ hostloop: ["Bash", "WebFetch", "NotebookEdit"],
1497
+ container: ["WebFetch"],
1498
+ // microvm is deliberately ABSENT, not an oversight: spawnMicroVm never receives `webFetchViaApi`
1499
+ // (execute.ts), so that tier still offers the built-in WebFetch. Listing it here would tell users to
1500
+ // re-record a cassette that faithfully describes what this build produces.
1501
+ };
1502
+ /** A cassette whose recorded init inventory names a built-in this build no longer offers at that tier.
1503
+ *
1504
+ * The gap this closes, hit for real: `example-pdf-skill` recorded `WebFetch` at `container` and asserted
1505
+ * `tool_not_called: WebFetch`. When the harness started modelling production's VM-loop swap, that tool
1506
+ * stopped existing there — so the assertion could never be violated and passed VACUOUSLY, while
1507
+ * `verify-cassettes` exited 0 throughout. Staleness keys on baseline/skillHash, so a fixture can describe
1508
+ * a tool surface the harness no longer produces and nothing notices.
1509
+ *
1510
+ * A NOTE, not a finding, deliberately: the swap is gate-conditional (`coworkWebFetchViaApi`), so a
1511
+ * cassette recorded with the gate off is legitimately different rather than stale. A false positive costs
1512
+ * a line of prose; a false FINDING would red a correct fixture. Mirrors computeDiscoverySurfaceNote. */
1513
+ function computeReplacedBuiltinNote(cassette) {
1514
+ const tools = recordedInitTools(cassette);
1515
+ if (!tools || tools.length === 0)
1516
+ return []; // no evidence — NOT "the tools are missing"
1517
+ const authored = cassette.scenario.fidelity;
1518
+ const tier = cassette.environment?.tier ?? cassette.effectiveFidelity ?? (authored !== "cowork" ? authored : undefined);
1519
+ if (!tier)
1520
+ return [];
1521
+ const replaced = REPLACED_BUILTINS_BY_TIER[tier];
1522
+ if (!replaced)
1523
+ return [];
1524
+ const stale = replaced.filter((name) => tools.includes(name));
1525
+ if (stale.length === 0)
1526
+ return [];
1527
+ return [
1528
+ `replaced-builtin: this cassette's init inventory names ${stale.join(", ")} at ${tier}, which this build ` +
1529
+ `no longer offers there — Bash/WebFetch are replaced by workspace MCP tools (mcp__workspace__*), and ` +
1530
+ `NotebookEdit is removed outright. An assertion naming one of them can no longer be violated: it would ` +
1531
+ `pass VACUOUSLY. Re-record if this scenario asserts on those names; harmless if it does not, and ` +
1532
+ `expected if the recording predates the swap or ran with coworkWebFetchViaApi off.`,
1533
+ ];
1534
+ }
1487
1535
  export function computeStaleness(cassette, cassetteDir, sessionOverride) {
1488
1536
  const tier = computeTierStaleness(cassette);
1489
1537
  const findings = [...tier.findings];
1490
- const notes = [...tier.notes, ...computeDiscoverySurfaceNote(cassette)];
1538
+ const notes = [...tier.notes, ...computeDiscoverySurfaceNote(cassette), ...computeReplacedBuiltinNote(cassette)];
1491
1539
  const fp = cassette.fingerprint;
1492
1540
  // BEFORE the fingerprint guard on purpose — same rationale as the tier check above: fingerprint-less
1493
1541
  // cassettes are the OLDEST, i.e. exactly the population the discovery-surface note targets.