cowork-harness 2.2.0 → 2.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +43 -17
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +15 -13
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -12
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +32 -1
  9. package/CHANGELOG.md +366 -0
  10. package/DESIGN.md +3 -3
  11. package/README.md +41 -11
  12. package/baselines/desktop-1.37937.1.json +816 -0
  13. package/baselines/prompts/cowork-system-prompt-fingerprints.json +27 -0
  14. package/dist/assert.js +105 -3
  15. package/dist/cli.js +15 -4
  16. package/dist/hostloop/workspace-handler.js +12 -2
  17. package/dist/run/cassette.js +304 -78
  18. package/dist/run/execute.js +135 -8
  19. package/dist/run/trace-view.js +39 -1
  20. package/dist/runtime/container.js +53 -2
  21. package/dist/runtime/hostloop.js +52 -8
  22. package/dist/sync/cowork-sync.js +233 -16
  23. package/dist/types.js +40 -10
  24. package/docs/README.md +3 -0
  25. package/docs/cassette.md +36 -11
  26. package/docs/debugging.md +25 -0
  27. package/docs/fidelity-gaps.md +133 -6
  28. package/docs/maintenance.md +15 -1
  29. package/docs/scenario.md +66 -17
  30. package/examples/replays/README.md +1 -1
  31. package/examples/replays/example-multiselect-gate.cassette.json +41 -46
  32. package/examples/replays/example-pdf-skill.cassette.json +164 -115
  33. package/examples/replays/hostloop-computer-links.cassette.json +62 -68
  34. package/examples/scenarios/example-pdf-skill.yaml +8 -1
  35. package/llms.txt +3 -1
  36. package/package.json +1 -1
  37. package/python/test_scenario_lint.py +39 -0
  38. package/schema/scenario.schema.json +29 -10
package/CHANGELOG.md CHANGED
@@ -4,6 +4,372 @@ All notable changes to this project are documented here. The format is based on
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
5
5
  [Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
6
6
 
7
+ ## [Unreleased]
8
+
9
+ ## [2.4.0] — 2026-08-27
10
+
11
+ ### Fidelity
12
+
13
+ - **Path resolution: the shell and the file tools use DIFFERENT roots, and the harness now models that.**
14
+ Measured on desktop-local Cowork 2026-08-27: `mcp__workspace__bash` starts every call at the bare
15
+ session root (`/sessions/<id>`), while the agent process sits at the outputs dir, so a bare `Write`
16
+ lands in `mnt/outputs` and is user-visible immediately. The two are different path spaces — a file the
17
+ shell creates with a relative path is *not* where a relative `Write` puts one.
18
+
19
+ `hostloop` ran its workspace bash at `${sessionRoot}/mnt/<firstFolder ?? outputs>`, collapsing the two.
20
+ The replaced derivation came from the asar's `cwd: c.vmCwd` spawn argument, which is **not
21
+ load-bearing on the cowork path** — only the `chat` branch prepends an explicit `cd ${vmCwd}`, which
22
+ would be redundant if the argument worked. It reproduced a prompt claim rather than an observed
23
+ behaviour. Both cwds now come from one function (`hostLoopCwds`) and are pinned together in a single
24
+ test: a single-value assertion cannot express "shell and file tools disagree, on purpose", which is the
25
+ contract every previous version of this bug flattened.
26
+
27
+ **What it changes for you:** a skill that writes deliverables from a shell script using relative paths
28
+ looked correct at `hostloop` and delivered nothing in production. It now fails here too.
29
+
30
+ - **`container` models production's VM-loop `web_fetch` swap.** When `coworkWebFetchViaApi` resolves on
31
+ (true in every shipped baseline), the tier registers a workspace SDK-MCP server exposing **web_fetch
32
+ only**, disallows the built-in `WebFetch`, and aliases the name to `mcp__workspace__web_fetch`
33
+ (`VM_LOOP_TOOL_ALIASES`). Bash is deliberately untouched — "Bash is the only tool that truly diverges
34
+ between loops" — which is why `container` keeps the built-in shell while `hostloop` replaces both.
35
+
36
+ The three parts ship together by necessity: disallowing the built-in without the alias turns a fidelity
37
+ fix into a regression, because the bare name stops resolving instead of landing on the workspace tool.
38
+ The tool is advertised but deliberately **not** pre-approved — production's VM-loop registration passes
39
+ the same approval hook the host loop does, so the call is gated at `can_use_tool`. Pre-approving it
40
+ would make a scripted `webfetch:<domain>` answer, and `decide: deny` on it, silently inert.
41
+
42
+ Its fetch is bound to the session egress allowlist and to URL provenance, exactly as the host loop's is.
43
+ Both matter: this fetch runs in the harness's own process, outside the container network namespace,
44
+ so the sidecar proxy never sees it and only these two gates constrain it. The handler's allowlist now
45
+ defaults to **deny-all** rather than `["*"]`, so a caller that forgets to pass one gets nothing through
46
+ instead of everything.
47
+
48
+ The workspace handler now gates **dispatch** on the same set it advertises, not just `tools/list`. An
49
+ unadvertised tool that still executed when named was a real hole: the VM-loop registration exposes
50
+ web_fetch only, and a `bash` call arriving there must be refused rather than quietly exec'd into the
51
+ container.
52
+
53
+ **`microvm` is unchanged** and still offers the built-in `WebFetch` — `spawnMicroVm` does not receive
54
+ the gate. **`chat` is unchanged too**: the swap applies to `run`/`record`, so `chat --fidelity container`
55
+ still offers the built-in and the two surfaces differ at the same declared tier.
56
+
57
+ ### Upgrade impact
58
+
59
+ - **At `fidelity: container` under `run`/`record`, `WebFetch` is no longer in the offered tool set.** An assertion naming it
60
+ no longer describes a callable tool. `tool_not_called: WebFetch` is the dangerous direction: it now
61
+ passes **vacuously** rather than failing loudly, so a scenario that was genuinely testing "this skill
62
+ does not fetch from the web" silently stops testing anything. Rename to
63
+ `tool_not_called: mcp__workspace__web_fetch`. The shipped `example-pdf-skill` scenario carried exactly
64
+ this defect and is fixed.
65
+
66
+ `verify-cassettes` grew a `replaced-builtin` note for this class: it reads a cassette's recorded init
67
+ inventory and reports built-in names the current build no longer offers at that tier. It is a **note,
68
+ not a finding** — the swap is gate-conditional, so a recording made with the gate off is legitimately
69
+ different, and an init event carrying no `tools` key is *no evidence* rather than a missing surface.
70
+ Neither should be told to re-record.
71
+
72
+ This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
73
+ no command, flag, schema or exit-code meaning changes. The tier's modelled tool inventory is a fidelity
74
+ property, and moving it toward production is the project's purpose — so it ships in a minor.
75
+
76
+ - **DEPRECATION: `fidelity:` becomes REQUIRED in the next major.** A scenario that omits it currently
77
+ defaults to `container`, which models the **VM loop** — but production runs the **host loop** (gate
78
+ `1143815894` is force-ON in every shipped baseline). So an omitted key silently measures the scenario
79
+ against a lane real users are not on, and the two lanes differ in exactly the ways that bite: where a
80
+ bare relative path lands, where the shell starts, and which tools are offered.
81
+
82
+ Both `run` and the skill's `scenario.py` lint now warn (`fidelity-defaulted`). The check reads the RAW
83
+ document, because Zod's `.default()` makes an omitted key indistinguishable from a deliberate
84
+ `fidelity: container` — and those two deserve different treatment. Naming the tier explicitly silences
85
+ it; `fidelity: container` remains a valid, non-warning choice.
86
+
87
+ ### Fixed
88
+
89
+ - **The semantic judge was told which authored files were never delivered.** The authored-file capture
90
+ deliberately includes the scratchpad, but production discards anything outside `mnt/` — "never reaches
91
+ the user or your file tools". Unlabelled, a rubric like "the report was written" graded TRUE on a file
92
+ the user never receives: a false green inside the one evaluator that reads free-form prose and cannot
93
+ infer the convention. Scratch files are now tagged `— SCRATCH, NOT delivered to the user`, with a note
94
+ explaining they are evidence of what the run DID and not that anything was delivered.
95
+
96
+ - **`sync` no longer refuses to write on the `subagentPromptServerOverride` gate.** Gate-ON only enables
97
+ the lookup; the payload that would actually override is delivered **per session by the server** and
98
+ appears in neither the asar, the fcache nor `config.json`. So gate state alone cannot separate
99
+ "override active" from "gate on, no payload, fallback still correct" — the guard blocked forever
100
+ rather than tripping, because it could never clear itself from its own inputs. It is now a
101
+ non-blocking note that says what it can and cannot know.
102
+
103
+ Settled by a live sub-agent probe instead: a real sub-agent's environment section matched the committed
104
+ paraphrase on all four load-bearing claims, so no override was reaching that account. That is evidence,
105
+ not proof — one account, one session, and a server rule can be segment-targeted. The note says so, and
106
+ says to re-probe if the sub-agent append matters to what you are shipping.
107
+
108
+ ### Documentation
109
+
110
+ - **[docs/scenario.md](./docs/scenario.md) now states where a relative path actually lands**, as a
111
+ measured table per tier, replacing a sentence that was simply wrong for the file tools. The guidance
112
+ that follows it is lane-dependent, because the correct answer inverts between lanes: on the desktop
113
+ host loop the file tools are already rooted at `outputs/`, so writing `outputs/x.md` doubles to
114
+ `outputs/outputs/x.md` and the user never sees it — a **bare filename** is right there. At
115
+ `container`/`microvm` a bare name lands in the scratchpad instead. Addressing a connected folder by
116
+ name never reaches it on either lane: it builds a same-named decoy inside `outputs`, reports success,
117
+ and gives no signal.
118
+
119
+ - **[docs/fidelity-gaps.md](./docs/fidelity-gaps.md)** gained the path-resolution split, and its
120
+ "VM tiers have no workspace tool aliases" entry is updated — closed at `container`, still open at
121
+ `microvm`.
122
+
123
+ - The skill's undelivered-deliverables guidance no longer hardcodes the literal prefix `outputs/`, which
124
+ was wrong on the lane production actually runs.
125
+
126
+ - **The product's vocabulary is mapped to this project's**, because one directory had four names and none
127
+ of them was the one Cowork's UI shows. "Working folder" is the user-visible roots (`outputs/` plus each
128
+ connected folder); "Scratchpad" is everything outside `mnt/`; `{{workspaceFolder}}` is a prompt token
129
+ that renders to the *first* user-visible root. Also recorded: Cowork's **Scratchpad panel is an activity
130
+ log, not a location listing** — it lists files as "wrote to" wherever they landed, so a file appearing
131
+ there is not evidence it was undelivered.
132
+
133
+ - **`Write`'s tool result echoes the raw path it was given and never absolutizes** (read from the agent
134
+ binary). Nothing in the harness parses a path out of a `Write` result; this is recorded so nothing
135
+ starts, since such an assertion would be reading something production does not emit. Cowork's own
136
+ chat-surface prompt claims the opposite, so the product's description of its own tool is wrong here.
137
+
138
+
139
+ ## [2.3.0] — 2026-08-26
140
+
141
+ ### Parity
142
+
143
+ - **Baseline `desktop-1.37937.1` (agent `2.1.246`).** `sync` had been refusing to write on three
144
+ `spawn.env` deltas; all three are now classified.
145
+
146
+ **Two new pinned keys.** The Cowork spawn sets `CLAUDE_CODE_PROMPT_CACHE_TTL="1h"` and
147
+ `CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL="5m"` unconditionally — no gate, no session or deployment
148
+ branch — so every first-party session receives them. They are **additive**:
149
+ `ENABLE_PROMPT_CACHING_1H="1"` is still set alongside. They read zero times in agent `2.1.241` and six
150
+ times each in `2.1.246`, so the contract went live one agent release after Desktop began setting it.
151
+
152
+ **`MCP_TOOL_TIMEOUT` is now classified per SITE, not per key.** Its first-party construction is
153
+ unchanged (still resolving to `180000`), but the third-party-only branch gained a second, settings-
154
+ conditional construction of the same key whose value expression the const resolver cannot reach.
155
+ Allowlisting the key — the obvious fix — would have been a silent contract loss: the allowlist is
156
+ checked *before* the pin list, so the key would have vanished from the generated env entirely, and it
157
+ is not a `REQUIRED_SPAWN_KEYS` member, so nothing would have hard-failed. Instead the 3p-only branch is
158
+ located by content and its inner keys are classified by name without resolving their values, which is
159
+ what the branch already meant. A brand-new key there still hard-fails.
160
+
161
+ **`CLAUDE_PREVIEW_CLASSIFIER_FLOOR` is now inert on the shipping agent** — recorded, not changed. Agent
162
+ `2.1.246` renamed the flag it reads to `CLAUDE_CHROME_CLASSIFIER_FLOOR` (and its consumer field
163
+ `previewClassifierFloorEnabled` → `chromeClassifierFloorEnabled`) while Desktop still sets only the old
164
+ name, so the classifier floor now falls through to its GrowthBook default. The key stays pinned: the
165
+ baseline records what the spawn constructs, and a Desktop-side rename must surface as a diff line
166
+ rather than as silence. Nothing in the harness reads it behaviourally. Together with the cache-TTL keys
167
+ above this is the same lesson pointing both ways — Desktop and the agent version the spawn env
168
+ independently, so "Desktop sets X" and "the agent reads X" are separately-dated claims.
169
+
170
+ Also found in the same pass and recorded in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md), with no
171
+ harness surface: the deliberately-unmodeled remote device-tool family gained `device_fs` and
172
+ `device_request_delete_permission` plus a folder-access announce mode — the Desktop half of the six
173
+ `cowork_*` risk categories that had appeared in agent `2.1.237`'s auto-mode rubric.
174
+
175
+ Also verified unchanged against the retained previous asar: the Cowork system prompt (the retained raw
176
+ files for 1.32885.1 through 1.37937.1 are byte-identical on disk), both sub-agent appends, `tools[]`,
177
+ the `canUseTool` chain, the mount modes, the egress contract, and the VM rootfs image.
178
+
179
+ - **All three committed example cassettes re-recorded against this baseline** — `example-pdf-skill`
180
+ (`container`), `example-multiselect-gate` (`protocol`) and `hostloop-computer-links` (`hostloop`).
181
+ The container one shows no behavioural change (transcript wording only). `verify-cassettes` is clean,
182
+ and the recorded MCP inventory is the Cowork lane's own servers (`cowork`, `plugins`, `skills`,
183
+ `workspace`) — no host inventory reached the fixtures, checked independently of the built-in scan
184
+ because the 2026-08-04 leak hid from a `mcp__` grep and surfaced only through NAME fields.
185
+
186
+ - **Live end-to-end pass re-run against this baseline.** `npm run test:live` — **4 suites, 19 assertions,
187
+ 19 green, 0 skipped** — across `protocol`, `container` and `hostloop`, on agent `2.1.246`. Every
188
+ `describe` in that lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated
189
+ case reports as skipped rather than passing vacuously; zero skips means the whole population executed.
190
+ `DESIGN.md`'s scope note is re-stamped accordingly and now records that no baseline is unverified for
191
+ want of a live run. Not claimed: the `boundary-check` sandbox proof and the example-scenario suite were
192
+ part of the previous stamp and were not run here.
193
+
194
+ ### Added
195
+
196
+ - **`question_context: {when_question?, matches}` — assert what a gate actually put in front of the user.**
197
+ A regex tested against a gate's founder-visible payload FIELD BY FIELD — the question label, every option
198
+ **label**, and every option **description**, each as a separate string. Matching per field rather than over
199
+ one joined blob is deliberate and load-bearing: a pattern cannot straddle two fields, so
200
+ `invoicing[\s\S]*Audit logging` will not match by stitching one option's description to the next option's
201
+ label — a "sentence" nobody was shown. The neighbouring transcript keys' docs teach `[\s\S]` for spanning
202
+ turns, so that is the habit an author brings here. `matches` is required NON-EMPTY: an empty pattern
203
+ compiles to `//i` and would green any run that fired a gate at all. `question_asked` matches question text only and `question_options`
204
+ compares labels only, so a sentence the model delivered inside an option's `description` was invisible to
205
+ every assertion key — a false-negative generator for any skill that puts context there, which the tool's
206
+ own schema invites. Measured on a consumer's paid run: a producer-authored sentence arrived verbatim in
207
+ the question, reworded inside it, and relocated into the proceed option's `description` across three runs
208
+ of one scenario; the third redded a lane on a run where the founder had in fact been told.
209
+
210
+ Evidence is the **ask-time** `AskUserQuestion` payload, never a `tool_result` — a skill's producer
211
+ typically also writes the same sentence into its own gate-state file, so `tool_result_matches` on that
212
+ phrase grades true whether or not the model ever surfaced it. Unlike `question_options`, omitting
213
+ `when_question` on a multi-gate run is **not** ambiguous: this key asks whether the text was shown at all.
214
+ Zero gates recorded fails; unreadable gate evidence fails evidence-unavailable, never vacuously.
215
+
216
+ ### Changed
217
+
218
+ - **`trace --view questions` now renders each gate's offered options — labels AND `description`s** — under
219
+ an `offered:` block, sub-question-labelled on a bundled gate. It previously printed the question label
220
+ alone, so the option `description` a skill routinely puts the deciding sentence in was reachable only by
221
+ hand-reading `events.jsonl`; a reader who found nothing in the view could reasonably conclude the text was
222
+ never delivered. The payload was always recorded — this was purely a rendering gap, and the same run dir
223
+ answers the question either way. The row (`--output-format json`) gains `subQuestions[]` carrying the
224
+ untruncated ask-time options; the text view caps each description at 240 chars, so nothing is lost, only
225
+ wrapped. This pairs with `question_context`: the view is how you *find* the text, that key is how you
226
+ *gate* on it.
227
+
228
+ - **The `tool_use` blindness of `transcript_contains`/`_not_contains`/`_matches`/`_not_matches` and
229
+ `computer_links_resolve`/`_if_present` is now documented and guarded.** `semantic_matches` has carried a
230
+ ⚠️ spelling out that its corpus excludes every `tool_use` — "a rubric claim about whether a tool was
231
+ called is unassertable" — while the six keys with the identical blindness said only "the assistant
232
+ transcript". A consumer wrote a `transcript_matches` against text living in a gate question; it could not
233
+ have matched at any phrasing, and the recording fail-closed after the spend. The caveat is now a property
234
+ of an enumerable set (`TOOL_USE_BLIND_KEYS`) enforced across every surface that documents a key — the
235
+ docs tables, the zod `.describe()` behind `assertions --list` and the generated JSON schema, and the
236
+ packaged skill reference — so a newly-added blind key cannot ship without it.
237
+ - **`question_asked`/`question_options`/`question_context` now warn that they match model-authored text.**
238
+ Gate question text and option labels are composed by the model and reworded run to run. The `choose:`
239
+ side already documented this (stable leading anchor, 1-based index); the assert side documented it
240
+ nowhere. Guarded by `MODEL_AUTHORED_TEXT_KEYS`.
241
+ - **A bad regex in a NESTED assertion field is now caught at load, not after the paid spawn.** The pre-compile
242
+ pass reached only top-level string keys, so every regex one level down — `artifact_text.matches`,
243
+ `artifact_text.not_matches`, `path_denied.path_matches`, `skill_tool_used.skill`/`.tool`,
244
+ `subagent_dispatch_healthy.type`, `subagent_output_contains.match`, `task_status.match`,
245
+ `question_options.when_question`, and the new `question_context.*` — was first compiled inside the
246
+ evaluator. All eleven are now validated at load, and `test/nested-regex-leaves.test.ts` reads `assert.ts`
247
+ and fails if the evaluator compiles a nested leaf the load-time table does not carry, so the gap cannot
248
+ reopen silently. `lint`'s double-quoted-regex warning also now covers nested `matches:` leaves.
249
+ - **`diff` with a single positional now names the missing operand** instead of printing bare usage. There is
250
+ deliberately no one-argument form: `diff` is polymorphic over baselines, run dirs and cassettes, so it
251
+ would need type dispatch plus a defined source for "the committed version".
252
+ - **The two cost keys are cross-referenced.** `RunResult.cost.usd` is one invocation's SDK
253
+ `total_cost_usd`; the critique report's `costUsd.totalUsd` aggregates the task turn, the reflection turn
254
+ and both evaluator passes. Reading the wrong one returns `undefined` rather than erroring, which reads as
255
+ "no cost recorded". Kept as two shapes on purpose — collapsing them would destroy the per-phase split.
256
+
257
+ ### Fixed
258
+ - **Two spawn-contract sentinels were weaker than their names implied; both now bite.**
259
+
260
+ `checkSpawnContractFacts` pinned the `allowedTools[]` built-in head and the built-in→`mcp__` boundary
261
+ but nothing between the boundary and the closing bracket — so `mcp__plugins__search_connectors` was
262
+ added to the array and both checks stayed green. The `mcp__` membership is now pinned as a set, and the
263
+ flag names the added and removed entries instead of reporting that something "moved". (That tool is
264
+ declared only on the third-party deployment, so the first-party inventory the harness serves is
265
+ unaffected — but the addition should not have been invisible.)
266
+
267
+ `checkMountModeFacts` asserted each read-only mount with a single `regex.test`, while `uploads`,
268
+ `.claude/skills` and `.projects/<uuid>` are each built at **two** sites — the VM-loop mount-set builder
269
+ and host-loop `computeBashMounts`. Either site satisfied the check, so a one-lane `ro`→`rw` flip — a
270
+ containment change on exactly one execution tier — passed green. It now compares the site count to the
271
+ count carrying `mode:"ro"` and names the lane count in the flag. Its fifth fact — the delete-deny
272
+ resolver `?"rwd":"rw"` — had the same shape and the same two-lane reality (it went 1 site to 2 in this
273
+ release) and now guards a floor on that count: a lane losing the resolver flags, a lane gaining one
274
+ does not.
275
+
276
+ - **All eight asar sentinels now carry a committed mutation case.** They are green on the previous
277
+ release's asar too, so a green proves nothing on its own unless the checker is known to bite; five of
278
+ them (`checkCodeTripwires`, `checkWebFetchFacts`, `checkEgressContractFacts`, `checkSyspromptMapFacts`,
279
+ `checkNormalizationSanity`) had no such case. Each new one changes an **inner** character of its
280
+ anchor, because a suffix rename still satisfies a substring regex — a mutation that cannot fail is not
281
+ evidence — and asserts that the mutation actually applied.
282
+
283
+ - **`record --dry-run <dir/>` no longer announces a WARN as a refusal.** The batch arm labelled every
284
+ advisory note `⚠ would-refuse (advisory)`, but `cassettePortabilityPreflight` can only ever return
285
+ `ok`/`warn` — it has no refuse path at all — and `hostInventoryPreflight` returns `warn` whenever the
286
+ target cassette already exists, which is every re-record corpus sitting at the default path. So the
287
+ preview told operators the real `record` would refuse runs it would in fact accept, on the one arm whose
288
+ design principle is that a guess must not gate. The label now follows the verdict kind
289
+ (`⚠ would-warn (advisory)`), the "ADVISORY, not this run's verdict" footer is unchanged, and exit codes
290
+ were never affected either way.
291
+
292
+ - **The 3p-branch rule refuses to blank W1.** W1 is the window every modeled first-party key is derived
293
+ from, so a branch marker appearing there hard-fails instead of blanking; deleting real pinned keys from
294
+ the derived env with nothing failing is the worse of the two outcomes.
295
+
296
+ - **`record --dry-run` now runs every pre-spend refusal the real record runs, and one refusal moved from
297
+ after the paid run to before it.** The rehearsal re-implemented the checks by hand, so it drifted:
298
+ `hostInventoryPreflight` shipped 2026-08-04 and the commit three days later titled *"make `--dry-run`
299
+ refuse what the real record refuses"* swept in the two checks returning `string | undefined` and missed
300
+ the one returning a `{kind}` verdict — for 19 days an operator could not discover that refusal without
301
+ spending. Separately, the slug-collision refusal (*"refusing to overwrite … it belongs to scenario X"*)
302
+ sat **after** `executeScenario`: you paid for the run and were then refused, though it is a pure function
303
+ of path + `--force` + the existing cassette's name. Both now run in one shared pre-spend block.
304
+
305
+ Refusals are also uniform now. `promptPolicyRejection` threw while `hostInventoryPreflight` called
306
+ `fail()`, which `process.exit`s — and the dir-batch loop catches a throw per item but cannot catch an
307
+ exit, so a host-inventory refusal mid-batch abandoned concurrent runs already paid for.
308
+
309
+ Under `--output-format json` a directory batch carries those advisory verdicts in a `notes[]` array, kept
310
+ separate from `refusals[]` so automation cannot read a guess as a binding verdict.
311
+
312
+ **On a directory, the path-dependent verdicts are ADVISORY (`⚠ would-refuse`) and do not affect the exit
313
+ code**, because a directory target takes no `--out`: the preview would have to guess the destination, and
314
+ a guess must not gate. Measured on a real consumer, gating on that guess would have refused 26 of 27
315
+ scenarios the real record accepts. Dry-run a single scenario file with the real flags for a binding
316
+ answer. `--quiet` suppresses the notes; it never suppresses a refusal.
317
+
318
+ Also new in the batch preview: the duplicate-cassette-path refusal the real batch already had.
319
+
320
+ ### Upgrade impact
321
+
322
+ - **`--allow-host-inventory-fixture` no longer waives a MEASURED host-inventory finding.** It was one flag
323
+ doing two jobs: bypassing the pre-flight refusal (a precondition the operator cannot check — "use only
324
+ when the session has no personal MCP servers or plugins") *and* downgrading the write-time scan's refusal
325
+ to a warning. So an operator who passed it to get past the undecidable precondition also switched off the
326
+ scan that would have caught a real leak. It is now the pre-flight bypass only; the finished recording is
327
+ still scanned, and a finding still refuses the write and quarantines the recording. Writing a flagged
328
+ recording needs the new, narrower `--allow-host-inventory-findings`. **A batch recorder that passes the
329
+ old flag on a new-fixture path will now abort on a genuine finding where it previously warned and wrote.**
330
+
331
+ A `--dry-run` that reports what *would* be captured was considered and declined: the inventory does not
332
+ exist until the agent has run, so a preview would either re-implement the scanner against a hypothetical
333
+ (a second oracle free to disagree with the real one) or require the spend it was meant to avoid. Reasoning
334
+ is in [docs/cassette.md](./docs/cassette.md).
335
+
336
+ This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
337
+ no command or flag is removed and no exit-code meaning changes — one flag's consent narrows, and the
338
+ capability it shed is reachable through the new one. So it ships in a minor.
339
+
340
+ ### Documentation
341
+
342
+ Five gaps found by asking a consumer which harness properties actually changed an outcome on a real
343
+ working day, then checking whether the docs said so. Four of the five were documented only as features,
344
+ never as the failure they prevent — the sentence a reader needs to recognise their own situation.
345
+
346
+ - **The blocking gate is now stated as a blocker.** `AskUserQuestion` *blocks*: it is a question to a
347
+ human and `claude -p` has no human, so a gated skill under a plain CLI run stalls or never reaches the
348
+ code behind the gate. That made half the skill untestable, not merely awkward to test — the README had
349
+ only "untestable headless unless something answers it", buried as the last of two afterthought bullets.
350
+ - **"Assert on the run, not the output"** — a new README section naming the class of claim the harness
351
+ exists for (`subagent_tool_absent`, `dispatch_count_max`, `no_delete_in_outputs`, `subagent_file_write`)
352
+ and why no output diff can reach it: a correct run and one that quietly handed a restricted sub-agent
353
+ shell access produce byte-identical files. `subagent_tool_absent` did not appear in the README at all.
354
+ - **The raw-`events.jsonl` escape hatch is documented**, with verified `jq` recipes, in
355
+ [docs/debugging.md](./docs/debugging.md). Every documented route into the event log went through
356
+ `trace`, whose views are a digest — so the run dir's *evidence* surface is wider than any view's
357
+ *observation* surface, and wider still than the assertion catalog. Concluding "it never happened" from
358
+ a view that doesn't render the field is a false negative, and the doc now says so and shows the read.
359
+ - **Sub-agent delivery has a route.** The tier-qualified outputs contract — the reason a hand-off path
360
+ that works on one loop lands in sandbox scratch on the other — was correctly documented in
361
+ [docs/subagents.md](./docs/subagents.md) but filed under "read on demand", reachable only by someone
362
+ who already knew the answer. It now has a "Common tasks" row keyed to the symptom.
363
+ - **Replay's cost claim carries a number.** "Zero spend" is the price; the wall-clock is what makes an
364
+ always-on per-PR gate obviously affordable, and no doc stated it (well under a second per cassette).
365
+
366
+ - **The project's provenance is stated on every surface the framing travels to.** `README.md` (which
367
+ renders on the npm page), `llms.txt`, both `.claude-plugin/marketplace.json` descriptions and `SKILL.md`
368
+ now say the same thing: an independent project, not affiliated with, endorsed by, or supported by
369
+ Anthropic; it bundles no Anthropic code; it is not Cowork. `SKILL.md` additionally tells the agent to say
370
+ so when a user asks what it is. Nothing enforces the four staying in step, so an edit can still drop it
371
+ from one of them unnoticed.
372
+
7
373
  ## [2.2.0] — 2026-08-25
8
374
 
9
375
  ### Upgrade impact
package/DESIGN.md CHANGED
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
47
47
  [docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
48
48
 
49
49
  - VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
50
- - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.237**, per `baselines/desktop-1.34493.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
50
+ - In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.246**, per `baselines/desktop-1.37937.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
51
51
  - Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
52
52
  - Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
53
53
 
@@ -174,9 +174,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
174
174
 
175
175
  > **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
176
176
 
177
- ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.234**, the native host app that `hostloop` runs is **2.1.234**, baseline **`desktop-1.32885.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-19, superseding the prior `1.32352.0` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
177
+ ### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
178
178
 
179
- > **Scope of that claim, stated plainly.** `2026-08-19 / desktop-1.32885.1` is the last baseline carrying a **full live end-to-end pass**, and it is no longer the newest committed baseline: **one** baseline has shipped since (`1.34493.1`), **one** of which moved the agent ELF, most recently to **2.1.237** — so the newest baseline is **not** live-verified, and this paragraph describes the 1.32885.1 pass only. The pass ran on agent `2.1.234` and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe and resume continuity on the native binary). It was one invocation of `npm run test:live` — **4 suites, 24 assertions, 20 green / 4 skipped** — plus the `boundary-check` sandbox proof (all six constraints enforced) and the example-scenario suite (6/6 across the same three tiers). Two caveats, so this is not read as more than it is: the four `live-outputs-delete` cases that **skipped** skipped because the agent issued no Bash call at all — it declined to run the pinned command, a model-behaviour miss the suite itself classifies as not a guard defect — so this is *every suite exercised*, not *every live assertion green*; and a live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — the `live-outputs-delete` skip set differs run to run for exactly that reason, which is why those cases skip loudly rather than fail. One such flake was root-caused rather than retried: the `hostloop` uploads-are-`Read`-able case counted bash calls, so its verdict turned on which exploration tool the model happened to pick (an `ls` of the uploads directory failed it; a `Glob` in the next run passed), even though the upload was `Read` directly and no outputs-delete fired either time. It now inspects what bash actually did and fails only when a command names the uploaded file — the workaround it exists to catch — so exploration is free and `cat`/`cp` of the upload still trips it. **One assertion has NOT passed against this baseline:** `live-outputs-delete`'s "a whole-line `#` comment is prose, not an executable delete" skipped in **all three** runs made against it — the full-lane pass plus two targeted re-runs of that suite. The other three cases that skipped in the full-lane run each passed in a re-run, which is the model variance described above; this one did not, and by the suite's own rule a skip that persists across runs means the model has stopped being willing to execute that pinned command, so the case needs its intent made unambiguous or needs retiring — it is a scenario-maintenance debt, not a guard defect, and it is recorded here rather than rounded away. Every other assertion in the lane has passed at least once against this baseline. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
179
+ > **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing in `baselines/` is currently unverified for want of a live run. The pass ran against agent **2.1.246** (the staged VM ELF and the native `.app` are both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
180
180
 
181
181
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
182
182
 
package/README.md CHANGED
@@ -11,6 +11,11 @@
11
11
  [![Built with Skill Creator Plus](https://img.shields.io/badge/Built_with-Skill_Creator_Plus-4ecdc4)](https://github.com/yaniv-golan/skill-creator-plus)
12
12
  [![Agent Skills compatible](https://img.shields.io/badge/Agent_Skills-compatible-4A90D9)](https://agentskills.io)
13
13
 
14
+ > **Unofficial.** An independent project, not affiliated with, endorsed by, or supported by Anthropic.
15
+ > "Claude" and "Claude Cowork" are Anthropic's. This harness emulates Cowork's *observable runtime
16
+ > contract* and drives Anthropic's own agent binary from your local Claude Desktop install — it
17
+ > bundles no Anthropic code, and it is not Cowork.
18
+
14
19
  Scriptable, CI-friendly test harness that reproduces **Claude Cowork's observable runtime contract** closely enough to test the skills you write — across many scenarios, headless, in CI — without the (locked) Desktop app. It reproduces not just Cowork's *behavior* but its *limitations*: sealed filesystem, default-deny egress, MCP-only cross-boundary — so a green test has cleared the constraints that break skills in Cowork. That is a far stronger signal than a bare `claude -p` run, and it is not a guarantee: this is an emulator of the contract, and the deliberate divergences are catalogued in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md).
15
20
 
16
21
  And because every run is recorded, you get the thing a transcript can't give you: **evidence of what the agent actually did, not what it said it did** — which skill was invoked (if any), which files a sub-agent really read, which hosts it reached, which options a person was really shown. See [Why not just `claude -p` or the Agent SDK?](#why-not-just-claude--p-or-the-agent-sdk).
@@ -83,8 +88,8 @@ Both are excellent, and if you're building an *agent* you should use them. This
83
88
 
84
89
  Two more that matter in practice:
85
90
 
86
- - **Token-free CI.** Record a run once, commit the cassette, and every PR re-runs it deterministically at **zero spend** — assertions, tool stream, gate answers and all. That's what makes an always-on gate affordable. See [cassette.md](./docs/cassette.md).
87
- - **Deterministic answers to questions and permission prompts.** A skill that asks the user something is untestable headless unless something answers it the way the UI would. The harness answers over the same `can_use_tool` control protocol Desktop uses, from your scenario's scripted `answers:` — or from a live decider when you're still discovering what it asks. See [scenario.md](./docs/scenario.md), [decider-dir.md](./docs/decider-dir.md).
91
+ - **A blocking gate you can actually answer.** `AskUserQuestion` *blocks*: it is a question to a human, and `claude -p` has no human. A gated skill under a plain CLI run either stalls at the gate or never reaches the code behind it — so the half of your skill that lives past the first question is untestable, not merely awkward to test. The harness answers over the same `can_use_tool` control protocol Desktop uses, from your scenario's scripted `answers:` — or from a live decider when you're still discovering what it asks — so the gate genuinely fires and is answered deterministically. `on_unanswered:` decides what an *un*scripted gate does (fail loud by default). See [scenario.md](./docs/scenario.md), [decider-dir.md](./docs/decider-dir.md).
92
+ - **Token-free CI.** Record a run once, commit the cassette, and every PR re-runs it deterministically at **zero spend** — assertions, tool stream, gate answers and all. A cassette replays in well under a second with no token, no Docker and no model call (the three shipped examples replay in ~0.6s total), which is what makes an always-on per-PR gate affordable. See [cassette.md](./docs/cassette.md).
88
93
 
89
94
  **What it doesn't do.** It runs and records; it does not design your experiment. Comparing "with skill" against "without skill" credibly — scrubbing tells, shuffling, judging blind, unblinding after grading — is still yours to build; the harness contributes the run execution and the control arm (`--ablate-skill`, one arm per invocation). And it emulates the *contract*, not the Desktop runtime: see [Limitations](#limitations) and [fidelity-gaps.md](./docs/fidelity-gaps.md) for what it deliberately does not reproduce.
90
95
 
@@ -115,7 +120,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
115
120
 
116
121
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
117
122
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
118
- > From a global install (`npm i -g "cowork-harness@^2.2.0"`), point at the package root instead:
123
+ > From a global install (`npm i -g "cowork-harness@^2.4.0"`), point at the package root instead:
119
124
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
120
125
  > (or copy the cassette into your own project and pass that path).
121
126
 
@@ -125,7 +130,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
125
130
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
126
131
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
127
132
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
128
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.2.0"`.
133
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.4.0"`.
129
134
 
130
135
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
131
136
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -150,7 +155,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
150
155
  claude plugin install cowork-harness@cowork-harness
151
156
  ```
152
157
 
153
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.2.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
158
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.4.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
154
159
 
155
160
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
156
161
 
@@ -175,7 +180,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
175
180
  ones that still need a source checkout. (The marketplace
176
181
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
177
182
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
178
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.2.0"` — see
183
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.4.0"` — see
179
184
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
180
185
 
181
186
  ### Prerequisites for anything above `protocol` fidelity
@@ -293,7 +298,7 @@ cowork-harness lint examples/scenarios/*.yaml --strict --min-severity WARN
293
298
  > | When | Assertions |
294
299
  > |---|---|
295
300
  > | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
296
- > | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
301
+ > | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
297
302
  > | Only if the cassette carries an `artifacts` manifest | `file_exists`, `artifact_text`, `user_visible_artifact`, `artifact_json`, `computer_links_resolve`, `computer_links_resolve_if_present`, `no_unexpected_files`, `input_unmodified` |
298
303
  > | Always skipped (live-only) | `file_absent`, `egress_*`, `expect_denied`, `no_delete_in_outputs`, `no_delete_in_mounts`, `self_heal_ran`, `transcript_no_host_path`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back` — keep these in a periodic live `run` |
299
304
  >
@@ -493,7 +498,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
493
498
 
494
499
  **Flags worth knowing** (the full list is always `<command> --help`):
495
500
  - `run`: a decider (`--decider-cmd <helper>`/`--decider-dir <dir>`, or a scenario's `on_unanswered: llm` — a scenario-YAML key, not a CLI flag; `run` rejects `--on-unanswered llm`) can answer unscripted gates; `--repeat N` (2-100) runs each scenario N times and aggregates a variance rollup instead of a single pass/fail (`--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd` tune the batch verdict/loop — and `--max-budget-usd` applies without `--repeat` too, as a pre-flight refusal from the scenario's cost history); `--matrix <matrix.yaml>` runs ONE scenario across the cross-product of baseline/model/skill_dir axes (worked example: `examples/matrices/csv-metrics-matrix.yaml`; `--max-cells`/`--concurrency` tune the cap/pool — any cell failing, assertion or infra, fails the run); `--compact`/`--demo` trim `run`/`skill` output for shareable screenshots/GIFs; `--label <tag>` (both `run` and `skill`) stamps a generation name for the iterate-across-fixes loop — surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and `status.json`, alongside the auto-recorded `skillCommit` (git provenance) and the authoritative `fingerprint.skillHash` a harvest step should pair critiques by.
496
- - `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `record` also **scans the finished recording** after redaction and before the write, and quarantines it to `<runs-root>/quarantine/` (with a `.findings.txt` naming what leaked) rather than writing a `host-inventory`/`machine-inventory` leak to a repo-visible path — `--allow-host-inventory-fixture` is the consent to write it anyway; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error). `replay` also prints an **advisory** note — never a failure — when the rootfs agent image differs from the one the cassette recorded (`environment.agentImage`): the image decides which capabilities exist, so replaying against a different one can move a verdict with nothing in the cassette having changed. It compares the registry digest first (the only identity comparable across machines) and falls back to the local image config id only when neither side has one; cassettes recorded before that field existed simply carry no image to compare.
501
+ - `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `record` also **scans the finished recording** after redaction and before the write, and quarantines it to `<runs-root>/quarantine/` (with a `.findings.txt` naming what leaked) rather than writing a `host-inventory`/`machine-inventory` leak to a repo-visible path — `--allow-host-inventory-fixture` bypasses only the *pre-flight* refusal (tier + destination path, checked before the spend) and deliberately leaves that scan in force — writing a recording the scan flagged needs the separate `--allow-host-inventory-findings`; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error). `replay` also prints an **advisory** note — never a failure — when the rootfs agent image differs from the one the cassette recorded (`environment.agentImage`): the image decides which capabilities exist, so replaying against a different one can move a verdict with nothing in the cassette having changed. It compares the registry digest first (the only identity comparable across machines) and falls back to the local image config id only when neither side has one; cassettes recorded before that field existed simply carry no image to compare.
497
502
  - `verify-cassettes`: privacy scan (email/currency/domain/path/machine-inventory) + staleness — exit 1 when verification RAN and found a real problem (a finding, a genuine drift, or scenario-prompt drift), exit 3 when it could NOT complete (an unverifiable-class staleness finding, a too-new cassette format, or a read error) — note that a cassette which fails SHAPE validation is still **privacy-scanned**, because the scan needs a readable transcript rather than a valid document, and each result carries `privacyScanned` saying whether it actually ran (`findings: []` with `privacyScanned: false` is an absence of evidence, not evidence of absence); whole-token allows via `--allow <regex>` (a pattern) / class-scoped `--allow-domain` / `--allow-email` / `--allow-path` / `--allow-machine-inventory` / `--allow-patterns-file <path>` (a **file** of patterns, one regex per line); `--skip-privacy` or `--skip-staleness` runs only part of the gate; a diverged scenario `prompt` vs. the cassette's frozen prompt is also a hard fail (its own `scenarioDrift` bucket), opt out with `--skip-scenario-drift`; `--margins` adds a per-cassette recorded-vs-budget report for count-bound assertions (a single-sample estimate — diagnostic only, never changes the gate verdict); `--allow-empty` makes an **existing but cassette-free directory** exit 0 instead of the default loud exit 2 — for a repo that deliberately commits no cassettes (a *missing* path still fails, so the flag can never green a typo'd path).
498
503
  - `stats`: reads `<runsRoot>/index.jsonl`, written automatically at every result; `--since`/`--baseline`/`--branch` filter; `--skill-hash <prefix>`/`--label <tag>` narrow to one generation of the iterate-across-fixes loop and `--group-by scenario|skill-hash|label|fidelity` splits per generation — or per effective fidelity tier — instead of aggregating across them (a window spanning >1 generation or >1 tier warns; under `skill-hash`/`label` grouping, rows lacking the field are excluded from grouping and counted, never bucketed blank — the `fidelity` key is total, so nothing is ever excluded under that grouping); `--runs` lists the individual runs behind each summary with their `skillHash`/`runLabel` (each `runs[]` entry also carries `fidelity` — the tier that run actually executed at, `effectiveFidelity ?? fidelity` — in the `--output-format json` envelope only; the text listing is unchanged); `--last <n>` windows per-group; `--reindex` rebuilds the index from the physical run-dir tree (the migration path for pre-index runs), reconstructing each critique's cost roll-up from its run dir along the way.
499
504
  - `diff`: `--changelog` renders known-field prose for a baseline diff; `--view tools|transcript|artifacts|meta` narrows a run/cassette diff to one section; normalization (default on) masks per-run noise (ids/timestamps/session markers/host paths) so two runs of the same scenario diff as identical — `--no-normalize` compares raw values.
@@ -573,6 +578,31 @@ egress:
573
578
 
574
579
  Multiple scenarios × sessions × platform baselines = your regression matrix. Drop YAML in `scenarios/` and CI runs them all.
575
580
 
581
+ ### Assert on the run, not the output
582
+
583
+ The assertions above that check a file or a phrase are the familiar half. The half that's hard to get any
584
+ other way asserts on **how the run behaved** — a property of the execution that leaves no trace in the
585
+ deliverable:
586
+
587
+ | Claim | Key |
588
+ |---|---|
589
+ | no sub-agent ever got `Bash` (a deliberately tool-restricted dispatch stayed restricted) | `subagent_tool_absent: Bash` |
590
+ | nothing was deleted from the outputs mount | `no_delete_in_outputs: true` |
591
+ | the dispatch count stayed bounded — no fan-out blow-up | `dispatch_count_max: <N>` |
592
+ | a sub-agent wrote the file, rather than the main loop doing it for them | `subagent_file_write: {path}` |
593
+ | the skill actually ran, rather than the model answering from its mounted source | `skill_triggered: <rx>` |
594
+ | the sandbox really blocked the network | `egress_denied: <host>` |
595
+
596
+ **You cannot diff your way to "no sub-agent used `Bash`."** A correct run and a run that quietly gave a
597
+ restricted sub-agent shell access produce byte-identical outputs; the difference exists only in the event
598
+ stream. This is the class of property that matters most for a skill whose sub-agents are deliberately
599
+ tool-restricted, or whose safety story is "it can't reach X" — and it is exactly what an output-diffing
600
+ test, or a human reading the final answer, will always report as fine.
601
+
602
+ The full catalog is `cowork-harness assertions --list`; [scenario.md → Which assertion for which question](./docs/scenario.md#which-assertion-for-which-question-goal--key)
603
+ is the goal-first chooser. Mind the two axes in each key's row: which **tier** it needs, and whether it
604
+ **survives `replay`** (several of the above are live-only — keep them in a periodic live `run`).
605
+
576
606
  ## Sandboxing: container vs. the real VM
577
607
 
578
608
  Local Cowork runs the agent in an **Apple Virtualization.framework microVM** (separate kernel). The harness's default `container` tier uses an OS container (shared kernel, namespaces/cgroups). For **testing skills you wrote**, that's faithful where it counts — same agent binary, same cowork mode, same mount layout, same egress allowlist, same permission protocol — because skill behavior is agent-loop + tool behavior, all kernel-invisible. The container is the right default precisely because it's CI-native; a VM needs nested virtualization most shared CI runners don't have.
@@ -760,7 +790,7 @@ jobs:
760
790
  - uses: actions/checkout@v4
761
791
  - name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
762
792
  run: |
763
- V=2.1.237 # match your scenario's pinned baseline's agentVersion
793
+ V=2.1.246 # match your scenario's pinned baseline's agentVersion
764
794
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
765
795
  chmod +x "$RUNNER_TEMP/claude-$V"
766
796
  echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
@@ -772,7 +802,7 @@ jobs:
772
802
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
773
803
  ```
774
804
 
775
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.2.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
805
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.4.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
776
806
 
777
807
  The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
778
808
 
@@ -940,6 +970,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
940
970
  ## Status
941
971
 
942
972
  The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
943
- **`desktop-1.34493.1`**. Release-by-release verification notes (what was re-verified against
973
+ **`desktop-1.37937.1`**. Release-by-release verification notes (what was re-verified against
944
974
  which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
945
975
  this section would otherwise duplicate lives in the sections above.