cowork-harness 2.1.0 → 2.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 2.1.0
7
- tracks-harness: cowork-harness 2.1.0 (baseline desktop-1.34493.1)
6
+ version: 2.2.0
7
+ tracks-harness: cowork-harness 2.2.0 (baseline desktop-1.34493.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.1.0` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.2.0` (baseline
26
26
  > `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.1.0"`. **Pin `@^2.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.2.0"`. **Pin `@^2.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "2.1.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "2.2.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
67
67
  GitHub-hosted runners, no token/Docker/agent:
68
68
 
69
69
  ```yaml
70
- - run: npm i -g "cowork-harness@^2.1.0"
70
+ - run: npm i -g "cowork-harness@^2.2.0"
71
71
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
72
72
  # no silent false-greens. WITHOUT --strict this
73
73
  # step cannot fail on a WARN-class rule (e.g.
@@ -340,7 +340,7 @@ jobs:
340
340
  with: { node-version: '24' }
341
341
  - uses: actions/setup-python@v5
342
342
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
343
- - run: npm i -g "cowork-harness@^2.1.0"
343
+ - run: npm i -g "cowork-harness@^2.2.0"
344
344
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
345
345
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
346
346
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -369,7 +369,7 @@ jobs:
369
369
  echo "live=true" >> "$GITHUB_OUTPUT"
370
370
  fi
371
371
  - if: steps.guard.outputs.live == 'true'
372
- run: npm i -g "cowork-harness@^2.1.0"
372
+ run: npm i -g "cowork-harness@^2.2.0"
373
373
  - if: steps.guard.outputs.live == 'true'
374
374
  run: cowork-harness run scenarios/ --output-format json
375
375
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`).
3
+ Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.1.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.2.0`
4
4
  (baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -342,8 +342,8 @@ same set live from the schema.
342
342
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
343
343
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
344
344
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
345
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
346
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
345
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
346
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
347
347
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
348
348
  | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
349
349
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 2.1.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -4,6 +4,117 @@ All notable changes to this project are documented here. The format is based on
4
4
  [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
5
5
  [Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
6
6
 
7
+ ## [2.2.0] — 2026-08-25
8
+
9
+ ### Upgrade impact
10
+
11
+ Two behaviour changes can turn a previously-green run red. Neither breaks a covered surface
12
+ ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) — both make an assertion report
13
+ what it always documented — so they ship in a minor:
14
+
15
+ - **`no_scratchpad_leak` at `container` can now FAIL.** It was measuring containment against a root in a
16
+ different path space, so every presented file classified `leaked: false` and the check passed
17
+ vacuously. A scenario whose skill genuinely leaves a presented file in the scratchpad will now red —
18
+ that is the leak the key exists to catch.
19
+ - **`baseline: desktop-1.11847.5` is now refused at `container`/`hostloop`/`microvm`.** It carries no
20
+ `spawn` block, so those tiers cannot reproduce Cowork's toolset — a run on it launched an agent with no
21
+ file or bash tools and still reported a verdict. Use a `sync`-recorded baseline, or `fidelity: protocol`.
22
+
23
+ ### Fixed
24
+
25
+ - **`present_files_called` no longer reports "the tool was never called" about a run that called it, and a
26
+ hostloop delivery scenario can be recorded with a host-path redaction policy at last.** The assertion
27
+ read presence off `RunResult.presentedFiles`, which is a *classification* of each presented path
28
+ (scratchpad? promoted? leaked?) and requires an absolute path to compute. At `hostloop` a presented path
29
+ is a real host path, so the shipped redaction policy rewrites it to
30
+ `[REDACTED:local-path:<hash>]/mnt/outputs/report.html` — correct, documented, and deliberately ordered so
31
+ the mount tail survives — and the classifier then drops every entry as un-normalizable. The list came
32
+ back empty and the assertion stated, as fact, that the tool had never been called.
33
+
34
+ Because `record` replays the base and redacted cassettes and refuses to write when the verdict differs,
35
+ this was not merely a wrong message: **no cassette asserting `present_files_called` could be recorded at
36
+ that tier at all**, so the assertion had never executed on the replay lane. The refusal was right — the
37
+ redacted cassette genuinely could not support the assert — but the defect it was reporting was in the
38
+ harness, not in the recording.
39
+
40
+ Presence now comes from `RunResult.presentFilesCalls`, a new count of the `present_files` invocations
41
+ that carried a well-formed `file_path`, taken from the tool_use input's shape and never from a path's
42
+ content, so redaction cannot alter it. Classification is untouched: `presentedFiles` still drops what it
43
+ cannot resolve, and `no_scratchpad_leak` still reads it. A run recorded before the field falls back to
44
+ the old `presentedFiles`-non-empty test, so no existing green changes.
45
+
46
+ A run whose every `present_files` call carried an unusable path now reports **cannot verify** rather than
47
+ "never called" — the tool *was* invoked, and the harness already knew it. (The malformed count is
48
+ deliberately kept out of that message: the record self-check normalizes `[REDACTED…]` tokens out of
49
+ failing messages but not digits, so an interpolated count that differed between the two replays would
50
+ refuse a cassette that is otherwise fine to write.)
51
+
52
+ Pinned as an invariant ([docs/invariants.md](./docs/invariants.md)) with the end-to-end record self-check
53
+ as its test anchor — the case that could not be recorded, and so had never run in CI.
54
+
55
+ - **Guest paths are built from the tree the harness stages, not from a baseline's recorded mount layout.**
56
+ `resolveMounts` returned `mountLayout.mntRoot` verbatim — or, when that field was absent and the recorded
57
+ `sessionRoot` already ended in `/mnt`, the session root itself. But the staged tree is always
58
+ `<sessionRoot>/mnt`: `stageWorkspace` creates it there and `dockerRunArgv` nests its read-only binds
59
+ there. On a baseline recording anything else, `--plugin-dir` therefore pointed one directory **above**
60
+ the staged plugin tree, so the plugin under test never loaded. `mntRoot` is now derived from the session
61
+ root, and a recorded layout the harness cannot stage is reported as a fidelity divergence at spawn
62
+ instead of silently composing a path no stager creates.
63
+
64
+ Guest paths now anchor on `sessionRoot` — the bind target — rather than `cwd`, which is only where the
65
+ agent's process starts (production's own working directory is a folder mount or `outputs`, so the two are
66
+ not interchangeable even though every synced baseline records them equal). `dockerRunArgv` takes an
67
+ explicit `agentCwd` for `-w`.
68
+
69
+ Two more derivations of the same rule are gone: `prompt.ts` had a private `/sessions/<id>` +`/mnt`
70
+ literal (so the prompt could describe a tree the runtimes had not staged), and **microvm** read the
71
+ agent's root from the baseline when lima structurally mounts it at `/sessions/<sessionId>` — which put
72
+ `CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages. A baseline recording a cwd that tier
73
+ cannot honour is now **refused**, not warned about: the guest `cd` would otherwise succeed at the wrong
74
+ directory.
75
+
76
+ - **A baseline with no `spawn` block is refused at the sandbox tiers.** That block carries the tool set,
77
+ pre-approvals, effort default and config-dir location, and the `?? []` fallbacks meant a run would launch
78
+ an agent with **no Read/Write/Bash/Skill/Task at all** and still report a verdict. `fidelity: protocol`
79
+ builds its own argv and is unaffected.
80
+
81
+ - **`no_scratchpad_leak` can see a container leak again — the session root was in the wrong path space.** The
82
+ root that `presentedFiles`' promoted/leaked classification is measured from was derived by the caller as
83
+ `<run-dir>/work/session`, a HOST path, and handed to every non-`protocol` tier. But the path space the
84
+ agent reports in is per-tier: at `container` (the default) it runs inside the sandbox and reports
85
+ `/sessions/<id>/…`, and only at `hostloop` does it run natively and report host paths. Measured against a
86
+ host root, no VM path is ever inside the root, so every presented file classified `leaked: false` —
87
+ including the handler's copy-failure branch, which returns the source path unchanged when a file is
88
+ blocked-extension, a directory, or absent. `no_scratchpad_leak` evaluates at `container` and nowhere else,
89
+ so the assertion that exists to catch that leak could not catch it, and `verdict.ts`'s delivery check read
90
+ the same `leaked: false` as a successful delivery.
91
+
92
+ Each runtime now reports the session root it actually launched the agent with, and the run consumes that
93
+ instead of deriving a second one — the two can no longer drift into different spaces. `chat` sets it on
94
+ both serving tiers as well; it never did, so a hostloop chat's `presentedFiles` was inverted in its result
95
+ file. `protocol` and `microvm` serve no `present_files` and keep the cwd fallback.
96
+
97
+ Independently, the classifier now fails CLOSED on a space mismatch: if the agent's cwd is not at or inside
98
+ the session root, the presented batch counts as malformed (`no_scratchpad_leak` → cannot verify) rather
99
+ than being graded against a root it cannot be compared to. `leaked: true` is not derivable in that state
100
+ either, and a `leaked: false` verdict there is exactly the vacuous pass the key exists to prevent.
101
+
102
+ ### Added
103
+
104
+ - **First live coverage for `present_files_called` / `no_scratchpad_leak`.** No scenario in the repo
105
+ asserted either key, which is how a session root in the wrong path space could record `leaked: false`
106
+ for every presented file without anything noticing. `e2e/scenarios/smoke-present-files.yaml` writes a
107
+ file OUTSIDE `mnt/` and delivers it, so the promotion is real and the pair is non-vacuous; it runs in
108
+ CI's live e2e loop. Measured on a live container run: `presentFilesCalls: 1`, promoted `true`, leaked
109
+ `false`.
110
+
111
+ - **`RunResult.presentFilesCalls`** — the count of `present_files` invocations that carried a well-formed
112
+ `file_path`, in `result.json` and [schema/run-result.json](./schema/run-result.json). Content-class, so a
113
+ replay re-drive reproduces it; read by `run`, `replay` and `verify-run` alike, and absent (not `0`) on a
114
+ result written before this release, which is what the assertion's fallback distinguishes. Use it to
115
+ answer "did the agent deliver anything?" from a result file without interpreting `presentedFiles`'
116
+ promoted/leaked classification.
117
+
7
118
  ## [2.1.0] — 2026-08-24
8
119
 
9
120
  ### Changed
package/README.md CHANGED
@@ -115,7 +115,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
115
115
 
116
116
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
117
117
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
118
- > From a global install (`npm i -g "cowork-harness@^2.1.0"`), point at the package root instead:
118
+ > From a global install (`npm i -g "cowork-harness@^2.2.0"`), point at the package root instead:
119
119
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
120
120
  > (or copy the cassette into your own project and pass that path).
121
121
 
@@ -125,7 +125,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
125
125
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
126
126
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
127
127
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
128
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.1.0"`.
128
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.2.0"`.
129
129
 
130
130
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
131
131
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -150,7 +150,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
150
150
  claude plugin install cowork-harness@cowork-harness
151
151
  ```
152
152
 
153
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.1.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
153
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.2.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
154
154
 
155
155
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
156
156
 
@@ -175,7 +175,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
175
175
  ones that still need a source checkout. (The marketplace
176
176
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
177
177
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
178
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.1.0"` — see
178
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.2.0"` — see
179
179
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
180
180
 
181
181
  ### Prerequisites for anything above `protocol` fidelity
@@ -772,7 +772,7 @@ jobs:
772
772
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
773
773
  ```
774
774
 
775
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.1.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
775
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.2.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
776
776
 
777
777
  The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
778
778
 
package/dist/assert.js CHANGED
@@ -1141,10 +1141,25 @@ function check(a, ctx) {
1141
1141
  results.push(fail("present_files_called: present_files is not served on `lane: remote` — a local MCP server cannot reach a remote Cowork session, which delivers via the agent-native SendUserFile instead (not modeled; see docs/fidelity-gaps.md) — cannot verify"));
1142
1142
  else if (ctx.effectiveFidelity !== "container" && ctx.effectiveFidelity !== "hostloop")
1143
1143
  results.push(fail(`present_files_called: present_files is served only on the container/hostloop tiers (this run: ${ctx.effectiveFidelity ?? "unknown"}) — cannot verify; use fidelity: container or hostloop for present_files-based delivery`));
1144
- else if (ctx.presentedFiles === undefined || ctx.presentedFiles.length === 0)
1145
- results.push(fail(`present_files_called: no file was delivered via present_files (the tool was never called)`));
1146
- else
1144
+ // PRESENCE comes from the invocation count, not from `presentedFiles`. The two answer different
1145
+ // questions, and only one of them survives redaction: `presentedFiles` entries are dropped when a
1146
+ // path can't be classified, which a host-path policy guarantees at hostloop (a real host path
1147
+ // redacts to `[REDACTED:…]/mnt/outputs/f`, and the classifier requires an absolute path). Reading
1148
+ // delivery off classification made a redacted recording report "the tool was never called" about a
1149
+ // run that called it three times — and, because record's redaction self-check replays both
1150
+ // cassettes and compares verdicts, no such cassette could be written at all.
1151
+ // `presentedFiles.length` stays as the fallback for a run recorded before the count existed.
1152
+ else if ((ctx.presentFilesCalls ?? 0) > 0 || (ctx.presentedFiles?.length ?? 0) > 0)
1147
1153
  results.push(ok());
1154
+ // Called, but every call's `files` was unusable — the tool WAS invoked, so "never called" would be a
1155
+ // factual claim the harness knows to be false. Mirrors no_scratchpad_leak's malformed branch above.
1156
+ // The count is deliberately NOT interpolated: record's self-check normalizes [REDACTED…] tokens out
1157
+ // of failing messages but not digits, so a count that differs between the base and redacted replays
1158
+ // would trip its message compare and refuse a record that is otherwise fine to write.
1159
+ else if (ctx.evidenceErrors?.presentFilesMalformed)
1160
+ results.push(fail(`present_files_called: present_files WAS called, but no call carried a usable file path (malformed input) — cannot verify delivery`));
1161
+ else
1162
+ results.push(fail(`present_files_called: no file was delivered via present_files (the tool was never called)`));
1148
1163
  }
1149
1164
  if (a.no_skill_triggered !== undefined) {
1150
1165
  const c = compileUserRegex(a.no_skill_triggered);
package/dist/baseline.js CHANGED
@@ -319,25 +319,47 @@ function latestBaselineFile() {
319
319
  files.sort(compareBaselineVersions);
320
320
  return join(BASELINES_DIR, files[files.length - 1]);
321
321
  }
322
+ /** The one segment the harness's staged session tree always adds under the session root. Every stager
323
+ * and every argv builder composes guest paths with it (`stage.ts`/`hostloop-stage.ts` create
324
+ * `<sessionHost>/mnt`, `dockerRunArgv` nests `:ro` binds at `<sessionRoot>/mnt/<mountPath>`), so a
325
+ * guest mnt root that is anything OTHER than `<sessionRoot>/mnt` cannot be produced. */
326
+ export const GUEST_MNT_SEGMENT = "mnt";
322
327
  /**
323
- * Expand the mount layout for a concrete session id.
324
- * cwd/sessionRoot = the session root (e.g. /sessions/<id>); mounts sit under mntRoot
325
- * (/sessions/<id>/mnt) and are returned as ABSOLUTE guest paths.
328
+ * Expand the mount layout for a concrete session id — the layout of the tree THIS HARNESS stages, which
329
+ * is what every guest path must be built from.
330
+ *
331
+ * `mntRoot` is DERIVED as `<sessionRoot>/mnt`, not read from `mountLayout.mntRoot`: the staged tree can
332
+ * only ever be there (see GUEST_MNT_SEGMENT), so honouring a recorded value that says otherwise emits
333
+ * guest paths pointing at directories no stager creates. A baseline recording a different mnt root is a
334
+ * FIDELITY divergence, surfaced by `recordedLayoutDivergence` at spawn, never a path this builds.
335
+ *
336
+ * Guest paths anchor on `sessionRoot` (the bind target), never on `cwd`: production's own working dir is
337
+ * a folder mount or `outputs` rather than the bare session root, so the two are not interchangeable even
338
+ * though every synced baseline currently records them equal. `cwd` is the agent's working directory and
339
+ * nothing else.
326
340
  */
327
341
  export function resolveMounts(baseline, sessionId, projectId = "proj1") {
328
342
  const subst = (s) => s.replace("{sessionId}", sessionId).replace("{projectId}", projectId);
329
343
  const cwd = subst(baseline.mountLayout.cwd);
330
344
  const sessionRoot = subst(baseline.mountLayout.sessionRoot);
331
- // TODO: container.ts:31 computes configGuest independently from sessionRoot and is NOT fixed here;
332
- // the legacy desktop-1.11847.5 baseline (sessionRoot ending in /mnt, no spawn block) still produces
333
- // a double-mnt configGuest path (/sessions/<id>/mnt/mnt/.claude) — that is a separate out-of-scope issue.
334
- const rawSessionRoot = baseline.mountLayout.sessionRoot;
335
- const mntRoot = subst(baseline.mountLayout.mntRoot ?? (rawSessionRoot.endsWith("/mnt") ? rawSessionRoot : `${rawSessionRoot}/mnt`));
336
- return {
337
- cwd,
338
- sessionRoot,
339
- mntRoot,
340
- configDir: `${mntRoot}/.claude`,
341
- mounts: baseline.mountLayout.mounts.map((m) => ({ ...m, mountPath: `${mntRoot}/${subst(m.mountPath)}` })),
342
- };
345
+ const mntRoot = `${sessionRoot}/${GUEST_MNT_SEGMENT}`;
346
+ return { cwd, sessionRoot, mntRoot };
347
+ }
348
+ /** Does this baseline RECORD a guest layout the harness cannot stage? Reads the recorded fields only
349
+ * (never the derived ones), so it stays a statement about the data: a `mntRoot` that is not
350
+ * `<sessionRoot>/mnt`, or a `sessionRoot` that already ends in the mnt segment — the shape that made
351
+ * `resolveMounts` return a root one level above the staged tree. Returns the divergence for the caller
352
+ * to surface, or undefined when the recording is reproducible. */
353
+ export function recordedLayoutDivergence(baseline) {
354
+ // Read STRUCTURALLY: a partially-constructed baseline (`{ spawn: {} }`) is normal at the argv seam, and
355
+ // this check must never be the thing that throws there. No recorded layout ⇒ nothing to diverge from.
356
+ const { sessionRoot, mntRoot } = baseline.mountLayout ?? {};
357
+ if (typeof sessionRoot !== "string")
358
+ return undefined;
359
+ const staged = `${sessionRoot}/${GUEST_MNT_SEGMENT}`;
360
+ if (mntRoot !== undefined && mntRoot !== staged)
361
+ return { recorded: mntRoot, staged };
362
+ if (sessionRoot.endsWith(`/${GUEST_MNT_SEGMENT}`))
363
+ return { recorded: sessionRoot, staged };
364
+ return undefined;
343
365
  }
package/dist/cli.js CHANGED
@@ -3769,6 +3769,7 @@ async function cmdVerifyRun(args) {
3769
3769
  fileToolAttempts: result.fileToolAttempts,
3770
3770
  pathDenials: result.pathDenials,
3771
3771
  presentedFiles: result.presentedFiles,
3772
+ presentFilesCalls: result.presentFilesCalls,
3772
3773
  evidenceErrors: result.evidenceErrors,
3773
3774
  effectiveFidelity: result.effectiveFidelity,
3774
3775
  // verify-run re-checks a kept run dir on the SAME machine that ran it — grouped with the live
package/dist/prompt.js CHANGED
@@ -2,6 +2,7 @@ import { warn } from "./io.js";
2
2
  import { readFileSync, existsSync } from "node:fs";
3
3
  import { join } from "node:path";
4
4
  import { fileURLToPath } from "node:url";
5
+ import { resolveMounts } from "./baseline.js";
5
6
  const BASELINES_DIR = join(fileURLToPath(new URL("..", import.meta.url)), "baselines");
6
7
  /** The {{placeholder}} names renderPrompts() substitutes. KEEP IN LOCKSTEP with the `tokens` map below. */
7
8
  export const MODELED_PLACEHOLDER_NAMES = new Set([
@@ -31,8 +32,10 @@ firstFolderMountPath, opts) {
31
32
  const spawn = baseline.spawn;
32
33
  if (!spawn)
33
34
  return {};
34
- const sessionRoot = `/sessions/${sessionId}`;
35
- const mntRoot = `${sessionRoot}/mnt`;
35
+ // Through resolveMounts, NOT a local `/sessions/<id>` literal: the prompt tells the agent where its
36
+ // files are, so it must name the same tree the runtimes stage and bind. A private derivation here was a
37
+ // fourth copy of this rule, and a fourth place for it to drift.
38
+ const { sessionRoot, mntRoot } = resolveMounts(baseline, sessionId);
36
39
  const workspaceFolder = firstFolderMountPath ? `${mntRoot}/${firstFolderMountPath}` : `${mntRoot}/outputs`;
37
40
  // {{currentDateTime}}/{{currentTimezone}} are deliberately render-time-impure (wall clock, host
38
41
  // TZ) — that's what the real Desktop builder substitutes, and the rendered append never enters a
@@ -1741,6 +1741,7 @@ function minimalRec() {
1741
1741
  fileToolAttempts: [],
1742
1742
  pathDenials: [],
1743
1743
  presentedFiles: [],
1744
+ presentFilesCalls: 0,
1744
1745
  webSearches: [],
1745
1746
  infraErrors: [],
1746
1747
  evidenceErrors: { taskTracking: 0, webSearchParse: 0, presentFilesMalformed: 0 },
@@ -3920,6 +3921,7 @@ function replayErrorResult(file) {
3920
3921
  fileToolAttempts: undefined, // no rec to read from on this early-bail lane
3921
3922
  pathDenials: undefined, // no rec to read from on this early-bail lane
3922
3923
  presentedFiles: undefined, // no rec to read from on this early-bail lane
3924
+ presentFilesCalls: undefined, // ditto
3923
3925
  preRunPaths: undefined,
3924
3926
  preRunLinkAware: undefined,
3925
3927
  preRunHashes: undefined,
@@ -5620,8 +5622,10 @@ export const ALWAYS_CONTENT_KEYS = [
5620
5622
  "task_status",
5621
5623
  "result",
5622
5624
  // content-class, NOT controlOut-gated: both the present_files tool_use and its own tool_result live
5623
- // in the ordinary events stream, so the re-drive reproduces `RunResult.presentedFiles` exactly like
5624
- // the other re-derived signals above (skill_triggered, redundantToolCalls, …).
5625
+ // in the ordinary events stream, so the re-drive reproduces the present_files signals like the other
5626
+ // re-derived ones above (skill_triggered, redundantToolCalls, …). `presentFilesCalls` (what
5627
+ // present_files_called reads) reproduces on every tier; `presentedFiles`' promoted/leaked booleans
5628
+ // reproduce at container, where cwd is the session root — no_scratchpad_leak's only tier.
5625
5629
  "no_scratchpad_leak",
5626
5630
  "present_files_called",
5627
5631
  // content-class, NOT controlOut-gated: fileToolAttempts re-derives from frozen tool_use blocks (the
@@ -6136,6 +6140,7 @@ export async function replayCassette(cassette, hooks = [], opts = {}) {
6136
6140
  // empty [] (nothing presented) vacuous-passes no_scratchpad_leak instead of reading as
6137
6141
  // evidence-unavailable.
6138
6142
  presentedFiles: rec.presentedFiles,
6143
+ presentFilesCalls: rec.presentFilesCalls,
6139
6144
  evidenceErrors: rec.evidenceErrors,
6140
6145
  effectiveFidelity: cassette.effectiveFidelity,
6141
6146
  // Replay has no live filesystem — computer_links_resolve normalizes both link shapes against the
@@ -6493,6 +6498,7 @@ export async function replayCassette(cassette, hooks = [], opts = {}) {
6493
6498
  // NOT reproduce — this one genuinely re-derives. Uncollapsed (an empty [] is the real "nothing
6494
6499
  // presented" signal no_scratchpad_leak's vacuous pass needs, matching live).
6495
6500
  presentedFiles: rec.presentedFiles,
6501
+ presentFilesCalls: rec.presentFilesCalls,
6496
6502
  preRunPaths: undefined,
6497
6503
  // Report the baseline semantics actually used during evaluation above (not undefined) so the returned
6498
6504
  // result doesn't misrepresent them. Same source of truth as the evaluate() ctx.
@@ -113,6 +113,7 @@ export function buildChatResult(record, opts) {
113
113
  fileToolAttempts: record.fileToolAttempts,
114
114
  pathDenials: record.pathDenials,
115
115
  presentedFiles: record.presentedFiles,
116
+ presentFilesCalls: record.presentFilesCalls,
116
117
  egress: opts.egress,
117
118
  resources,
118
119
  stderrLogPath: join(opts.outDir, "agent.stderr.log"),
package/dist/run/chat.js CHANGED
@@ -417,6 +417,7 @@ export async function cmdChat(args) {
417
417
  },
418
418
  };
419
419
  const run = new Run(agent, decider, [renderer, tripwireHook], sessionId);
420
+ run.setSessionRoot(hl.sessionRoot); // HOST tree — without it cwd (mnt/outputs) stands in for the root
420
421
  stopHeartbeat = startHeartbeat(renderer, renderPlan, start);
421
422
  if (viaApiOn) {
422
423
  run.enableWebFetchGate();
@@ -450,6 +451,7 @@ export async function cmdChat(args) {
450
451
  const decider = Chain(new ScriptedDecider([]), new PermissionDefaultDecider("cowork"), new PromptDecider(ask));
451
452
  const renderer = makeRenderer(renderPlan);
452
453
  const run = new Run(agent, decider, [renderer], sessionId);
454
+ run.setSessionRoot(ct.sessionRoot); // VM path — same space the agent reports (see execute.ts)
453
455
  stopHeartbeat = startHeartbeat(renderer, renderPlan, start);
454
456
  record = await run.drive(withSeedPrompt(seedPrompt, ttyTurns(rl)), chatDriveOpts(prompts, { tier: "container", sdkMcp: ct.sdkMcp }));
455
457
  }
@@ -692,6 +692,10 @@ export async function executeScenario(scenario, opts = {}) {
692
692
  const prompts = renderPrompts(baseline, session, sessionId, plan.mounts.find((m) => m.kind === "folder")?.mountPath, hostLoopOpts);
693
693
  promptFidelityWarnings = prompts.fidelityWarnings; // hoist out so RunResult construction (after try) can access it
694
694
  let sdkMcp;
695
+ // The session root as reported BY THE SPAWN that just happened — the dir whose `mnt/` is the
696
+ // user-visible workspace, in the same path space the agent reports its own paths in. Only the two
697
+ // tiers that serve present_files supply one; the others keep the cwd fallback (see setSessionRoot).
698
+ let spawnedSessionRoot;
695
699
  if (effectiveFidelity === "hostloop") {
696
700
  const hl = spawnHostLoop(scenario, baseline, plan, outDir, sessionId, {
697
701
  systemPromptAppend: prompts.systemPromptAppend,
@@ -712,6 +716,7 @@ export async function executeScenario(scenario, opts = {}) {
712
716
  hostloopPathGateFired = hl.pathGateFired;
713
717
  hostloopInfraErrors = hl.infraErrors;
714
718
  hostloopMarkTearingDown = hl.markTearingDown;
719
+ spawnedSessionRoot = hl.sessionRoot; // HOST tree — the native agent runs there
715
720
  logHostWriteNotice(plan.mounts.filter((mt) => mt.kind === "folder").map((mt) => ({ from: mt.hostPath, mode: mt.mode })), warn);
716
721
  if (scenario.assert.some((a) => a.transcript_no_host_path === true) && !opts.compact)
717
722
  warn(`::warning:: [hostloop] scenario asserts transcript_no_host_path — hostloop's native file tools legitimately ` +
@@ -729,6 +734,7 @@ export async function executeScenario(scenario, opts = {}) {
729
734
  child = ct.child;
730
735
  containerName = ct.containerName; // so the Ctrl-C / finally reap removes the agent container by name
731
736
  sdkMcp = ct.sdkMcp; // cowork/present_files + the skills/plugins discovery servers (combineSdkMcp)
737
+ spawnedSessionRoot = ct.sessionRoot; // VM path (`/sessions/<id>`) — what the agent inside reports
732
738
  }
733
739
  else if (effectiveFidelity === "microvm") {
734
740
  child = spawnMicroVm(scenario, baseline, plan, outDir, sessionId, {
@@ -771,12 +777,18 @@ export async function executeScenario(scenario, opts = {}) {
771
777
  const decider = effectiveFidelity === "hostloop" ? Chain(makeHostLoopCanUseToolGate(), policyDecider) : policyDecider;
772
778
  const run = new Run(sessionT, decider, opts.hooks ?? [], sessionId, dialogTimeoutMs ?? undefined, scenario.timeout_ms);
773
779
  run.seedApprovedDomains(session.web_fetch.approved_domains); // test convenience: pre-approved web_fetch hosts
774
- // The session root — the dir whose `mnt/` IS the user-visible workspace. Given explicitly because
775
- // the agent's own cwd only coincides with it on some tiers (at hostloop the agent runs inside
776
- // mnt/outputs), and present_files' promoted/leaked classification is measured from it.
777
- // `protocol` has no session/mnt layout at all, so it stays unset and keeps the cwd fallback.
778
- if (effectiveFidelity !== "protocol")
779
- run.setSessionRoot(join(outDir, "work", "session"));
780
+ // The session root — the dir whose `mnt/` IS the user-visible workspace, and what present_files'
781
+ // promoted/leaked classification is measured from. Taken from the SPAWN, never re-derived here: the
782
+ // root and the agent's reported paths must be in the SAME path space, and they are not the same space
783
+ // on every tier. A host path was passed unconditionally once; at container the agent reports VM paths
784
+ // (`/sessions/<id>/…`), so nothing was ever inside the root, every presented file classified
785
+ // `leaked: false`, and `no_scratchpad_leak` — which evaluates at container and nowhere else — passed
786
+ // vacuously over a real copy-failure leak.
787
+ //
788
+ // Unset on the tiers that serve no present_files (`protocol`, `microvm`), where the cwd fallback is
789
+ // already the session root and there is no delivery to classify.
790
+ if (spawnedSessionRoot !== undefined)
791
+ run.setSessionRoot(spawnedSessionRoot);
780
792
  // fill the provenance bundle (backed by Run's tracker + recorded approval) BEFORE drive().
781
793
  // Host-loop only, and only when the web_fetch-via-API gate is on; otherwise the handler stays
782
794
  // allowlist-only (ref.current undefined). Run seeds the set from turns + tool_results.
@@ -1157,6 +1169,7 @@ export async function executeScenario(scenario, opts = {}) {
1157
1169
  // Always defined live — an empty array is the real "nothing presented" signal no_scratchpad_leak's
1158
1170
  // vacuous pass needs, distinct from replay's evidence-unavailable undefined on an older cassette.
1159
1171
  presentedFiles: record.presentedFiles,
1172
+ presentFilesCalls: record.presentFilesCalls,
1160
1173
  evidenceErrors: record.evidenceErrors,
1161
1174
  effectiveFidelity,
1162
1175
  // Live lane (this run's own machine) — host-shaped computer:// links (hostloop) are checked
@@ -1404,6 +1417,7 @@ export async function executeScenario(scenario, opts = {}) {
1404
1417
  fileToolAttempts: record.fileToolAttempts, // uncollapsed — content-class, same as toolResults/decisions above
1405
1418
  pathDenials: record.pathDenials, // uncollapsed — content-class, same as fileToolAttempts above
1406
1419
  presentedFiles: record.presentedFiles, // uncollapsed — an empty [] is the real "nothing presented" signal no_scratchpad_leak's vacuous pass needs
1420
+ presentFilesCalls: record.presentFilesCalls,
1407
1421
  // The pre-spawn baseline no_unexpected_files diffs against (same single read the evaluate ctx got).
1408
1422
  // undefined = the run didn't capture (key not asserted, microvm, pre-seam) — the assertion then
1409
1423
  // fails evidence-unavailable, loud.
@@ -1788,6 +1802,7 @@ export function buildPartialResult(args) {
1788
1802
  fileToolAttempts: record.fileToolAttempts, // uncollapsed — content-class, same as toolResults/decisions above
1789
1803
  pathDenials: record.pathDenials, // uncollapsed — content-class, same as fileToolAttempts above
1790
1804
  presentedFiles: record.presentedFiles, // uncollapsed — an empty [] is the real "nothing presented" signal no_scratchpad_leak's vacuous pass needs
1805
+ presentFilesCalls: record.presentFilesCalls,
1791
1806
  preRunPaths: readPreRunManifest(args.outDir),
1792
1807
  preRunLinkAware: readPreRunManifestLinkAware(args.outDir),
1793
1808
  preRunHashes: readPreRunManifestHashes(args.outDir),
package/dist/run/run.js CHANGED
@@ -267,6 +267,7 @@ export class Run {
267
267
  fileToolAttempts: [],
268
268
  pathDenials: [],
269
269
  presentedFiles: [],
270
+ presentFilesCalls: 0,
270
271
  webSearches: [],
271
272
  infraErrors: [],
272
273
  evidenceErrors: { taskTracking: 0, webSearchParse: 0, presentFilesMalformed: 0 },
@@ -488,6 +489,12 @@ export class Run {
488
489
  const pf = presentFilesInput(ev.input);
489
490
  this.pendingPresentFiles.set(ev.toolUseId, pf.files);
490
491
  this.rec.evidenceErrors.presentFilesMalformed += pf.malformed;
492
+ // Presence, counted here and NOT from `presentedFiles` below — see the field's comment.
493
+ // Gated on a well-formed file (not merely on the call) so the count keeps the DELIVERY
494
+ // meaning `present_files_called` documents: a call whose `files` were unusable delivered
495
+ // nothing, leaves this at 0, and is reported through the malformed counter instead.
496
+ if (pf.files.length > 0)
497
+ this.rec.presentFilesCalls++;
491
498
  }
492
499
  this.toolLog.push({ name: ev.name, input: ev.input, synthetic: ev.synthetic, parentToolUseId: ev.parentToolUseId }); // still logged for provenance/trace
493
500
  break;
@@ -932,6 +939,18 @@ export class Run {
932
939
  // genuine leak as fine. This bug over-reports; that one would under-report, and a false green is the
933
940
  // worse failure.
934
941
  const root = this.sessionRoot !== undefined ? posixPath.normalize(this.sessionRoot) : cwd;
942
+ // SPACE CHECK, before any classification: the agent's cwd must sit AT or INSIDE the session root. That
943
+ // holds on every tier that serves present_files — at container cwd IS the root, at hostloop it is
944
+ // `<root>/mnt/<outputs|folder>` — so a cwd outside the root means the two are in different path spaces
945
+ // (a host root against VM-reported paths, say) and every containment test below is meaningless. Count
946
+ // the batch malformed instead of grading it: nothing would be under the root, so the classification
947
+ // would silently read `leaked: false` for a genuine leak, which is exactly the vacuous pass
948
+ // `no_scratchpad_leak` exists to prevent. Deliberately NOT "no presented path is under the root" —
949
+ // a hostloop delivery out of a connected folder legitimately sits outside the session tree.
950
+ if (root !== undefined && cwd !== undefined && cwd !== root && !cwd.startsWith(`${root}/`)) {
951
+ this.rec.evidenceErrors.presentFilesMalformed += froms.length;
952
+ return;
953
+ }
935
954
  const isScratchpad = (p) => root !== undefined && p.startsWith(`${root}/`) && !p.startsWith(`${root}/mnt/`);
936
955
  for (let i = 0; i < froms.length; i++) {
937
956
  const rawTo = tos[i];
@@ -1,3 +1,5 @@
1
+ import { recordedLayoutDivergence } from "../baseline.js";
2
+ import { warn } from "../io.js";
1
3
  import { DEFAULT_MAX_THINKING_TOKENS } from "../types.js";
2
4
  import { SECRET_ENV_KEYS } from "./host-env.js";
3
5
  /**
@@ -9,6 +11,23 @@ import { SECRET_ENV_KEYS } from "./host-env.js";
9
11
  */
10
12
  export function baseAgentArgs(baseline, plan, opts) {
11
13
  const spawn = baseline.spawn;
14
+ // A baseline with NO `spawn` block cannot be spawned faithfully at a sandbox tier: the tool set,
15
+ // pre-approvals, effort default and config-dir location all come from it, and the `?? []` fallbacks
16
+ // below would silently emit an agent with no Read/Write/Bash/Skill/Task at all — a run that cannot
17
+ // execute the skill under test while still reporting a verdict. Refuse instead. (`protocol` builds its
18
+ // own argv and inherits the host CLI's toolset, so it is unaffected and stays usable.)
19
+ if (!spawn)
20
+ throw new Error(`baseline "${baseline.appVersion}" has no \`spawn\` block, so a sandbox tier cannot reproduce Cowork's ` +
21
+ `toolset, pre-approvals or config-dir layout — the agent would launch with none of its file/bash tools. ` +
22
+ `Use a baseline recorded by \`sync\`, or run at \`fidelity: protocol\` (which builds its own argv).`);
23
+ // FIDELITY, not correctness: guest paths are always staged at `<sessionRoot>/mnt` (see
24
+ // GUEST_MNT_SEGMENT), so a baseline recording a different mnt root is reproduced approximately. Say so
25
+ // once, here, rather than letting the recorded value build a path no stager creates.
26
+ const divergence = recordedLayoutDivergence(baseline);
27
+ if (divergence)
28
+ warn(`::warning:: baseline "${baseline.appVersion}" records a guest mnt root this harness cannot stage ` +
29
+ `(recorded ${divergence.recorded}, staged ${divergence.staged}) — guest paths use the staged layout, ` +
30
+ `so mount-path fidelity at this tier is approximate.\n`);
12
31
  // Real Cowork ALWAYS emits `--effort`, for every model class (picker, no-picker, regex-default,
13
32
  // unknown) — falling back to the baseline's synced medium default when the session left it unset
14
33
  // (per-model validation of an EXPLICIT value already ran in buildLaunchPlan's validateEffort; the
@@ -207,7 +226,7 @@ export function dockerRunArgv(i) {
207
226
  i.network,
208
227
  ...(i.lockdown ? HARDENING : []),
209
228
  "-w",
210
- i.sessionRoot,
229
+ i.agentCwd ?? i.sessionRoot,
211
230
  // Render SECRET values by NAME only (`-e KEY`) so the token never lands in `docker run`'s
212
231
  // argv (visible via ps / /proc/<pid>/cmdline). Docker inherits the value from its own env — the
213
232
  // harness process env, where runtimeAuthEnv read it. Non-secret env keeps the explicit KEY=value.
@@ -23,8 +23,11 @@ import { resolveAgentImage, resolveContainerRuntime } from "./agent-image.js";
23
23
  */
24
24
  export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opts = {}) {
25
25
  const m = resolveMounts(baseline, sessionId, "proj1");
26
- const sessionRoot = m.cwd; // /sessions/<id>
27
- const mntRoot = m.mntRoot; // /sessions/<id>/mnt
26
+ // BIND TARGET (and the anchor for every guest path), vs the agent's working dir. Equal on every synced
27
+ // baseline; kept distinct because production's cwd is a folder mount or `outputs`, not the session root.
28
+ const sessionRoot = m.sessionRoot; // /sessions/<id>
29
+ const agentCwd = m.cwd;
30
+ const mntRoot = m.mntRoot; // <sessionRoot>/mnt — the tree stageWorkspace creates
28
31
  const configGuest = `${sessionRoot}/${baseline.spawn?.configDirInGuest ?? "mnt/.claude"}`;
29
32
  const AGENT_IN = "/usr/local/bin/claude";
30
33
  // Name by the per-invocation runToken (NOT sessionId) so a --resume after a failed run doesn't collide
@@ -88,6 +91,7 @@ export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opt
88
91
  lockdown: (process.env.COWORK_LOCKDOWN ?? "on") !== "off",
89
92
  name: containerName,
90
93
  sessionRoot,
94
+ agentCwd,
91
95
  sessionHost,
92
96
  agentHost,
93
97
  agentIn: AGENT_IN,
@@ -130,5 +134,9 @@ export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opt
130
134
  handle: makePluginsHandler({ mountedPlugins }),
131
135
  };
132
136
  const sdkMcp = combineSdkMcp(...(coworkBundle ? [coworkBundle] : []), skillsBundle, pluginsBundle);
133
- return { child, containerName, sdkMcp };
137
+ // `sessionRoot` is the VM path the agent sees (`-w` above, and the cowork handler's own
138
+ // `sessionRootVm`). Returned so the caller classifies present_files against the root THIS spawn used,
139
+ // instead of re-deriving one — the two lived in different path spaces (host vs VM) once, which made
140
+ // every container leak read as `leaked: false`.
141
+ return { child, containerName, sdkMcp, sessionRoot };
134
142
  }
@@ -363,7 +363,11 @@ export function spawnHostLoop(_scenario, baseline, plan, outDir, sessionId, opts
363
363
  handle: makePluginsHandler({ mountedPlugins }),
364
364
  };
365
365
  const sdkMcp = combineSdkMcp(workspaceBundle, ...(coworkBundle ? [coworkBundle] : []), skillsBundle, pluginsBundle);
366
- return { child, sdkMcp, hooks, pathGateFired, containerName, hostEgress, infraErrors, markTearingDown };
366
+ // `sessionRoot` here is the HOST tree (`sessionHost`), not the VM path: the agent runs natively on the
367
+ // host at this tier (see the `spawn(agentNativeHost, …, { cwd: hostOutputsDir })` above), so the paths
368
+ // it reports — and the ones its present_files handler validates — are host paths. Returned for the same
369
+ // reason as container's: the caller must not re-derive it.
370
+ return { child, sdkMcp, hooks, pathGateFired, containerName, hostEgress, infraErrors, markTearingDown, sessionRoot: sessionHost };
367
371
  }
368
372
  /** The two infra-error emitters a host-loop run needs, sharing one sink and one events.jsonl writer.
369
373
  * Every row carries its `source` so the verdict can tell a dead supervisor apart from a failed command,
@@ -5,6 +5,12 @@ import { join, dirname, basename } from "node:path";
5
5
  import { createHash } from "node:crypto";
6
6
  /** Host dir mounted writable into the VM at /sessions (the staging area; per-session subdirs). */
7
7
  export const VM_WORK_HOST = join(homedir(), ".cowork-harness", "vm-work");
8
+ /** Guest mount point of `VM_WORK_HOST` — the literal in the lima template below. The microVM's guest
9
+ * session root is `${VM_GUEST_SESSIONS_ROOT}/<sessionId>` STRUCTURALLY: lima mounts the work root here
10
+ * and the per-session dirs live inside it, so this tier cannot honour a baseline that records the agent
11
+ * running anywhere else. Exported so the runtime anchors on the same string the template mounts, and so
12
+ * a test can check that pairing without booting a VM. */
13
+ export const VM_GUEST_SESSIONS_ROOT = "/sessions";
8
14
  /**
9
15
  * L2 microVM provisioning via Lima with `vmType: vz` — Apple Virtualization.framework,
10
16
  * the SAME hypervisor Claude Cowork uses. This gives a real Linux kernel (VM-grade
@@ -141,7 +147,7 @@ mounts:
141
147
  # the agent persists its session (enabling --resume). Lima creates the mountpoint, writable by the
142
148
  # mounting user — no guest /sessions permission problem.
143
149
  - location: "${VM_WORK_HOST}"
144
- mountPoint: "/sessions"
150
+ mountPoint: "${VM_GUEST_SESSIONS_ROOT}"
145
151
  writable: true
146
152
  provision:
147
153
  - mode: system
@@ -1,7 +1,7 @@
1
1
  import { spawn } from "node:child_process";
2
2
  import { cpSync, existsSync, mkdirSync, rmSync } from "node:fs";
3
3
  import { dirname, join } from "node:path";
4
- import { limaPath, vmInit, applyGuestFirewall, vmGatewayIp, VM_WORK_HOST } from "./lima.js";
4
+ import { limaPath, vmInit, applyGuestFirewall, vmGatewayIp, VM_WORK_HOST, VM_GUEST_SESSIONS_ROOT } from "./lima.js";
5
5
  import { resolveMounts } from "../baseline.js";
6
6
  import { spawnEnv, baseAgentArgs } from "./argv.js";
7
7
  import { stageWorkspace } from "./stage.js";
@@ -62,11 +62,29 @@ function parseEnvPortMicroVm(name, defaultValue) {
62
62
  *
63
63
  * Egress: a host-side allowlist proxy + a guest default-deny iptables firewall.
64
64
  */
65
+ /** The microVM's guest session root, which is STRUCTURAL rather than baseline-derived: lima mounts the
66
+ * work root at `VM_GUEST_SESSIONS_ROOT` with a per-`sessionId` dir inside it, so the agent runs at
67
+ * `<VM_GUEST_SESSIONS_ROOT>/<sessionId>` whatever the baseline records. Reading it from
68
+ * `mountLayout.cwd` instead put `CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages whenever
69
+ * the two disagreed.
70
+ *
71
+ * A recorded cwd this tier cannot honour THROWS rather than warns: the guest `cd` would otherwise
72
+ * succeed at the wrong directory, and a silently-wrong cwd is exactly what the loud `cd` prevents.
73
+ *
74
+ * Pure and exported so this pairing is checkable without booting a VM (`vmInit` needs `limactl`). */
75
+ export function microvmGuestSessionRoot(baseline, sessionId) {
76
+ const sessionVm = `${VM_GUEST_SESSIONS_ROOT}/${sessionId}`;
77
+ const recordedCwd = resolveMounts(baseline, sessionId, "proj1").cwd;
78
+ if (recordedCwd !== sessionVm)
79
+ throw new Error(`baseline "${baseline.appVersion}" records the agent's cwd as ${recordedCwd}, but the microVM mounts its ` +
80
+ `session tree at ${sessionVm} and cannot place the agent anywhere else — run this baseline at ` +
81
+ `fidelity: container/hostloop, or use one whose mountLayout.cwd matches the microVM mount point.`);
82
+ return sessionVm;
83
+ }
65
84
  export function spawnMicroVm(_scenario, baseline, plan, outDir, sessionId, opts = {}) {
66
85
  const { instance } = vmInit(baseline);
67
- const m = resolveMounts(baseline, sessionId, "proj1");
68
- const sessionVm = m.cwd; // /sessions/<id>
69
- const mntVm = m.mntRoot; // /sessions/<id>/mnt
86
+ const sessionVm = microvmGuestSessionRoot(baseline, sessionId);
87
+ const mntVm = `${sessionVm}/mnt`; // the staged tree — see GUEST_MNT_SEGMENT
70
88
  const configVm = `${sessionVm}/${baseline.spawn?.configDirInGuest ?? "mnt/.claude"}`;
71
89
  // Stage into the lima-mounted work dir (host VM_WORK_HOST -> guest /cowork-work) via the shared
72
90
  // helper. It honors plan.resume for ALL of .claude + mounts + mcp.json (the .claude guard
package/dist/types.js CHANGED
@@ -456,7 +456,7 @@ export const Assertion = z.strictObject({
456
456
  present_files_called: z
457
457
  .literal(true)
458
458
  .optional()
459
- .describe("at least one file was actually delivered via the present_files tool (presentedFiles is non-empty). The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid"),
459
+ .describe("at least one file was actually delivered via the present_files tool (at least one call carried a well-formed file_path). Presence is read from the INVOCATION count, not from the classified presentedFiles list, so it is unaffected by a redaction policy that rewrites host paths; a run that called the tool but whose every call carried an unusable path reports cannot-verify, never 'the tool was never called'. The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid"),
460
460
  egress_denied: z.string().optional().describe("egress to this host was denied"),
461
461
  egress_allowed: z.string().optional().describe("egress to this host was allowed"),
462
462
  // Only `true` is accepted: `false` is rejected as a footgun. The assertion is presence-semantic — authoring
package/docs/cassette.md CHANGED
@@ -512,8 +512,8 @@ the rules and CI-placement rationale (why each category behaves this way), see
512
512
  | `path_denied` | **`fidelity: hostloop` only** — a path denial matched all given matchers (`tool`/`path_matches`/`source`/`agent_scope`) — replay: needs `controlOut`; any other tier FAILS "cannot verify" |
513
513
  | `no_path_denied` | **`fidelity: hostloop` only** — no path denial was recorded at all — replay: needs `controlOut`. **Only `true` is valid**; any other tier FAILS "cannot verify" |
514
514
  | `result` | run ended with `success` or `error` |
515
- | `no_scratchpad_leak` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuous pass if nothing was presented; content-class: both the tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay; evidence-unavailable only when `presentedFiles` is absent (an older run predating the feature); **container tier only** — `present_files` *is* served on hostloop, but that branch passes a validated path through without promoting, so there is no scratch→outputs copy to leak; `microvm`/`protocol` do not serve the tool at all |
516
- | `present_files_called` | at least one file was delivered via `present_files` (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak`; content-class (re-derives identically on replay); **container + hostloop tiers** |
515
+ | `no_scratchpad_leak` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuous pass if nothing was presented; content-class: both the tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only; evidence-unavailable only when `presentedFiles` is absent (an older run predating the feature); **container tier only** — `present_files` *is* served on hostloop, but that branch passes a validated path through without promoting, so there is no scratch→outputs copy to leak; `microvm`/`protocol` do not serve the tool at all |
516
+ | `present_files_called` | at least one file was delivered via `present_files` (`RunResult.presentFilesCalls` > 0 — the invocation count, which a host-path redaction policy cannot alter) — the presence companion to `no_scratchpad_leak`; content-class (re-derives identically on replay); **container + hostloop tiers** |
517
517
  | `allow_permissive_auto_allow` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
518
518
  | `allow_missing_capability` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
519
519
  | `allow_l0_plugin_divergence` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
@@ -22,6 +22,9 @@ duplicate it here; this table is for finding things fast, not for the full ratio
22
22
  | A cassette carries a `sessionFingerprint` (v9+, content-SHAPE hash of the AUTHORED, pre-resolution session — connected folders, plugins, skills, mcp and egress config, web_fetch approved domains, plus projects and agent_env when set; authored so the hash stays relocatable across checkouts) checked ONLY by `verify-cassettes`, deliberately excluded from `computeStaleness`/`checkStaleness` so it never affects the default replay verdict | `src/run/cassette.ts` (`sessionFingerprintDrift`) | `test/session-fingerprint.test.ts` (`describe("buildSessionFingerprint (function-level)")` — covers `buildSessionFingerprint` + `sessionFingerprintDrift`) |
23
23
  | A corrupt `timeline.jsonl` header must read as "no timeline" (`undefined`) at every consumer, never a present-empty timeline — a header-corrupt read that folds to `[]`/`undefined` events bakes a novel "ran, no activity" shape into a recorded cassette instead of the honest evidence-unavailable state | `src/run/execute.ts`, `src/run/chat-result.ts`, `src/run/cassette.ts` (each checks `!timelineRaw.headerCorrupt` alongside `malformedLines === 0` before trusting a `readTimeline()` result) | `test/central-cluster-guards.test.ts` |
24
24
  | `evidenceErrors.egressParse` must be included in the presence gate that decides whether `RunResult.evidenceErrors` serializes at all — omitting it silently absorbs the one counter that exists specifically to surface dropped egress-proxy lines | `src/run/run.ts` (`evidenceErrorsForResult`) | `test/central-cluster-guards.test.ts` (`describe("#39 egressParse reaches result.json (presence-gate fix, not just parseEgressLine)")`) |
25
+ | Every GUEST path must be composed from the tree the harness stages — `<sessionRoot>/mnt` — and never from a baseline's recorded `mountLayout.mntRoot`. A recorded mnt root that differs is a fidelity divergence reported at spawn, not a path to build: honouring one put `--plugin-dir` a directory above the staged plugin tree, so the plugin under test never loaded. Guest paths anchor on `sessionRoot` (the bind target), never `cwd` (where the agent merely starts); at microvm the root is lima's mount point and a baseline recording another cwd is REFUSED, since the guest `cd` would otherwise succeed at the wrong dir. A baseline with no `spawn` block is refused at the sandbox tiers outright — its toolset and pre-approvals are that block | `src/baseline.ts` (`GUEST_MNT_SEGMENT`, `resolveMounts`, `recordedLayoutDivergence`), `src/runtime/argv.ts` (the spawn-block refusal + divergence warning), `src/runtime/microvm.ts` (`microvmGuestSessionRoot`), `src/prompt.ts` (routed through `resolveMounts` instead of a private derivation) | `test/guest-path-layout.test.ts` (a table over every shipped baseline, oracle taken from the STAGING side, plus a synthetic spawnable-but-divergent baseline — the shipped divergent one has no `spawn` block, so only the synthetic case exercises the path composition) |
26
+ | The session root handed to `Run.setSessionRoot` must come from the SPAWN that started the agent, in the same path space the agent reports its own paths in — host at hostloop (native agent), the VM `/sessions/<id>` at container — never re-derived by the caller. A host root measured against VM-reported paths puts nothing inside the root, so every presented file classifies `leaked: false` and `no_scratchpad_leak` (which evaluates at container and nowhere else) passes vacuously over a real copy-failure leak | `src/runtime/container.ts` + `src/runtime/hostloop.ts` (each returns the `sessionRoot` it used), `src/run/execute.ts` + `src/run/chat.ts` (pass it through), `src/run/run.ts` (`notePresentedFiles`' cwd-at-or-inside-root space check counts the batch malformed rather than grading it in the wrong space) | `test/session-root-path-space.test.ts` (per-tier geometry, the fail-closed space mismatch, live-vs-replay agreement, and the spawn-reports-its-own-root seam) |
27
+ | `present_files_called` must read presence from `RunResult.presentFilesCalls` (the invocation count, derived from the tool_use input's SHAPE), never from the classified `presentedFiles` list — that list drops any path it cannot resolve, which a host-path redaction policy guarantees at `hostloop`, so reading classification claimed "the tool was never called" about a run that called it AND made every such scenario unrecordable (record refuses a cassette whose verdict redaction changed) | `src/assert.ts` (the `present_files_called` branch), `src/run/run.ts` (`presentFilesCalls`, counted at the `present_files` tool_use) | `test/present-files-redaction-invariant.test.ts` (drives the real `assertRedactionVerdictPreserved` over a hostloop cassette redacted with the shipped policy; the CONTROL case pins that the un-redacted side genuinely passes, so a shared tier-gate failure can't green it) |
25
28
  | An anonymous sub-agent dispatch's synthesized `toolUseId` must derive from the assistant message's own (stable, unique) id, never a process-lifetime counter — a counter-derived id isn't reproducible across record→replay and can collide across messages in a single run | `src/agent/session.ts` (`parseMessage`) | `test/session-parse-guards.test.ts` (the regression test asserting synthesized dispatch id uniqueness across messages, in the `parseMessage` describe block) |
26
29
  | The pre-commit cassette gate must fail CLOSED: any `verify-cassettes` outcome that is not a proven clean `0` blocks the commit, and an exit 3 blocks unless its cause is *staleness* specifically. `ci.yml` triggers on `push: [main]` and `pull_request`, but the documented local workflow lands with `merge --ff-only` into `main` and pushes afterwards — so for the maintainer, the only person who records host-inheriting cassettes, CI is not a pre-publication gate and the hook is not one layer of two. It is the only gate, and anything it waves through reaches public history | `.githooks/pre-commit` (allowlist on `hook_status`; missing `dist/cli.js` blocks rather than warning; exit 3 split by cause via `--output-format json`) | `test/cassette-gate.test.ts` |
27
30
  | Staging a new/updated `baselines/desktop-*.json` must re-stamp or re-record the committed `examples/replays/*.cassette.json` fixtures against it in the same commit — `latest` resolves to the newest baseline file (`src/baseline.ts`, `latestBaselineFile`), so any staged baseline can move `latest` and instantly stale every example cassette's `fingerprint.baseline`; caught 3 times (3431e09, 9eaba8d, the 0.29.0 cycle), always late (on the release PR's first CI run) because these commits sit unpushed for a while | `.githooks/pre-commit` (runs `verify-cassettes` whenever a `baselines/desktop-*.json`, a `*.cassette.json`, or any staged `.json` carrying the `"generator": "cowork-harness"` marker is staged) | `test/cassette-gate.test.ts` — drives the hook in a scratch repo through a stubbed CLI, and separately pins the stub's exit-code contract against the real binary; CI's `Cassette privacy + staleness scan` and `Cassette scan covers every TRACKED cassette` steps (`ci.yml`) are the non-local backstop |
@@ -23,9 +23,11 @@ forward from the previous baseline untouched)
23
23
  ```
24
24
 
25
25
  > **`mountLayout.mounts[].mode` is documentary, and older baselines carry a stale `projects` row.**
26
- > Nothing reads that array at run time: `resolveMounts()` spreads it, but container/microvm/hostloop each
27
- > take only `cwd` and `mntRoot` from the result, and the one `mode === "r"` filter reads the launch plan's
28
- > mounts, not the baseline's. The `projects` row reads `mode: "rw"` in every baseline before
26
+ > Nothing reads that array at run time: `resolveMounts()` does not return it at all, and
27
+ > container/microvm/hostloop take only `cwd`, `sessionRoot` and `mntRoot` from the result — where
28
+ > `mntRoot` is DERIVED as `<sessionRoot>/mnt` (the only tree the stagers create) rather than read from
29
+ > `mountLayout.mntRoot`; a recorded value that disagrees is reported as a fidelity divergence at spawn.
30
+ > The one `mode === "r"` filter reads the launch plan's mounts, not the baseline's. The `projects` row reads `mode: "rw"` in every baseline before
29
31
  > `desktop-1.25927.0`, which is **not uniformly wrong**: below `MOUNT_BARE_NAME_MIN_VERSION` (1.14271.0)
30
32
  > `.projects/<name>` really was the connected-folder namespace, and folders are resolver-driven `rw`. From
31
33
  > that boundary on, folders moved to `mnt/<basename>` and `.projects/<uuid>` became the project-attachment
package/docs/scenario.md CHANGED
@@ -470,8 +470,8 @@ whether it **survives `replay`**. Both are in the key's row below, and the repla
470
470
  | `all_tasks_completed: true` | every task in the run's task list reached status `completed` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); **only `true` is valid**; also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry) |
471
471
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions; also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry) |
472
472
  | `task_status: {match, status}` | a task whose subject or id matches the `match` regex reached `status` — also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry), mirroring `all_tasks_completed`/`task_count_min` |
473
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so this is meaningfully replay-checkable (the re-drive reproduces it); fails as **evidence unavailable** when `presentedFiles` telemetry is absent (an old run predating this key); **container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through unchanged), so there is no scratch→outputs copy for this key to check — that's not a detection gap, though: hostloop's cwd already *is* the outputs dir, so a delivered file is visible there immediately (see `user_visible_artifact`'s footgun note above). On microvm and protocol, `present_files` isn't served at all, so there's no delivery record for this key to check — cannot-verify. Use `container` for present_files-based delivery you want this key to verify, or write directly to `outputs/`; **the tool name is lane-specific** — `present_files` is the desktop-local lane's tool (the one this harness emulates) while remote Cowork delivers via the agent-native `SendUserFile`, so a skill should describe the delivery outcome rather than naming either tool ([fidelity-gaps.md](./fidelity-gaps.md), "File delivery"); this key asserts the harness-side delivery record either way; **only `true` is valid** |
474
- | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak; **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note above: the tool name differs on remote Cowork; **only `true` is valid** |
473
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so this is meaningfully replay-checkable at the tier it evaluates on — the re-drive reproduces the classification at container, where the agent's cwd IS the session root the live lane measures from (at hostloop a re-drive has only the recorded cwd, `mnt/outputs`, so the booleans are not equivalent there); fails as **evidence unavailable** when `presentedFiles` telemetry is absent (an old run predating this key); **container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through unchanged), so there is no scratch→outputs copy for this key to check — that's not a detection gap, though: hostloop's cwd already *is* the outputs dir, so a delivered file is visible there immediately (see `user_visible_artifact`'s footgun note above). On microvm and protocol, `present_files` isn't served at all, so there's no delivery record for this key to check — cannot-verify. Use `container` for present_files-based delivery you want this key to verify, or write directly to `outputs/`; **the tool name is lane-specific** — `present_files` is the desktop-local lane's tool (the one this harness emulates) while remote Cowork delivers via the agent-native `SendUserFile`, so a skill should describe the delivery outcome rather than naming either tool ([fidelity-gaps.md](./fidelity-gaps.md), "File delivery"); this key asserts the harness-side delivery record either way; **only `true` is valid** |
474
+ | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (at least one call carried a well-formed `file_path`, counted at the invocation — **not** read off the classified `presentedFiles` list, so a redaction policy that rewrites host paths cannot turn a real delivery into "never called"; a run whose every call carried an unusable path reports cannot-verify) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak; **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note above: the tool name differs on remote Cowork; **only `true` is valid** |
475
475
  | `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD — fails as **evidence unavailable** when cost telemetry is absent (an old run predating this key). **Live lane only in spirit**: on replay this asserts the *frozen recording's* cost, not fresh spend — a cost regression is caught by a live run, not a token-free replay |
476
476
  | `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache-read/creation tokens excluded — priced separately). Same replay caveat as `max_cost_usd`: asserts the recording, not fresh spend |
477
477
  | `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`, sub-agent tools excluded) ≤ N — unlike the cost/token keys, this **is** meaningfully replay-checkable (the re-drive recomputes `toolCounts` deterministically from the recorded events) |
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^2.1.0"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^2.2.0"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "2.1.0",
3
+ "version": "2.2.0",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -137,6 +137,10 @@
137
137
  },
138
138
  "description": "present_files promotions ({from, to, promoted, leaked}) \u2014 scratch\u2192outputs deliverable surfacing (container tier)."
139
139
  },
140
+ "presentFilesCalls": {
141
+ "type": "integer",
142
+ "description": "How many present_files calls carried at least one well-formed file_path \u2014 the presence evidence present_files_called reads, counted from the tool_use input's shape so it survives redaction (presentedFiles entries are dropped when a path can't be classified, which a host-path policy guarantees at hostloop). Absent on pre-field results."
143
+ },
140
144
  "mode": {
141
145
  "type": "string",
142
146
  "enum": ["run", "chat"],
@@ -480,7 +480,7 @@
480
480
  "const": true
481
481
  },
482
482
  "present_files_called": {
483
- "description": "at least one file was actually delivered via the present_files tool (presentedFiles is non-empty). The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid",
483
+ "description": "at least one file was actually delivered via the present_files tool (at least one call carried a well-formed file_path). Presence is read from the INVOCATION count, not from the classified presentedFiles list, so it is unaffected by a redaction policy that rewrites host paths; a run that called the tool but whose every call carried an unusable path reports cannot-verify, never 'the tool was never called'. The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid",
484
484
  "type": "boolean",
485
485
  "const": true
486
486
  },