cowork-harness 2.1.0 → 2.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -3
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +111 -0
- package/README.md +5 -5
- package/dist/assert.js +18 -3
- package/dist/baseline.js +37 -15
- package/dist/cli.js +1 -0
- package/dist/prompt.js +5 -2
- package/dist/run/cassette.js +8 -2
- package/dist/run/chat-result.js +1 -0
- package/dist/run/chat.js +2 -0
- package/dist/run/execute.js +21 -6
- package/dist/run/run.js +19 -0
- package/dist/runtime/argv.js +20 -1
- package/dist/runtime/container.js +11 -3
- package/dist/runtime/hostloop.js +5 -1
- package/dist/runtime/lima.js +7 -1
- package/dist/runtime/microvm.js +22 -4
- package/dist/types.js +1 -1
- package/docs/cassette.md +2 -2
- package/docs/invariants.md +3 -0
- package/docs/maintenance.md +5 -3
- package/docs/scenario.md +2 -2
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/run-result.json +4 -0
- package/schema/scenario.schema.json +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.
|
|
7
|
-
tracks-harness: cowork-harness 2.
|
|
6
|
+
version: 2.2.0
|
|
7
|
+
tracks-harness: cowork-harness 2.2.0 (baseline desktop-1.34493.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.2.0` (baseline
|
|
26
26
|
> `desktop-1.34493.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.2.0"`. **Pin `@^2.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
44
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
45
45
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "2.
|
|
20
|
+
(e.g. `version: "2.2.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.2.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -340,7 +340,7 @@ jobs:
|
|
|
340
340
|
with: { node-version: '24' }
|
|
341
341
|
- uses: actions/setup-python@v5
|
|
342
342
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
343
|
-
- run: npm i -g "cowork-harness@^2.
|
|
343
|
+
- run: npm i -g "cowork-harness@^2.2.0"
|
|
344
344
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
345
345
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
346
346
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -369,7 +369,7 @@ jobs:
|
|
|
369
369
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
370
370
|
fi
|
|
371
371
|
- if: steps.guard.outputs.live == 'true'
|
|
372
|
-
run: npm i -g "cowork-harness@^2.
|
|
372
|
+
run: npm i -g "cowork-harness@^2.2.0"
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
374
|
run: cowork-harness run scenarios/ --output-format json
|
|
375
375
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.
|
|
3
|
+
Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.2.0`
|
|
4
4
|
(baseline `desktop-1.34493.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -342,8 +342,8 @@ same set live from the schema.
|
|
|
342
342
|
| `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
|
|
343
343
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
|
|
344
344
|
| `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
|
|
345
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives
|
|
346
|
-
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.
|
|
345
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
|
|
346
|
+
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
|
|
347
347
|
| `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
|
|
348
348
|
| `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
|
|
349
349
|
| `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.
|
|
5
|
+
Tracks `cowork-harness 2.2.0` (baseline `desktop-1.34493.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,117 @@ All notable changes to this project are documented here. The format is based on
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
|
|
5
5
|
[Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
|
|
6
6
|
|
|
7
|
+
## [2.2.0] — 2026-08-25
|
|
8
|
+
|
|
9
|
+
### Upgrade impact
|
|
10
|
+
|
|
11
|
+
Two behaviour changes can turn a previously-green run red. Neither breaks a covered surface
|
|
12
|
+
([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) — both make an assertion report
|
|
13
|
+
what it always documented — so they ship in a minor:
|
|
14
|
+
|
|
15
|
+
- **`no_scratchpad_leak` at `container` can now FAIL.** It was measuring containment against a root in a
|
|
16
|
+
different path space, so every presented file classified `leaked: false` and the check passed
|
|
17
|
+
vacuously. A scenario whose skill genuinely leaves a presented file in the scratchpad will now red —
|
|
18
|
+
that is the leak the key exists to catch.
|
|
19
|
+
- **`baseline: desktop-1.11847.5` is now refused at `container`/`hostloop`/`microvm`.** It carries no
|
|
20
|
+
`spawn` block, so those tiers cannot reproduce Cowork's toolset — a run on it launched an agent with no
|
|
21
|
+
file or bash tools and still reported a verdict. Use a `sync`-recorded baseline, or `fidelity: protocol`.
|
|
22
|
+
|
|
23
|
+
### Fixed
|
|
24
|
+
|
|
25
|
+
- **`present_files_called` no longer reports "the tool was never called" about a run that called it, and a
|
|
26
|
+
hostloop delivery scenario can be recorded with a host-path redaction policy at last.** The assertion
|
|
27
|
+
read presence off `RunResult.presentedFiles`, which is a *classification* of each presented path
|
|
28
|
+
(scratchpad? promoted? leaked?) and requires an absolute path to compute. At `hostloop` a presented path
|
|
29
|
+
is a real host path, so the shipped redaction policy rewrites it to
|
|
30
|
+
`[REDACTED:local-path:<hash>]/mnt/outputs/report.html` — correct, documented, and deliberately ordered so
|
|
31
|
+
the mount tail survives — and the classifier then drops every entry as un-normalizable. The list came
|
|
32
|
+
back empty and the assertion stated, as fact, that the tool had never been called.
|
|
33
|
+
|
|
34
|
+
Because `record` replays the base and redacted cassettes and refuses to write when the verdict differs,
|
|
35
|
+
this was not merely a wrong message: **no cassette asserting `present_files_called` could be recorded at
|
|
36
|
+
that tier at all**, so the assertion had never executed on the replay lane. The refusal was right — the
|
|
37
|
+
redacted cassette genuinely could not support the assert — but the defect it was reporting was in the
|
|
38
|
+
harness, not in the recording.
|
|
39
|
+
|
|
40
|
+
Presence now comes from `RunResult.presentFilesCalls`, a new count of the `present_files` invocations
|
|
41
|
+
that carried a well-formed `file_path`, taken from the tool_use input's shape and never from a path's
|
|
42
|
+
content, so redaction cannot alter it. Classification is untouched: `presentedFiles` still drops what it
|
|
43
|
+
cannot resolve, and `no_scratchpad_leak` still reads it. A run recorded before the field falls back to
|
|
44
|
+
the old `presentedFiles`-non-empty test, so no existing green changes.
|
|
45
|
+
|
|
46
|
+
A run whose every `present_files` call carried an unusable path now reports **cannot verify** rather than
|
|
47
|
+
"never called" — the tool *was* invoked, and the harness already knew it. (The malformed count is
|
|
48
|
+
deliberately kept out of that message: the record self-check normalizes `[REDACTED…]` tokens out of
|
|
49
|
+
failing messages but not digits, so an interpolated count that differed between the two replays would
|
|
50
|
+
refuse a cassette that is otherwise fine to write.)
|
|
51
|
+
|
|
52
|
+
Pinned as an invariant ([docs/invariants.md](./docs/invariants.md)) with the end-to-end record self-check
|
|
53
|
+
as its test anchor — the case that could not be recorded, and so had never run in CI.
|
|
54
|
+
|
|
55
|
+
- **Guest paths are built from the tree the harness stages, not from a baseline's recorded mount layout.**
|
|
56
|
+
`resolveMounts` returned `mountLayout.mntRoot` verbatim — or, when that field was absent and the recorded
|
|
57
|
+
`sessionRoot` already ended in `/mnt`, the session root itself. But the staged tree is always
|
|
58
|
+
`<sessionRoot>/mnt`: `stageWorkspace` creates it there and `dockerRunArgv` nests its read-only binds
|
|
59
|
+
there. On a baseline recording anything else, `--plugin-dir` therefore pointed one directory **above**
|
|
60
|
+
the staged plugin tree, so the plugin under test never loaded. `mntRoot` is now derived from the session
|
|
61
|
+
root, and a recorded layout the harness cannot stage is reported as a fidelity divergence at spawn
|
|
62
|
+
instead of silently composing a path no stager creates.
|
|
63
|
+
|
|
64
|
+
Guest paths now anchor on `sessionRoot` — the bind target — rather than `cwd`, which is only where the
|
|
65
|
+
agent's process starts (production's own working directory is a folder mount or `outputs`, so the two are
|
|
66
|
+
not interchangeable even though every synced baseline records them equal). `dockerRunArgv` takes an
|
|
67
|
+
explicit `agentCwd` for `-w`.
|
|
68
|
+
|
|
69
|
+
Two more derivations of the same rule are gone: `prompt.ts` had a private `/sessions/<id>` +`/mnt`
|
|
70
|
+
literal (so the prompt could describe a tree the runtimes had not staged), and **microvm** read the
|
|
71
|
+
agent's root from the baseline when lima structurally mounts it at `/sessions/<sessionId>` — which put
|
|
72
|
+
`CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages. A baseline recording a cwd that tier
|
|
73
|
+
cannot honour is now **refused**, not warned about: the guest `cd` would otherwise succeed at the wrong
|
|
74
|
+
directory.
|
|
75
|
+
|
|
76
|
+
- **A baseline with no `spawn` block is refused at the sandbox tiers.** That block carries the tool set,
|
|
77
|
+
pre-approvals, effort default and config-dir location, and the `?? []` fallbacks meant a run would launch
|
|
78
|
+
an agent with **no Read/Write/Bash/Skill/Task at all** and still report a verdict. `fidelity: protocol`
|
|
79
|
+
builds its own argv and is unaffected.
|
|
80
|
+
|
|
81
|
+
- **`no_scratchpad_leak` can see a container leak again — the session root was in the wrong path space.** The
|
|
82
|
+
root that `presentedFiles`' promoted/leaked classification is measured from was derived by the caller as
|
|
83
|
+
`<run-dir>/work/session`, a HOST path, and handed to every non-`protocol` tier. But the path space the
|
|
84
|
+
agent reports in is per-tier: at `container` (the default) it runs inside the sandbox and reports
|
|
85
|
+
`/sessions/<id>/…`, and only at `hostloop` does it run natively and report host paths. Measured against a
|
|
86
|
+
host root, no VM path is ever inside the root, so every presented file classified `leaked: false` —
|
|
87
|
+
including the handler's copy-failure branch, which returns the source path unchanged when a file is
|
|
88
|
+
blocked-extension, a directory, or absent. `no_scratchpad_leak` evaluates at `container` and nowhere else,
|
|
89
|
+
so the assertion that exists to catch that leak could not catch it, and `verdict.ts`'s delivery check read
|
|
90
|
+
the same `leaked: false` as a successful delivery.
|
|
91
|
+
|
|
92
|
+
Each runtime now reports the session root it actually launched the agent with, and the run consumes that
|
|
93
|
+
instead of deriving a second one — the two can no longer drift into different spaces. `chat` sets it on
|
|
94
|
+
both serving tiers as well; it never did, so a hostloop chat's `presentedFiles` was inverted in its result
|
|
95
|
+
file. `protocol` and `microvm` serve no `present_files` and keep the cwd fallback.
|
|
96
|
+
|
|
97
|
+
Independently, the classifier now fails CLOSED on a space mismatch: if the agent's cwd is not at or inside
|
|
98
|
+
the session root, the presented batch counts as malformed (`no_scratchpad_leak` → cannot verify) rather
|
|
99
|
+
than being graded against a root it cannot be compared to. `leaked: true` is not derivable in that state
|
|
100
|
+
either, and a `leaked: false` verdict there is exactly the vacuous pass the key exists to prevent.
|
|
101
|
+
|
|
102
|
+
### Added
|
|
103
|
+
|
|
104
|
+
- **First live coverage for `present_files_called` / `no_scratchpad_leak`.** No scenario in the repo
|
|
105
|
+
asserted either key, which is how a session root in the wrong path space could record `leaked: false`
|
|
106
|
+
for every presented file without anything noticing. `e2e/scenarios/smoke-present-files.yaml` writes a
|
|
107
|
+
file OUTSIDE `mnt/` and delivers it, so the promotion is real and the pair is non-vacuous; it runs in
|
|
108
|
+
CI's live e2e loop. Measured on a live container run: `presentFilesCalls: 1`, promoted `true`, leaked
|
|
109
|
+
`false`.
|
|
110
|
+
|
|
111
|
+
- **`RunResult.presentFilesCalls`** — the count of `present_files` invocations that carried a well-formed
|
|
112
|
+
`file_path`, in `result.json` and [schema/run-result.json](./schema/run-result.json). Content-class, so a
|
|
113
|
+
replay re-drive reproduces it; read by `run`, `replay` and `verify-run` alike, and absent (not `0`) on a
|
|
114
|
+
result written before this release, which is what the assertion's fallback distinguishes. Use it to
|
|
115
|
+
answer "did the agent deliver anything?" from a result file without interpreting `presentedFiles`'
|
|
116
|
+
promoted/leaked classification.
|
|
117
|
+
|
|
7
118
|
## [2.1.0] — 2026-08-24
|
|
8
119
|
|
|
9
120
|
### Changed
|
package/README.md
CHANGED
|
@@ -115,7 +115,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
115
115
|
|
|
116
116
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
117
117
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
118
|
-
> From a global install (`npm i -g "cowork-harness@^2.
|
|
118
|
+
> From a global install (`npm i -g "cowork-harness@^2.2.0"`), point at the package root instead:
|
|
119
119
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
120
120
|
> (or copy the cassette into your own project and pass that path).
|
|
121
121
|
|
|
@@ -125,7 +125,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
125
125
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
126
126
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
127
127
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
128
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.
|
|
128
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.2.0"`.
|
|
129
129
|
|
|
130
130
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
131
131
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -150,7 +150,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
150
150
|
claude plugin install cowork-harness@cowork-harness
|
|
151
151
|
```
|
|
152
152
|
|
|
153
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.
|
|
153
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.2.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
154
154
|
|
|
155
155
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
156
156
|
|
|
@@ -175,7 +175,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
175
175
|
ones that still need a source checkout. (The marketplace
|
|
176
176
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
177
177
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
178
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.
|
|
178
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.2.0"` — see
|
|
179
179
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
180
180
|
|
|
181
181
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -772,7 +772,7 @@ jobs:
|
|
|
772
772
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
773
773
|
```
|
|
774
774
|
|
|
775
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.
|
|
775
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.2.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
776
776
|
|
|
777
777
|
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
778
778
|
|
package/dist/assert.js
CHANGED
|
@@ -1141,10 +1141,25 @@ function check(a, ctx) {
|
|
|
1141
1141
|
results.push(fail("present_files_called: present_files is not served on `lane: remote` — a local MCP server cannot reach a remote Cowork session, which delivers via the agent-native SendUserFile instead (not modeled; see docs/fidelity-gaps.md) — cannot verify"));
|
|
1142
1142
|
else if (ctx.effectiveFidelity !== "container" && ctx.effectiveFidelity !== "hostloop")
|
|
1143
1143
|
results.push(fail(`present_files_called: present_files is served only on the container/hostloop tiers (this run: ${ctx.effectiveFidelity ?? "unknown"}) — cannot verify; use fidelity: container or hostloop for present_files-based delivery`));
|
|
1144
|
-
|
|
1145
|
-
|
|
1146
|
-
|
|
1144
|
+
// PRESENCE comes from the invocation count, not from `presentedFiles`. The two answer different
|
|
1145
|
+
// questions, and only one of them survives redaction: `presentedFiles` entries are dropped when a
|
|
1146
|
+
// path can't be classified, which a host-path policy guarantees at hostloop (a real host path
|
|
1147
|
+
// redacts to `[REDACTED:…]/mnt/outputs/f`, and the classifier requires an absolute path). Reading
|
|
1148
|
+
// delivery off classification made a redacted recording report "the tool was never called" about a
|
|
1149
|
+
// run that called it three times — and, because record's redaction self-check replays both
|
|
1150
|
+
// cassettes and compares verdicts, no such cassette could be written at all.
|
|
1151
|
+
// `presentedFiles.length` stays as the fallback for a run recorded before the count existed.
|
|
1152
|
+
else if ((ctx.presentFilesCalls ?? 0) > 0 || (ctx.presentedFiles?.length ?? 0) > 0)
|
|
1147
1153
|
results.push(ok());
|
|
1154
|
+
// Called, but every call's `files` was unusable — the tool WAS invoked, so "never called" would be a
|
|
1155
|
+
// factual claim the harness knows to be false. Mirrors no_scratchpad_leak's malformed branch above.
|
|
1156
|
+
// The count is deliberately NOT interpolated: record's self-check normalizes [REDACTED…] tokens out
|
|
1157
|
+
// of failing messages but not digits, so a count that differs between the base and redacted replays
|
|
1158
|
+
// would trip its message compare and refuse a record that is otherwise fine to write.
|
|
1159
|
+
else if (ctx.evidenceErrors?.presentFilesMalformed)
|
|
1160
|
+
results.push(fail(`present_files_called: present_files WAS called, but no call carried a usable file path (malformed input) — cannot verify delivery`));
|
|
1161
|
+
else
|
|
1162
|
+
results.push(fail(`present_files_called: no file was delivered via present_files (the tool was never called)`));
|
|
1148
1163
|
}
|
|
1149
1164
|
if (a.no_skill_triggered !== undefined) {
|
|
1150
1165
|
const c = compileUserRegex(a.no_skill_triggered);
|
package/dist/baseline.js
CHANGED
|
@@ -319,25 +319,47 @@ function latestBaselineFile() {
|
|
|
319
319
|
files.sort(compareBaselineVersions);
|
|
320
320
|
return join(BASELINES_DIR, files[files.length - 1]);
|
|
321
321
|
}
|
|
322
|
+
/** The one segment the harness's staged session tree always adds under the session root. Every stager
|
|
323
|
+
* and every argv builder composes guest paths with it (`stage.ts`/`hostloop-stage.ts` create
|
|
324
|
+
* `<sessionHost>/mnt`, `dockerRunArgv` nests `:ro` binds at `<sessionRoot>/mnt/<mountPath>`), so a
|
|
325
|
+
* guest mnt root that is anything OTHER than `<sessionRoot>/mnt` cannot be produced. */
|
|
326
|
+
export const GUEST_MNT_SEGMENT = "mnt";
|
|
322
327
|
/**
|
|
323
|
-
* Expand the mount layout for a concrete session id
|
|
324
|
-
*
|
|
325
|
-
*
|
|
328
|
+
* Expand the mount layout for a concrete session id — the layout of the tree THIS HARNESS stages, which
|
|
329
|
+
* is what every guest path must be built from.
|
|
330
|
+
*
|
|
331
|
+
* `mntRoot` is DERIVED as `<sessionRoot>/mnt`, not read from `mountLayout.mntRoot`: the staged tree can
|
|
332
|
+
* only ever be there (see GUEST_MNT_SEGMENT), so honouring a recorded value that says otherwise emits
|
|
333
|
+
* guest paths pointing at directories no stager creates. A baseline recording a different mnt root is a
|
|
334
|
+
* FIDELITY divergence, surfaced by `recordedLayoutDivergence` at spawn, never a path this builds.
|
|
335
|
+
*
|
|
336
|
+
* Guest paths anchor on `sessionRoot` (the bind target), never on `cwd`: production's own working dir is
|
|
337
|
+
* a folder mount or `outputs` rather than the bare session root, so the two are not interchangeable even
|
|
338
|
+
* though every synced baseline currently records them equal. `cwd` is the agent's working directory and
|
|
339
|
+
* nothing else.
|
|
326
340
|
*/
|
|
327
341
|
export function resolveMounts(baseline, sessionId, projectId = "proj1") {
|
|
328
342
|
const subst = (s) => s.replace("{sessionId}", sessionId).replace("{projectId}", projectId);
|
|
329
343
|
const cwd = subst(baseline.mountLayout.cwd);
|
|
330
344
|
const sessionRoot = subst(baseline.mountLayout.sessionRoot);
|
|
331
|
-
|
|
332
|
-
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
};
|
|
345
|
+
const mntRoot = `${sessionRoot}/${GUEST_MNT_SEGMENT}`;
|
|
346
|
+
return { cwd, sessionRoot, mntRoot };
|
|
347
|
+
}
|
|
348
|
+
/** Does this baseline RECORD a guest layout the harness cannot stage? Reads the recorded fields only
|
|
349
|
+
* (never the derived ones), so it stays a statement about the data: a `mntRoot` that is not
|
|
350
|
+
* `<sessionRoot>/mnt`, or a `sessionRoot` that already ends in the mnt segment — the shape that made
|
|
351
|
+
* `resolveMounts` return a root one level above the staged tree. Returns the divergence for the caller
|
|
352
|
+
* to surface, or undefined when the recording is reproducible. */
|
|
353
|
+
export function recordedLayoutDivergence(baseline) {
|
|
354
|
+
// Read STRUCTURALLY: a partially-constructed baseline (`{ spawn: {} }`) is normal at the argv seam, and
|
|
355
|
+
// this check must never be the thing that throws there. No recorded layout ⇒ nothing to diverge from.
|
|
356
|
+
const { sessionRoot, mntRoot } = baseline.mountLayout ?? {};
|
|
357
|
+
if (typeof sessionRoot !== "string")
|
|
358
|
+
return undefined;
|
|
359
|
+
const staged = `${sessionRoot}/${GUEST_MNT_SEGMENT}`;
|
|
360
|
+
if (mntRoot !== undefined && mntRoot !== staged)
|
|
361
|
+
return { recorded: mntRoot, staged };
|
|
362
|
+
if (sessionRoot.endsWith(`/${GUEST_MNT_SEGMENT}`))
|
|
363
|
+
return { recorded: sessionRoot, staged };
|
|
364
|
+
return undefined;
|
|
343
365
|
}
|
package/dist/cli.js
CHANGED
|
@@ -3769,6 +3769,7 @@ async function cmdVerifyRun(args) {
|
|
|
3769
3769
|
fileToolAttempts: result.fileToolAttempts,
|
|
3770
3770
|
pathDenials: result.pathDenials,
|
|
3771
3771
|
presentedFiles: result.presentedFiles,
|
|
3772
|
+
presentFilesCalls: result.presentFilesCalls,
|
|
3772
3773
|
evidenceErrors: result.evidenceErrors,
|
|
3773
3774
|
effectiveFidelity: result.effectiveFidelity,
|
|
3774
3775
|
// verify-run re-checks a kept run dir on the SAME machine that ran it — grouped with the live
|
package/dist/prompt.js
CHANGED
|
@@ -2,6 +2,7 @@ import { warn } from "./io.js";
|
|
|
2
2
|
import { readFileSync, existsSync } from "node:fs";
|
|
3
3
|
import { join } from "node:path";
|
|
4
4
|
import { fileURLToPath } from "node:url";
|
|
5
|
+
import { resolveMounts } from "./baseline.js";
|
|
5
6
|
const BASELINES_DIR = join(fileURLToPath(new URL("..", import.meta.url)), "baselines");
|
|
6
7
|
/** The {{placeholder}} names renderPrompts() substitutes. KEEP IN LOCKSTEP with the `tokens` map below. */
|
|
7
8
|
export const MODELED_PLACEHOLDER_NAMES = new Set([
|
|
@@ -31,8 +32,10 @@ firstFolderMountPath, opts) {
|
|
|
31
32
|
const spawn = baseline.spawn;
|
|
32
33
|
if (!spawn)
|
|
33
34
|
return {};
|
|
34
|
-
|
|
35
|
-
|
|
35
|
+
// Through resolveMounts, NOT a local `/sessions/<id>` literal: the prompt tells the agent where its
|
|
36
|
+
// files are, so it must name the same tree the runtimes stage and bind. A private derivation here was a
|
|
37
|
+
// fourth copy of this rule, and a fourth place for it to drift.
|
|
38
|
+
const { sessionRoot, mntRoot } = resolveMounts(baseline, sessionId);
|
|
36
39
|
const workspaceFolder = firstFolderMountPath ? `${mntRoot}/${firstFolderMountPath}` : `${mntRoot}/outputs`;
|
|
37
40
|
// {{currentDateTime}}/{{currentTimezone}} are deliberately render-time-impure (wall clock, host
|
|
38
41
|
// TZ) — that's what the real Desktop builder substitutes, and the rendered append never enters a
|
package/dist/run/cassette.js
CHANGED
|
@@ -1741,6 +1741,7 @@ function minimalRec() {
|
|
|
1741
1741
|
fileToolAttempts: [],
|
|
1742
1742
|
pathDenials: [],
|
|
1743
1743
|
presentedFiles: [],
|
|
1744
|
+
presentFilesCalls: 0,
|
|
1744
1745
|
webSearches: [],
|
|
1745
1746
|
infraErrors: [],
|
|
1746
1747
|
evidenceErrors: { taskTracking: 0, webSearchParse: 0, presentFilesMalformed: 0 },
|
|
@@ -3920,6 +3921,7 @@ function replayErrorResult(file) {
|
|
|
3920
3921
|
fileToolAttempts: undefined, // no rec to read from on this early-bail lane
|
|
3921
3922
|
pathDenials: undefined, // no rec to read from on this early-bail lane
|
|
3922
3923
|
presentedFiles: undefined, // no rec to read from on this early-bail lane
|
|
3924
|
+
presentFilesCalls: undefined, // ditto
|
|
3923
3925
|
preRunPaths: undefined,
|
|
3924
3926
|
preRunLinkAware: undefined,
|
|
3925
3927
|
preRunHashes: undefined,
|
|
@@ -5620,8 +5622,10 @@ export const ALWAYS_CONTENT_KEYS = [
|
|
|
5620
5622
|
"task_status",
|
|
5621
5623
|
"result",
|
|
5622
5624
|
// content-class, NOT controlOut-gated: both the present_files tool_use and its own tool_result live
|
|
5623
|
-
// in the ordinary events stream, so the re-drive reproduces
|
|
5624
|
-
//
|
|
5625
|
+
// in the ordinary events stream, so the re-drive reproduces the present_files signals like the other
|
|
5626
|
+
// re-derived ones above (skill_triggered, redundantToolCalls, …). `presentFilesCalls` (what
|
|
5627
|
+
// present_files_called reads) reproduces on every tier; `presentedFiles`' promoted/leaked booleans
|
|
5628
|
+
// reproduce at container, where cwd is the session root — no_scratchpad_leak's only tier.
|
|
5625
5629
|
"no_scratchpad_leak",
|
|
5626
5630
|
"present_files_called",
|
|
5627
5631
|
// content-class, NOT controlOut-gated: fileToolAttempts re-derives from frozen tool_use blocks (the
|
|
@@ -6136,6 +6140,7 @@ export async function replayCassette(cassette, hooks = [], opts = {}) {
|
|
|
6136
6140
|
// empty [] (nothing presented) vacuous-passes no_scratchpad_leak instead of reading as
|
|
6137
6141
|
// evidence-unavailable.
|
|
6138
6142
|
presentedFiles: rec.presentedFiles,
|
|
6143
|
+
presentFilesCalls: rec.presentFilesCalls,
|
|
6139
6144
|
evidenceErrors: rec.evidenceErrors,
|
|
6140
6145
|
effectiveFidelity: cassette.effectiveFidelity,
|
|
6141
6146
|
// Replay has no live filesystem — computer_links_resolve normalizes both link shapes against the
|
|
@@ -6493,6 +6498,7 @@ export async function replayCassette(cassette, hooks = [], opts = {}) {
|
|
|
6493
6498
|
// NOT reproduce — this one genuinely re-derives. Uncollapsed (an empty [] is the real "nothing
|
|
6494
6499
|
// presented" signal no_scratchpad_leak's vacuous pass needs, matching live).
|
|
6495
6500
|
presentedFiles: rec.presentedFiles,
|
|
6501
|
+
presentFilesCalls: rec.presentFilesCalls,
|
|
6496
6502
|
preRunPaths: undefined,
|
|
6497
6503
|
// Report the baseline semantics actually used during evaluation above (not undefined) so the returned
|
|
6498
6504
|
// result doesn't misrepresent them. Same source of truth as the evaluate() ctx.
|
package/dist/run/chat-result.js
CHANGED
|
@@ -113,6 +113,7 @@ export function buildChatResult(record, opts) {
|
|
|
113
113
|
fileToolAttempts: record.fileToolAttempts,
|
|
114
114
|
pathDenials: record.pathDenials,
|
|
115
115
|
presentedFiles: record.presentedFiles,
|
|
116
|
+
presentFilesCalls: record.presentFilesCalls,
|
|
116
117
|
egress: opts.egress,
|
|
117
118
|
resources,
|
|
118
119
|
stderrLogPath: join(opts.outDir, "agent.stderr.log"),
|
package/dist/run/chat.js
CHANGED
|
@@ -417,6 +417,7 @@ export async function cmdChat(args) {
|
|
|
417
417
|
},
|
|
418
418
|
};
|
|
419
419
|
const run = new Run(agent, decider, [renderer, tripwireHook], sessionId);
|
|
420
|
+
run.setSessionRoot(hl.sessionRoot); // HOST tree — without it cwd (mnt/outputs) stands in for the root
|
|
420
421
|
stopHeartbeat = startHeartbeat(renderer, renderPlan, start);
|
|
421
422
|
if (viaApiOn) {
|
|
422
423
|
run.enableWebFetchGate();
|
|
@@ -450,6 +451,7 @@ export async function cmdChat(args) {
|
|
|
450
451
|
const decider = Chain(new ScriptedDecider([]), new PermissionDefaultDecider("cowork"), new PromptDecider(ask));
|
|
451
452
|
const renderer = makeRenderer(renderPlan);
|
|
452
453
|
const run = new Run(agent, decider, [renderer], sessionId);
|
|
454
|
+
run.setSessionRoot(ct.sessionRoot); // VM path — same space the agent reports (see execute.ts)
|
|
453
455
|
stopHeartbeat = startHeartbeat(renderer, renderPlan, start);
|
|
454
456
|
record = await run.drive(withSeedPrompt(seedPrompt, ttyTurns(rl)), chatDriveOpts(prompts, { tier: "container", sdkMcp: ct.sdkMcp }));
|
|
455
457
|
}
|
package/dist/run/execute.js
CHANGED
|
@@ -692,6 +692,10 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
692
692
|
const prompts = renderPrompts(baseline, session, sessionId, plan.mounts.find((m) => m.kind === "folder")?.mountPath, hostLoopOpts);
|
|
693
693
|
promptFidelityWarnings = prompts.fidelityWarnings; // hoist out so RunResult construction (after try) can access it
|
|
694
694
|
let sdkMcp;
|
|
695
|
+
// The session root as reported BY THE SPAWN that just happened — the dir whose `mnt/` is the
|
|
696
|
+
// user-visible workspace, in the same path space the agent reports its own paths in. Only the two
|
|
697
|
+
// tiers that serve present_files supply one; the others keep the cwd fallback (see setSessionRoot).
|
|
698
|
+
let spawnedSessionRoot;
|
|
695
699
|
if (effectiveFidelity === "hostloop") {
|
|
696
700
|
const hl = spawnHostLoop(scenario, baseline, plan, outDir, sessionId, {
|
|
697
701
|
systemPromptAppend: prompts.systemPromptAppend,
|
|
@@ -712,6 +716,7 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
712
716
|
hostloopPathGateFired = hl.pathGateFired;
|
|
713
717
|
hostloopInfraErrors = hl.infraErrors;
|
|
714
718
|
hostloopMarkTearingDown = hl.markTearingDown;
|
|
719
|
+
spawnedSessionRoot = hl.sessionRoot; // HOST tree — the native agent runs there
|
|
715
720
|
logHostWriteNotice(plan.mounts.filter((mt) => mt.kind === "folder").map((mt) => ({ from: mt.hostPath, mode: mt.mode })), warn);
|
|
716
721
|
if (scenario.assert.some((a) => a.transcript_no_host_path === true) && !opts.compact)
|
|
717
722
|
warn(`::warning:: [hostloop] scenario asserts transcript_no_host_path — hostloop's native file tools legitimately ` +
|
|
@@ -729,6 +734,7 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
729
734
|
child = ct.child;
|
|
730
735
|
containerName = ct.containerName; // so the Ctrl-C / finally reap removes the agent container by name
|
|
731
736
|
sdkMcp = ct.sdkMcp; // cowork/present_files + the skills/plugins discovery servers (combineSdkMcp)
|
|
737
|
+
spawnedSessionRoot = ct.sessionRoot; // VM path (`/sessions/<id>`) — what the agent inside reports
|
|
732
738
|
}
|
|
733
739
|
else if (effectiveFidelity === "microvm") {
|
|
734
740
|
child = spawnMicroVm(scenario, baseline, plan, outDir, sessionId, {
|
|
@@ -771,12 +777,18 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
771
777
|
const decider = effectiveFidelity === "hostloop" ? Chain(makeHostLoopCanUseToolGate(), policyDecider) : policyDecider;
|
|
772
778
|
const run = new Run(sessionT, decider, opts.hooks ?? [], sessionId, dialogTimeoutMs ?? undefined, scenario.timeout_ms);
|
|
773
779
|
run.seedApprovedDomains(session.web_fetch.approved_domains); // test convenience: pre-approved web_fetch hosts
|
|
774
|
-
// The session root — the dir whose `mnt/` IS the user-visible workspace
|
|
775
|
-
//
|
|
776
|
-
//
|
|
777
|
-
//
|
|
778
|
-
|
|
779
|
-
|
|
780
|
+
// The session root — the dir whose `mnt/` IS the user-visible workspace, and what present_files'
|
|
781
|
+
// promoted/leaked classification is measured from. Taken from the SPAWN, never re-derived here: the
|
|
782
|
+
// root and the agent's reported paths must be in the SAME path space, and they are not the same space
|
|
783
|
+
// on every tier. A host path was passed unconditionally once; at container the agent reports VM paths
|
|
784
|
+
// (`/sessions/<id>/…`), so nothing was ever inside the root, every presented file classified
|
|
785
|
+
// `leaked: false`, and `no_scratchpad_leak` — which evaluates at container and nowhere else — passed
|
|
786
|
+
// vacuously over a real copy-failure leak.
|
|
787
|
+
//
|
|
788
|
+
// Unset on the tiers that serve no present_files (`protocol`, `microvm`), where the cwd fallback is
|
|
789
|
+
// already the session root and there is no delivery to classify.
|
|
790
|
+
if (spawnedSessionRoot !== undefined)
|
|
791
|
+
run.setSessionRoot(spawnedSessionRoot);
|
|
780
792
|
// fill the provenance bundle (backed by Run's tracker + recorded approval) BEFORE drive().
|
|
781
793
|
// Host-loop only, and only when the web_fetch-via-API gate is on; otherwise the handler stays
|
|
782
794
|
// allowlist-only (ref.current undefined). Run seeds the set from turns + tool_results.
|
|
@@ -1157,6 +1169,7 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
1157
1169
|
// Always defined live — an empty array is the real "nothing presented" signal no_scratchpad_leak's
|
|
1158
1170
|
// vacuous pass needs, distinct from replay's evidence-unavailable undefined on an older cassette.
|
|
1159
1171
|
presentedFiles: record.presentedFiles,
|
|
1172
|
+
presentFilesCalls: record.presentFilesCalls,
|
|
1160
1173
|
evidenceErrors: record.evidenceErrors,
|
|
1161
1174
|
effectiveFidelity,
|
|
1162
1175
|
// Live lane (this run's own machine) — host-shaped computer:// links (hostloop) are checked
|
|
@@ -1404,6 +1417,7 @@ export async function executeScenario(scenario, opts = {}) {
|
|
|
1404
1417
|
fileToolAttempts: record.fileToolAttempts, // uncollapsed — content-class, same as toolResults/decisions above
|
|
1405
1418
|
pathDenials: record.pathDenials, // uncollapsed — content-class, same as fileToolAttempts above
|
|
1406
1419
|
presentedFiles: record.presentedFiles, // uncollapsed — an empty [] is the real "nothing presented" signal no_scratchpad_leak's vacuous pass needs
|
|
1420
|
+
presentFilesCalls: record.presentFilesCalls,
|
|
1407
1421
|
// The pre-spawn baseline no_unexpected_files diffs against (same single read the evaluate ctx got).
|
|
1408
1422
|
// undefined = the run didn't capture (key not asserted, microvm, pre-seam) — the assertion then
|
|
1409
1423
|
// fails evidence-unavailable, loud.
|
|
@@ -1788,6 +1802,7 @@ export function buildPartialResult(args) {
|
|
|
1788
1802
|
fileToolAttempts: record.fileToolAttempts, // uncollapsed — content-class, same as toolResults/decisions above
|
|
1789
1803
|
pathDenials: record.pathDenials, // uncollapsed — content-class, same as fileToolAttempts above
|
|
1790
1804
|
presentedFiles: record.presentedFiles, // uncollapsed — an empty [] is the real "nothing presented" signal no_scratchpad_leak's vacuous pass needs
|
|
1805
|
+
presentFilesCalls: record.presentFilesCalls,
|
|
1791
1806
|
preRunPaths: readPreRunManifest(args.outDir),
|
|
1792
1807
|
preRunLinkAware: readPreRunManifestLinkAware(args.outDir),
|
|
1793
1808
|
preRunHashes: readPreRunManifestHashes(args.outDir),
|
package/dist/run/run.js
CHANGED
|
@@ -267,6 +267,7 @@ export class Run {
|
|
|
267
267
|
fileToolAttempts: [],
|
|
268
268
|
pathDenials: [],
|
|
269
269
|
presentedFiles: [],
|
|
270
|
+
presentFilesCalls: 0,
|
|
270
271
|
webSearches: [],
|
|
271
272
|
infraErrors: [],
|
|
272
273
|
evidenceErrors: { taskTracking: 0, webSearchParse: 0, presentFilesMalformed: 0 },
|
|
@@ -488,6 +489,12 @@ export class Run {
|
|
|
488
489
|
const pf = presentFilesInput(ev.input);
|
|
489
490
|
this.pendingPresentFiles.set(ev.toolUseId, pf.files);
|
|
490
491
|
this.rec.evidenceErrors.presentFilesMalformed += pf.malformed;
|
|
492
|
+
// Presence, counted here and NOT from `presentedFiles` below — see the field's comment.
|
|
493
|
+
// Gated on a well-formed file (not merely on the call) so the count keeps the DELIVERY
|
|
494
|
+
// meaning `present_files_called` documents: a call whose `files` were unusable delivered
|
|
495
|
+
// nothing, leaves this at 0, and is reported through the malformed counter instead.
|
|
496
|
+
if (pf.files.length > 0)
|
|
497
|
+
this.rec.presentFilesCalls++;
|
|
491
498
|
}
|
|
492
499
|
this.toolLog.push({ name: ev.name, input: ev.input, synthetic: ev.synthetic, parentToolUseId: ev.parentToolUseId }); // still logged for provenance/trace
|
|
493
500
|
break;
|
|
@@ -932,6 +939,18 @@ export class Run {
|
|
|
932
939
|
// genuine leak as fine. This bug over-reports; that one would under-report, and a false green is the
|
|
933
940
|
// worse failure.
|
|
934
941
|
const root = this.sessionRoot !== undefined ? posixPath.normalize(this.sessionRoot) : cwd;
|
|
942
|
+
// SPACE CHECK, before any classification: the agent's cwd must sit AT or INSIDE the session root. That
|
|
943
|
+
// holds on every tier that serves present_files — at container cwd IS the root, at hostloop it is
|
|
944
|
+
// `<root>/mnt/<outputs|folder>` — so a cwd outside the root means the two are in different path spaces
|
|
945
|
+
// (a host root against VM-reported paths, say) and every containment test below is meaningless. Count
|
|
946
|
+
// the batch malformed instead of grading it: nothing would be under the root, so the classification
|
|
947
|
+
// would silently read `leaked: false` for a genuine leak, which is exactly the vacuous pass
|
|
948
|
+
// `no_scratchpad_leak` exists to prevent. Deliberately NOT "no presented path is under the root" —
|
|
949
|
+
// a hostloop delivery out of a connected folder legitimately sits outside the session tree.
|
|
950
|
+
if (root !== undefined && cwd !== undefined && cwd !== root && !cwd.startsWith(`${root}/`)) {
|
|
951
|
+
this.rec.evidenceErrors.presentFilesMalformed += froms.length;
|
|
952
|
+
return;
|
|
953
|
+
}
|
|
935
954
|
const isScratchpad = (p) => root !== undefined && p.startsWith(`${root}/`) && !p.startsWith(`${root}/mnt/`);
|
|
936
955
|
for (let i = 0; i < froms.length; i++) {
|
|
937
956
|
const rawTo = tos[i];
|
package/dist/runtime/argv.js
CHANGED
|
@@ -1,3 +1,5 @@
|
|
|
1
|
+
import { recordedLayoutDivergence } from "../baseline.js";
|
|
2
|
+
import { warn } from "../io.js";
|
|
1
3
|
import { DEFAULT_MAX_THINKING_TOKENS } from "../types.js";
|
|
2
4
|
import { SECRET_ENV_KEYS } from "./host-env.js";
|
|
3
5
|
/**
|
|
@@ -9,6 +11,23 @@ import { SECRET_ENV_KEYS } from "./host-env.js";
|
|
|
9
11
|
*/
|
|
10
12
|
export function baseAgentArgs(baseline, plan, opts) {
|
|
11
13
|
const spawn = baseline.spawn;
|
|
14
|
+
// A baseline with NO `spawn` block cannot be spawned faithfully at a sandbox tier: the tool set,
|
|
15
|
+
// pre-approvals, effort default and config-dir location all come from it, and the `?? []` fallbacks
|
|
16
|
+
// below would silently emit an agent with no Read/Write/Bash/Skill/Task at all — a run that cannot
|
|
17
|
+
// execute the skill under test while still reporting a verdict. Refuse instead. (`protocol` builds its
|
|
18
|
+
// own argv and inherits the host CLI's toolset, so it is unaffected and stays usable.)
|
|
19
|
+
if (!spawn)
|
|
20
|
+
throw new Error(`baseline "${baseline.appVersion}" has no \`spawn\` block, so a sandbox tier cannot reproduce Cowork's ` +
|
|
21
|
+
`toolset, pre-approvals or config-dir layout — the agent would launch with none of its file/bash tools. ` +
|
|
22
|
+
`Use a baseline recorded by \`sync\`, or run at \`fidelity: protocol\` (which builds its own argv).`);
|
|
23
|
+
// FIDELITY, not correctness: guest paths are always staged at `<sessionRoot>/mnt` (see
|
|
24
|
+
// GUEST_MNT_SEGMENT), so a baseline recording a different mnt root is reproduced approximately. Say so
|
|
25
|
+
// once, here, rather than letting the recorded value build a path no stager creates.
|
|
26
|
+
const divergence = recordedLayoutDivergence(baseline);
|
|
27
|
+
if (divergence)
|
|
28
|
+
warn(`::warning:: baseline "${baseline.appVersion}" records a guest mnt root this harness cannot stage ` +
|
|
29
|
+
`(recorded ${divergence.recorded}, staged ${divergence.staged}) — guest paths use the staged layout, ` +
|
|
30
|
+
`so mount-path fidelity at this tier is approximate.\n`);
|
|
12
31
|
// Real Cowork ALWAYS emits `--effort`, for every model class (picker, no-picker, regex-default,
|
|
13
32
|
// unknown) — falling back to the baseline's synced medium default when the session left it unset
|
|
14
33
|
// (per-model validation of an EXPLICIT value already ran in buildLaunchPlan's validateEffort; the
|
|
@@ -207,7 +226,7 @@ export function dockerRunArgv(i) {
|
|
|
207
226
|
i.network,
|
|
208
227
|
...(i.lockdown ? HARDENING : []),
|
|
209
228
|
"-w",
|
|
210
|
-
i.sessionRoot,
|
|
229
|
+
i.agentCwd ?? i.sessionRoot,
|
|
211
230
|
// Render SECRET values by NAME only (`-e KEY`) so the token never lands in `docker run`'s
|
|
212
231
|
// argv (visible via ps / /proc/<pid>/cmdline). Docker inherits the value from its own env — the
|
|
213
232
|
// harness process env, where runtimeAuthEnv read it. Non-secret env keeps the explicit KEY=value.
|
|
@@ -23,8 +23,11 @@ import { resolveAgentImage, resolveContainerRuntime } from "./agent-image.js";
|
|
|
23
23
|
*/
|
|
24
24
|
export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opts = {}) {
|
|
25
25
|
const m = resolveMounts(baseline, sessionId, "proj1");
|
|
26
|
-
|
|
27
|
-
|
|
26
|
+
// BIND TARGET (and the anchor for every guest path), vs the agent's working dir. Equal on every synced
|
|
27
|
+
// baseline; kept distinct because production's cwd is a folder mount or `outputs`, not the session root.
|
|
28
|
+
const sessionRoot = m.sessionRoot; // /sessions/<id>
|
|
29
|
+
const agentCwd = m.cwd;
|
|
30
|
+
const mntRoot = m.mntRoot; // <sessionRoot>/mnt — the tree stageWorkspace creates
|
|
28
31
|
const configGuest = `${sessionRoot}/${baseline.spawn?.configDirInGuest ?? "mnt/.claude"}`;
|
|
29
32
|
const AGENT_IN = "/usr/local/bin/claude";
|
|
30
33
|
// Name by the per-invocation runToken (NOT sessionId) so a --resume after a failed run doesn't collide
|
|
@@ -88,6 +91,7 @@ export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opt
|
|
|
88
91
|
lockdown: (process.env.COWORK_LOCKDOWN ?? "on") !== "off",
|
|
89
92
|
name: containerName,
|
|
90
93
|
sessionRoot,
|
|
94
|
+
agentCwd,
|
|
91
95
|
sessionHost,
|
|
92
96
|
agentHost,
|
|
93
97
|
agentIn: AGENT_IN,
|
|
@@ -130,5 +134,9 @@ export function spawnContainer(_scenario, baseline, plan, outDir, sessionId, opt
|
|
|
130
134
|
handle: makePluginsHandler({ mountedPlugins }),
|
|
131
135
|
};
|
|
132
136
|
const sdkMcp = combineSdkMcp(...(coworkBundle ? [coworkBundle] : []), skillsBundle, pluginsBundle);
|
|
133
|
-
|
|
137
|
+
// `sessionRoot` is the VM path the agent sees (`-w` above, and the cowork handler's own
|
|
138
|
+
// `sessionRootVm`). Returned so the caller classifies present_files against the root THIS spawn used,
|
|
139
|
+
// instead of re-deriving one — the two lived in different path spaces (host vs VM) once, which made
|
|
140
|
+
// every container leak read as `leaked: false`.
|
|
141
|
+
return { child, containerName, sdkMcp, sessionRoot };
|
|
134
142
|
}
|
package/dist/runtime/hostloop.js
CHANGED
|
@@ -363,7 +363,11 @@ export function spawnHostLoop(_scenario, baseline, plan, outDir, sessionId, opts
|
|
|
363
363
|
handle: makePluginsHandler({ mountedPlugins }),
|
|
364
364
|
};
|
|
365
365
|
const sdkMcp = combineSdkMcp(workspaceBundle, ...(coworkBundle ? [coworkBundle] : []), skillsBundle, pluginsBundle);
|
|
366
|
-
|
|
366
|
+
// `sessionRoot` here is the HOST tree (`sessionHost`), not the VM path: the agent runs natively on the
|
|
367
|
+
// host at this tier (see the `spawn(agentNativeHost, …, { cwd: hostOutputsDir })` above), so the paths
|
|
368
|
+
// it reports — and the ones its present_files handler validates — are host paths. Returned for the same
|
|
369
|
+
// reason as container's: the caller must not re-derive it.
|
|
370
|
+
return { child, sdkMcp, hooks, pathGateFired, containerName, hostEgress, infraErrors, markTearingDown, sessionRoot: sessionHost };
|
|
367
371
|
}
|
|
368
372
|
/** The two infra-error emitters a host-loop run needs, sharing one sink and one events.jsonl writer.
|
|
369
373
|
* Every row carries its `source` so the verdict can tell a dead supervisor apart from a failed command,
|
package/dist/runtime/lima.js
CHANGED
|
@@ -5,6 +5,12 @@ import { join, dirname, basename } from "node:path";
|
|
|
5
5
|
import { createHash } from "node:crypto";
|
|
6
6
|
/** Host dir mounted writable into the VM at /sessions (the staging area; per-session subdirs). */
|
|
7
7
|
export const VM_WORK_HOST = join(homedir(), ".cowork-harness", "vm-work");
|
|
8
|
+
/** Guest mount point of `VM_WORK_HOST` — the literal in the lima template below. The microVM's guest
|
|
9
|
+
* session root is `${VM_GUEST_SESSIONS_ROOT}/<sessionId>` STRUCTURALLY: lima mounts the work root here
|
|
10
|
+
* and the per-session dirs live inside it, so this tier cannot honour a baseline that records the agent
|
|
11
|
+
* running anywhere else. Exported so the runtime anchors on the same string the template mounts, and so
|
|
12
|
+
* a test can check that pairing without booting a VM. */
|
|
13
|
+
export const VM_GUEST_SESSIONS_ROOT = "/sessions";
|
|
8
14
|
/**
|
|
9
15
|
* L2 microVM provisioning via Lima with `vmType: vz` — Apple Virtualization.framework,
|
|
10
16
|
* the SAME hypervisor Claude Cowork uses. This gives a real Linux kernel (VM-grade
|
|
@@ -141,7 +147,7 @@ mounts:
|
|
|
141
147
|
# the agent persists its session (enabling --resume). Lima creates the mountpoint, writable by the
|
|
142
148
|
# mounting user — no guest /sessions permission problem.
|
|
143
149
|
- location: "${VM_WORK_HOST}"
|
|
144
|
-
mountPoint: "
|
|
150
|
+
mountPoint: "${VM_GUEST_SESSIONS_ROOT}"
|
|
145
151
|
writable: true
|
|
146
152
|
provision:
|
|
147
153
|
- mode: system
|
package/dist/runtime/microvm.js
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
import { spawn } from "node:child_process";
|
|
2
2
|
import { cpSync, existsSync, mkdirSync, rmSync } from "node:fs";
|
|
3
3
|
import { dirname, join } from "node:path";
|
|
4
|
-
import { limaPath, vmInit, applyGuestFirewall, vmGatewayIp, VM_WORK_HOST } from "./lima.js";
|
|
4
|
+
import { limaPath, vmInit, applyGuestFirewall, vmGatewayIp, VM_WORK_HOST, VM_GUEST_SESSIONS_ROOT } from "./lima.js";
|
|
5
5
|
import { resolveMounts } from "../baseline.js";
|
|
6
6
|
import { spawnEnv, baseAgentArgs } from "./argv.js";
|
|
7
7
|
import { stageWorkspace } from "./stage.js";
|
|
@@ -62,11 +62,29 @@ function parseEnvPortMicroVm(name, defaultValue) {
|
|
|
62
62
|
*
|
|
63
63
|
* Egress: a host-side allowlist proxy + a guest default-deny iptables firewall.
|
|
64
64
|
*/
|
|
65
|
+
/** The microVM's guest session root, which is STRUCTURAL rather than baseline-derived: lima mounts the
|
|
66
|
+
* work root at `VM_GUEST_SESSIONS_ROOT` with a per-`sessionId` dir inside it, so the agent runs at
|
|
67
|
+
* `<VM_GUEST_SESSIONS_ROOT>/<sessionId>` whatever the baseline records. Reading it from
|
|
68
|
+
* `mountLayout.cwd` instead put `CLAUDE_CONFIG_DIR` and `--mcp-config` at paths nothing stages whenever
|
|
69
|
+
* the two disagreed.
|
|
70
|
+
*
|
|
71
|
+
* A recorded cwd this tier cannot honour THROWS rather than warns: the guest `cd` would otherwise
|
|
72
|
+
* succeed at the wrong directory, and a silently-wrong cwd is exactly what the loud `cd` prevents.
|
|
73
|
+
*
|
|
74
|
+
* Pure and exported so this pairing is checkable without booting a VM (`vmInit` needs `limactl`). */
|
|
75
|
+
export function microvmGuestSessionRoot(baseline, sessionId) {
|
|
76
|
+
const sessionVm = `${VM_GUEST_SESSIONS_ROOT}/${sessionId}`;
|
|
77
|
+
const recordedCwd = resolveMounts(baseline, sessionId, "proj1").cwd;
|
|
78
|
+
if (recordedCwd !== sessionVm)
|
|
79
|
+
throw new Error(`baseline "${baseline.appVersion}" records the agent's cwd as ${recordedCwd}, but the microVM mounts its ` +
|
|
80
|
+
`session tree at ${sessionVm} and cannot place the agent anywhere else — run this baseline at ` +
|
|
81
|
+
`fidelity: container/hostloop, or use one whose mountLayout.cwd matches the microVM mount point.`);
|
|
82
|
+
return sessionVm;
|
|
83
|
+
}
|
|
65
84
|
export function spawnMicroVm(_scenario, baseline, plan, outDir, sessionId, opts = {}) {
|
|
66
85
|
const { instance } = vmInit(baseline);
|
|
67
|
-
const
|
|
68
|
-
const
|
|
69
|
-
const mntVm = m.mntRoot; // /sessions/<id>/mnt
|
|
86
|
+
const sessionVm = microvmGuestSessionRoot(baseline, sessionId);
|
|
87
|
+
const mntVm = `${sessionVm}/mnt`; // the staged tree — see GUEST_MNT_SEGMENT
|
|
70
88
|
const configVm = `${sessionVm}/${baseline.spawn?.configDirInGuest ?? "mnt/.claude"}`;
|
|
71
89
|
// Stage into the lima-mounted work dir (host VM_WORK_HOST -> guest /cowork-work) via the shared
|
|
72
90
|
// helper. It honors plan.resume for ALL of .claude + mounts + mcp.json (the .claude guard
|
package/dist/types.js
CHANGED
|
@@ -456,7 +456,7 @@ export const Assertion = z.strictObject({
|
|
|
456
456
|
present_files_called: z
|
|
457
457
|
.literal(true)
|
|
458
458
|
.optional()
|
|
459
|
-
.describe("at least one file was actually delivered via the present_files tool (presentedFiles is
|
|
459
|
+
.describe("at least one file was actually delivered via the present_files tool (at least one call carried a well-formed file_path). Presence is read from the INVOCATION count, not from the classified presentedFiles list, so it is unaffected by a redaction policy that rewrites host paths; a run that called the tool but whose every call carried an unusable path reports cannot-verify, never 'the tool was never called'. The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid"),
|
|
460
460
|
egress_denied: z.string().optional().describe("egress to this host was denied"),
|
|
461
461
|
egress_allowed: z.string().optional().describe("egress to this host was allowed"),
|
|
462
462
|
// Only `true` is accepted: `false` is rejected as a footgun. The assertion is presence-semantic — authoring
|
package/docs/cassette.md
CHANGED
|
@@ -512,8 +512,8 @@ the rules and CI-placement rationale (why each category behaves this way), see
|
|
|
512
512
|
| `path_denied` | **`fidelity: hostloop` only** — a path denial matched all given matchers (`tool`/`path_matches`/`source`/`agent_scope`) — replay: needs `controlOut`; any other tier FAILS "cannot verify" |
|
|
513
513
|
| `no_path_denied` | **`fidelity: hostloop` only** — no path denial was recorded at all — replay: needs `controlOut`. **Only `true` is valid**; any other tier FAILS "cannot verify" |
|
|
514
514
|
| `result` | run ended with `success` or `error` |
|
|
515
|
-
| `no_scratchpad_leak` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuous pass if nothing was presented; content-class: both the tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives
|
|
516
|
-
| `present_files_called` | at least one file was delivered via `present_files` (`RunResult.
|
|
515
|
+
| `no_scratchpad_leak` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuous pass if nothing was presented; content-class: both the tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only; evidence-unavailable only when `presentedFiles` is absent (an older run predating the feature); **container tier only** — `present_files` *is* served on hostloop, but that branch passes a validated path through without promoting, so there is no scratch→outputs copy to leak; `microvm`/`protocol` do not serve the tool at all |
|
|
516
|
+
| `present_files_called` | at least one file was delivered via `present_files` (`RunResult.presentFilesCalls` > 0 — the invocation count, which a host-path redaction policy cannot alter) — the presence companion to `no_scratchpad_leak`; content-class (re-derives identically on replay); **container + hostloop tiers** |
|
|
517
517
|
| `allow_permissive_auto_allow` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
|
|
518
518
|
| `allow_missing_capability` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
|
|
519
519
|
| `allow_l0_plugin_divergence` | verdict modifier — kept on replay → no-op pass (the live signal it suppresses is zeroed) |
|
package/docs/invariants.md
CHANGED
|
@@ -22,6 +22,9 @@ duplicate it here; this table is for finding things fast, not for the full ratio
|
|
|
22
22
|
| A cassette carries a `sessionFingerprint` (v9+, content-SHAPE hash of the AUTHORED, pre-resolution session — connected folders, plugins, skills, mcp and egress config, web_fetch approved domains, plus projects and agent_env when set; authored so the hash stays relocatable across checkouts) checked ONLY by `verify-cassettes`, deliberately excluded from `computeStaleness`/`checkStaleness` so it never affects the default replay verdict | `src/run/cassette.ts` (`sessionFingerprintDrift`) | `test/session-fingerprint.test.ts` (`describe("buildSessionFingerprint (function-level)")` — covers `buildSessionFingerprint` + `sessionFingerprintDrift`) |
|
|
23
23
|
| A corrupt `timeline.jsonl` header must read as "no timeline" (`undefined`) at every consumer, never a present-empty timeline — a header-corrupt read that folds to `[]`/`undefined` events bakes a novel "ran, no activity" shape into a recorded cassette instead of the honest evidence-unavailable state | `src/run/execute.ts`, `src/run/chat-result.ts`, `src/run/cassette.ts` (each checks `!timelineRaw.headerCorrupt` alongside `malformedLines === 0` before trusting a `readTimeline()` result) | `test/central-cluster-guards.test.ts` |
|
|
24
24
|
| `evidenceErrors.egressParse` must be included in the presence gate that decides whether `RunResult.evidenceErrors` serializes at all — omitting it silently absorbs the one counter that exists specifically to surface dropped egress-proxy lines | `src/run/run.ts` (`evidenceErrorsForResult`) | `test/central-cluster-guards.test.ts` (`describe("#39 egressParse reaches result.json (presence-gate fix, not just parseEgressLine)")`) |
|
|
25
|
+
| Every GUEST path must be composed from the tree the harness stages — `<sessionRoot>/mnt` — and never from a baseline's recorded `mountLayout.mntRoot`. A recorded mnt root that differs is a fidelity divergence reported at spawn, not a path to build: honouring one put `--plugin-dir` a directory above the staged plugin tree, so the plugin under test never loaded. Guest paths anchor on `sessionRoot` (the bind target), never `cwd` (where the agent merely starts); at microvm the root is lima's mount point and a baseline recording another cwd is REFUSED, since the guest `cd` would otherwise succeed at the wrong dir. A baseline with no `spawn` block is refused at the sandbox tiers outright — its toolset and pre-approvals are that block | `src/baseline.ts` (`GUEST_MNT_SEGMENT`, `resolveMounts`, `recordedLayoutDivergence`), `src/runtime/argv.ts` (the spawn-block refusal + divergence warning), `src/runtime/microvm.ts` (`microvmGuestSessionRoot`), `src/prompt.ts` (routed through `resolveMounts` instead of a private derivation) | `test/guest-path-layout.test.ts` (a table over every shipped baseline, oracle taken from the STAGING side, plus a synthetic spawnable-but-divergent baseline — the shipped divergent one has no `spawn` block, so only the synthetic case exercises the path composition) |
|
|
26
|
+
| The session root handed to `Run.setSessionRoot` must come from the SPAWN that started the agent, in the same path space the agent reports its own paths in — host at hostloop (native agent), the VM `/sessions/<id>` at container — never re-derived by the caller. A host root measured against VM-reported paths puts nothing inside the root, so every presented file classifies `leaked: false` and `no_scratchpad_leak` (which evaluates at container and nowhere else) passes vacuously over a real copy-failure leak | `src/runtime/container.ts` + `src/runtime/hostloop.ts` (each returns the `sessionRoot` it used), `src/run/execute.ts` + `src/run/chat.ts` (pass it through), `src/run/run.ts` (`notePresentedFiles`' cwd-at-or-inside-root space check counts the batch malformed rather than grading it in the wrong space) | `test/session-root-path-space.test.ts` (per-tier geometry, the fail-closed space mismatch, live-vs-replay agreement, and the spawn-reports-its-own-root seam) |
|
|
27
|
+
| `present_files_called` must read presence from `RunResult.presentFilesCalls` (the invocation count, derived from the tool_use input's SHAPE), never from the classified `presentedFiles` list — that list drops any path it cannot resolve, which a host-path redaction policy guarantees at `hostloop`, so reading classification claimed "the tool was never called" about a run that called it AND made every such scenario unrecordable (record refuses a cassette whose verdict redaction changed) | `src/assert.ts` (the `present_files_called` branch), `src/run/run.ts` (`presentFilesCalls`, counted at the `present_files` tool_use) | `test/present-files-redaction-invariant.test.ts` (drives the real `assertRedactionVerdictPreserved` over a hostloop cassette redacted with the shipped policy; the CONTROL case pins that the un-redacted side genuinely passes, so a shared tier-gate failure can't green it) |
|
|
25
28
|
| An anonymous sub-agent dispatch's synthesized `toolUseId` must derive from the assistant message's own (stable, unique) id, never a process-lifetime counter — a counter-derived id isn't reproducible across record→replay and can collide across messages in a single run | `src/agent/session.ts` (`parseMessage`) | `test/session-parse-guards.test.ts` (the regression test asserting synthesized dispatch id uniqueness across messages, in the `parseMessage` describe block) |
|
|
26
29
|
| The pre-commit cassette gate must fail CLOSED: any `verify-cassettes` outcome that is not a proven clean `0` blocks the commit, and an exit 3 blocks unless its cause is *staleness* specifically. `ci.yml` triggers on `push: [main]` and `pull_request`, but the documented local workflow lands with `merge --ff-only` into `main` and pushes afterwards — so for the maintainer, the only person who records host-inheriting cassettes, CI is not a pre-publication gate and the hook is not one layer of two. It is the only gate, and anything it waves through reaches public history | `.githooks/pre-commit` (allowlist on `hook_status`; missing `dist/cli.js` blocks rather than warning; exit 3 split by cause via `--output-format json`) | `test/cassette-gate.test.ts` |
|
|
27
30
|
| Staging a new/updated `baselines/desktop-*.json` must re-stamp or re-record the committed `examples/replays/*.cassette.json` fixtures against it in the same commit — `latest` resolves to the newest baseline file (`src/baseline.ts`, `latestBaselineFile`), so any staged baseline can move `latest` and instantly stale every example cassette's `fingerprint.baseline`; caught 3 times (3431e09, 9eaba8d, the 0.29.0 cycle), always late (on the release PR's first CI run) because these commits sit unpushed for a while | `.githooks/pre-commit` (runs `verify-cassettes` whenever a `baselines/desktop-*.json`, a `*.cassette.json`, or any staged `.json` carrying the `"generator": "cowork-harness"` marker is staged) | `test/cassette-gate.test.ts` — drives the hook in a scratch repo through a stubbed CLI, and separately pins the stub's exit-code contract against the real binary; CI's `Cassette privacy + staleness scan` and `Cassette scan covers every TRACKED cassette` steps (`ci.yml`) are the non-local backstop |
|
package/docs/maintenance.md
CHANGED
|
@@ -23,9 +23,11 @@ forward from the previous baseline untouched)
|
|
|
23
23
|
```
|
|
24
24
|
|
|
25
25
|
> **`mountLayout.mounts[].mode` is documentary, and older baselines carry a stale `projects` row.**
|
|
26
|
-
> Nothing reads that array at run time: `resolveMounts()`
|
|
27
|
-
> take only `cwd` and `mntRoot` from the result
|
|
28
|
-
>
|
|
26
|
+
> Nothing reads that array at run time: `resolveMounts()` does not return it at all, and
|
|
27
|
+
> container/microvm/hostloop take only `cwd`, `sessionRoot` and `mntRoot` from the result — where
|
|
28
|
+
> `mntRoot` is DERIVED as `<sessionRoot>/mnt` (the only tree the stagers create) rather than read from
|
|
29
|
+
> `mountLayout.mntRoot`; a recorded value that disagrees is reported as a fidelity divergence at spawn.
|
|
30
|
+
> The one `mode === "r"` filter reads the launch plan's mounts, not the baseline's. The `projects` row reads `mode: "rw"` in every baseline before
|
|
29
31
|
> `desktop-1.25927.0`, which is **not uniformly wrong**: below `MOUNT_BARE_NAME_MIN_VERSION` (1.14271.0)
|
|
30
32
|
> `.projects/<name>` really was the connected-folder namespace, and folders are resolver-driven `rw`. From
|
|
31
33
|
> that boundary on, folders moved to `mnt/<basename>` and `.projects/<uuid>` became the project-attachment
|
package/docs/scenario.md
CHANGED
|
@@ -470,8 +470,8 @@ whether it **survives `replay`**. Both are in the key's row below, and the repla
|
|
|
470
470
|
| `all_tasks_completed: true` | every task in the run's task list reached status `completed` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); **only `true` is valid**; also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry) |
|
|
471
471
|
| `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions; also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry) |
|
|
472
472
|
| `task_status: {match, status}` | a task whose subject or id matches the `match` regex reached `status` — also fails **evidence unavailable** ("malformed") when any TaskCreate result was unparseable (corrupt task telemetry), mirroring `all_tasks_completed`/`task_count_min` |
|
|
473
|
-
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so this is meaningfully replay-checkable
|
|
474
|
-
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`presentedFiles`
|
|
473
|
+
| `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so this is meaningfully replay-checkable at the tier it evaluates on — the re-drive reproduces the classification at container, where the agent's cwd IS the session root the live lane measures from (at hostloop a re-drive has only the recorded cwd, `mnt/outputs`, so the booleans are not equivalent there); fails as **evidence unavailable** when `presentedFiles` telemetry is absent (an old run predating this key); **container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through unchanged), so there is no scratch→outputs copy for this key to check — that's not a detection gap, though: hostloop's cwd already *is* the outputs dir, so a delivered file is visible there immediately (see `user_visible_artifact`'s footgun note above). On microvm and protocol, `present_files` isn't served at all, so there's no delivery record for this key to check — cannot-verify. Use `container` for present_files-based delivery you want this key to verify, or write directly to `outputs/`; **the tool name is lane-specific** — `present_files` is the desktop-local lane's tool (the one this harness emulates) while remote Cowork delivers via the agent-native `SendUserFile`, so a skill should describe the delivery outcome rather than naming either tool ([fidelity-gaps.md](./fidelity-gaps.md), "File delivery"); this key asserts the harness-side delivery record either way; **only `true` is valid** |
|
|
474
|
+
| `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (at least one call carried a well-formed `file_path`, counted at the invocation — **not** read off the classified `presentedFiles` list, so a redaction policy that rewrites host paths cannot turn a real delivery into "never called"; a run whose every call carried an unusable path reports cannot-verify) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak; **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note above: the tool name differs on remote Cowork; **only `true` is valid** |
|
|
475
475
|
| `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD — fails as **evidence unavailable** when cost telemetry is absent (an old run predating this key). **Live lane only in spirit**: on replay this asserts the *frozen recording's* cost, not fresh spend — a cost regression is caught by a live run, not a token-free replay |
|
|
476
476
|
| `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache-read/creation tokens excluded — priced separately). Same replay caveat as `max_cost_usd`: asserts the recording, not fresh spend |
|
|
477
477
|
| `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`, sub-agent tools excluded) ≤ N — unlike the cost/token keys, this **is** meaningfully replay-checkable (the re-drive recomputes `toolCounts` deterministically from the recorded events) |
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^2.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@^2.2.0"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.2.0",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
package/schema/run-result.json
CHANGED
|
@@ -137,6 +137,10 @@
|
|
|
137
137
|
},
|
|
138
138
|
"description": "present_files promotions ({from, to, promoted, leaked}) \u2014 scratch\u2192outputs deliverable surfacing (container tier)."
|
|
139
139
|
},
|
|
140
|
+
"presentFilesCalls": {
|
|
141
|
+
"type": "integer",
|
|
142
|
+
"description": "How many present_files calls carried at least one well-formed file_path \u2014 the presence evidence present_files_called reads, counted from the tool_use input's shape so it survives redaction (presentedFiles entries are dropped when a path can't be classified, which a host-path policy guarantees at hostloop). Absent on pre-field results."
|
|
143
|
+
},
|
|
140
144
|
"mode": {
|
|
141
145
|
"type": "string",
|
|
142
146
|
"enum": ["run", "chat"],
|
|
@@ -480,7 +480,7 @@
|
|
|
480
480
|
"const": true
|
|
481
481
|
},
|
|
482
482
|
"present_files_called": {
|
|
483
|
-
"description": "at least one file was actually delivered via the present_files tool (presentedFiles is
|
|
483
|
+
"description": "at least one file was actually delivered via the present_files tool (at least one call carried a well-formed file_path). Presence is read from the INVOCATION count, not from the classified presentedFiles list, so it is unaffected by a redaction policy that rewrites host paths; a run that called the tool but whose every call carried an unusable path reports cannot-verify, never 'the tool was never called'. The presence companion to no_scratchpad_leak (which passes vacuously when nothing was presented, and stays container-only) — pair them to require a delivery AND require it not to leak; CONTAINER + HOSTLOOP TIERS — the harness serves present_files at both, mirroring real Cowork advertising the tool in both its VM and host-loop modes; every other tier is still a harness coverage gap (see docs/fidelity-gaps.md, 'File delivery'). present_files is the DESKTOP-LOCAL lane's tool name; remote Cowork uses the agent-native SendUserFile (docs/fidelity-gaps.md, 'File delivery') — this key asserts the harness-side delivery record either way; only `true` is valid",
|
|
484
484
|
"type": "boolean",
|
|
485
485
|
"const": true
|
|
486
486
|
},
|