cowork-harness 2.2.0 → 2.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +43 -17
- package/.claude/skills/cowork-harness/references/ci-recipe.md +15 -13
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +14 -12
- package/.claude/skills/cowork-harness/references/task-recipes.md +2 -2
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +32 -1
- package/CHANGELOG.md +366 -0
- package/DESIGN.md +3 -3
- package/README.md +41 -11
- package/baselines/desktop-1.37937.1.json +816 -0
- package/baselines/prompts/cowork-system-prompt-fingerprints.json +27 -0
- package/dist/assert.js +105 -3
- package/dist/cli.js +15 -4
- package/dist/hostloop/workspace-handler.js +12 -2
- package/dist/run/cassette.js +304 -78
- package/dist/run/execute.js +135 -8
- package/dist/run/trace-view.js +39 -1
- package/dist/runtime/container.js +53 -2
- package/dist/runtime/hostloop.js +52 -8
- package/dist/sync/cowork-sync.js +233 -16
- package/dist/types.js +40 -10
- package/docs/README.md +3 -0
- package/docs/cassette.md +36 -11
- package/docs/debugging.md +25 -0
- package/docs/fidelity-gaps.md +133 -6
- package/docs/maintenance.md +15 -1
- package/docs/scenario.md +66 -17
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-multiselect-gate.cassette.json +41 -46
- package/examples/replays/example-pdf-skill.cassette.json +164 -115
- package/examples/replays/hostloop-computer-links.cassette.json +62 -68
- package/examples/scenarios/example-pdf-skill.yaml +8 -1
- package/llms.txt +3 -1
- package/package.json +1 -1
- package/python/test_scenario_lint.py +39 -0
- package/schema/scenario.schema.json +29 -10
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,372 @@ All notable changes to this project are documented here. The format is based on
|
|
|
4
4
|
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/). The project uses
|
|
5
5
|
[Semantic Versioning](https://semver.org/); as of 1.0.0, a backwards-incompatible change to a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)) requires a major bump.
|
|
6
6
|
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
## [2.4.0] — 2026-08-27
|
|
10
|
+
|
|
11
|
+
### Fidelity
|
|
12
|
+
|
|
13
|
+
- **Path resolution: the shell and the file tools use DIFFERENT roots, and the harness now models that.**
|
|
14
|
+
Measured on desktop-local Cowork 2026-08-27: `mcp__workspace__bash` starts every call at the bare
|
|
15
|
+
session root (`/sessions/<id>`), while the agent process sits at the outputs dir, so a bare `Write`
|
|
16
|
+
lands in `mnt/outputs` and is user-visible immediately. The two are different path spaces — a file the
|
|
17
|
+
shell creates with a relative path is *not* where a relative `Write` puts one.
|
|
18
|
+
|
|
19
|
+
`hostloop` ran its workspace bash at `${sessionRoot}/mnt/<firstFolder ?? outputs>`, collapsing the two.
|
|
20
|
+
The replaced derivation came from the asar's `cwd: c.vmCwd` spawn argument, which is **not
|
|
21
|
+
load-bearing on the cowork path** — only the `chat` branch prepends an explicit `cd ${vmCwd}`, which
|
|
22
|
+
would be redundant if the argument worked. It reproduced a prompt claim rather than an observed
|
|
23
|
+
behaviour. Both cwds now come from one function (`hostLoopCwds`) and are pinned together in a single
|
|
24
|
+
test: a single-value assertion cannot express "shell and file tools disagree, on purpose", which is the
|
|
25
|
+
contract every previous version of this bug flattened.
|
|
26
|
+
|
|
27
|
+
**What it changes for you:** a skill that writes deliverables from a shell script using relative paths
|
|
28
|
+
looked correct at `hostloop` and delivered nothing in production. It now fails here too.
|
|
29
|
+
|
|
30
|
+
- **`container` models production's VM-loop `web_fetch` swap.** When `coworkWebFetchViaApi` resolves on
|
|
31
|
+
(true in every shipped baseline), the tier registers a workspace SDK-MCP server exposing **web_fetch
|
|
32
|
+
only**, disallows the built-in `WebFetch`, and aliases the name to `mcp__workspace__web_fetch`
|
|
33
|
+
(`VM_LOOP_TOOL_ALIASES`). Bash is deliberately untouched — "Bash is the only tool that truly diverges
|
|
34
|
+
between loops" — which is why `container` keeps the built-in shell while `hostloop` replaces both.
|
|
35
|
+
|
|
36
|
+
The three parts ship together by necessity: disallowing the built-in without the alias turns a fidelity
|
|
37
|
+
fix into a regression, because the bare name stops resolving instead of landing on the workspace tool.
|
|
38
|
+
The tool is advertised but deliberately **not** pre-approved — production's VM-loop registration passes
|
|
39
|
+
the same approval hook the host loop does, so the call is gated at `can_use_tool`. Pre-approving it
|
|
40
|
+
would make a scripted `webfetch:<domain>` answer, and `decide: deny` on it, silently inert.
|
|
41
|
+
|
|
42
|
+
Its fetch is bound to the session egress allowlist and to URL provenance, exactly as the host loop's is.
|
|
43
|
+
Both matter: this fetch runs in the harness's own process, outside the container network namespace,
|
|
44
|
+
so the sidecar proxy never sees it and only these two gates constrain it. The handler's allowlist now
|
|
45
|
+
defaults to **deny-all** rather than `["*"]`, so a caller that forgets to pass one gets nothing through
|
|
46
|
+
instead of everything.
|
|
47
|
+
|
|
48
|
+
The workspace handler now gates **dispatch** on the same set it advertises, not just `tools/list`. An
|
|
49
|
+
unadvertised tool that still executed when named was a real hole: the VM-loop registration exposes
|
|
50
|
+
web_fetch only, and a `bash` call arriving there must be refused rather than quietly exec'd into the
|
|
51
|
+
container.
|
|
52
|
+
|
|
53
|
+
**`microvm` is unchanged** and still offers the built-in `WebFetch` — `spawnMicroVm` does not receive
|
|
54
|
+
the gate. **`chat` is unchanged too**: the swap applies to `run`/`record`, so `chat --fidelity container`
|
|
55
|
+
still offers the built-in and the two surfaces differ at the same declared tier.
|
|
56
|
+
|
|
57
|
+
### Upgrade impact
|
|
58
|
+
|
|
59
|
+
- **At `fidelity: container` under `run`/`record`, `WebFetch` is no longer in the offered tool set.** An assertion naming it
|
|
60
|
+
no longer describes a callable tool. `tool_not_called: WebFetch` is the dangerous direction: it now
|
|
61
|
+
passes **vacuously** rather than failing loudly, so a scenario that was genuinely testing "this skill
|
|
62
|
+
does not fetch from the web" silently stops testing anything. Rename to
|
|
63
|
+
`tool_not_called: mcp__workspace__web_fetch`. The shipped `example-pdf-skill` scenario carried exactly
|
|
64
|
+
this defect and is fixed.
|
|
65
|
+
|
|
66
|
+
`verify-cassettes` grew a `replaced-builtin` note for this class: it reads a cassette's recorded init
|
|
67
|
+
inventory and reports built-in names the current build no longer offers at that tier. It is a **note,
|
|
68
|
+
not a finding** — the swap is gate-conditional, so a recording made with the gate off is legitimately
|
|
69
|
+
different, and an init event carrying no `tools` key is *no evidence* rather than a missing surface.
|
|
70
|
+
Neither should be told to re-record.
|
|
71
|
+
|
|
72
|
+
This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
|
|
73
|
+
no command, flag, schema or exit-code meaning changes. The tier's modelled tool inventory is a fidelity
|
|
74
|
+
property, and moving it toward production is the project's purpose — so it ships in a minor.
|
|
75
|
+
|
|
76
|
+
- **DEPRECATION: `fidelity:` becomes REQUIRED in the next major.** A scenario that omits it currently
|
|
77
|
+
defaults to `container`, which models the **VM loop** — but production runs the **host loop** (gate
|
|
78
|
+
`1143815894` is force-ON in every shipped baseline). So an omitted key silently measures the scenario
|
|
79
|
+
against a lane real users are not on, and the two lanes differ in exactly the ways that bite: where a
|
|
80
|
+
bare relative path lands, where the shell starts, and which tools are offered.
|
|
81
|
+
|
|
82
|
+
Both `run` and the skill's `scenario.py` lint now warn (`fidelity-defaulted`). The check reads the RAW
|
|
83
|
+
document, because Zod's `.default()` makes an omitted key indistinguishable from a deliberate
|
|
84
|
+
`fidelity: container` — and those two deserve different treatment. Naming the tier explicitly silences
|
|
85
|
+
it; `fidelity: container` remains a valid, non-warning choice.
|
|
86
|
+
|
|
87
|
+
### Fixed
|
|
88
|
+
|
|
89
|
+
- **The semantic judge was told which authored files were never delivered.** The authored-file capture
|
|
90
|
+
deliberately includes the scratchpad, but production discards anything outside `mnt/` — "never reaches
|
|
91
|
+
the user or your file tools". Unlabelled, a rubric like "the report was written" graded TRUE on a file
|
|
92
|
+
the user never receives: a false green inside the one evaluator that reads free-form prose and cannot
|
|
93
|
+
infer the convention. Scratch files are now tagged `— SCRATCH, NOT delivered to the user`, with a note
|
|
94
|
+
explaining they are evidence of what the run DID and not that anything was delivered.
|
|
95
|
+
|
|
96
|
+
- **`sync` no longer refuses to write on the `subagentPromptServerOverride` gate.** Gate-ON only enables
|
|
97
|
+
the lookup; the payload that would actually override is delivered **per session by the server** and
|
|
98
|
+
appears in neither the asar, the fcache nor `config.json`. So gate state alone cannot separate
|
|
99
|
+
"override active" from "gate on, no payload, fallback still correct" — the guard blocked forever
|
|
100
|
+
rather than tripping, because it could never clear itself from its own inputs. It is now a
|
|
101
|
+
non-blocking note that says what it can and cannot know.
|
|
102
|
+
|
|
103
|
+
Settled by a live sub-agent probe instead: a real sub-agent's environment section matched the committed
|
|
104
|
+
paraphrase on all four load-bearing claims, so no override was reaching that account. That is evidence,
|
|
105
|
+
not proof — one account, one session, and a server rule can be segment-targeted. The note says so, and
|
|
106
|
+
says to re-probe if the sub-agent append matters to what you are shipping.
|
|
107
|
+
|
|
108
|
+
### Documentation
|
|
109
|
+
|
|
110
|
+
- **[docs/scenario.md](./docs/scenario.md) now states where a relative path actually lands**, as a
|
|
111
|
+
measured table per tier, replacing a sentence that was simply wrong for the file tools. The guidance
|
|
112
|
+
that follows it is lane-dependent, because the correct answer inverts between lanes: on the desktop
|
|
113
|
+
host loop the file tools are already rooted at `outputs/`, so writing `outputs/x.md` doubles to
|
|
114
|
+
`outputs/outputs/x.md` and the user never sees it — a **bare filename** is right there. At
|
|
115
|
+
`container`/`microvm` a bare name lands in the scratchpad instead. Addressing a connected folder by
|
|
116
|
+
name never reaches it on either lane: it builds a same-named decoy inside `outputs`, reports success,
|
|
117
|
+
and gives no signal.
|
|
118
|
+
|
|
119
|
+
- **[docs/fidelity-gaps.md](./docs/fidelity-gaps.md)** gained the path-resolution split, and its
|
|
120
|
+
"VM tiers have no workspace tool aliases" entry is updated — closed at `container`, still open at
|
|
121
|
+
`microvm`.
|
|
122
|
+
|
|
123
|
+
- The skill's undelivered-deliverables guidance no longer hardcodes the literal prefix `outputs/`, which
|
|
124
|
+
was wrong on the lane production actually runs.
|
|
125
|
+
|
|
126
|
+
- **The product's vocabulary is mapped to this project's**, because one directory had four names and none
|
|
127
|
+
of them was the one Cowork's UI shows. "Working folder" is the user-visible roots (`outputs/` plus each
|
|
128
|
+
connected folder); "Scratchpad" is everything outside `mnt/`; `{{workspaceFolder}}` is a prompt token
|
|
129
|
+
that renders to the *first* user-visible root. Also recorded: Cowork's **Scratchpad panel is an activity
|
|
130
|
+
log, not a location listing** — it lists files as "wrote to" wherever they landed, so a file appearing
|
|
131
|
+
there is not evidence it was undelivered.
|
|
132
|
+
|
|
133
|
+
- **`Write`'s tool result echoes the raw path it was given and never absolutizes** (read from the agent
|
|
134
|
+
binary). Nothing in the harness parses a path out of a `Write` result; this is recorded so nothing
|
|
135
|
+
starts, since such an assertion would be reading something production does not emit. Cowork's own
|
|
136
|
+
chat-surface prompt claims the opposite, so the product's description of its own tool is wrong here.
|
|
137
|
+
|
|
138
|
+
|
|
139
|
+
## [2.3.0] — 2026-08-26
|
|
140
|
+
|
|
141
|
+
### Parity
|
|
142
|
+
|
|
143
|
+
- **Baseline `desktop-1.37937.1` (agent `2.1.246`).** `sync` had been refusing to write on three
|
|
144
|
+
`spawn.env` deltas; all three are now classified.
|
|
145
|
+
|
|
146
|
+
**Two new pinned keys.** The Cowork spawn sets `CLAUDE_CODE_PROMPT_CACHE_TTL="1h"` and
|
|
147
|
+
`CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL="5m"` unconditionally — no gate, no session or deployment
|
|
148
|
+
branch — so every first-party session receives them. They are **additive**:
|
|
149
|
+
`ENABLE_PROMPT_CACHING_1H="1"` is still set alongside. They read zero times in agent `2.1.241` and six
|
|
150
|
+
times each in `2.1.246`, so the contract went live one agent release after Desktop began setting it.
|
|
151
|
+
|
|
152
|
+
**`MCP_TOOL_TIMEOUT` is now classified per SITE, not per key.** Its first-party construction is
|
|
153
|
+
unchanged (still resolving to `180000`), but the third-party-only branch gained a second, settings-
|
|
154
|
+
conditional construction of the same key whose value expression the const resolver cannot reach.
|
|
155
|
+
Allowlisting the key — the obvious fix — would have been a silent contract loss: the allowlist is
|
|
156
|
+
checked *before* the pin list, so the key would have vanished from the generated env entirely, and it
|
|
157
|
+
is not a `REQUIRED_SPAWN_KEYS` member, so nothing would have hard-failed. Instead the 3p-only branch is
|
|
158
|
+
located by content and its inner keys are classified by name without resolving their values, which is
|
|
159
|
+
what the branch already meant. A brand-new key there still hard-fails.
|
|
160
|
+
|
|
161
|
+
**`CLAUDE_PREVIEW_CLASSIFIER_FLOOR` is now inert on the shipping agent** — recorded, not changed. Agent
|
|
162
|
+
`2.1.246` renamed the flag it reads to `CLAUDE_CHROME_CLASSIFIER_FLOOR` (and its consumer field
|
|
163
|
+
`previewClassifierFloorEnabled` → `chromeClassifierFloorEnabled`) while Desktop still sets only the old
|
|
164
|
+
name, so the classifier floor now falls through to its GrowthBook default. The key stays pinned: the
|
|
165
|
+
baseline records what the spawn constructs, and a Desktop-side rename must surface as a diff line
|
|
166
|
+
rather than as silence. Nothing in the harness reads it behaviourally. Together with the cache-TTL keys
|
|
167
|
+
above this is the same lesson pointing both ways — Desktop and the agent version the spawn env
|
|
168
|
+
independently, so "Desktop sets X" and "the agent reads X" are separately-dated claims.
|
|
169
|
+
|
|
170
|
+
Also found in the same pass and recorded in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md), with no
|
|
171
|
+
harness surface: the deliberately-unmodeled remote device-tool family gained `device_fs` and
|
|
172
|
+
`device_request_delete_permission` plus a folder-access announce mode — the Desktop half of the six
|
|
173
|
+
`cowork_*` risk categories that had appeared in agent `2.1.237`'s auto-mode rubric.
|
|
174
|
+
|
|
175
|
+
Also verified unchanged against the retained previous asar: the Cowork system prompt (the retained raw
|
|
176
|
+
files for 1.32885.1 through 1.37937.1 are byte-identical on disk), both sub-agent appends, `tools[]`,
|
|
177
|
+
the `canUseTool` chain, the mount modes, the egress contract, and the VM rootfs image.
|
|
178
|
+
|
|
179
|
+
- **All three committed example cassettes re-recorded against this baseline** — `example-pdf-skill`
|
|
180
|
+
(`container`), `example-multiselect-gate` (`protocol`) and `hostloop-computer-links` (`hostloop`).
|
|
181
|
+
The container one shows no behavioural change (transcript wording only). `verify-cassettes` is clean,
|
|
182
|
+
and the recorded MCP inventory is the Cowork lane's own servers (`cowork`, `plugins`, `skills`,
|
|
183
|
+
`workspace`) — no host inventory reached the fixtures, checked independently of the built-in scan
|
|
184
|
+
because the 2026-08-04 leak hid from a `mcp__` grep and surfaced only through NAME fields.
|
|
185
|
+
|
|
186
|
+
- **Live end-to-end pass re-run against this baseline.** `npm run test:live` — **4 suites, 19 assertions,
|
|
187
|
+
19 green, 0 skipped** — across `protocol`, `container` and `hostloop`, on agent `2.1.246`. Every
|
|
188
|
+
`describe` in that lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated
|
|
189
|
+
case reports as skipped rather than passing vacuously; zero skips means the whole population executed.
|
|
190
|
+
`DESIGN.md`'s scope note is re-stamped accordingly and now records that no baseline is unverified for
|
|
191
|
+
want of a live run. Not claimed: the `boundary-check` sandbox proof and the example-scenario suite were
|
|
192
|
+
part of the previous stamp and were not run here.
|
|
193
|
+
|
|
194
|
+
### Added
|
|
195
|
+
|
|
196
|
+
- **`question_context: {when_question?, matches}` — assert what a gate actually put in front of the user.**
|
|
197
|
+
A regex tested against a gate's founder-visible payload FIELD BY FIELD — the question label, every option
|
|
198
|
+
**label**, and every option **description**, each as a separate string. Matching per field rather than over
|
|
199
|
+
one joined blob is deliberate and load-bearing: a pattern cannot straddle two fields, so
|
|
200
|
+
`invoicing[\s\S]*Audit logging` will not match by stitching one option's description to the next option's
|
|
201
|
+
label — a "sentence" nobody was shown. The neighbouring transcript keys' docs teach `[\s\S]` for spanning
|
|
202
|
+
turns, so that is the habit an author brings here. `matches` is required NON-EMPTY: an empty pattern
|
|
203
|
+
compiles to `//i` and would green any run that fired a gate at all. `question_asked` matches question text only and `question_options`
|
|
204
|
+
compares labels only, so a sentence the model delivered inside an option's `description` was invisible to
|
|
205
|
+
every assertion key — a false-negative generator for any skill that puts context there, which the tool's
|
|
206
|
+
own schema invites. Measured on a consumer's paid run: a producer-authored sentence arrived verbatim in
|
|
207
|
+
the question, reworded inside it, and relocated into the proceed option's `description` across three runs
|
|
208
|
+
of one scenario; the third redded a lane on a run where the founder had in fact been told.
|
|
209
|
+
|
|
210
|
+
Evidence is the **ask-time** `AskUserQuestion` payload, never a `tool_result` — a skill's producer
|
|
211
|
+
typically also writes the same sentence into its own gate-state file, so `tool_result_matches` on that
|
|
212
|
+
phrase grades true whether or not the model ever surfaced it. Unlike `question_options`, omitting
|
|
213
|
+
`when_question` on a multi-gate run is **not** ambiguous: this key asks whether the text was shown at all.
|
|
214
|
+
Zero gates recorded fails; unreadable gate evidence fails evidence-unavailable, never vacuously.
|
|
215
|
+
|
|
216
|
+
### Changed
|
|
217
|
+
|
|
218
|
+
- **`trace --view questions` now renders each gate's offered options — labels AND `description`s** — under
|
|
219
|
+
an `offered:` block, sub-question-labelled on a bundled gate. It previously printed the question label
|
|
220
|
+
alone, so the option `description` a skill routinely puts the deciding sentence in was reachable only by
|
|
221
|
+
hand-reading `events.jsonl`; a reader who found nothing in the view could reasonably conclude the text was
|
|
222
|
+
never delivered. The payload was always recorded — this was purely a rendering gap, and the same run dir
|
|
223
|
+
answers the question either way. The row (`--output-format json`) gains `subQuestions[]` carrying the
|
|
224
|
+
untruncated ask-time options; the text view caps each description at 240 chars, so nothing is lost, only
|
|
225
|
+
wrapped. This pairs with `question_context`: the view is how you *find* the text, that key is how you
|
|
226
|
+
*gate* on it.
|
|
227
|
+
|
|
228
|
+
- **The `tool_use` blindness of `transcript_contains`/`_not_contains`/`_matches`/`_not_matches` and
|
|
229
|
+
`computer_links_resolve`/`_if_present` is now documented and guarded.** `semantic_matches` has carried a
|
|
230
|
+
⚠️ spelling out that its corpus excludes every `tool_use` — "a rubric claim about whether a tool was
|
|
231
|
+
called is unassertable" — while the six keys with the identical blindness said only "the assistant
|
|
232
|
+
transcript". A consumer wrote a `transcript_matches` against text living in a gate question; it could not
|
|
233
|
+
have matched at any phrasing, and the recording fail-closed after the spend. The caveat is now a property
|
|
234
|
+
of an enumerable set (`TOOL_USE_BLIND_KEYS`) enforced across every surface that documents a key — the
|
|
235
|
+
docs tables, the zod `.describe()` behind `assertions --list` and the generated JSON schema, and the
|
|
236
|
+
packaged skill reference — so a newly-added blind key cannot ship without it.
|
|
237
|
+
- **`question_asked`/`question_options`/`question_context` now warn that they match model-authored text.**
|
|
238
|
+
Gate question text and option labels are composed by the model and reworded run to run. The `choose:`
|
|
239
|
+
side already documented this (stable leading anchor, 1-based index); the assert side documented it
|
|
240
|
+
nowhere. Guarded by `MODEL_AUTHORED_TEXT_KEYS`.
|
|
241
|
+
- **A bad regex in a NESTED assertion field is now caught at load, not after the paid spawn.** The pre-compile
|
|
242
|
+
pass reached only top-level string keys, so every regex one level down — `artifact_text.matches`,
|
|
243
|
+
`artifact_text.not_matches`, `path_denied.path_matches`, `skill_tool_used.skill`/`.tool`,
|
|
244
|
+
`subagent_dispatch_healthy.type`, `subagent_output_contains.match`, `task_status.match`,
|
|
245
|
+
`question_options.when_question`, and the new `question_context.*` — was first compiled inside the
|
|
246
|
+
evaluator. All eleven are now validated at load, and `test/nested-regex-leaves.test.ts` reads `assert.ts`
|
|
247
|
+
and fails if the evaluator compiles a nested leaf the load-time table does not carry, so the gap cannot
|
|
248
|
+
reopen silently. `lint`'s double-quoted-regex warning also now covers nested `matches:` leaves.
|
|
249
|
+
- **`diff` with a single positional now names the missing operand** instead of printing bare usage. There is
|
|
250
|
+
deliberately no one-argument form: `diff` is polymorphic over baselines, run dirs and cassettes, so it
|
|
251
|
+
would need type dispatch plus a defined source for "the committed version".
|
|
252
|
+
- **The two cost keys are cross-referenced.** `RunResult.cost.usd` is one invocation's SDK
|
|
253
|
+
`total_cost_usd`; the critique report's `costUsd.totalUsd` aggregates the task turn, the reflection turn
|
|
254
|
+
and both evaluator passes. Reading the wrong one returns `undefined` rather than erroring, which reads as
|
|
255
|
+
"no cost recorded". Kept as two shapes on purpose — collapsing them would destroy the per-phase split.
|
|
256
|
+
|
|
257
|
+
### Fixed
|
|
258
|
+
- **Two spawn-contract sentinels were weaker than their names implied; both now bite.**
|
|
259
|
+
|
|
260
|
+
`checkSpawnContractFacts` pinned the `allowedTools[]` built-in head and the built-in→`mcp__` boundary
|
|
261
|
+
but nothing between the boundary and the closing bracket — so `mcp__plugins__search_connectors` was
|
|
262
|
+
added to the array and both checks stayed green. The `mcp__` membership is now pinned as a set, and the
|
|
263
|
+
flag names the added and removed entries instead of reporting that something "moved". (That tool is
|
|
264
|
+
declared only on the third-party deployment, so the first-party inventory the harness serves is
|
|
265
|
+
unaffected — but the addition should not have been invisible.)
|
|
266
|
+
|
|
267
|
+
`checkMountModeFacts` asserted each read-only mount with a single `regex.test`, while `uploads`,
|
|
268
|
+
`.claude/skills` and `.projects/<uuid>` are each built at **two** sites — the VM-loop mount-set builder
|
|
269
|
+
and host-loop `computeBashMounts`. Either site satisfied the check, so a one-lane `ro`→`rw` flip — a
|
|
270
|
+
containment change on exactly one execution tier — passed green. It now compares the site count to the
|
|
271
|
+
count carrying `mode:"ro"` and names the lane count in the flag. Its fifth fact — the delete-deny
|
|
272
|
+
resolver `?"rwd":"rw"` — had the same shape and the same two-lane reality (it went 1 site to 2 in this
|
|
273
|
+
release) and now guards a floor on that count: a lane losing the resolver flags, a lane gaining one
|
|
274
|
+
does not.
|
|
275
|
+
|
|
276
|
+
- **All eight asar sentinels now carry a committed mutation case.** They are green on the previous
|
|
277
|
+
release's asar too, so a green proves nothing on its own unless the checker is known to bite; five of
|
|
278
|
+
them (`checkCodeTripwires`, `checkWebFetchFacts`, `checkEgressContractFacts`, `checkSyspromptMapFacts`,
|
|
279
|
+
`checkNormalizationSanity`) had no such case. Each new one changes an **inner** character of its
|
|
280
|
+
anchor, because a suffix rename still satisfies a substring regex — a mutation that cannot fail is not
|
|
281
|
+
evidence — and asserts that the mutation actually applied.
|
|
282
|
+
|
|
283
|
+
- **`record --dry-run <dir/>` no longer announces a WARN as a refusal.** The batch arm labelled every
|
|
284
|
+
advisory note `⚠ would-refuse (advisory)`, but `cassettePortabilityPreflight` can only ever return
|
|
285
|
+
`ok`/`warn` — it has no refuse path at all — and `hostInventoryPreflight` returns `warn` whenever the
|
|
286
|
+
target cassette already exists, which is every re-record corpus sitting at the default path. So the
|
|
287
|
+
preview told operators the real `record` would refuse runs it would in fact accept, on the one arm whose
|
|
288
|
+
design principle is that a guess must not gate. The label now follows the verdict kind
|
|
289
|
+
(`⚠ would-warn (advisory)`), the "ADVISORY, not this run's verdict" footer is unchanged, and exit codes
|
|
290
|
+
were never affected either way.
|
|
291
|
+
|
|
292
|
+
- **The 3p-branch rule refuses to blank W1.** W1 is the window every modeled first-party key is derived
|
|
293
|
+
from, so a branch marker appearing there hard-fails instead of blanking; deleting real pinned keys from
|
|
294
|
+
the derived env with nothing failing is the worse of the two outcomes.
|
|
295
|
+
|
|
296
|
+
- **`record --dry-run` now runs every pre-spend refusal the real record runs, and one refusal moved from
|
|
297
|
+
after the paid run to before it.** The rehearsal re-implemented the checks by hand, so it drifted:
|
|
298
|
+
`hostInventoryPreflight` shipped 2026-08-04 and the commit three days later titled *"make `--dry-run`
|
|
299
|
+
refuse what the real record refuses"* swept in the two checks returning `string | undefined` and missed
|
|
300
|
+
the one returning a `{kind}` verdict — for 19 days an operator could not discover that refusal without
|
|
301
|
+
spending. Separately, the slug-collision refusal (*"refusing to overwrite … it belongs to scenario X"*)
|
|
302
|
+
sat **after** `executeScenario`: you paid for the run and were then refused, though it is a pure function
|
|
303
|
+
of path + `--force` + the existing cassette's name. Both now run in one shared pre-spend block.
|
|
304
|
+
|
|
305
|
+
Refusals are also uniform now. `promptPolicyRejection` threw while `hostInventoryPreflight` called
|
|
306
|
+
`fail()`, which `process.exit`s — and the dir-batch loop catches a throw per item but cannot catch an
|
|
307
|
+
exit, so a host-inventory refusal mid-batch abandoned concurrent runs already paid for.
|
|
308
|
+
|
|
309
|
+
Under `--output-format json` a directory batch carries those advisory verdicts in a `notes[]` array, kept
|
|
310
|
+
separate from `refusals[]` so automation cannot read a guess as a binding verdict.
|
|
311
|
+
|
|
312
|
+
**On a directory, the path-dependent verdicts are ADVISORY (`⚠ would-refuse`) and do not affect the exit
|
|
313
|
+
code**, because a directory target takes no `--out`: the preview would have to guess the destination, and
|
|
314
|
+
a guess must not gate. Measured on a real consumer, gating on that guess would have refused 26 of 27
|
|
315
|
+
scenarios the real record accepts. Dry-run a single scenario file with the real flags for a binding
|
|
316
|
+
answer. `--quiet` suppresses the notes; it never suppresses a refusal.
|
|
317
|
+
|
|
318
|
+
Also new in the batch preview: the duplicate-cassette-path refusal the real batch already had.
|
|
319
|
+
|
|
320
|
+
### Upgrade impact
|
|
321
|
+
|
|
322
|
+
- **`--allow-host-inventory-fixture` no longer waives a MEASURED host-inventory finding.** It was one flag
|
|
323
|
+
doing two jobs: bypassing the pre-flight refusal (a precondition the operator cannot check — "use only
|
|
324
|
+
when the session has no personal MCP servers or plugins") *and* downgrading the write-time scan's refusal
|
|
325
|
+
to a warning. So an operator who passed it to get past the undecidable precondition also switched off the
|
|
326
|
+
scan that would have caught a real leak. It is now the pre-flight bypass only; the finished recording is
|
|
327
|
+
still scanned, and a finding still refuses the write and quarantines the recording. Writing a flagged
|
|
328
|
+
recording needs the new, narrower `--allow-host-inventory-findings`. **A batch recorder that passes the
|
|
329
|
+
old flag on a new-fixture path will now abort on a genuine finding where it previously warned and wrote.**
|
|
330
|
+
|
|
331
|
+
A `--dry-run` that reports what *would* be captured was considered and declined: the inventory does not
|
|
332
|
+
exist until the agent has run, so a preview would either re-implement the scanner against a hypothetical
|
|
333
|
+
(a second oracle free to disagree with the real one) or require the spend it was meant to avoid. Reasoning
|
|
334
|
+
is in [docs/cassette.md](./docs/cassette.md).
|
|
335
|
+
|
|
336
|
+
This does not break a covered surface ([SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)):
|
|
337
|
+
no command or flag is removed and no exit-code meaning changes — one flag's consent narrows, and the
|
|
338
|
+
capability it shed is reachable through the new one. So it ships in a minor.
|
|
339
|
+
|
|
340
|
+
### Documentation
|
|
341
|
+
|
|
342
|
+
Five gaps found by asking a consumer which harness properties actually changed an outcome on a real
|
|
343
|
+
working day, then checking whether the docs said so. Four of the five were documented only as features,
|
|
344
|
+
never as the failure they prevent — the sentence a reader needs to recognise their own situation.
|
|
345
|
+
|
|
346
|
+
- **The blocking gate is now stated as a blocker.** `AskUserQuestion` *blocks*: it is a question to a
|
|
347
|
+
human and `claude -p` has no human, so a gated skill under a plain CLI run stalls or never reaches the
|
|
348
|
+
code behind the gate. That made half the skill untestable, not merely awkward to test — the README had
|
|
349
|
+
only "untestable headless unless something answers it", buried as the last of two afterthought bullets.
|
|
350
|
+
- **"Assert on the run, not the output"** — a new README section naming the class of claim the harness
|
|
351
|
+
exists for (`subagent_tool_absent`, `dispatch_count_max`, `no_delete_in_outputs`, `subagent_file_write`)
|
|
352
|
+
and why no output diff can reach it: a correct run and one that quietly handed a restricted sub-agent
|
|
353
|
+
shell access produce byte-identical files. `subagent_tool_absent` did not appear in the README at all.
|
|
354
|
+
- **The raw-`events.jsonl` escape hatch is documented**, with verified `jq` recipes, in
|
|
355
|
+
[docs/debugging.md](./docs/debugging.md). Every documented route into the event log went through
|
|
356
|
+
`trace`, whose views are a digest — so the run dir's *evidence* surface is wider than any view's
|
|
357
|
+
*observation* surface, and wider still than the assertion catalog. Concluding "it never happened" from
|
|
358
|
+
a view that doesn't render the field is a false negative, and the doc now says so and shows the read.
|
|
359
|
+
- **Sub-agent delivery has a route.** The tier-qualified outputs contract — the reason a hand-off path
|
|
360
|
+
that works on one loop lands in sandbox scratch on the other — was correctly documented in
|
|
361
|
+
[docs/subagents.md](./docs/subagents.md) but filed under "read on demand", reachable only by someone
|
|
362
|
+
who already knew the answer. It now has a "Common tasks" row keyed to the symptom.
|
|
363
|
+
- **Replay's cost claim carries a number.** "Zero spend" is the price; the wall-clock is what makes an
|
|
364
|
+
always-on per-PR gate obviously affordable, and no doc stated it (well under a second per cassette).
|
|
365
|
+
|
|
366
|
+
- **The project's provenance is stated on every surface the framing travels to.** `README.md` (which
|
|
367
|
+
renders on the npm page), `llms.txt`, both `.claude-plugin/marketplace.json` descriptions and `SKILL.md`
|
|
368
|
+
now say the same thing: an independent project, not affiliated with, endorsed by, or supported by
|
|
369
|
+
Anthropic; it bundles no Anthropic code; it is not Cowork. `SKILL.md` additionally tells the agent to say
|
|
370
|
+
so when a user asks what it is. Nothing enforces the four staying in step, so an edit can still drop it
|
|
371
|
+
from one of them unnoticed.
|
|
372
|
+
|
|
7
373
|
## [2.2.0] — 2026-08-25
|
|
8
374
|
|
|
9
375
|
### Upgrade impact
|
package/DESIGN.md
CHANGED
|
@@ -47,7 +47,7 @@ in how a file reaches the user, which is what changes skill behaviour: see
|
|
|
47
47
|
[docs/scenario.md](./docs/scenario.md)'s `lane:` key for holding a run to either contract. The local lane:
|
|
48
48
|
|
|
49
49
|
- VM bundle: `~/Library/Application Support/Claude/vm_bundles/claudevm.bundle/` (`rootfs.img`, `sessiondata.img`, `efivars.fd`, `machineIdentifier`, `gvisorMacAddress`, `vmIP`); a warm pool at `vm_bundles/warm/<sha>/`.
|
|
50
|
-
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.
|
|
50
|
+
- In-VM agent: `~/Library/Application Support/Claude/claude-code-vm/<ver>/claude` (currently **2.1.246**, per `baselines/desktop-1.37937.1.json`, an **ELF aarch64** binary), spawned by the host **in cowork mode via the `CLAUDE_CODE_IS_COWORK=1` env var** — *not* a `--cowork` flag (that flag is plugin-scope and the staged agent rejects it; see the control-protocol note below). Each baseline records this ELF's `sha256` (`agentBinary.sha256`/`shaProvenance`), and the resolver integrity-checks the binary it's about to run against that hash by default (opt out `COWORK_HARNESS_VERIFY_AGENT_SHA=0`), so "the same pinned agent" is enforced, not just asserted. Old versions are re-downloadable + verifiable from the official release channel (see `docs/maintenance.md`).
|
|
51
51
|
- Network: `vm_network_mode: "gvisor"`, egress through a userspace netstack with a **compiled domain allowlist**; off-list partners rejected (`partner rejected: entry not on compiled allowlist`).
|
|
52
52
|
- Control plane: Electron renderer→main typed IPC on channels named `$eipc_message$_<per-build-UUID>_$_claude.web_$_<Class>_$_<method>`, every handler validating `event.senderFrame.url` against a trusted-origin allowlist. The session manager is `LocalAgentModeSessions` (80 methods: `start`, `sendMessage`, `setDraftSessionFolders`, `onToolPermissionRequest`, `respondToToolPermission`, `getTranscript`, `onEvent`, …), bridged to the renderer as `window.cowork`.
|
|
53
53
|
|
|
@@ -174,9 +174,9 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
174
174
|
|
|
175
175
|
> **Machine-readable form:** the five shapes below are schema'd as `schema/protocol.v1.json`, with a golden vector pack at `fixtures/protocol/v1/` — see [docs/protocol.md](./docs/protocol.md) for scope, versioning, and how to conformance-test against them.
|
|
176
176
|
|
|
177
|
-
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.
|
|
177
|
+
### Control protocol — VERIFIED end-to-end against the live host CLI (macOS. The staged in-VM agent that L1/L2 run is **2.1.246**, the native host app that `hostloop` runs is **2.1.246**, baseline **`desktop-1.37937.1`** — a fresh live end-to-end pass across the `protocol`, `container`, and `hostloop` tiers was run against this baseline on 2026-08-26, superseding the prior `1.32885.1` pin. Scope and caveats of that pass are in the note directly below; read it before citing this heading.) The subsequent `desktop-1.20186.1` baseline is a patch-only Desktop release (egress allowlist, spawn config, and the Cowork system-prompt fingerprint unchanged from 1.20186.0; the staged VM ELF re-synced 2.1.202 → 2.1.205) — the live pass is deliberately not restamped onto it.
|
|
178
178
|
|
|
179
|
-
> **Scope of that claim, stated plainly.** `2026-08-
|
|
179
|
+
> **Scope of that claim, stated plainly.** `2026-08-26 / desktop-1.37937.1` is the baseline carrying the latest **full live end-to-end pass**, and it is also the newest committed baseline — **no baselines have shipped since**, so nothing in `baselines/` is currently unverified for want of a live run. The pass ran against agent **2.1.246** (the staged VM ELF and the native `.app` are both at that version) and covered all three tiers: `protocol` (the matrix runner's own live e2e plus its unanswered-gate regression pin), `container` (the spawn contract against the staged binary, resume continuity across the container boundary, and the outputs-delete guard), and `hostloop` (the sub-agent relative-`Write` acceptance probe, resume continuity on the native binary, critique at the unpinned tier, uploads readability, sub-agent WebSearch capture, and the discovery-server declaration check). It was one invocation of `npm run test:live` — **4 suites, 19 assertions, 19 green / 0 skipped**. **Nothing was gated out**, which is the part worth stating: every `describe` in this lane is `skipIf`-gated on Docker, the staged-binary version and the token, so a gated case reports as SKIPPED rather than passing vacuously — vitest reported zero skips, and the 19 that ran are exactly the 19 the suite enumerates. **Scope-out, so this is not read as more than it is.** (a) The assertion population is 19 here against the 24 recorded for the 1.32885.1 pass; cases have been retired and consolidated since (the `live-outputs-delete` whole-line-`#`-comment case was retired in 1.25.0 after its pinned command stopped being executed by the model), so the two counts are not comparable and the drop is not coverage lost in this pass. (b) The `boundary-check` sandbox proof and the example-scenario suite were part of the 1.32885.1 stamp and were **NOT** run here — this paragraph claims `npm run test:live` only. (c) A live pass verifies observed behaviour, not the whole spawn contract by construction. **These suites are model-dependent, so a single red is evidence of model variance until a re-run says otherwise, not of a regression** — which is why the gated cases skip loudly rather than fail. Separately and not a live matter: the three committed example cassettes (`example-pdf-skill` at `container`, `example-multiselect-gate` at `protocol`, `hostloop-computer-links` at `hostloop`) were re-recorded against this baseline in the same change, so nothing in `examples/replays/` is stale against it. Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
180
180
|
|
|
181
181
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
182
182
|
|
package/README.md
CHANGED
|
@@ -11,6 +11,11 @@
|
|
|
11
11
|
[](https://github.com/yaniv-golan/skill-creator-plus)
|
|
12
12
|
[](https://agentskills.io)
|
|
13
13
|
|
|
14
|
+
> **Unofficial.** An independent project, not affiliated with, endorsed by, or supported by Anthropic.
|
|
15
|
+
> "Claude" and "Claude Cowork" are Anthropic's. This harness emulates Cowork's *observable runtime
|
|
16
|
+
> contract* and drives Anthropic's own agent binary from your local Claude Desktop install — it
|
|
17
|
+
> bundles no Anthropic code, and it is not Cowork.
|
|
18
|
+
|
|
14
19
|
Scriptable, CI-friendly test harness that reproduces **Claude Cowork's observable runtime contract** closely enough to test the skills you write — across many scenarios, headless, in CI — without the (locked) Desktop app. It reproduces not just Cowork's *behavior* but its *limitations*: sealed filesystem, default-deny egress, MCP-only cross-boundary — so a green test has cleared the constraints that break skills in Cowork. That is a far stronger signal than a bare `claude -p` run, and it is not a guarantee: this is an emulator of the contract, and the deliberate divergences are catalogued in [docs/fidelity-gaps.md](./docs/fidelity-gaps.md).
|
|
15
20
|
|
|
16
21
|
And because every run is recorded, you get the thing a transcript can't give you: **evidence of what the agent actually did, not what it said it did** — which skill was invoked (if any), which files a sub-agent really read, which hosts it reached, which options a person was really shown. See [Why not just `claude -p` or the Agent SDK?](#why-not-just-claude--p-or-the-agent-sdk).
|
|
@@ -83,8 +88,8 @@ Both are excellent, and if you're building an *agent* you should use them. This
|
|
|
83
88
|
|
|
84
89
|
Two more that matter in practice:
|
|
85
90
|
|
|
86
|
-
- **
|
|
87
|
-
- **
|
|
91
|
+
- **A blocking gate you can actually answer.** `AskUserQuestion` *blocks*: it is a question to a human, and `claude -p` has no human. A gated skill under a plain CLI run either stalls at the gate or never reaches the code behind it — so the half of your skill that lives past the first question is untestable, not merely awkward to test. The harness answers over the same `can_use_tool` control protocol Desktop uses, from your scenario's scripted `answers:` — or from a live decider when you're still discovering what it asks — so the gate genuinely fires and is answered deterministically. `on_unanswered:` decides what an *un*scripted gate does (fail loud by default). See [scenario.md](./docs/scenario.md), [decider-dir.md](./docs/decider-dir.md).
|
|
92
|
+
- **Token-free CI.** Record a run once, commit the cassette, and every PR re-runs it deterministically at **zero spend** — assertions, tool stream, gate answers and all. A cassette replays in well under a second with no token, no Docker and no model call (the three shipped examples replay in ~0.6s total), which is what makes an always-on per-PR gate affordable. See [cassette.md](./docs/cassette.md).
|
|
88
93
|
|
|
89
94
|
**What it doesn't do.** It runs and records; it does not design your experiment. Comparing "with skill" against "without skill" credibly — scrubbing tells, shuffling, judging blind, unblinding after grading — is still yours to build; the harness contributes the run execution and the control arm (`--ablate-skill`, one arm per invocation). And it emulates the *contract*, not the Desktop runtime: see [Limitations](#limitations) and [fidelity-gaps.md](./docs/fidelity-gaps.md) for what it deliberately does not reproduce.
|
|
90
95
|
|
|
@@ -115,7 +120,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
115
120
|
|
|
116
121
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
117
122
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
118
|
-
> From a global install (`npm i -g "cowork-harness@^2.
|
|
123
|
+
> From a global install (`npm i -g "cowork-harness@^2.4.0"`), point at the package root instead:
|
|
119
124
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
120
125
|
> (or copy the cassette into your own project and pass that path).
|
|
121
126
|
|
|
@@ -125,7 +130,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
125
130
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
126
131
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
127
132
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
128
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.
|
|
133
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.4.0"`.
|
|
129
134
|
|
|
130
135
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
131
136
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -150,7 +155,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
150
155
|
claude plugin install cowork-harness@cowork-harness
|
|
151
156
|
```
|
|
152
157
|
|
|
153
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.
|
|
158
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.4.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
154
159
|
|
|
155
160
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
156
161
|
|
|
@@ -175,7 +180,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
175
180
|
ones that still need a source checkout. (The marketplace
|
|
176
181
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
177
182
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
178
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.
|
|
183
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.4.0"` — see
|
|
179
184
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
180
185
|
|
|
181
186
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -293,7 +298,7 @@ cowork-harness lint examples/scenarios/*.yaml --strict --min-severity WARN
|
|
|
293
298
|
> | When | Assertions |
|
|
294
299
|
> |---|---|
|
|
295
300
|
> | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
|
|
296
|
-
> | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
|
|
301
|
+
> | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
|
|
297
302
|
> | Only if the cassette carries an `artifacts` manifest | `file_exists`, `artifact_text`, `user_visible_artifact`, `artifact_json`, `computer_links_resolve`, `computer_links_resolve_if_present`, `no_unexpected_files`, `input_unmodified` |
|
|
298
303
|
> | Always skipped (live-only) | `file_absent`, `egress_*`, `expect_denied`, `no_delete_in_outputs`, `no_delete_in_mounts`, `self_heal_ran`, `transcript_no_host_path`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back` — keep these in a periodic live `run` |
|
|
299
304
|
>
|
|
@@ -493,7 +498,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
493
498
|
|
|
494
499
|
**Flags worth knowing** (the full list is always `<command> --help`):
|
|
495
500
|
- `run`: a decider (`--decider-cmd <helper>`/`--decider-dir <dir>`, or a scenario's `on_unanswered: llm` — a scenario-YAML key, not a CLI flag; `run` rejects `--on-unanswered llm`) can answer unscripted gates; `--repeat N` (2-100) runs each scenario N times and aggregates a variance rollup instead of a single pass/fail (`--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd` tune the batch verdict/loop — and `--max-budget-usd` applies without `--repeat` too, as a pre-flight refusal from the scenario's cost history); `--matrix <matrix.yaml>` runs ONE scenario across the cross-product of baseline/model/skill_dir axes (worked example: `examples/matrices/csv-metrics-matrix.yaml`; `--max-cells`/`--concurrency` tune the cap/pool — any cell failing, assertion or infra, fails the run); `--compact`/`--demo` trim `run`/`skill` output for shareable screenshots/GIFs; `--label <tag>` (both `run` and `skill`) stamps a generation name for the iterate-across-fixes loop — surfaced in `result.json` (`runLabel`), the run-index row, `inspect`, and `status.json`, alongside the auto-recorded `skillCommit` (git provenance) and the authoritative `fingerprint.skillHash` a harvest step should pair critiques by.
|
|
496
|
-
- `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `record` also **scans the finished recording** after redaction and before the write, and quarantines it to `<runs-root>/quarantine/` (with a `.findings.txt` naming what leaked) rather than writing a `host-inventory`/`machine-inventory` leak to a repo-visible path — `--allow-host-inventory-fixture`
|
|
501
|
+
- `record`/`replay`: `replay --explain` prints the evidence trail behind every passing assert (the flagship false-green tool — text mode; `--output-format json` already carries `assertions[].evidence`); `--decider-llm`/`--decider-dir` answer gates live during recording; `record <dir/>` is itself a first-class batch input (recording a whole directory of scenarios), and `--rerecord-stale` re-records everything stale in one pass — `--concurrency <N>` bounds the parallelism for either; `--assert-from <scenario.yaml>`/`--reassert` re-check the on-disk `assert:` instead of the frozen one; `--strict`/`--fail-on-skill-drift` control staleness handling on replay; `--no-redact` skips record-time redaction; `--allow-failing` relaxes the post-run verdict gate; `--dry-run` resolves without recording — it runs the REAL loader, so it is also the token-free way to check whether a scenario still loads (exit 2 on a schema error for a single file; a directory reports each broken file and exits 1); `--max-budget-usd <x>` refuses before spending when prior-run history says this scenario — or, on a batch, the whole batch — has cost more than x (at `--concurrency 1` a running total also stops the batch once x is reached; above that it is a pre-flight estimate only, and says so); `--force` overrides the refusal to overwrite a default-path cassette that belongs to a *different* scenario (a slug collision) — not a general overwrite-anything flag; `record` also **scans the finished recording** after redaction and before the write, and quarantines it to `<runs-root>/quarantine/` (with a `.findings.txt` naming what leaked) rather than writing a `host-inventory`/`machine-inventory` leak to a repo-visible path — `--allow-host-inventory-fixture` bypasses only the *pre-flight* refusal (tier + destination path, checked before the spend) and deliberately leaves that scan in force — writing a recording the scan flagged needs the separate `--allow-host-inventory-findings`; `replay --best-effort-future-cassette` lets a cassette from a newer format version replay anyway (warn instead of the default hard error). `replay` also prints an **advisory** note — never a failure — when the rootfs agent image differs from the one the cassette recorded (`environment.agentImage`): the image decides which capabilities exist, so replaying against a different one can move a verdict with nothing in the cassette having changed. It compares the registry digest first (the only identity comparable across machines) and falls back to the local image config id only when neither side has one; cassettes recorded before that field existed simply carry no image to compare.
|
|
497
502
|
- `verify-cassettes`: privacy scan (email/currency/domain/path/machine-inventory) + staleness — exit 1 when verification RAN and found a real problem (a finding, a genuine drift, or scenario-prompt drift), exit 3 when it could NOT complete (an unverifiable-class staleness finding, a too-new cassette format, or a read error) — note that a cassette which fails SHAPE validation is still **privacy-scanned**, because the scan needs a readable transcript rather than a valid document, and each result carries `privacyScanned` saying whether it actually ran (`findings: []` with `privacyScanned: false` is an absence of evidence, not evidence of absence); whole-token allows via `--allow <regex>` (a pattern) / class-scoped `--allow-domain` / `--allow-email` / `--allow-path` / `--allow-machine-inventory` / `--allow-patterns-file <path>` (a **file** of patterns, one regex per line); `--skip-privacy` or `--skip-staleness` runs only part of the gate; a diverged scenario `prompt` vs. the cassette's frozen prompt is also a hard fail (its own `scenarioDrift` bucket), opt out with `--skip-scenario-drift`; `--margins` adds a per-cassette recorded-vs-budget report for count-bound assertions (a single-sample estimate — diagnostic only, never changes the gate verdict); `--allow-empty` makes an **existing but cassette-free directory** exit 0 instead of the default loud exit 2 — for a repo that deliberately commits no cassettes (a *missing* path still fails, so the flag can never green a typo'd path).
|
|
498
503
|
- `stats`: reads `<runsRoot>/index.jsonl`, written automatically at every result; `--since`/`--baseline`/`--branch` filter; `--skill-hash <prefix>`/`--label <tag>` narrow to one generation of the iterate-across-fixes loop and `--group-by scenario|skill-hash|label|fidelity` splits per generation — or per effective fidelity tier — instead of aggregating across them (a window spanning >1 generation or >1 tier warns; under `skill-hash`/`label` grouping, rows lacking the field are excluded from grouping and counted, never bucketed blank — the `fidelity` key is total, so nothing is ever excluded under that grouping); `--runs` lists the individual runs behind each summary with their `skillHash`/`runLabel` (each `runs[]` entry also carries `fidelity` — the tier that run actually executed at, `effectiveFidelity ?? fidelity` — in the `--output-format json` envelope only; the text listing is unchanged); `--last <n>` windows per-group; `--reindex` rebuilds the index from the physical run-dir tree (the migration path for pre-index runs), reconstructing each critique's cost roll-up from its run dir along the way.
|
|
499
504
|
- `diff`: `--changelog` renders known-field prose for a baseline diff; `--view tools|transcript|artifacts|meta` narrows a run/cassette diff to one section; normalization (default on) masks per-run noise (ids/timestamps/session markers/host paths) so two runs of the same scenario diff as identical — `--no-normalize` compares raw values.
|
|
@@ -573,6 +578,31 @@ egress:
|
|
|
573
578
|
|
|
574
579
|
Multiple scenarios × sessions × platform baselines = your regression matrix. Drop YAML in `scenarios/` and CI runs them all.
|
|
575
580
|
|
|
581
|
+
### Assert on the run, not the output
|
|
582
|
+
|
|
583
|
+
The assertions above that check a file or a phrase are the familiar half. The half that's hard to get any
|
|
584
|
+
other way asserts on **how the run behaved** — a property of the execution that leaves no trace in the
|
|
585
|
+
deliverable:
|
|
586
|
+
|
|
587
|
+
| Claim | Key |
|
|
588
|
+
|---|---|
|
|
589
|
+
| no sub-agent ever got `Bash` (a deliberately tool-restricted dispatch stayed restricted) | `subagent_tool_absent: Bash` |
|
|
590
|
+
| nothing was deleted from the outputs mount | `no_delete_in_outputs: true` |
|
|
591
|
+
| the dispatch count stayed bounded — no fan-out blow-up | `dispatch_count_max: <N>` |
|
|
592
|
+
| a sub-agent wrote the file, rather than the main loop doing it for them | `subagent_file_write: {path}` |
|
|
593
|
+
| the skill actually ran, rather than the model answering from its mounted source | `skill_triggered: <rx>` |
|
|
594
|
+
| the sandbox really blocked the network | `egress_denied: <host>` |
|
|
595
|
+
|
|
596
|
+
**You cannot diff your way to "no sub-agent used `Bash`."** A correct run and a run that quietly gave a
|
|
597
|
+
restricted sub-agent shell access produce byte-identical outputs; the difference exists only in the event
|
|
598
|
+
stream. This is the class of property that matters most for a skill whose sub-agents are deliberately
|
|
599
|
+
tool-restricted, or whose safety story is "it can't reach X" — and it is exactly what an output-diffing
|
|
600
|
+
test, or a human reading the final answer, will always report as fine.
|
|
601
|
+
|
|
602
|
+
The full catalog is `cowork-harness assertions --list`; [scenario.md → Which assertion for which question](./docs/scenario.md#which-assertion-for-which-question-goal--key)
|
|
603
|
+
is the goal-first chooser. Mind the two axes in each key's row: which **tier** it needs, and whether it
|
|
604
|
+
**survives `replay`** (several of the above are live-only — keep them in a periodic live `run`).
|
|
605
|
+
|
|
576
606
|
## Sandboxing: container vs. the real VM
|
|
577
607
|
|
|
578
608
|
Local Cowork runs the agent in an **Apple Virtualization.framework microVM** (separate kernel). The harness's default `container` tier uses an OS container (shared kernel, namespaces/cgroups). For **testing skills you wrote**, that's faithful where it counts — same agent binary, same cowork mode, same mount layout, same egress allowlist, same permission protocol — because skill behavior is agent-loop + tool behavior, all kernel-invisible. The container is the right default precisely because it's CI-native; a VM needs nested virtualization most shared CI runners don't have.
|
|
@@ -760,7 +790,7 @@ jobs:
|
|
|
760
790
|
- uses: actions/checkout@v4
|
|
761
791
|
- name: Stage the agent binary (official channel, sha256-verified — see docs/maintenance.md)
|
|
762
792
|
run: |
|
|
763
|
-
V=2.1.
|
|
793
|
+
V=2.1.246 # match your scenario's pinned baseline's agentVersion
|
|
764
794
|
curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
|
|
765
795
|
chmod +x "$RUNNER_TEMP/claude-$V"
|
|
766
796
|
echo "COWORK_AGENT_BINARY=$RUNNER_TEMP/claude-$V" >> "$GITHUB_ENV"
|
|
@@ -772,7 +802,7 @@ jobs:
|
|
|
772
802
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
773
803
|
```
|
|
774
804
|
|
|
775
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.
|
|
805
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.4.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
776
806
|
|
|
777
807
|
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
778
808
|
|
|
@@ -940,6 +970,6 @@ inputs/outputs. Human-readable terminal text is explicitly **not** part of the c
|
|
|
940
970
|
## Status
|
|
941
971
|
|
|
942
972
|
The latest shipped baseline — what `baseline: latest` resolves to (`cowork-harness list`) — is
|
|
943
|
-
**`desktop-1.
|
|
973
|
+
**`desktop-1.37937.1`**. Release-by-release verification notes (what was re-verified against
|
|
944
974
|
which live agent/asar) are recorded in [CHANGELOG.md](./CHANGELOG.md); the feature catalogue
|
|
945
975
|
this section would otherwise duplicate lives in the sections above.
|