cowork-harness 3.0.0 → 3.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +10 -4
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +57 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +4 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +14 -6
- package/AGENTS.md +1 -1
- package/CHANGELOG.md +159 -0
- package/CONTRIBUTING.md +35 -0
- package/README.md +86 -703
- package/SPEC.md +2 -1
- package/dist/cli.js +61 -5
- package/dist/run/cassette.js +113 -11
- package/dist/run/chat-result.js +2 -0
- package/dist/run/chat.js +6 -0
- package/dist/run/execute.js +23 -2
- package/dist/run/hook-events.js +38 -5
- package/dist/run/model-provenance.js +100 -0
- package/dist/run/run.js +12 -0
- package/dist/run/verdict.js +26 -0
- package/docs/README.md +9 -6
- package/docs/boundary.md +7 -0
- package/docs/cassette.md +10 -3
- package/docs/ci.md +114 -0
- package/docs/cli.md +533 -0
- package/docs/companion-skill.md +58 -0
- package/docs/debugging.md +10 -7
- package/docs/fidelity-gaps.md +65 -1
- package/docs/gotchas.md +27 -2
- package/docs/invariants.md +1 -1
- package/docs/maintenance.md +28 -0
- package/docs/scenario.md +7 -1
- package/docs/session.md +5 -3
- package/docs/stats.md +1 -1
- package/examples/README.md +3 -3
- package/examples/replays/README.md +1 -1
- package/examples/sessions/default.yaml +1 -1
- package/llms.txt +3 -0
- package/package.json +1 -1
- package/schema/cassette.v12.json +15 -3
- package/schema/run-result.json +30 -1
- package/scripts/bump-version.ts +12 -1
- package/scripts/check-versions.ts +1 -1
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 3.
|
|
7
|
-
tracks-harness: cowork-harness 3.
|
|
6
|
+
version: 3.1.0
|
|
7
|
+
tracks-harness: cowork-harness 3.1.0 (baseline desktop-1.40609.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.1.0` (baseline
|
|
29
29
|
> `desktop-1.40609.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.1.0"`. **Pin `@^3.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -568,6 +568,12 @@ Recognize these before "fixing" a non-bug:
|
|
|
568
568
|
would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
|
|
569
569
|
with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
|
|
570
570
|
a harness coverage gap (see the *File delivery* section of fidelity-gaps).
|
|
571
|
+
- **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
|
|
572
|
+
`model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
|
|
573
|
+
this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
|
|
574
|
+
transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
|
|
575
|
+
by a different model than the scenario names, so treat a green as evidence about the fallback model.
|
|
576
|
+
|
|
571
577
|
- **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
|
|
572
578
|
connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
|
|
573
579
|
approval, not just outputs — a connected folder shows the identical default — so this run diverged from
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "3.
|
|
20
|
+
(e.g. `version: "3.1.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^3.
|
|
70
|
+
- run: npm i -g "cowork-harness@^3.1.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^3.
|
|
345
|
+
- run: npm i -g "cowork-harness@^3.1.0"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^3.
|
|
374
|
+
run: npm i -g "cowork-harness@^3.1.0"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 3.
|
|
3
|
+
Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -20,7 +20,8 @@ Self-contained reference. Tracks `cowork-harness 3.0.0` (baseline `desktop-1.406
|
|
|
20
20
|
- A `hostloop` scenario with a **writable** connected folder (`mode: rw`/`rwd`) needs `allow_host_writes:
|
|
21
21
|
true` — with no container around the native file tools, that combination gives the agent genuine,
|
|
22
22
|
software-checked-only host filesystem access. Read-only folders and folder-less runs need no opt-in.
|
|
23
|
-
- A `protocol` scenario staging a plugin that declares runnable hooks
|
|
23
|
+
- A `protocol` scenario staging a plugin that declares runnable hooks — in `<plugin>/hooks/hooks.json`
|
|
24
|
+
or the plugin manifest's `hooks` key, both live — needs `allow_host_hooks: true`
|
|
24
25
|
(`--allow-host-hooks` on `chat`/`skill`). L0 passes `--plugin-dir`, so the CLI runs those
|
|
25
26
|
hooks as native host processes under your account with no sandbox. Plugins without hooks need no opt-in.
|
|
26
27
|
Needs cowork-harness >= 3.0.0 — an older CLI rejects the key outright (`Unrecognized key`, exit 2), it
|
|
@@ -155,6 +156,60 @@ hands `fn` exactly this dict.
|
|
|
155
156
|
- `agent` is **retired** — `on_unanswered: agent` is rejected by the schema. The enum is
|
|
156
157
|
`["fail", "prompt", "llm", "first"]`.
|
|
157
158
|
|
|
159
|
+
### A script-running skill trips `permissive_auto_allow` — at `protocol` specifically
|
|
160
|
+
|
|
161
|
+
The default `permission_parity: cowork` auto-allows an unscripted, off-registry tool ask — and records
|
|
162
|
+
it, because real Cowork would have BLOCKED for the user. `computeVerdict` then FAILS the run: a green
|
|
163
|
+
carrying one would not be a faithful pass. Only `Read`, `Glob` and `Grep` are default-allow; **`Bash` is
|
|
164
|
+
off-registry**.
|
|
165
|
+
|
|
166
|
+
Two things decide whether you ever see it, and neither is obvious. The harness only decides how to
|
|
167
|
+
ANSWER a permission ask; whether the agent ASKS is the agent binary's own logic, and that varies by
|
|
168
|
+
both the command and the tier. All rows below are measured, same scenario shape, same session:
|
|
169
|
+
|
|
170
|
+
| Tier | Bash command | tool used | ask | verdict |
|
|
171
|
+
|---|---|---|---|---|
|
|
172
|
+
| `protocol` | `echo hello` | `Bash` | none | ✓ green |
|
|
173
|
+
| `protocol` | `python3 -c "print(42)"` | `Bash` | **one** | ✗ `permissive_auto_allow` |
|
|
174
|
+
| `container` | `python3 -c "print(42)"` | `Bash` | none | ✓ green |
|
|
175
|
+
| `hostloop` | `python3 -c "print(42)"` | `mcp__workspace__bash` | none | ✓ green |
|
|
176
|
+
|
|
177
|
+
So the guard is **not** a general hazard for script-running skills — it is a `protocol` one. The host
|
|
178
|
+
CLI that L0 spawns asks for a command like `python3 -c …` and not for `echo`; the staged agent at
|
|
179
|
+
`container` asks for neither, and `hostloop` routes shell through `mcp__workspace__bash` and so is not
|
|
180
|
+
even the same tool. A skill that runs `python3 ${CLAUDE_SKILL_DIR}/scripts/…` therefore goes red at
|
|
181
|
+
`protocol` on an otherwise-correct run, while the same skill is green on the sandboxed tiers — and a
|
|
182
|
+
hello-world probe is green everywhere and teaches the wrong expectation.
|
|
183
|
+
|
|
184
|
+
Three ways through, in preference order: script the gate (`--answer` / `answers:`), which is what the
|
|
185
|
+
warning tells you and keeps the run deterministic; set `permission_parity: strict` to deny instead of
|
|
186
|
+
allow, if refusal is what you want to test; or assert `allow_permissive_auto_allow: true` when the
|
|
187
|
+
permissive behaviour is deliberately what the scenario is about.
|
|
188
|
+
|
|
189
|
+
> **What is NOT established.** Why the host CLI asks for one command and not another was observed, not
|
|
190
|
+
> traced — do not infer a rule for which commands ask from these rows. `microvm` was not measured.
|
|
191
|
+
|
|
192
|
+
### `baseline agent binary not found` — a Desktop update pruned the pinned ELF
|
|
193
|
+
|
|
194
|
+
A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
|
|
195
|
+
a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
|
|
196
|
+
current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
|
|
197
|
+
|
|
198
|
+
Prefer **repinning `baseline:` to an installed version** for anything
|
|
199
|
+
reproducibility-bound — you keep an exact pin and move it deliberately. `baseline: latest` never rots
|
|
200
|
+
but silently drifts, so two runs weeks apart are not comparable; a pin and `latest` have opposite
|
|
201
|
+
failure modes.
|
|
202
|
+
|
|
203
|
+
To find which versions you actually have, do NOT use `cowork-harness list` — it enumerates the baseline
|
|
204
|
+
definitions shipped with the harness, which are present regardless of what Desktop pruned from this
|
|
205
|
+
machine, so a pruned pin lists as healthy. Test the staged binary: read `agentBinary.stagedPath` from
|
|
206
|
+
the baseline JSON, expand the leading `~`, and check it is a FILE — the pruned case leaves the empty
|
|
207
|
+
directory behind, so a directory test passes on exactly the case that fails.
|
|
208
|
+
|
|
209
|
+
`COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` runs the newest sibling instead of the pinned
|
|
210
|
+
binary and downgrades the sha check to advisory — that is the substitution the hard failure exists to
|
|
211
|
+
prevent, so use it to unblock once, never in CI.
|
|
212
|
+
|
|
158
213
|
### Determinism contract
|
|
159
214
|
|
|
160
215
|
- `fail` — the default for `run`. On an unscripted gate it hard-errors; the error names the exact
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.1.0`
|
|
4
4
|
(baseline `desktop-1.40609.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -100,7 +100,8 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
|
|
|
100
100
|
# Read-only folders and folder-less runs need no opt-in.
|
|
101
101
|
|
|
102
102
|
allow_host_hooks: true # OPTIONAL — required to run `protocol` fidelity when a staged plugin
|
|
103
|
-
# declares runnable hooks
|
|
103
|
+
# declares runnable hooks — either `<plugin>/hooks/hooks.json` OR
|
|
104
|
+
# the manifest's `hooks` key: L0 passes
|
|
104
105
|
# --plugin-dir, so the CLI executes those hooks as NATIVE HOST
|
|
105
106
|
# processes under your account, with no container sandbox. A plugin
|
|
106
107
|
# that declares no hooks needs no opt-in, and a misplaced root-level
|
|
@@ -420,6 +421,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
|
|
|
420
421
|
| `infra_error` | fail | A supervising process died mid-run (VM/egress sidecar) — the run's evidence is contaminated, not author-suppressible |
|
|
421
422
|
| `stalled` | fail | The run ended on an unanswered question, or on a trailing-`?` final turn with no tool work after the last gate (opt out: `allow_stall`) |
|
|
422
423
|
| `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
|
|
424
|
+
| `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
|
|
423
425
|
| `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
|
|
424
426
|
| `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
|
|
425
427
|
| `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 3.
|
|
5
|
+
Tracks `cowork-harness 3.1.0` (baseline `desktop-1.40609.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -46,6 +46,11 @@ what production would do. Two layers of defense:
|
|
|
46
46
|
statically knowable and only produce a non-failing informational note.
|
|
47
47
|
- **Manual (one-liner):** `grep -h '"effectiveFidelity"' cassettes/*.cassette.json | sort | uniq -c`
|
|
48
48
|
shows the tier distribution of the fleet at a glance.
|
|
49
|
+
- **Do NOT re-record for a `[note]` alone.** A `·`/`[note]` row is informational and the run still exits
|
|
50
|
+
**0** — `session-fingerprint: predates \`model\` coverage` and `prompt-assets: cassette predates …` both
|
|
51
|
+
mean "this cassette was recorded before that field existed, and everything the hash *does* cover still
|
|
52
|
+
matches". Re-recording buys the new coverage and nothing else, so on a fleet of heavy cassettes it is a
|
|
53
|
+
real bill for no verdict change. Re-record when a `✗` says to (baseline moved, skill drift, tier moved).
|
|
49
54
|
|
|
50
55
|
### Cassette anatomy (what you're looking at when you open one)
|
|
51
56
|
|
|
@@ -66,7 +71,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v12.json`](htt
|
|
|
66
71
|
| `preRunOrigin` | How that pre-run baseline was obtained — `local-walk` (real), `remote-unavailable` or `local-unreadable`. Only `local-walk` supports a verdict: replay fails `no_unexpected_files` as evidence-unavailable on the other two rather than passing vacuously |
|
|
67
72
|
| `scenarioSource` | Relative path to the authored YAML this was recorded from |
|
|
68
73
|
| `authoring` | Present iff a live decider answered ≥1 gate during recording (`nonDeterministic: true`) |
|
|
69
|
-
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
74
|
+
| `sessionFingerprint` | Optional even on v9+ (the minimum readable version): hash of the session's content-relevant SHAPE (model/folders/plugins/skills/mcp/egress/web_fetch, plus projects and agent_env when set). Checked ONLY by `verify-cassettes`, never the default replay verdict; absent → not checked |
|
|
70
75
|
| `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
|
|
71
76
|
| `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
|
|
72
77
|
| `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
|
|
@@ -241,10 +246,13 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
241
246
|
unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
|
|
242
247
|
dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
|
|
243
248
|
whose source was never committed names a generation that is unrecoverable, which makes the
|
|
244
|
-
comparison uninterpretable rather than merely noisy. And with no `model:` in the session
|
|
245
|
-
`--model` on the
|
|
246
|
-
before/after can silently straddle two models
|
|
247
|
-
|
|
249
|
+
comparison uninterpretable rather than merely noisy. And with no `model:` in the session and no
|
|
250
|
+
`--model` on the command (every lane takes it), each run uses whatever the staged agent binary
|
|
251
|
+
defaults to, so a before/after can silently straddle two models — the run warns when nothing pinned
|
|
252
|
+
one. Read `result.json` back to confirm: `modelSource` says whether anything pinned the model at all,
|
|
253
|
+
and `modelPinHonored` whether the pin survived (**absent means unverifiable, not "yes"**). `models`
|
|
254
|
+
lists what served the run — ignore any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent
|
|
255
|
+
fabricated locally, not a model).
|
|
248
256
|
3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
249
257
|
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
250
258
|
content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
|
package/AGENTS.md
CHANGED
|
@@ -22,7 +22,7 @@ protocol layer or run-loop bookkeeping in the CLI.
|
|
|
22
22
|
Don't add a test that needs a live model or Docker to the default suite; that's the `pytest -m cowork` /
|
|
23
23
|
`npm run test:live` lane. Python fast lane (from `python/`): `pytest -m 'not cowork'`.
|
|
24
24
|
- CLI binary `cowork-harness`; env vars `COWORK_HARNESS_*` (+ `COWORK_AGENT_BINARY` / `COWORK_AGENT_IMAGE`) —
|
|
25
|
-
see README's [Reproducibility knobs](./
|
|
25
|
+
see README's [Reproducibility knobs](./docs/cli.md#reproducibility-knobs) for the full env-var list.
|
|
26
26
|
Node ≥ 22.
|
|
27
27
|
- `cowork-harness sync` is **local-only** (needs Desktop + `app.asar`; not on CI). The committed
|
|
28
28
|
`baselines/*.json` are CI's source of truth — never hand-edit release facts into source; they come from
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,165 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [3.1.0] — 2026-08-31
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **The pre-coverage cassette note was a paragraph repeated per file.** It fires on every cassette
|
|
14
|
+
recorded before `model` joined the session shape — for a consumer with a dozen, a dozen identical
|
|
15
|
+
paragraphs burying the findings that are genuinely per-file. Cut to one line (370 chars to 157), matching
|
|
16
|
+
the prompt-assets note beside it; the reasoning stays in `docs/fidelity-gaps.md`. Notes are deliberately
|
|
17
|
+
**not** aggregated in `verify-cassettes` — it is a per-file audit, and collapsing them would destroy the
|
|
18
|
+
attribution — so length was the part to fix. The note has always been non-failing: a pre-coverage
|
|
19
|
+
cassette stays clean, and `replay --strict` exits 0.
|
|
20
|
+
- **The unpinned-model warning told `run` users the flag lived on other lanes.** It named the `skill`,
|
|
21
|
+
`probe-dispatch` and `chat` lanes for a flag `run` and `record` now accept — caught by running it, not
|
|
22
|
+
by reading it.
|
|
23
|
+
|
|
24
|
+
- **`run --model ""` started a spending run instead of failing.** SPEC §CB-2 requires an empty or
|
|
25
|
+
whitespace model value to be a usage error, never a silently-propagated empty string. `record` inherits
|
|
26
|
+
that from the shared flag parser; `run` hand-rolls its own loop (the session-less lanes take `--model`
|
|
27
|
+
out of the leftovers themselves, so the shared parser cannot consume it) and did not enforce it.
|
|
28
|
+
|
|
29
|
+
### Added
|
|
30
|
+
|
|
31
|
+
- **`--model <id>` on `run` and `record`.** The `skill`, `probe-dispatch` and `chat` lanes already had it;
|
|
32
|
+
the two lanes that produce cassettes and CI verdicts did not, so the only way to pin them was editing the
|
|
33
|
+
session file. Precedence is explicit: an `--model` flag or a `--matrix` `models:` axis (a per-invocation
|
|
34
|
+
act) outranks the session's `model:`, while `COWORK_HARNESS_MODEL` only fills a gap — a machine-scoped
|
|
35
|
+
variable must never silently outrank a model the scenario declares, or the run's model becomes a
|
|
36
|
+
property of the shell it was launched from.
|
|
37
|
+
- **Model provenance on every run, and a warning when nothing pins the model.** `--model` reaches the
|
|
38
|
+
agent only when a session declares `model:`, so an unset session inherits whatever the local CLI would
|
|
39
|
+
pick — and because the agent selects part of its system prompt by model capability, that moves the
|
|
40
|
+
*instructions* it is given, not just answer quality. `--effort` is not symmetric: it is always emitted
|
|
41
|
+
from the baseline's `spawn.effortDefault`, so an unset session pinned the effort and left the model it
|
|
42
|
+
applies to floating. Runs now warn when no model resolves (`chat` in its own words), naming the key and
|
|
43
|
+
the flag. Omitting it is deprecated and becomes an error in the next major — the same path
|
|
44
|
+
`Scenario.fidelity` took, and deliberately not a hard fail today: the repo chose deprecation for a
|
|
45
|
+
default with a larger blast radius, and a harder gate for a lesser field would be incoherent with that.
|
|
46
|
+
- **A family pin is checked as family membership, not equality.** `--model opus` is resolved by the agent
|
|
47
|
+
to a concrete id (`claude-opus-5`), never echoed back, and which member an account resolves it to is not
|
|
48
|
+
something the harness can know — so comparing the two as strings would report a pinned-and-honored run as
|
|
49
|
+
`modelPinHonored: false`. `opus`/`sonnet`/`haiku`/`fable` now compare by family and still fail on a
|
|
50
|
+
wrong-family model; `best` and `opusplan` name no comparable model and stay unverifiable.
|
|
51
|
+
- **`RunResult.modelFallbacks`, `modelPinHonored` and `modelSource`.** Fallbacks come from the agent's own
|
|
52
|
+
`system`/`model_fallback` event rather than a diff of `models[]` — the event names the trigger, which is
|
|
53
|
+
what separates a retired or blocked pin (every run falls back the same way until the id changes) from a
|
|
54
|
+
transient overload. A new `model_fallback` warn signal surfaces it in the verdict; it cannot fire on a
|
|
55
|
+
healthy run, so it adds no volume to a currently-green one. `modelPinHonored` is three-state on purpose:
|
|
56
|
+
`true`, `false`, and **absent for unverifiable** — nothing pinned, or no live model evidence. The
|
|
57
|
+
natural two-state shape reports "we could not tell" as "the pin held", which is a false green, so the
|
|
58
|
+
no-evidence cases are asserted explicitly rather than left to a truthiness check.
|
|
59
|
+
- **A cassette records the model it was recorded with**, in its `environment` block: the id the agent
|
|
60
|
+
reported (so it survives a mid-run fallback) plus the source. A cassette recorded with nothing pinned
|
|
61
|
+
says `unresolved`, which is the difference between "recorded on opus-5 by choice" and "on whatever that
|
|
62
|
+
laptop happened to pick".
|
|
63
|
+
- **The session fingerprint now covers the pinned model.** Editing `model:` after recording previously
|
|
64
|
+
moved nothing and `verify-cassettes` said nothing, even though the edit changes what the agent was told.
|
|
65
|
+
A cassette recorded before this coverage reports **unverifiable** rather than clean — its hash cannot
|
|
66
|
+
tell "the field was never covered" from "the pin changed since", and claiming otherwise in the remedy
|
|
67
|
+
would reintroduce the false green the coverage exists to close.
|
|
68
|
+
|
|
69
|
+
## [3.0.1] — 2026-08-30
|
|
70
|
+
|
|
71
|
+
### Changed
|
|
72
|
+
|
|
73
|
+
- **The README now shows the product, not just the argument for it.** Three independent reviews found
|
|
74
|
+
the same gap: for a test harness, the file a user authors is the most persuasive thing the project
|
|
75
|
+
owns, and the router split had left it with zero examples. Adds a worked scenario (prompt, scripted
|
|
76
|
+
answers, assertions) with its verdict output — the scenario is extracted and linted in place, so it is
|
|
77
|
+
a real file rather than plausible-looking YAML — plus a "What that catches" section built from an
|
|
78
|
+
actual run where a named skill was offered, declined, and the green answer looked fine anyway
|
|
79
|
+
(`skillsInvoked: []`, `toolCounts: {}`) — framed as a closed loop, since the author fixed and re-verified
|
|
80
|
+
it with the same instrument roughly ninety minutes later. An outcome example is a claim about a MOMENT;
|
|
81
|
+
the better the tool works the faster its own examples get fixed out from under it, so it needs a tense. Also names two things the docs asserted but never taught:
|
|
82
|
+
diffing the same scenario across two fidelity tiers as a discovery technique, and proving an assertion
|
|
83
|
+
can fail before trusting it. The requirements block moves below the `claude -p` argument — three
|
|
84
|
+
reviewers independently reported reading install prerequisites before any reason to want them.
|
|
85
|
+
|
|
86
|
+
- **The README is now a router, not the whole manual.** It was 975 lines carrying three unrelated
|
|
87
|
+
audiences at once; it is now **282** and branches to per-audience pages: **`docs/cli.md`** (install,
|
|
88
|
+
prerequisites, commands, the two files, run output, env knobs), **`docs/companion-skill.md`** (install
|
|
89
|
+
+ orientation; usage stays in `SKILL.md`), and **`docs/ci.md`** (the token-free gate, the packaged
|
|
90
|
+
Action, the live lane — written to stand alone). The harness's own pipeline and contributor suite moved
|
|
91
|
+
to `CONTRIBUTING.md`, which keeps `docs/ci.md` about *consuming* the Action rather than about this repo.
|
|
92
|
+
README keeps what all three audiences share: the fidelity-tier vocabulary, architecture, limitations,
|
|
93
|
+
the docs index, and versioning. Content moved verbatim — this is a relocation, not a rewrite.
|
|
94
|
+
|
|
95
|
+
Nothing is dropped: `Sandboxing` folded into `docs/boundary.md` and `Maintenance` into
|
|
96
|
+
`docs/maintenance.md`; `Discovery` was already covered in `docs/discovery.md`, so it was deleted rather
|
|
97
|
+
than duplicated. Every moved relative link was rewritten for its new depth, and every deep link into a
|
|
98
|
+
moved section was repointed across `docs/`, `examples/README.md`, `llms.txt` and the docs index.
|
|
99
|
+
|
|
100
|
+
### Fixed
|
|
101
|
+
|
|
102
|
+
- **`bump-version` did not know about the router-split pages.** `docs/cli.md`, `docs/ci.md` and
|
|
103
|
+
`docs/companion-skill.md` carry `cowork-harness@^X.Y.Z` install floors that `check:versions` enforces,
|
|
104
|
+
but the bump tool's target list predated them — so `npm run bump` rewrote every other floor and then
|
|
105
|
+
failed its own post-write lockstep check. Caught at release time by that self-check rather than
|
|
106
|
+
shipping a repo with mismatched floors. All three are now registered, and the test pinning that list
|
|
107
|
+
carries the reason.
|
|
108
|
+
|
|
109
|
+
|
|
110
|
+
- **22 dead anchor links introduced by the README router split, and the guard gap that let them ship.**
|
|
111
|
+
Moving 11 `##` sections out of README left every `](#slug)` link pointing at them dangling — 19 in
|
|
112
|
+
README (including two badges and the two nav lines the page opened with), plus one each in `AGENTS.md`,
|
|
113
|
+
`docs/ci.md` and `docs/companion-skill.md`. GitHub fails these silently: the page simply does not
|
|
114
|
+
scroll. Every file-path link still resolved and every existing guard stayed green, because the anchor
|
|
115
|
+
suite validated links *into* README and *into* sibling docs but never a page's links to its **own**
|
|
116
|
+
headings. `test/repo-docs-anchors.test.ts` now covers that third case across root pages and `docs/`,
|
|
117
|
+
with a canary against a checker that extracts nothing, and was confirmed to fail against the exact
|
|
118
|
+
regression that shipped. The two nav lines are deleted rather than repointed — they were a
|
|
119
|
+
table-of-contents for a 975-line page, and "Pick your path" plus the Documentation table now do that job.
|
|
120
|
+
|
|
121
|
+
Found by two independent reviewers; the second caught the three non-README instances a README-scoped
|
|
122
|
+
fix would have missed.
|
|
123
|
+
|
|
124
|
+
Follow-up from the same review: three cross-page links still read "…\[Prerequisites](./docs/cli.md#…)
|
|
125
|
+
**below**" — resolving correctly while the sentence lied, because Prerequisites had moved to another
|
|
126
|
+
page. Reworded, and guarded: a cross-page link followed closely by "below"/"above" now fails the suite.
|
|
127
|
+
This half is the nastier one — the link works, so every link checker stays green forever.
|
|
128
|
+
|
|
129
|
+
- **The `protocol` host-hook consent gate was defeatable by a spelling choice.** `allow_host_hooks`
|
|
130
|
+
(3.0.0) refuses a spawn until the operator consents to a staged plugin's hooks running as native host
|
|
131
|
+
processes — but detection keyed only on a file named `hooks.json`, and a plugin may equally declare its
|
|
132
|
+
hooks in the manifest's `hooks` key (`.claude-plugin/plugin.json`, or a bare `plugin.json`), which is
|
|
133
|
+
the more common spelling. Such a plugin sailed past the gate: reproduced end-to-end, the hook executed
|
|
134
|
+
under the operator's own account while the run went green, no consent was asked and no disclosure
|
|
135
|
+
printed. The same blind predicate feeds the tier-independent disclosure, so those hooks also ran
|
|
136
|
+
unannounced at `hostloop`. Detection now returns the UNION of both channels, which fixes the gate and
|
|
137
|
+
the disclosure together since both call through it. The `hooks/hooks.json` placement carve-out is
|
|
138
|
+
deliberately NOT extended to the manifest — that carve-out exists because a misplaced `hooks.json` is
|
|
139
|
+
inert, and a manifest-declared hook is live wherever the manifest sits.
|
|
140
|
+
|
|
141
|
+
Reported by a consumer during a 3.0.0 adoption pass, with the defect proven from committed cassettes:
|
|
142
|
+
10 of theirs carried a `SessionStart:startup` hook_started/hook_response pair from a manifest-only
|
|
143
|
+
declaration, including one recorded at `container`.
|
|
144
|
+
|
|
145
|
+
### Documentation
|
|
146
|
+
|
|
147
|
+
- **The pruned-agent-binary failure now has a documented remedy, with the tradeoffs named.** A Desktop
|
|
148
|
+
update deletes the prior version's staged ELF while often leaving an empty version directory, so a
|
|
149
|
+
scenario pinning that agent version dies with `baseline agent binary not found`. The code anticipates
|
|
150
|
+
this by name, but the remedy appeared in no skill surface at all — and an installed companion skill
|
|
151
|
+
ships without repo `docs/`, so the agent most likely to hit it had nothing to read. Now in
|
|
152
|
+
`docs/gotchas.md` and the skill's own reference, stating that the three remedies are NOT equivalent:
|
|
153
|
+
repin for a reproducibility-bound suite, `latest` for a one-off (a pin rots silently, `latest` drifts
|
|
154
|
+
silently — opposite failure modes), and `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` only to unblock once,
|
|
155
|
+
never in CI, since it runs the newest sibling and downgrades the sha check to advisory.
|
|
156
|
+
|
|
157
|
+
- **`permissive_auto_allow` and script-running skills is a `protocol`-specific hazard, not a general one.**
|
|
158
|
+
`Bash` is off-registry (only `Read`/`Glob`/`Grep` are default-allow), but the harness only decides how to
|
|
159
|
+
ANSWER a permission ask — whether the agent ASKS varies by command AND tier. Measured: at `protocol`, `echo
|
|
160
|
+
hello` produces no ask (green) while `python3 -c "print(42)"` produces one and fails the guard; the
|
|
161
|
+
identical command at `container` produces no ask (green), and `hostloop` routes shell through
|
|
162
|
+
`mcp__workspace__bash` entirely. So a skill running `python3 ${CLAUDE_SKILL_DIR}/scripts/…` goes red at L0
|
|
163
|
+
and green on the sandboxed tiers, while a hello-world probe is green everywhere and teaches the wrong
|
|
164
|
+
expectation. `references/fidelity-and-answers.md` carries the measured table and the three ways through; why
|
|
165
|
+
the host CLI asks for one command and not another is explicitly left untraced, and `microvm` is named as
|
|
166
|
+
unmeasured.
|
|
167
|
+
|
|
9
168
|
## [3.0.0] — 2026-08-29
|
|
10
169
|
|
|
11
170
|
### Breaking
|
package/CONTRIBUTING.md
CHANGED
|
@@ -108,3 +108,38 @@ When a Desktop release moves something `sync` doesn't read, it reports an `unkno
|
|
|
108
108
|
## Reporting issues
|
|
109
109
|
|
|
110
110
|
Use the issue templates. For anything security/sandbox related, see [SECURITY.md](./SECURITY.md).
|
|
111
|
+
|
|
112
|
+
## This repo's own CI pipeline
|
|
113
|
+
|
|
114
|
+
Contributor-facing. For *consuming* the harness in your own CI — the token-free gate and the packaged
|
|
115
|
+
Action — see [docs/ci.md](./docs/ci.md).
|
|
116
|
+
|
|
117
|
+
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
118
|
+
|
|
119
|
+
| Stage | Runs | Needs | Gates |
|
|
120
|
+
|---|---|---|---|
|
|
121
|
+
| **build** | format check · version-lockstep guard · typecheck · source guards · build · CLI smoke · token-free `replay` · `verify-cassettes` · `lint` | nothing | every push/PR |
|
|
122
|
+
| **test** | the unit suite (vitest), sharded 4-way | nothing | every push/PR |
|
|
123
|
+
| **floor** | the unit suite once, unsharded, on Node 22 — the version `engines.node` declares — so the floor is exercised rather than asserted (other jobs run Node 24, the Active LTS line) | nothing | every push/PR; gates the merge context |
|
|
124
|
+
| **action-self-test** | packs this commit and runs the packaged `uses: ./` Action across its full case set — pass (committed example cassette), usage-error fail (nonexistent path), assertion fail (checks the reporter renders a ❌ row), `lint`, and `analyze-skill`; `ci.yml` currently carries 11 `command:` invocations, so re-count here rather than trusting this sentence | nothing | every push/PR |
|
|
125
|
+
| **python** | `pytest` helper self-checks (`python/`, run with `-m 'not cowork'` — the token-free subset; the Docker/token `@pytest.mark.cowork` tests are excluded) | nothing (token-free assertions only) | every push/PR |
|
|
126
|
+
| **boundary** | builds the pinned agent image, brings up the default-deny network, runs `boundary-check`, then `npm run test:live` (live contract tests that guard the binary-resolution assumptions, no token needed) | Docker, arm64 runner | proves the sandbox enforces Cowork's limits — **no API key** |
|
|
127
|
+
| **image-recipe** | compiles `docker/Dockerfile.agent` in **both** variants (lean/core and `COWORK_FULL_PARITY=1`) on an arm64 runner, so a recipe change cannot reach the merge gate uncompiled | Docker, arm64 runner | every push/PR; gates the merge context |
|
|
128
|
+
| **scenarios** | the live scenario suite (mixed `protocol` + `container` fidelity across `examples/scenarios/`), plus the `e2e/scenarios/*.yaml` smoke set (this repo's own L0/L1/hostloop self-tests, `microvm` excluded — needs a real VM); uploads transcripts/egress logs as artifacts; relies on a runner-local staged agent binary (no in-workflow download step). | `ANTHROPIC_API_KEY` | fork PRs: the whole job is skipped (`if:` guard); same-repo without a key: warns and exits 0 |
|
|
129
|
+
| **parity-drift** | reminder to re-`sync` when Desktop updates | nothing | **goes red** if the newest committed baseline is > 90 days old, but sits outside `ci-green`'s `needs:` list, so a red run does not block a merge |
|
|
130
|
+
|
|
131
|
+
This ordering means cheap checks fail fast, the **boundary parity gate runs without secrets** (so forks get it too), and expensive live runs only happen when a key is present.
|
|
132
|
+
|
|
133
|
+
|
|
134
|
+
## The harness's own suite
|
|
135
|
+
|
|
136
|
+
```bash
|
|
137
|
+
npm run ci # typecheck + build + test (run format:check separately; NOT the same set as CI's `build` job — see CONTRIBUTING.md)
|
|
138
|
+
npm test # vitest: decider, egress allowlist, launch plan, example validation
|
|
139
|
+
cowork-harness boundary-check # self-verify the sandbox (needs Docker; not part of `npm run ci`)
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
Unit tests cover the scripted-answer logic, the egress allowlist matcher, the session→launch-plan materialization (mounts + discovery settings + env-strip), and a **schema guard** that fails if any shipped baseline/session/scenario stops validating. Add a test alongside any new schema field or `Decider` rule — see [CONTRIBUTING.md](./CONTRIBUTING.md).
|
|
143
|
+
|
|
144
|
+
> Copy your starting scenarios/sessions from **`examples/`**. The **`e2e/`** directory is the harness's *own* fidelity self-tests (smoke scenarios per tier) — not a template to copy.
|
|
145
|
+
|