cowork-harness 3.10.0 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +74 -1124
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +159 -0
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +60 -0
  4. package/.claude/skills/cowork-harness/references/authoring.md +258 -0
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +41 -32
  6. package/.claude/skills/cowork-harness/references/critique.md +4 -3
  7. package/.claude/skills/cowork-harness/references/debugging.md +145 -0
  8. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +8 -5
  9. package/.claude/skills/cowork-harness/references/gotchas.md +285 -0
  10. package/.claude/skills/cowork-harness/references/measurement.md +56 -0
  11. package/.claude/skills/cowork-harness/references/run-record-replay.md +339 -0
  12. package/.claude/skills/cowork-harness/references/scenario-schema.md +37 -172
  13. package/.claude/skills/cowork-harness/references/task-recipes.md +36 -23
  14. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +10 -0
  15. package/.claude/skills/cowork-harness/scripts/scenario.py +471 -32
  16. package/.env.example +4 -0
  17. package/AGENTS.md +6 -0
  18. package/CHANGELOG.md +422 -0
  19. package/DESIGN.md +2 -2
  20. package/README.md +9 -9
  21. package/RELEASING.md +2 -4
  22. package/SPEC.md +177 -32
  23. package/baselines/desktop-2.9939.4.json +1116 -0
  24. package/dist/assert.js +37 -10
  25. package/dist/cli.js +354 -147
  26. package/dist/critique/command.js +61 -27
  27. package/dist/critique/evaluator.js +9 -4
  28. package/dist/decide/external-channel.js +41 -12
  29. package/dist/decide/llm-transport.js +2 -0
  30. package/dist/decide/semantic-judge.js +7 -2
  31. package/dist/redactable-literal.js +43 -0
  32. package/dist/run/analyze-skill.js +5 -3
  33. package/dist/run/budget.js +8 -2
  34. package/dist/run/cassette.js +560 -102
  35. package/dist/run/chat-result.js +3 -2
  36. package/dist/run/chat.js +37 -21
  37. package/dist/run/command-globals.js +136 -0
  38. package/dist/run/doctor.js +6 -4
  39. package/dist/run/envelope.js +36 -79
  40. package/dist/run/execute.js +174 -83
  41. package/dist/run/lint-load.js +3 -1
  42. package/dist/run/migrate-run-dir.js +3 -1
  43. package/dist/run/model-provenance.js +30 -17
  44. package/dist/run/run.js +50 -5
  45. package/dist/run/runs-gc.js +4 -2
  46. package/dist/run/scaffold.js +2 -0
  47. package/dist/run/scenario-tool.js +122 -23
  48. package/dist/run/skill-flag-surface.js +9 -1
  49. package/dist/run/tier-vacuous-tools.js +12 -0
  50. package/dist/run/timeline-fold.js +48 -19
  51. package/dist/run/trace-view.js +109 -18
  52. package/dist/run/verdict.js +8 -4
  53. package/dist/session.js +87 -68
  54. package/dist/spawn-guard.js +13 -0
  55. package/dist/staging/resolve.js +11 -9
  56. package/dist/tool-call-assert.js +237 -0
  57. package/dist/types.js +64 -6
  58. package/docs/cassette.md +46 -8
  59. package/docs/chat.md +6 -2
  60. package/docs/ci.md +17 -16
  61. package/docs/cli.md +24 -19
  62. package/docs/companion-skill.md +5 -3
  63. package/docs/critique.md +11 -8
  64. package/docs/debugging.md +7 -4
  65. package/docs/decider-dir.md +21 -4
  66. package/docs/fidelity-gaps.md +41 -17
  67. package/docs/gotchas.md +2 -2
  68. package/docs/invariants.md +1 -0
  69. package/docs/maintenance.md +1 -1
  70. package/docs/plugin-root.md +32 -6
  71. package/docs/run-status.md +1 -0
  72. package/docs/scenario.md +41 -23
  73. package/docs/session.md +8 -8
  74. package/docs/stats.md +5 -0
  75. package/examples/replays/README.md +5 -5
  76. package/examples/replays/example-multiselect-gate.cassette.json +48 -48
  77. package/examples/replays/example-pdf-skill.cassette.json +92 -95
  78. package/examples/replays/hostloop-computer-links.cassette.json +62 -61
  79. package/examples/sessions/default.yaml +1 -1
  80. package/llms.txt +1 -1
  81. package/package.json +1 -1
  82. package/python/test_scenario_lint.py +434 -21
  83. package/schema/cassette.v13.json +330 -0
  84. package/schema/critique-report.json +1 -1
  85. package/schema/run-result.json +36 -1
  86. package/schema/scenario.schema.json +190 -9
  87. package/schema/verify-cassettes.json +2 -2
  88. package/scripts/bump-version.ts +99 -93
  89. package/scripts/check-claims.ts +32 -19
  90. package/scripts/check-surface.ts +58 -10
  91. package/scripts/check-versions.ts +22 -15
  92. package/scripts/gen-schema.ts +9 -1
@@ -1,10 +1,10 @@
1
1
  ---
2
2
  name: cowork-harness
3
- description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
3
+ description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.10.0
7
- tracks-harness: cowork-harness 3.10.0 (baseline desktop-2.9939.2)
6
+ version: 4.0.0
7
+ tracks-harness: cowork-harness 4.0.0 (baseline desktop-2.9939.4)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,11 +22,12 @@ Anthropic. Say so if a user asks what it is.
22
22
  The single most important idea: **a green run is not automatically a correct run.** The harness has
23
23
  several ways to no-op a check while still producing a green run (skip an assertion on replay — now
24
24
  flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
25
- allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
- the highest-value part. Read it.
25
+ allowlist). This skill exists mostly to keep you out of those traps — the *Invariants* below and the
26
+ full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
+ Read them.
27
28
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.10.0` (baseline
29
- > `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.0.0` (baseline
30
+ > `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
32
 
32
33
  ## Preflight — make sure the harness can actually run
@@ -35,14 +36,14 @@ The 10-second inner loop, once the CLI is on PATH:
35
36
 
36
37
  ```bash
37
38
  cowork-harness doctor # prerequisites OK? (Docker, agent, token, baseline)
38
- cowork-harness skill ./my-skill "do X" # run the skill once against the staged agent
39
+ cowork-harness skill ./my-skill "do X" --model claude-sonnet-5 # run once (or set COWORK_HARNESS_MODEL)
39
40
  ```
40
41
 
41
42
  Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
42
43
 
43
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.10.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.10.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.10.0"`. **Pin `@^3.10.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.0.0"`. **Pin `@^4.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
47
 
47
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -50,20 +51,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
50
51
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
51
52
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
52
53
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
53
- - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
54
+ - **Point at another `.env`** with `--dotenv <path>`, and relocate run output with `--run-dir <path>`; both work before or after the subcommand.
54
55
 
55
56
  ## Orient — the three loops
56
57
 
57
- Everything you do with the harness is one of **three loops**, and the rest of this skill is organized
58
- into three Parts to match: **author** a scenario (Part I), **run / record / lock** it into a
59
- reproducible regression (Part II), and **debug** a run that misbehaved or greened when it shouldn't
60
- (Part III — reachable straight from here in one hop).
58
+ Everything you do with the harness is one of **three loops**, and the detail lives in one reference
59
+ file per loop: **author** a scenario ([`references/authoring.md`](references/authoring.md)), **run /
60
+ record / lock** it into a reproducible regression ([`references/run-record-replay.md`](references/run-record-replay.md)),
61
+ and **debug** a run that misbehaved or greened when it shouldn't ([`references/debugging.md`](references/debugging.md)).
61
62
 
62
63
  Pick the entry point you need. The first three are the everyday path — a quick liveness check, the
63
64
  CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that hang off them:
64
65
 
65
- - **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>"`. Fastest; no
66
- scenario file.
66
+ - **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>" --model <id>`. Fastest;
67
+ no scenario file.
67
68
  - **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
68
69
  This is the CI-grade path and most of this skill.
69
70
  - **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
@@ -72,7 +73,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
72
73
  `cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
73
74
  then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
74
75
  the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
75
- output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
76
+ output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
76
77
  [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
77
78
  **"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
78
79
  grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
@@ -87,9 +88,9 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
87
88
  the cost, and it answers that question directly. Report and evidence-package shapes:
88
89
  `references/critique.md`.
89
90
  - **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
90
- TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
91
+ TTY, **not** an asserted test — see *Debugging with `chat`* in `references/debugging.md`).
91
92
  **"Interactive" splits two ways — don't take the wrong branch.** Want to answer gates yourself *and*
92
- still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path* below),
93
+ still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path* in `references/authoring.md`),
93
94
  **not** `chat`. Reach for `chat` only when you are exploring by hand and do NOT want a verdict.
94
95
 
95
96
  > **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
@@ -102,1109 +103,58 @@ lint-skill · analyze-skill · probe-dispatch ·
102
103
  verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
103
104
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
104
105
 
105
- **Two different `scaffold` tools — don't confuse them.** The native `cowork-harness scaffold <run-id>`
106
- above turns an already-*recorded* run into a scenario (needs a run to exist first). The bundled
107
- `scripts/scenario.py scaffold --name … --skill …` — see *Scaffold a valid scenario, then lint before
108
- you push* in **Part I** — builds a scenario from flags alone, no run required. Passing that section's
109
- flag set to the native command fails with `unknown flag: --name` (exit 2).
110
-
111
- ## Part I — AUTHOR a scenario
112
-
113
- Everything below composes one deterministic, asserted `scenarios/*.yaml`: the session/scenario split,
114
- how the skill mounts, the fidelity tier, the answer path, the two assertion axes, `web_fetch`
115
- provenance, and the scaffold/lint tools that keep the YAML honest.
116
-
117
- ### Two files: session vs scenario
118
-
119
- - **`sessions/*.yaml`** — pre-prompt setup: `model`, mounts (`folders`), and discovery
120
- (marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario that
121
- omits `session:` gets an all-defaults **inline** session (not a file on disk).
122
- - **`scenarios/*.yaml`** — the test: `prompt`, scripted `answers:`, and `assert:`.
123
-
124
- This split matters: release ground truth (`baseline:` / `baselines/`, produced by `sync`) is
125
- **separate** from authored setup (`session:` / `sessions/`). "profile" is retired vocabulary — do
126
- not use it. See `references/scenario-schema.md` for every field.
127
-
128
- ### Discovery: how the skill-under-test gets mounted
129
-
130
- The skill is **copied fresh into the sandbox each run**. Wire it via `plugins.local_plugins` +
131
- `plugins.enabled: [<plugin>@local]` in the session (or `--marketplace` / `--plugin` flags on
132
- `skill`). A missing mount source is now a **hard error** (`mount source(s) not found …`); set
133
- `COWORK_HARNESS_SOFT_MISSING=1` to fall back to warn-and-exclude. Mount names are always derived from
134
- the folder basename (collision-resolved); there is no `to:` override. See `references/scenario-schema.md`.
135
-
136
- > **`git add` a brand-new skill before testing it.** Inside a git repo the harness stages the
137
- > **git-tracked** files (the fidelity boundary — real Cowork installs from a repo and sees only committed
138
- > files). *Tracked* means **in the git index** (committed **or** `git add`-staged); the **content** staged
139
- > is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
140
- > commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
141
- > record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
142
- > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
143
- > installed" then did the work itself — a green-looking run where the skill never loaded. That now
144
- > **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
145
- > `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
146
- > to copy untracked files (won't reflect what ships). A folder **outside** any repo is copied raw (no guard).
147
-
148
- ### Choose a fidelity tier
149
-
150
- | Tier | What it gives you | Use when |
151
- |---|---|---|
152
- | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
153
- | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
154
- | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
155
- | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
156
-
157
- Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
158
- (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
159
- `references/fidelity-and-answers.md`.
160
-
161
- **Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
162
- `/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
163
- **remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
164
- different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
165
- reproduces it and none can — that container is not something a local tool can stand up. Which lane a
166
- real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
167
- So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
168
- asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
169
- `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
170
- rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
171
- above).
172
-
173
- ### Choose an answer path (gates: AskUserQuestion + tool-permission)
174
-
175
- Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
176
- live model into answering flags the run `nonDeterministic` — keep those out of deterministic
177
- regressions.
178
-
179
- <!-- answer-channels:begin -->
180
- **Pick by asking one question about your situation**, not by scanning a table — the channels are not
181
- interchangeable and the wrong one either masks a gate or can't run at all:
182
-
183
- ```
184
- Will this run be re-executed UNATTENDED? (CI, a committed cassette, --repeat, --matrix)
185
- │
186
- ├─ YES ──► scripted `answers:` / `--answer` / `--answer-policy` + `on_unanswered: fail`
187
- │ The ONLY reproducible channel. Non-negotiable for CI and committed cassettes.
188
- │ Labels reworded every run? STAY HERE: pin a stable leading SUBSTRING
189
- │ (uniqueness-guarded, fails loud) or a positional `choose`. Both keep determinism.
190
- │
191
- └─ NO — a discovery / validation run. Who holds the context to answer?
192
- │
193
- ├─ a model, steered by one line of intent
194
- │ ──► `--decider-llm --intent "<…>"` [skill · record]
195
- │ NOT on `run` — there the spelling is the scenario-YAML `on_unanswered: llm`.
196
- │ Can false-green an oracle-less semantic gate.
197
- │
198
- ├─ deterministic logic you can write down
199
- │ ──► `--decider-cmd '<helper>'` [skill · run]
200
- │ Determinism is your helper's, not the harness's. NOT on `record`.
201
- │
202
- ├─ YOU, the driving agent, holding the task context
203
- │ ──► `--decider-dir <FRESH, EMPTY dir>` [skill · run · record]
204
- │ + `cowork-harness gates <dir> --follow` (arm a Monitor here)
205
- │ + `cowork-harness answer <dir> --gate N --choose "<label>"`
206
- │ Its ONE unique property: it needs no advance knowledge of the option SET.
207
- │ (Label *text* drift alone does not need this — substring anchors handle that.)
208
- │
209
- └─ a human at a keyboard, and you are NOT producing a test
210
- ──► `cowork-harness chat` (TTY; no pass/fail verdict — see below)
211
- ```
212
-
213
- | Channel | Deterministic? | Don't use it when |
214
- |---|---|---|
215
- | Scripted | ✅ the CI/agent default | you cannot know the option set in advance |
216
- | `--decider-llm` / `on_unanswered: llm` | ❌ nonDeterministic | the gate has no oracle a model could judge |
217
- | `--decider-cmd` | delegated to your helper | the logic needs task context code doesn't have |
218
- | `--decider-dir` | ❌ nonDeterministic | nobody is present to drive it — it BLOCKS per gate |
219
- | `on_unanswered: first` | ❌ nonDeterministic | the answer matters — it *masks* the gate |
220
-
221
- **Cost of `--decider-dir`, stated plainly:** flags the run `nonDeterministic`; needs a live driver + a
222
- Monitor, so it is **unusable unattended**; blocks at each gate, strictly serial; needs a fresh empty dir
223
- per run (a dirty one is refused); rejected with `--repeat`, `--on-unanswered`, `--decider-cmd`, and with
224
- `--matrix --concurrency > 1`; and a cassette recorded this way carries a **re-record cost** — regenerating
225
- it needs the driver present again.
226
-
227
- **Rehearse it in ~2s before wiring it into a real run** — `cowork-harness decide --decider-dir <dir>` fires
228
- one sample gate through the same channel, then blocks (10-min backstop) until you answer it with the two
229
- commands above. It is the cheapest way to see the protocol work. Full recipe, including the multiSelect
230
- wire shape and the `gates --follow` Monitor loop:
231
- [`docs/decider-dir.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/decider-dir.md)
232
- (repo-only — an npm install ships it at `node_modules/cowork-harness/docs/decider-dir.md`). <!-- npm-only-ok -->
233
-
234
- **It is a FEEDER for the scripted default, not a rival.** `record --decider-dir` is a first-class way to
235
- *produce* a cassette: the non-reproducibility is spent once at authoring time and the cassette replays
236
- deterministically forever. The loop is **discover → transcribe → script** — answer live, then paste the
237
- run's echoed `--answer "<q>=<choice>"` footer lines into the scenario's `answers:` so re-records go back to
238
- being unattended. Skip the transcribe step only for one-off/exploratory runs.
239
- <!-- answer-channels:end -->
240
-
241
- **For a QUESTION gate, never hand-write the `req-N.json`/`resp-N.json` files.** `gates` and `answer` wrap
242
- the protocol — the atomic temp+rename, the `{id, answers}` envelope, the multiSelect array shape.
243
- Hand-rolling a Monitor over the raw files is the single most common mistake on this channel.
244
-
245
- `answer` writes `{id, answers}` and nothing else, so it covers **question gates only**. The channel also
246
- carries **permission**, **dialog** and **elicit** gates, whose replies need `{behavior}` / `{action}` — for
247
- those, write `resp-N.json` yourself, following the `reply_with` template the gate's own `req-N.json`
248
- advertises (it spells out the exact shape, e.g. `{"id":"…","behavior":"allow|deny"}`).
249
-
250
- Exact accepted values (teach precisely): `--on-unanswered` takes `fail|prompt|first` on `skill`,
251
- only `fail|first` on `run`. **`llm` is NOT an `--on-unanswered` value** — the bare flag
252
- `--on-unanswered llm` is rejected (use `--decider-llm`); the YAML spelling is `on_unanswered: llm`.
253
- The word `agent` is **retired** — do not write `on_unanswered: agent` (the schema rejects it).
254
- `--on-unanswered` also conflicts with `--decider-dir`/`--decider-cmd`/`--decider-llm` (the channel or
255
- model IS the terminal, so a policy alongside it never applies) — pass one, not both. On `record`, a
256
- scenario setting `on_unanswered: prompt` is rejected too: the YAML field outranks the flag, and a TTY
257
- wait can't produce a deterministic committed fixture.
258
- `--on-unanswered first` is itself flagged `nonDeterministic` — it is *not* a deterministic stand-in
259
- for scripted answers. See `references/fidelity-and-answers.md`.
260
-
261
- **Which gates to anchor (re-record robustness).** The model rewords option labels (and sometimes the
262
- question) every run, so a brittle exact-label `choose:` is itself a re-record-fragility source — it drifts and
263
- forces a re-record. The practical rule: **label-anchor only the gates whose choice drives an `assert:`** (or
264
- materially changes behavior); for gates whose answer is immaterial to your assertions, `on_unanswered: first`
265
- is the more re-record-robust choice — accept the `nonDeterministic` flag rather than trade it for a flaky
266
- anchor. (When label *order* is stable but the text drifts, a positional `choose` is the middle option — the
267
- linter flags positional `choose` as order-dependent, so use it deliberately.) The caution stands: `first`
268
- *masks* an unanswered gate, so don't use it for a gate you actually need answered a specific way.
269
-
270
- **Drifting label TEXT and an unknowable option SET are different problems — don't reach past the cheap
271
- fix.** Text that rewords while the choices stay the same is a *scripted* problem with a deterministic
272
- answer: a uniqueness-guarded leading substring, or a positional `choose`. Only when you cannot know what
273
- the options will *be* — they're generated per input document, so no anchor can be written in advance — does
274
- the answer move to a live channel (`--decider-dir` if you're driving, `--decider-llm` if nobody is).
275
-
276
- #### External deciders and the "first" shorthand
277
-
278
- When using `--decider-cmd` or `--decider-dir`, the helper's output is passed through
279
- `coerceLabel` **with the "first" shorthand disabled**. This means a helper that returns the literal
280
- string `"first"` must match an actual label named `"first"` — it is **not** coerced to option 1.
281
- This prevents a helper bug (accidentally emitting `"first"`) from silently green-ing option 1.
282
-
283
- The `"first"` shorthand remains active only for the built-in `--on-unanswered first` path. If you
284
- write an external helper, return a label name or option index — never the bare word `"first"` unless
285
- your gate actually has a label called `"first"`.
286
-
287
- ### Assertions: two orthogonal axes
288
-
289
- Conflating these is the **biggest landmine**. An assertion key has two independent properties:
290
-
291
- - **Axis A — robust to LLM phrasing drift?** Structural/boundary keys (`subagent_dispatched`,
292
- `egress_*`, `file_exists`, `user_visible_artifact`, `result`) are robust. Free-text content is
293
- not: match prose with `transcript_matches` / `transcript_contains` (stable lexical markers only —
294
- not semantic content the model paraphrases, which re-records red); check structured JSON with YAML
295
- `artifact_json` (or the [pytest lane](https://github.com/yaniv-golan/cowork-harness/blob/main/python/README.md) for complex predicates), not via a transcript substring.
296
- - **Axis B — survives `replay`?** *Independent of Axis A.* On the token-free `replay` lane, only
297
- **content keys** evaluate; filesystem / egress keys are skipped (live-only) — loudly, via an
298
- `::warning::` annotation, not a silent no-op. A key
299
- being "robust" says nothing about whether it runs on your replay gate.
300
-
301
- Getting Axis B wrong means a check that **does nothing in CI** — the harness warns loudly when it skips
302
- (an `::warning::` annotation, not a silent no-op — see the Axis B bullet above), and the bundled linter
303
- catches it before you push — run it (see *Scaffold a valid scenario, then lint before you push* below).
304
-
305
- See `references/scenario-schema.md` for the full assertion catalog with each key's replay class.
306
-
307
- #### Which assertion for which question (goal → key)
308
-
309
- Beyond the outcome/content keys most scenarios reach for first (`result`, `transcript_*`,
310
- `file_exists`/`user_visible_artifact`, `artifact_json`), the harness surfaces the agent's *behavior*
311
- — tool health, sub-agent work, panels, skill attribution, resources — as assertable keys. Reach for
312
- them by what you're trying to prove:
313
-
314
- | You want to check that… | Reach for |
106
+ ## Invariants — how a green run lies
107
+
108
+ Each of these has produced a green run that tested nothing. The full catalog, with the reasoning
109
+ behind each, is [`references/gotchas.md`](references/gotchas.md).
110
+
111
+ 1. **`result: success` is not "the task completed".** It means the agent didn't error. Assert the
112
+ deliverable (`file_exists` / `artifact_json` / `transcript_matches`). A `skill`-lane `PASS` only means
113
+ no guard fired: read `skillsInvoked`, `models` and `ablated` before concluding anything from it.
114
+ 2. **`replay` skips live-only keys.** Filesystem and egress keys are skipped on replay (loudly), so a
115
+ mixed item like `{result, egress_denied}` greens on its content half. Keep one concern per `assert:`
116
+ item, put live-only checks on a live gate, and run `cowork-harness lint`.
117
+ 3. **Only scripted answers reproduce.** Scripted `answers:` + `on_unanswered: fail` is the CI channel.
118
+ `first`, an LLM decider and `--decider-dir` all flag the run `nonDeterministic`, and `first` masks
119
+ the gate it answers.
120
+ 4. **`replay` evaluates the FROZEN scenario.** Editing `scenarios/*.yaml` changes nothing on a plain
121
+ `replay`: re-check an `assert:` edit with `replay --assert-from <file>`, and re-record for any other key.
122
+ 5. **Some keys pass on absence.** `gate_answers_delivered` passes when no gate fired — pair it with
123
+ `gate_answer_count_min: 1`. `tool_called` proves a tool ran, not that it was attempted.
124
+ 6. **An untracked skill mounts empty.** `git add` a new skill before testing it, and commit before
125
+ recording the cassette that locks it.
126
+ 7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
127
+ (`container` serves `mcp__workspace__web_fetch`, not `WebFetch`), and every tier models Cowork's
128
+ desktop-local lane only.
129
+ 8. **A WARN signal never blocks a green.** Read the verdict signals after every run
130
+ (`prompt_asset_missing`, `undelivered_deliverables`, `model_fallback`, …).
131
+
132
+ ## Short workflows
133
+
134
+ - **Author, then lock:** `cowork-harness scaffold --name … --prompt …` → `cowork-harness lint scenarios/` →
135
+ `cowork-harness record <file.yaml> --dry-run` (free) → `record` once, with `--out` at a tracked path
136
+ (a cassette cannot be moved) → `replay` on the PR gate.
137
+ - **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
138
+ `answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
139
+ - **Debug:** the triage table in `references/debugging.md` → `inspect`, `trace --view …`,
140
+ `verify-run`, `diff`. For a green you don't trust: `replay --explain`, then the gotchas.
141
+ - **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
142
+ arm yourself, with the model pinned and a recoverable source frozen.
143
+
144
+ ## References — where the detail lives
145
+
146
+ | File | Read it for |
315
147
  |---|---|
316
- | the skill didn't error out of a tool | `tool_no_error: <regex>`, `max_tool_errors: <N>` |
317
- | it didn't waste repeated identical calls | `max_redundant_tool_calls: <N>` |
318
- | a deliverable reached the user | `user_visible_artifact: <path>` (+ `no_scratchpad_leak: true` if it delivers via `present_files` — **`container` only**) |
319
- | an internal name/path did **not** leak into a delivered file | `artifact_text: {artifact, not_contains}` — `artifact_json`'s companion for non-JSON bodies; literal path, no glob, so one entry per delivered surface |
320
- | a named path must **not** exist after the run | `file_absent: <path>` (**live/verify-run only**) — do NOT invert `no_unexpected_files`: that is an allowlist over *newly created* files and needs a pre-run manifest |
321
- | a to-do workflow finished | `all_tasks_completed: true`, `task_status: {match, status}` |
322
- | a skill / connector / tool was **offered** | `skill_available`, `connector_available`, `tool_available` (all `<regex>`) |
323
- | a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
324
- | a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
325
- | a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
326
- | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
327
- | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
328
- | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
329
- | the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
330
- | the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
331
- | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
332
- | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
333
- | a context compaction happened | `compaction_occurred: true` |
334
-
335
- Every one of these still obeys the two axes above — several are live-only or need a `controlOut`
336
- cassette on replay, so check the catalog's replay class before putting one on a PR gate.
337
- `cowork-harness assertions --list` prints the full, always-current key set with one-line semantics
338
- straight from the schema — treat it (and the catalog) as the source of truth; this map is a
339
- goal-oriented index into it, not a second catalog.
340
-
341
- ### web_fetch (fail-closed, two-path)
342
-
343
- `web_fetch` behaves unlike `curl`. A URL is gated by **provenance**, not the egress allowlist:
344
-
345
- - A URL is *provenanced* iff it appeared in the **prompt** or a **prior `web_fetch` result**. To
346
- make a fetch succeed, put the URL in the prompt.
347
- - **Provenanced** → fetches (still SSRF-guarded per redirect hop); the egress hostname allowlist is
348
- **not consulted**.
349
- - **Not provenanced** → raises a per-domain approval gate (`webfetch:<domain>`) that is
350
- **fail-closed** (it is *not* auto-allowed; `--on-unanswered first` won't allow it). Answer it with
351
- a scripted rule (`when_tool: "webfetch:<domain>"` + `grant: domain|once`), a session
352
- `web_fetch.approved_domains`, or a live decider.
353
-
354
- Surprise to remember: adding a host to `egress.extra_allow` is a **no-op** for a provenanced fetch.
355
- Full model in `references/scenario-schema.md`.
356
-
357
- ### Scaffold a valid scenario, then lint before you push
358
-
359
- Don't hand-write the YAML from memory — that's how invented keys (`assertions:` vs `assert:`,
360
- `json_file`, `answer_policy`) creep in. Start from the bundled generator, which emits the
361
- known-good skeleton (right tier, scripted `answers:` + `on_unanswered: fail`, content assertions
362
- separated from live-only ones, one concern per item) and **self-lints its own output**. The
363
- generator is the bundled `scripts/scenario.py` — installed as a plugin, point `S` at
364
- `${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`; from a repo checkout, use the literal path below:
365
-
366
- ```bash
367
- S=".claude/skills/cowork-harness/scripts/scenario.py"
368
- python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
369
- --prompt "Generate the weekly report to outputs/report.md." \
370
- --content 'weekly report' --artifact outputs/report.md \
371
- --egress-allowed api.weather.example.com --out scenarios/report-check.yaml
372
- ```
373
-
374
- Then lint every scenario — it encodes the no-silent-false-green invariants. Use the CLI wrapper
375
- `cowork-harness lint`: it runs the bundled `scenario.py lint` **and** the harness's own scenario loader,
376
- so a file `run`/`record` would refuse fails lint too (running `scenario.py lint` directly skips the loader):
377
-
378
- ```bash
379
- cowork-harness lint scenarios/*.yaml
380
- ```
381
-
382
- `lint` flags: filesystem/egress-only assertions on a `replay` gate (silent no-op), bad regex
383
- quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `hostloop`/`protocol`
384
- (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
385
- baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
386
- `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
387
- unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
388
- `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
389
- to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
390
- `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
391
- and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
392
- (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
393
- emits a scenario `lint` would reject.
394
-
395
- **`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
396
- the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
397
- a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
398
- naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
399
- does not check what depends on the machine the run happens on (the session file and its mounts, an
400
- absolute `baseline:` path, environment variables). A session or matrix YAML in a linted directory is not
401
- a scenario and is reported as one that does not load — keep those out of the linted set. `python3
402
- scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
403
- `cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
404
- on a schema error; a directory reports each `✗ broken:` file and exits 1). **Read the exit code, not just
405
- its sign:** `record <file>` — with or without
406
- `--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
407
- pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
408
- as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
409
- but **`replay` does not**: a frozen top-level key it doesn't recognize (e.g. `lane:` recorded pre-1.16.0) is
410
- silently ignored and can flip a lane-sensitive verdict green; only frozen **assertion** keys stay
411
- hard-rejected there. Full split + the v11 version-regime:
412
- [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md#unknown-keys-the-loader-is-strict-lint-is-lenient).
413
-
414
- ## Part II — RUN, RECORD & LOCK
415
-
416
- You have an authored scenario. This Part runs it, reads the verdict, locks it into a
417
- byte-deterministic cassette, checks a background run's liveness, and places the assertions in the
418
- right CI lane.
419
-
420
- ### Run, then lock determinism
421
-
422
- Read the verdict and the inline failing transcript. To pin a flaky-because-stochastic gate, paste
423
- the echoed `--answer "<q>=<choice>"` footer lines back into the scenario's `answers:` for a
424
- deterministic re-run. Use `cowork-harness trace <id>` to digest a run. If only an *assertion* is wrong (the
425
- run itself was fine), `cowork-harness verify-run <run-dir> <scenario.yaml>` re-checks the `assert:` block against
426
- a **kept** run dir (`--keep`, or a `--session-id` run) with no live re-record — tokens-free, ~1s per iteration.
427
- When the scenario declares `answers:`, verify-run **also** checks they still match the run's actual gates (a
428
- reworded gate or a `choose:` the run never offered fails here in ~1s instead of on a paid re-record). Or skip
429
- the discovery/encode/record dance entirely and answer gates **live during the recording** with
430
- `record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
431
- `run` takes no `--dry-run`: to check that a scenario **loads** without spending, `cowork-harness lint
432
- <file.yaml>` runs the real loader (and resolves a named `baseline:`); `cowork-harness record <file.yaml>
433
- --dry-run` runs the real loader too AND the same scenario-level
434
- refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
435
- cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
436
- guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
437
- one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
438
- portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
439
- the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
440
- takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
441
- contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
442
- real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
443
- reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
444
-
445
- **Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
446
- concluded the free pre-flight was unavailable:
447
- - **"Does my whole corpus still load?"** → `cowork-harness lint scenarios/` answers it (every file the
448
- loader rejects is an ERROR, and so is a `baseline:` naming no shipped baseline), or the **directory** arm
449
- (`record scenarios/ --dry-run --quiet`, the CI shape in `references/ci-recipe.md`) when you also want the
450
- pre-spend refusals. The directory arm reports every offender in one pass, and the
451
- destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
452
- are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
453
- knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
454
- with no `prompt:` key reports as `· skipped:` rather than broken (so a renamed or mis-indented
455
- `prompt:` reads as "not a scenario" and the batch still exits 0), exit 1 covers **refused**
456
- (path-independent only — prompt policy, assert contradiction, duplicate target) as well as broken, and
457
- `--quiet` suppresses the advisory notes entirely — it keeps `✗ broken:` / `✗ refused:` / `· skipped:`,
458
- which is what you want in CI but means the notes are not a thing you will see there.
459
- - **"Would THIS record be refused?"** → the **single-file** arm **with the flags and `--out` the real
460
- record will get**. The destination it evaluates is `--out` if given, else `cassettes/<slug>.cassette.json`
461
- *relative to your cwd* — so previewing from the repo root a record that really runs from a subdirectory
462
- asks about a path that may not even exist, and a scenario whose cassette IS committed can come back
463
- refused. Point it at the real destination and the answer is binding.
464
-
465
- Neither question is answered by passing `--allow-host-inventory-fixture` to get past the refusal: that
466
- flag is consent for a recording you intend to make, and reaching for it as a load-check habit is how it
467
- stops meaning anything.
468
-
469
- **Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
470
- Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
471
- default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
472
- That choice is permanent: the cassette rewrites `scenario.session` and `scenarioSource` **relative to
473
- its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
474
- into another repo — leaves those unresolvable and
475
- `verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
476
- re-record at the new location — or point `replay`/`verify-cassettes` at the session with `--session <file>`, which resolves it without a re-record. Since 2.0.0 a bare `replay` FAILS on this class rather than warning. **`record` now says so BEFORE it spends:** a pre-flight — at the same
477
- pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
478
- warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
479
- outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
480
- cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
481
- missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
482
- (`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25 below).
483
- The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
484
- redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
485
- for a cassette that cannot verify staleness from its own location — recoverable only by passing
486
- `--session <file>` on every invocation thereafter.
487
-
488
- **Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
489
- scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
490
- (and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
491
- `events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
492
- `description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
493
- only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
494
- was never delivered — the views are a digest, and the record is wider (`jq` recipes in
495
- [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
496
- your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
497
- skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
498
- current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
499
- deliberately. (Same fail-closed family: corrupt gate evidence — unparseable `events.jsonl` lines, or fewer
500
- gates than `trace.json` recorded questions — a structurally invalid `result.json`, a `command:"replay"`
501
- result (a replay is a re-check of a recorded cassette, not run evidence — verify the original live run dir),
502
- and a `mode:"chat"` result (chat carries no assertions or verdict by contract) also refuse rather than
503
- certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
504
-
505
- Run artifacts are written to `~/.cowork-harness/runs/…` by default — **outside any working tree**, so a run
506
- launched from a repo root never drops sensitive skill inputs/outputs into it. Pass `--run-dir <path>` (or set
507
- `COWORK_HARNESS_RUNS_DIR`) to relocate; in CI point it at a workspace path so an artifact-upload step can
508
- collect the runs.
509
-
510
- #### Validate a skill against real documents (not a cassette)
511
-
512
- The loops above build **deterministic regressions**. A different job — drive a skill against *real* input
513
- documents to judge whether it actually does the work (extraction, analysis), with no intent to record a
514
- cassette — has its own recipe:
515
-
516
- 1. **Explore with the LLM decider.** `cowork-harness skill <dir> --decider-llm --intent "<one line of what
517
- this run is testing>"` lets a model (Sonnet default) answer each gate steered by your intent. The model replies with
518
- the option **number** and the harness maps it to the exact label (so it can't whiff by mis-typing the
519
- label text); an out-of-set answer fails loud. This is exploration, **not** a deterministic regression —
520
- the run is flagged non-deterministic and a green here is not a scripted pass. The answering model
521
- defaults to a Sonnet id (a weaker model tends to prose-decline an ambiguous judgment gate → fail-loud);
522
- override it with `--decider-model <id>` — a cheaper model (e.g. Haiku) for simple gates to cut cost,
523
- or Opus for the hardest judgment gates; it won't make an under-specified gate deterministic. A live
524
- decider can false-green a semantic assertion on an oracle-less gate — see `references/fidelity-and-answers.md`.
525
- 2. **Script the load-bearing gates — especially binary confirm gates.** Once you know which gates fire
526
- (`trace <run-dir> --view questions`), pin the ones whose choice drives the outcome with
527
- `--answer "<q>=<label>"` / `--answer-policy <yaml>`. When a skill **re-words its option labels run-to-run**
528
- (LLM-authored gates), pin a **stable leading substring** instead of the full label — `--answer
529
- "<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
530
- **fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
531
- CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
532
- 3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
533
- **salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
534
- missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
535
- 4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
536
- produced — the artifacts plus a shallow field preview of each JSON artifact (e.g. the extracted figures).
537
- It works on a salvaged partial run too. (A partial run is marked `PARTIAL`; `verify-run` and `scaffold`
538
- refuse to treat its half-finished output as a passing result.)
539
- 5. **For image-only / scanned PDFs, use the full-parity image.** The default agent image omits OCR and
540
- PDF-table tooling; if a **scenario** sets `requires_capabilities` (a scenario field — not skill
541
- frontmatter) and the image provably omits one, the harness **aborts before the paid run (exit 3)** —
542
- unless the scenario asserts `allow_missing_capability: true`, which downgrades it to a notice and
543
- proceeds. Rebuild with `--build-arg COWORK_FULL_PARITY=1` and point `COWORK_AGENT_IMAGE` at it for those
544
- skills.
545
- 6. **Iterate across fixes — verify before you trust, and don't cross-pair generations.** A green run is
546
- not a correct run, and a skill's self-reported finding is not real until its cited evidence is found in
547
- the run's own output. Ground each finding against `result.json` (`finalMessage` = the skill's own
548
- answer/critique; `toolResults` = tool outputs) and the tool-call stream via
549
- `cowork-harness trace <run-dir> --output-format json` — add `--full-results` so a successful call's full
550
- input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
551
- pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
552
- produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
553
- kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
554
- `run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
555
- version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
556
- generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
557
- `--label <tag>`); an un-split window spanning more than one generation now warns.
558
- **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
559
- skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
560
- critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
561
- pairing step there silently groups on an absent key instead of erroring — check the field is present, or
562
- require ≥ 1.5.0. See [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)
563
- (repo-only) for the full loop.
564
-
565
- #### Interpreting verdict signals
566
-
567
- The run verdict may include `WARN`-severity signals in addition to pass/fail. One to watch for:
568
-
569
- - **`prompt_asset_missing`** — the run proceeded but a prompt asset referenced by the scenario was
570
- not found. The model ran against an incomplete prompt. This is a `WARN`, not a hard failure, so
571
- the run can still green. If you see it, fix the asset path — a green with a missing asset is
572
- not a valid pass.
573
-
574
- **False negatives — signals that are tier/image artifacts, not skill defects.** Some fail-severity
575
- signals read like a skill gap but are really a property of the reduced test image or the fidelity tier.
576
- Recognize these before "fixing" a non-bug:
577
-
578
- - **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
579
- rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
580
- `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
581
- (`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
582
- Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
583
- behind that sentence). The message says so ("likely a FALSE
584
- NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
585
- `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
586
- `allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
587
- or a declared `requires_capabilities` the tier can't provide, both lanes — an unknown family name
588
- hard-fails rather than silently passing.) **On an open-ended `skill` run** (no `assert:` block to carry
589
- the modifier), pass **`--allow-missing-capability`** — the CLI equivalent of the assertion.
590
- - **`ended_with_question`** (`WARN`, live lane) — a heuristic: the agent's final answer contains a
591
- question and the run wrote **no deliverable to `outputs/`** — it may have ended on a request for input
592
- instead of finishing. Warn-only; the fix is scripting/steering the answer (`answer:` / `--answer` / a
593
- decider, or `--decider-llm --intent`), not editing the skill's prose. The strict, fail-severity sibling
594
- `stalled` already catches a *trailing*-`?` final turn that did no tool work after the last gate; this
595
- covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
596
- the final message before acting — a legitimate question-posing answer that wrote a file never fires.
597
- Assert `allow_stall: true` if ending on a question is the intended terminal state.
598
- - **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
599
- and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
600
- they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
601
- does not get them. It fires with no assertion written — `present_files_called` covers the positive case
602
- only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
603
- **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
604
- scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
605
- **The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
606
- can see them, and give the file tools an **absolute** path under the outputs directory the agent's prompt
607
- names. On the desktop-local host-loop lane (what production runs), against Desktop **2.7032.0 and later**,
608
- the agent process runs outside the session (`/var/empty`), so a relative `Read`/`Write`/`Edit` — a bare
609
- filename or `outputs/x.md` alike — is **refused** ("File is in a directory that is denied by your
610
- permission settings."); only a pathless or relative `Grep`/`Glob` is redirected to outputs. (Before
611
- 2.7032.0 the file tools were rooted at `outputs/`, so a bare filename landed there and `outputs/x.md`
612
- doubled to `outputs/outputs/x.md`.) At `fidelity: container`/`microvm` (VM-loop, the harness default) the
613
- base is the session root, so a bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an
614
- explicit delivery. Addressing
615
- a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
616
- decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
617
- [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
618
- nothing is delivered by location there, so only an explicit delivery counts. Assert
619
- **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
620
- downloaded inputs) rather than a delivery gap.
621
- - **`delivery_unobservable`** (`WARN`, `lane: remote` only) — the run produced file(s) but the harness
622
- serves **no delivery tool on that lane**, so whether they reached the user is unanswerable. This is the
623
- honest cannot-verify companion to `undelivered_deliverables`: reporting every remote file as undelivered
624
- would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
625
- with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
626
- a harness coverage gap (see the *File delivery* section of fidelity-gaps).
627
- - **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
628
- `model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
629
- this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
630
- transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
631
- by a different model than the scenario names, so treat a green as evidence about the fallback model.
632
-
633
- - **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
634
- connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
635
- approval, not just outputs — a connected folder shows the identical default — so this run diverged from
636
- what production would have allowed. `WARN` rather than `FAIL` because the harness **detects** post-hoc
637
- what production **enforces**: by the time the scan sees it, the agent already proceeded where it would
638
- have hit `EPERM`, so failing the run would overstate what a post-hoc scan knows. Author
639
- `no_delete_in_mounts: true` to hard-fail on it, or `allow_delete_in: ["<mount>"]` to waive that mount
640
- (detection still runs and the hit is still recorded — the waiver is a verdict decision).
641
- - **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
642
- paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
643
- `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
644
- `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
645
- run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
646
- it's valid.
647
- - **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
648
- infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
649
- rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
650
- the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
651
- a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
652
- `Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
653
- failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
654
- suspiciously empty.
655
- - **`outputs_delete_unconfirmed`** (`WARN`) — a delete-shaped command near `mnt/outputs` that nothing
656
- confirms: the per-turn filesystem diff shows no output present at turn start was deleted, and no flagged
657
- delete has an `outputs/` path as its own operand. The classic case is a Python variable named `rm` in a
658
- `python3 -c` body that also reads a report from outputs: `rm = json.load(open(".../outputs/r.json"))`.
659
- A file the turn created and then deleted is invisible to the diff, so real deletes land here too:
660
- - a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`);
661
- - a `cd` then a relative path;
662
- - chained variables (`A=…; B=$A/x; rm "$B"`);
663
- - a Python path held in a variable set on another line (`p = …` then `os.remove(p)`);
664
- - wrapper flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`);
665
- - calls outside the modelled set, such as Node's `fs.promises.rm(…)`.
666
-
667
- Read the command before dismissing it. A literal-path delete (`rm -f mnt/outputs/x`,
668
- `os.remove(".../outputs/x")`) still fails `outputs_delete`. So do two non-deletes: quoted text where a
669
- delete command with an outputs operand follows a separator, subshell or keyword
670
- (`echo 'note; rm mnt/outputs/x'` — the classifier does not track quotes), and a heredoc that *writes* a
671
- script instead of running it. A statement over 4 KiB or a command over 16 KiB is judged by the stricter
672
- original rule, so a huge one-line body with a variable named `rm` fails again, and a command whose variable
673
- expansion would exceed the scanner's work budget (about a hundred distinct variables in one 10 KB line) is
674
- not expanded — every mount it names literally counts as deleted in. Waive any of these with
675
- `allow_outputs_delete`. The warn is raised even when `no_delete_in_outputs` is authored (the assertion
676
- passes; this warn is how the hit stays visible in text output).
677
- - **`outputs_diff_unavailable`** (`WARN`) — the outputs filesystem diff could not verify this turn and the
678
- text scan saw nothing, so a delete by a script file or a non-bash tool would have gone unseen.
679
- - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
680
- `RunResult.scan` is undefined and the host-path guard and the outputs-delete **text scan did not run this
681
- run** (the outputs filesystem diff still did, and a delete it proves still fails). Not a
682
- pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
683
-
684
- The full 22-code signal table (severity + per-signal opt-out) is in
685
- [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
686
- the fuller narrative.
687
-
688
- ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
689
-
690
- A single green proves the run passed **once**. Two questions need more than that, and both have a
691
- discipline that is cheap to follow and expensive to skip.
692
-
693
- **"Did it pass, or pass once?"** → `--repeat N` (2-100, on `skill` AND `run`) samples the same
694
- skill+prompt N times and prints a variance rollup instead of a single verdict. `--min-pass-rate` sets
695
- the batch threshold, `--stop-on-diverge` stops the moment flakiness is proven, `--max-budget-usd` caps
696
- spend.
697
-
698
- **"Does the skill actually help?"** → `--ablate-skill` runs the prompt with every skill/plugin
699
- discovery source removed, so the agent answers from its own priors. **It is ONE arm, not a paired
700
- experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
701
- the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
702
- and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
703
- The rollup says so on its verdict line: `repeat "<skill>": PASS [ABLATED — control arm] — 5/5 passed`.
704
- Every ablated run is stamped `ablated: true` in `result.json` and carries `ablated=true` on its
705
- `[provenance]` footer line; a run that isn't stamped is a real run.
706
- What the harness gives you here is the run execution and the control arm — designing the comparison
707
- (scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
708
-
709
- **Measurement hygiene — four things that silently invalidate a batch:**
710
-
711
- 1. **Pin the model.** With no `model:` in the session (or `--model` on the `skill` lane) the run uses
712
- whatever the staged agent binary defaults to — not a harness constant, and it can move under a
713
- baseline bump. Read `result.json`'s `models` back before believing any cross-run comparison — and when
714
- you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
715
- fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
716
- array purely by whether such a turn occurred.
717
- 2. **Commit the skill first.** `fingerprint.skillHash` is content-exact, so an edit mid-batch silently
718
- splits your dataset into two generations — and a hash whose source was never committed identifies a
719
- generation that is unrecoverable. `stats --group-by skill-hash` separates them after the fact;
720
- nothing recovers the source.
721
- 3. **Check which arm you actually ran** before analysing anything: `ablated` and
722
- `context.availableSkills` in each `result.json`.
723
- 4. **Read `skillsInvoked`.** A rep where the skill never triggered is a measurement of the model, not
724
- of your skill — discard or re-run it.
725
-
726
- ### Checking whether a background run is alive
727
-
728
- Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
729
- processes in your OWN PID namespace, which is frequently NOT the harness process's namespace (e.g. when
730
- you're a sandboxed subagent). An empty `ps aux` match tells you nothing about whether the run is still
731
- going.
732
-
733
- Use **`cowork-harness status <dir> [--follow]`** instead — reads `<outDir>/status.json`, a file the
734
- harness writes/updates throughout the run's lifecycle (including a crash-safety net for a thrown
735
- error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
736
- either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
737
- checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
738
- run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
739
- which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
740
- modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
741
- `<dir>` also accepts the run-dir root
742
- passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
743
- newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
744
- rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
745
- payload; `cowork-harness status --help` has the flags.)
746
-
747
- **Poll with `--follow`, not with a shell loop over `status`'s stdout.** The one-shot text form prints to
748
- **stderr** and writes nothing to stdout; `--output-format json` (one envelope) and `--follow` (one JSON
749
- line per status change) are the **stdout** forms. A poll that greps `status`'s stdout therefore matches
750
- nothing, exits 1, and returns instantly against a run with minutes left to go — a silent false "done":
751
-
752
- ```bash
753
- # WRONG — stdout is empty, so grep exits 1, `!` inverts it, and the loop never sleeps.
754
- until ! cowork-harness status "$D" | grep -q '● running'; do sleep 30; done
755
-
756
- # RIGHT — the harness owns the poll loop and exits when the run reaches a terminal state.
757
- cowork-harness status "$D" --follow
758
- ```
759
-
760
- **A multi-minute `record`/`run` outlives a short-lived wrapper.** Don't launch a long record from a
761
- subagent that returns before it finishes — the returning agent tears down its process tree and kills the
762
- in-flight run mid-artifact-write. Run it foreground, or detached from any process that will exit first.
763
- (The `status.json` liveness above is exactly what surfaces such a teardown as `"error"`/`stale` rather
764
- than a stuck `"running"`.)
765
-
766
- ### Place assertions in the right CI lane
767
-
768
- CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
769
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
770
- PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
771
- the four-stage pipeline.
772
-
773
- ## Part III — Debug
774
-
775
- A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
776
- afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
777
- re-run and hope. Start at the triage below, then use the observability output and, when you need to
778
- reproduce interactively, `chat`.
779
-
780
- > **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
781
- > `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
782
- > narrower record — `critique-evidence-package.txt`, what a grade was actually computed against — none of
783
- > the five tools above surface it; see `references/critique.md`.
784
-
785
- ### Triage — a run misbehaved, or a green looks wrong
786
-
787
- <!-- BEGIN triage-canonical -->
788
- Two situations need different tools — figure out which one you're in first, then reach for the tool
789
- instead of re-running and hoping. The run already wrote its evidence to a kept run dir (`--keep` prints
790
- the path; `trace <run-id>` finds it), and every tool below reads that evidence **token-free** — no
791
- Docker, no re-record.
792
-
793
- | Situation | Symptom | Reach for (in order) |
794
- |---|---|---|
795
- | **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
796
- | **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `replay --mutate` — perturbs a CAPPED SAMPLE of recorded JSON values (10/file, 50 total) and reports which perturbations NOTHING caught; the report names the sample size, so read it as a sample not a total (reporting only; never moves the verdict/exit code) · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
797
-
798
- A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
799
- agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
800
- already does.
801
- <!-- END triage-canonical -->
802
-
803
- **Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
804
- **deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
805
- so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
806
- `references/fidelity-and-answers.md` (shipped); the specific deltas vs. real Cowork and the sandbox
807
- boundary model live in [`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md) / [`docs/boundary.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/boundary.md) (repo-only, not in the installed
808
- payload). If the behavior is on that gap list, it's expected — stop debugging your skill.
809
-
810
- ### Inspecting a run's observability output
811
-
812
- A verdict is only the top of what a run records, and the run dir persists after the verdict
813
- (`~/.cowork-harness/runs/…`). Beyond pass/fail, every `run`/`skill`/`chat` writes a `result.json` and a
814
- trace you read back without a re-record — the debugging loop is *localize the failure from that
815
- already-written evidence*, not re-run-and-hope. Use them to diagnose a failure (and, secondarily, to
816
- decide which assertions from *Assertions: two orthogonal axes* are worth adding):
817
-
818
- - **`cowork-harness trace <run-dir> --view <view>`** — focuses one of the run's rollups (the per-tool
819
- call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
820
- bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
821
- the current list rather than relying on a fixed enumeration here.
822
- - **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
823
- is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
824
- cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
825
- location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
826
- session end), `present_files` is NOT served, and `user_visible_artifact` /
827
- `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
828
- to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
829
- `lane: remote` scenario still runs locally.
830
- - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
831
- `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
832
- plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
833
- `--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
834
- tier — instead of aggregating across them (a window spanning >1 generation warns, so an un-split A/B
835
- average announces itself rather than passing as one number;
836
- >1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
837
- individual runs behind each summary with their `skillHash`/`runLabel`, so
838
- you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
839
- - **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
840
- contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
841
- walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
842
- `total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
843
- `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
844
- `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
845
- `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
846
- `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
847
- `context` (tools/mcpServers/availableSkills), `tasks`,
848
- `workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
849
- (`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
850
- evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
851
- `gateProvenance` (per-gate `scripted`/`decided(llm|external)`/`first-option`/`prompt` with a
852
- `bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
853
- `egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
854
- cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
855
- `outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
856
- `jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
857
- `{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
858
- semantics: [`docs/cli.md` → What you get out](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md#what-you-get-out-inspectable-output) (repo-only); [`schema/run-result.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/run-result.json) is the
859
- machine source.)
860
- - **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
861
- **`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
862
- a re-record rarely tells you more than the captured stderr already does. Also check
863
- **`resultErrorKind`** (`"transport" | "agent" | "usage_limit"`) before spending another paid run: a
864
- `"usage_limit"` failure is a quota exhaustion, not a skill bug — retry after the limit resets rather
865
- than debugging; `"transport"`/`"agent"` means something actually broke, worth localizing before
866
- re-running.
867
- - **Attributing cost to sub-agent work.** `subagents[]` gives the dispatch tree — each sub-agent's
868
- `dispatchModel`/`resolvedModel`, `toolsUsed`, `prompt`/`output`, and `attributedSkillId` — but **not** its own token/cost;
869
- aggregate cost is per-**model** in `modelUsage` (and `trace --view usage`), not per-sub-agent. So a
870
- cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
871
- per-model usage — the harness doesn't line-item each sub-agent's tokens.
872
- - **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
873
- `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
874
- `trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
875
- **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
876
- **absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
877
- collection) is evidence **UNAVAILABLE**, not an empty run — `trace --view files` reports a loud
878
- UNAVAILABLE marker (`workspaceFilesRecorded: false` in JSON, and no phantom "removed" diff rows) and
879
- `inspect` prints `artifacts: UNAVAILABLE` (`artifactsRecorded: false`) instead of `artifacts (0):`.
880
-
881
- ### Debugging with `chat`
882
-
883
- `cowork-harness chat` opens an interactive multi-turn REPL against a live Cowork session. It is
884
- **not** an asserted test — no `assert:` block, no cassette. Use it to explore behavior, reproduce a
885
- bug interactively, or test a prompt before committing it to a scenario.
886
-
887
- Each session still writes an informational `result.json` (`mode: "chat"`, no `assertions`) plus a
888
- trace and index row under its run dir — the same telemetry (tool durations, model usage, resources,
889
- etc.) that `run`/`skill` produce — so `cowork-harness trace <chat-run-dir>` / `stats` work on a chat
890
- session too, even though it never yields a verdict.
891
-
892
- **`--plugin <dir>` flag (repeatable).** Load additional skill folders alongside the primary session
893
- plugin. Each `--plugin <dir>` appends the folder to `local_plugins`. Useful when the skill-under-test
894
- depends on a sibling plugin:
895
-
896
- ```bash
897
- cowork-harness chat ./skills/report-gen --plugin ./skills/shared-utils
898
- ```
899
-
900
- **Note:** `--raw` mode (native `docker run -it`) can't honor the harness-managed flags, so `--upload`,
901
- `--folder`, `--plugin`, and `--fidelity` are **rejected** with a usage error if combined with `--raw`;
902
- only `--model` is carried through.
903
-
904
- **`/help` in the REPL.** Type `/help` at the prompt to see available commands:
905
-
906
- ```
907
- Commands: /exit /quit /help
908
- ```
909
-
910
- The startup banner now reads `type your message (/help for commands)` as a reminder. `/exit` and
911
- `/quit` both terminate the session.
912
-
913
- ## Gotchas — the "✓ passed ≠ correct" landmines
914
-
915
- Stated as *symptom → why → fix*. This is the **workflow/record/answer-path** view — the broader of the
916
- two lists, but **not a strict superset**: `references/scenario-schema.md`'s *Authoring gotcha list*
917
- carries a few assertion-level landmines this one omits (`transcript_no_host_path`'s scan width,
918
- `egress.extra_allow`'s no-op on the provenanced `web_fetch` path, `replay_protocol_fidelity` not being
919
- authorable). Reach for this list when debugging a run's behavior, that one while authoring `assert:`.
920
- **The two lists are numbered independently** — a bare "gotcha N" means the list you are reading.
921
-
922
- 1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
923
- `replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
924
- `artifact_text`, `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
925
- *mixed* item like
926
- `{result, egress_denied}` greens on `result` while its `egress_denied` half is dropped. (`record`
927
- snapshots an `artifacts` manifest, which makes
928
- `file_exists`/`user_visible_artifact`/`artifact_json`/`artifact_text`/`computer_links_resolve`
929
- replay-checkable — but the live-only egress keys stay skipped, and `file_absent` is never
930
- replay-checkable at all: proving absence needs an exhaustive, healthy walk a manifest does not record.) *Fix:* put egress/live-only checks on
931
- a live gate; keep one concern per `assert:` item; run the linter. The harness warns loudly on skip.
932
-
933
- 2. **A steered gate answer never reached the model.** *Why:* `serializeDecision` must emit
934
- `updatedInput: { questions, answers }`; a header-only gate (empty `question`) can never be keyed.
935
- *Fix:* give every gate a non-empty `question`. (multiSelect gates ARE supported on **every** answer
936
- channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
937
- reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
938
- Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
939
- string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
940
- `questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
941
- old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
942
- on unobserved delivery (absence of evidence is failure, not neutral).
943
-
944
- 3. **A multi-key `assert:` item is an AND.** A single list item with more than one key passes iff
945
- **every** key passes. *Fix:* one concern per item unless you genuinely mean conjunction (and a
946
- mixed-class conjunction still loses its filesystem half on replay — see gotcha 1 below).
947
-
948
- 4. **`tool_called` doesn't mean "attempted".** Tool counts are authoritative and de-duped: a tool
949
- that was *requested then denied* does **not** register as called. *Fix:* don't assert `tool_called`
950
- to prove an attempt; it proves the tool actually ran.
951
-
952
- 5. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
953
- sub-agent used other tools. `subagent_dispatched` / `subagent_output_contains` match on dispatch
954
- type (`dispatchAgentType`), the binary-*resolved* type (`resolvedAgentType`), *or* the dispatch
955
- **description** — so a type-less dispatch that resolved to e.g. `general-purpose` is still
956
- selectable, by either the resolved type or the description. A `Task` dispatch that carries NO
957
- `subagent_type` at all falls back to the built-in `general-purpose` agent with a **wildcard tool
958
- surface** (`tools:["*"]`, including workspace bash) — faithful production behavior, and it fires
959
- routinely. The harness warns loudly on this fallback and records `subagents[].dispatchTypeOmitted`;
960
- an *explicit* `subagent_type: "general-purpose"` is a deliberate author choice and does not warn.
961
- Implication: `subagent_tool_absent` on a type-less dispatch is weaker evidence (wildcard surface) —
962
- pin `subagent_type` explicitly when you need a tight tool-absence guarantee.
963
-
964
- **Cross-tier "no shell" caveat.** On `hostloop`, native `Bash` calls route through the
965
- `mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
966
- `mcp__workspace__*` to hold across every tier.
967
-
968
- 6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
969
- reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
970
- run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
971
- per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
972
- (gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
973
- — repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
974
- `dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
975
- a fan-out you don't want.
976
-
977
- 7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
978
- assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
979
-
980
- 8. **Read-only mounts are enforced; delete-deny is a HARNESS gap — production DOES enforce it.**
981
- `mode:r` mounts get a real `:ro` bind (a write fails in-guest). But `rw` vs `rwd`
982
- (write-but-no-delete on `outputs/` / connected folders) is *not* mount-enforced **in the harness** —
983
- `rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`. **Real Cowork enforces it live:**
984
- outputs is a FUSE mount, and `unlink`/`rmdir` fail `Operation not permitted`; a skill must request
985
- approval via `allow_cowork_file_delete` (which re-mounts the folder `rwd` mid-session) to delete.
986
- **Only unlinking is denied.** Emptying a file in place — `truncate -s 0`, `> file`, `shred` without
987
- `-u` — and renaming *within* outputs both SUCCEED in production, so the harness does not flag them
988
- either. Renaming a file OUT of outputs fails (`EXDEV`, then `EPERM` on the copy-then-unlink
989
- fallback), so that stays a delete. Two consequences: a skill should not stage disposable scratch
990
- under `outputs/` (in production, cleanup there costs an approval prompt), and a skill's
991
- "catch-EPERM-then-request-approval" branch cannot be exercised at any harness tier (the `rm` just
992
- succeeds here). Do not read this gotcha as "delete-deny may not be real in production" — it is real.
993
- If a scenario's deletion IS intended, assert `allow_outputs_delete: true` rather than dropping
994
- `no_delete_in_outputs` — omitting it does not permit anything.
995
-
996
- 9. **Keep `.env` out of any mounted folder** — it is copied into the sandbox and the token could
997
- leak. Put it at a working-dir or install root (token resolution: env > `--dotenv` > `./.env` >
998
- install `.env`). **Inverse footgun — running from a git worktree:** a worktree's `./.env` is gitignored, so
999
- it's **absent** there and you'll get "no model credentials." *Fix:* pass `--dotenv <main-checkout>/.env`
1000
- (or set the env var) — that's exactly what `--dotenv` is for.
1001
-
1002
- 10. **A base64 artifact that was scrubbed at record time will fail artifact assertions at replay.**
1003
- When `record` detects a secret embedded in a base64 artifact, it replaces the entire artifact
1004
- body with `[REDACTED:base64]` and emits a `::warning::`. Any `artifact_json` or content
1005
- assertion targeting that artifact will fail at replay because the body no longer matches. *Fix:*
1006
- do not let secrets flow into artifacts; if the artifact is intentionally opaque, drop the
1007
- content assertion and gate on `file_exists` on the live lane instead.
1008
-
1009
- 11. **An external decider returning `"first"` does not select option 1.** The `"first"` keyword
1010
- shorthand is disabled for `--decider-cmd` / `--decider-dir` helpers (see *Choose an answer path*
1011
- → External deciders). If your helper
1012
- accidentally emits `"first"` and no label named `"first"` exists, the gate fails — it does
1013
- **not** silently pick the first option. This is intentional: a helper bug should fail loud, not
1014
- green wrong. *Fix:* have helpers return a label name or numeric index.
1015
-
1016
- 12. **`prompt_asset_missing` is a WARN, not a hard failure — greens can hide it.** The
1017
- `prompt_asset_missing` verdict signal (see *Interpreting verdict signals*) does not block a green verdict. Scan the verdict
1018
- signals section after every run; a run that greened with this signal ran against an incomplete
1019
- prompt. *Fix:* treat `prompt_asset_missing` as a blocking error in CI by checking the signals
1020
- array.
1021
- 13. **`result: success` means the agent didn't error, NOT that the task completed — always assert on
1022
- artifacts/content.**
1023
- - A turn that ends on a plain-text re-ask ("which file did you mean?") still reports
1024
- `result: success`.
1025
- - The harness catches this with a **`stalled`** verdict signal: a run that ends on a question and
1026
- did **no productive work after its last gate** — both the no-gate case ("which file?" with no
1027
- tool calls) AND the *answered-gate-then-re-ask* case (the agent answers an `AskUserQuestion`,
1028
- then asks again in plain text and stops). Suppress with `allow_stall: true` if ending on a
1029
- question is intended.
1030
- - The signal is a **tool-position heuristic**, not deliverable detection, so it is imprecise both
1031
- ways:
1032
- - **False negative:** a post-gate tool *call* clears the flag whether it **succeeded or
1033
- errored** — an agent that ran a tool after the gate and still stalled is not caught.
1034
- - **False positive:** a deliverable written *before* a final confirmation gate does **not**
1035
- clear it, so a write-then-confirm-then-question run is flagged — use `allow_stall: true` for a
1036
- deliberate confirm-terminal skill.
1037
- - The broad guard is therefore YOUR assertions — assert the deliverable (`file_exists` /
1038
- `artifact_json` / `transcript_matches`), never just `result: success`.
1039
- - `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
1040
- stalling *after* one is answered — two different failure modes.
1041
- - **Free-text aside:** the scripted key for a "type-it-in-notes" option is **`answer:`** — an
1042
- arbitrary string delivered verbatim, bypassing label validation by author intent (Cowork
1043
- auto-provides an "Other" free-text path on every gate). Mutually exclusive with `choose:`; setting
1044
- both fails loud. What has no scripted equivalent is the `OTHER:` *directive*
1045
- (it works only on the LLM-decider path, not scripted `choose:`, and only on
1046
- **single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
1047
- options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
1048
- LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
1049
- `OTHER:` is marked `[via Other free-text]` in its `gateProvenance` rationale.
1050
- 14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
1051
- but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
1052
- index lands on a different option (a silent re-record flake). Prefer an exact label when order is
1053
- stable; `lint` flags positional `choose` with an advisory. Unstable option order is also what the
1054
- **user** sees — a reordered gate puts a different choice in the default slot — so pin what was shown
1055
- with `question_options`, rather than only hardening the answer rule against it.
1056
- 15. **A scripted `choose:` matching no offered option HARD-fails the run — `on_unanswered: first` does NOT
1057
- backstop it.** This is distinct from an *unanswered* gate (no rule matched → falls to `on_unanswered`): a
1058
- rule that DID match the gate but whose `choose:` names a label the gate never offered (the model reworded
1059
- it) is treated as an authoring bug and fails loud — `first`/`llm` won't absorb it. The error now prints the
1060
- **offered options** (and a closest-match suggestion), so fix the anchor from the error alone — no need to
1061
- dig through `events.jsonl`. (This is exactly the drift `verify-run` answer-coverage catches in ~1s; use it
1062
- before a paid record.)
1063
- 16. **Batch record keeps going — you don't need a one-at-a-time wrapper.** `record <dir>` and `record <dir>
1064
- --rerecord-stale` run **every** scenario, collect failures, and report them at the end (non-zero exit on
1065
- any failure) — a failing scenario does NOT abort the batch. So a single `cowork-harness record cassettes/
1066
- --rerecord-stale` surfaces ALL stale anchors in one pass (add `--concurrency <N>` to parallelize); a shell
1067
- wrapper that loops one cassette at a time with `set -e` defeats this and rediscovers stale anchors serially.
1068
- Two durability properties make the batch safe to trust: each cassette is written **atomically** (a
1069
- same-directory temp file + rename), so an interrupted or OOM-killed batch never leaves a partial/corrupt
1070
- cassette — a failed scenario simply produces none; and under `--concurrency <N>` each scenario runs **fully
1071
- isolated** (its own egress sidecar network + proxy, its own per-session run dir), so parallel records don't
1072
- cross-talk — the concurrency bound exists only for the Docker address pool + API rate limits, not correctness.
1073
-
1074
- 17. **Editing `scenarios/*.yaml` does NOT change a plain `replay` — the WHOLE scenario is frozen, not just
1075
- `assert:`.** *Why:* a cassette captures every key (`lane:`, `fidelity:`, `baseline:`, `prompt:`, `skills:` …)
1076
- and `replay` evaluates all of them from that frozen copy — byte-deterministic, ignoring the working tree (so
1077
- a committed cassette can't silently re-interpret against an uncommitted YAML). **Only `assert:`
1078
- (+`expect_denied:`) can be opted back to disk; any other edited key reaches a replay only by re-recording.**
1079
- This is *loud* rather than a *silent* no-op: plain `replay` prints a `::notice::` when a sibling's
1080
- `assert:`/`prompt:` differs, and when the sibling **fails to load** at all (a typo'd or too-new key) — and
1081
- points you at the fix.
1082
- *Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
1083
- `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
1084
- `prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
1085
- fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
1086
- are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the `session`
1087
- (model / data mounts / discovery) is NOT drift-checked or fingerprinted, so a **model change** between record
1088
- and re-assert is undetected — the notice flags this; re-record if the session changed. `verify-run` reads
1089
- on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
1090
-
1091
- 18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
1092
- bundle several sub-questions into a single gate; the assertion counts each sub-question, so a
1093
- 3-sub-question bundle counts as 3, not 1. `trace --view questions` shows the same per-gate
1094
- sub-question count and a matching footer total — read that off instead of the tool-call count when
1095
- sizing the budget.
1096
-
1097
- 19. **`gate_answers_delivered` passes vacuously when no gate fires — pair it, or drop it.** Whether a
1098
- gate fires is model-dependent, so `gate_answers_delivered: true` alone can't catch "the gate never
1099
- fired at all". If the scenario is meant to gate, pair it with `gate_answer_count_min: 1` (a floor of
1100
- `0` witnesses nothing). If the scenario is gate-clean by design, drop the key — it asserts nothing
1101
- there — and declare `questions_count_max: 0`, which fails loudly if a gate ever appears. Asserting
1102
- `questions_count_max: 0` alongside a gate-presence key is unsatisfiable: `run`/`skill`/`record`
1103
- refuse it before spending.
1104
-
1105
- 20. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
1106
- read-only folder's files as path + hash only (`truncated: true`, no `body`) — it's an input the agent
1107
- read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass against it on replay
1108
- (the hash-only entry still materializes a placeholder); `artifact_json`/`artifact_text` report a clear
1109
- evidence-unavailable on every lane (live/verify-run/replay agree — no green-record/red-replay). This is
1110
- also why a `mode: r` input never trips the `binary` privacy finding or needs `--allow` — only a
1111
- *committed* body is scanned. `scaffold` won't emit `file_exists` for one either (it's not in
1112
- `RunResult.artifacts`). A `mode: rw`/`rwd` folder's contents are captured with a full body, same as
1113
- `outputs/`.
1114
-
1115
- 21. **A `fidelity: cowork` cassette can go stale in a way `skill`/`format` drift won't catch.** Its recorded
1116
- `effectiveFidelity` field pins which concrete tier (`hostloop` or `container`) the baseline resolved to
1117
- AT RECORD TIME. If a later Desktop baseline flips that resolution, `verify-cassettes` reports it as a
1118
- `resolved-tier` finding (re-record — the recording now exercises the wrong tier); a cassette with no
1119
- `effectiveFidelity` at all, or an unloadable pinned `baseline:`, reports `unverifiable-tier` instead
1120
- (couldn't check — also re-record). Both are `fidelity: cowork`-only; an explicit-tier scenario never
1121
- produces them. (Details: [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § tier staleness — repo-only.)
1122
-
1123
- 22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
1124
- `manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
1125
- assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
1126
- yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
1127
- fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
1128
- suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
1129
- rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
1130
- about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
1131
- `lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
1132
- `--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
1133
- 23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
1134
- *Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
1135
- servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
1136
- fail your gate. *Fix:* nothing, unless the scenario asserts `tool_available` on
1137
- `mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
1138
- re-recording would never produce those tools anyway.
1139
-
1140
- 24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
1141
- lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
1142
- emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
1143
- Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
1144
- optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
1145
- other — and probing a remote session makes this harness look like it emulates the wrong tool under the
1146
- wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
1147
- file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
1148
- `present_files_called` assertion keys are harness-side names and stay valid either way.
1149
- ([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
1150
- → "File delivery" has the binary-verified detail; repo-only.)
1151
-
1152
- 25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
1153
- --allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
1154
- (`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
1155
- `record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
1156
- fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
1157
- scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
1158
- pass it. Writing a recording the scan DID flag is the separate `record
1159
- --allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
1160
- exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
1161
- destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
1162
- before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
1163
- **quarantined** to `<runs-root>/quarantine/` with a `.findings.txt` naming what leaked, and the command
1164
- fails without writing the path you asked for (the recording is not discarded — you paid for it).
1165
- `verify-cassettes --allow-host-inventory <regex>` is unrelated: a per-finding suppressor for the
1166
- scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
1167
- wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
1168
-
1169
- 26. **A `skill`-lane `PASS` does not mean the skill ran, or that the run was the one you wanted.** *Why:*
1170
- an open-ended `skill` run has no `assert:` block, so its verdict reports only that **no guard fired**
1171
- (no error, stall, host-path leak, `outputs/` delete, permissive auto-allow or capability gap). On
1172
- `run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
1173
- means N runs cleared the guards — it says nothing about which model served them, whether the skill
1174
- was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
1175
- carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
1176
- `models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
1177
- exactly like skill output is not evidence: the skill's own source is mounted where the model can
1178
- read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
1179
- directly, with `skillActivity` empty.
1180
-
1181
- 27. **`allow_stall: true` is a scenario assertion, so the `skill` lane cannot use it.** *Why:* the
1182
- `stalled` guard fires when a run's final message ends in `?` with no productive tool call after the
1183
- last gate — which includes a complete answer that closes by *offering* a follow-up ("want me to run
1184
- this through a structured pass?"). The documented opt-out lives in an `assert:` block, and an
1185
- open-ended `skill` run has none, so the failure message names a remedy that lane can't perform.
1186
- *Fix:* on `skill`, read the final message before believing `stalled`, or move the check to a
1187
- `run` scenario where `allow_stall: true` is authorable.
1188
-
1189
- For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
1190
- files in `references/` (the gotchas above are the full list; the references repeat only the
1191
- assertion/replay-relevant ones).
1192
-
1193
- ## References
1194
-
1195
- - `references/task-recipes.md` — end-to-end recipes for the four jobs fleet owners actually hit:
1196
- evolve a cassette's `assert:` (usually no re-record), audit a fleet for tier drift, set up
1197
- redaction before the first hostloop/protocol record, derive budget assertions without a
1198
- two-pass record. Start here when the question is "how do I do X", not "what does flag Y mean".
1199
- - `references/scenario-schema.md` — scenario/session YAML schema, full assertion catalog (with each
1200
- key's replay class), the web_fetch model, and an assertion/replay-scoped gotcha subset (the full
1201
- landmine catalog lives in this SKILL's Gotchas section above).
1202
- - `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
1203
- - `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
1204
- GitHub Actions pipeline.
1205
- - `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
1206
- no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
1207
- (validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
1208
- - Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
1209
- alive* (Part II) above; the fuller recipe is in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) (repo-only, not shipped with the
1210
- installed skill).
148
+ | [`references/authoring.md`](references/authoring.md) | session vs scenario, discovery, fidelity tier, the answer-channel decision tree, `web_fetch`, scaffold + lint |
149
+ | [`references/assertions-guide.md`](references/assertions-guide.md) | the two assertion axes, the goal → key map |
150
+ | [`references/run-record-replay.md`](references/run-record-replay.md) | run / `verify-run`, recording and cassette placement, real-document validation, verdict signals (incl. where a relative or `outputs/` path lands per tier), background-run liveness, CI lanes |
151
+ | [`references/measurement.md`](references/measurement.md) | `--repeat`, `--ablate-skill`, measurement hygiene |
152
+ | [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
153
+ | [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
154
+ | [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality |
155
+ | [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
156
+ | [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
157
+ | [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
158
+ | [`references/ci-recipe.md`](references/ci-recipe.md) | the GitHub Action, replay-vs-live lanes, the four-stage pipeline |
159
+ | [`references/critique.md`](references/critique.md) | `critique` report and evidence-package shapes |
160
+ | `scripts/scenario.py` | `scaffold`, `lint`, `lint-skill`, `resolve-agent-types <plugin-dir>` (validates a pinned `subagent_type` against `plugin.json` + `agents/*.md`) |