cowork-harness 3.9.0 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (119) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +74 -1085
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +159 -0
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +60 -0
  4. package/.claude/skills/cowork-harness/references/authoring.md +258 -0
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +50 -39
  6. package/.claude/skills/cowork-harness/references/critique.md +13 -7
  7. package/.claude/skills/cowork-harness/references/debugging.md +145 -0
  8. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +12 -7
  9. package/.claude/skills/cowork-harness/references/gotchas.md +285 -0
  10. package/.claude/skills/cowork-harness/references/measurement.md +56 -0
  11. package/.claude/skills/cowork-harness/references/run-record-replay.md +339 -0
  12. package/.claude/skills/cowork-harness/references/scenario-schema.md +41 -174
  13. package/.claude/skills/cowork-harness/references/task-recipes.md +38 -24
  14. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +10 -0
  15. package/.claude/skills/cowork-harness/scripts/scenario.py +530 -32
  16. package/.cowork-redact.json +18 -7
  17. package/.env.example +4 -0
  18. package/AGENTS.md +7 -1
  19. package/CHANGELOG.md +792 -0
  20. package/CONTRIBUTING.md +7 -5
  21. package/DESIGN.md +2 -2
  22. package/README.md +20 -9
  23. package/RELEASING.md +2 -4
  24. package/SPEC.md +203 -34
  25. package/baselines/desktop-2.9939.4.json +1116 -0
  26. package/dist/agent/session.js +7 -0
  27. package/dist/assert.js +61 -18
  28. package/dist/baseline.js +6 -0
  29. package/dist/cli.js +365 -151
  30. package/dist/critique/command.js +223 -149
  31. package/dist/critique/evaluator.js +9 -4
  32. package/dist/critique/limitations.js +1 -1
  33. package/dist/critique/mount-check.js +52 -0
  34. package/dist/critique/package-evidence.js +33 -17
  35. package/dist/decide/decider.js +4 -8
  36. package/dist/decide/external-channel.js +57 -28
  37. package/dist/decide/llm-transport.js +2 -0
  38. package/dist/decide/semantic-judge.js +7 -2
  39. package/dist/egress/sidecar.js +18 -24
  40. package/dist/errors.js +38 -5
  41. package/dist/hostloop/pretooluse-path-hook.js +64 -0
  42. package/dist/hostloop/process-cwd.js +124 -0
  43. package/dist/redact.js +10 -2
  44. package/dist/redactable-literal.js +43 -0
  45. package/dist/run/analyze-skill.js +5 -3
  46. package/dist/run/budget.js +8 -2
  47. package/dist/run/cassette.js +567 -102
  48. package/dist/run/chat-result.js +4 -2
  49. package/dist/run/chat.js +40 -22
  50. package/dist/run/command-globals.js +136 -0
  51. package/dist/run/doctor.js +8 -6
  52. package/dist/run/envelope.js +36 -79
  53. package/dist/run/execute.js +773 -145
  54. package/dist/run/lint-load.js +235 -0
  55. package/dist/run/migrate-run-dir.js +3 -1
  56. package/dist/run/model-provenance.js +30 -17
  57. package/dist/run/outputs-delete-tier.js +59 -0
  58. package/dist/run/pre-run-manifest.js +57 -4
  59. package/dist/run/run.js +103 -11
  60. package/dist/run/runs-gc.js +4 -2
  61. package/dist/run/scaffold.js +2 -0
  62. package/dist/run/scenario-tool.js +193 -30
  63. package/dist/run/skill-flag-surface.js +9 -1
  64. package/dist/run/tier-vacuous-tools.js +12 -0
  65. package/dist/run/timeline-fold.js +48 -19
  66. package/dist/run/trace-view.js +109 -18
  67. package/dist/run/verdict.js +55 -10
  68. package/dist/runtime/agent-image.js +20 -1
  69. package/dist/runtime/argv.js +1 -0
  70. package/dist/runtime/hostloop-prompt.js +1 -1
  71. package/dist/runtime/hostloop.js +61 -16
  72. package/dist/runtime/microvm.js +111 -4
  73. package/dist/scan.js +69 -8
  74. package/dist/secrets.js +1 -1
  75. package/dist/session.js +87 -68
  76. package/dist/spawn-guard.js +13 -0
  77. package/dist/staging/resolve.js +11 -9
  78. package/dist/termination.js +143 -0
  79. package/dist/tool-call-assert.js +237 -0
  80. package/dist/types.js +66 -8
  81. package/docs/README.md +4 -4
  82. package/docs/boundary.md +7 -6
  83. package/docs/cassette.md +119 -19
  84. package/docs/chat.md +6 -2
  85. package/docs/ci.md +18 -17
  86. package/docs/cli.md +120 -35
  87. package/docs/companion-skill.md +5 -3
  88. package/docs/critique.md +48 -26
  89. package/docs/debugging.md +7 -4
  90. package/docs/decider-dir.md +24 -5
  91. package/docs/fidelity-gaps.md +84 -72
  92. package/docs/gotchas.md +3 -3
  93. package/docs/invariants.md +6 -5
  94. package/docs/maintenance.md +4 -1
  95. package/docs/plugin-root.md +32 -6
  96. package/docs/run-status.md +13 -4
  97. package/docs/scenario.md +123 -69
  98. package/docs/session.md +14 -13
  99. package/docs/stats.md +5 -0
  100. package/docs/subagents.md +14 -11
  101. package/examples/README.md +2 -1
  102. package/examples/replays/README.md +5 -5
  103. package/examples/replays/example-multiselect-gate.cassette.json +48 -48
  104. package/examples/replays/example-pdf-skill.cassette.json +92 -95
  105. package/examples/replays/hostloop-computer-links.cassette.json +58 -58
  106. package/examples/sessions/default.yaml +1 -1
  107. package/llms.txt +4 -4
  108. package/package.json +1 -1
  109. package/python/test_scenario_lint.py +532 -21
  110. package/schema/cassette.v13.json +330 -0
  111. package/schema/critique-report.json +1 -1
  112. package/schema/run-result.json +71 -4
  113. package/schema/scenario.schema.json +193 -12
  114. package/schema/verify-cassettes.json +2 -2
  115. package/scripts/bump-version.ts +99 -93
  116. package/scripts/check-claims.ts +32 -19
  117. package/scripts/check-surface.ts +58 -10
  118. package/scripts/check-versions.ts +68 -19
  119. package/scripts/gen-schema.ts +15 -6
@@ -1,10 +1,10 @@
1
1
  ---
2
2
  name: cowork-harness
3
- description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
3
+ description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.9.0
7
- tracks-harness: cowork-harness 3.9.0 (baseline desktop-2.9939.2)
6
+ version: 4.0.0
7
+ tracks-harness: cowork-harness 4.0.0 (baseline desktop-2.9939.4)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,11 +22,12 @@ Anthropic. Say so if a user asks what it is.
22
22
  The single most important idea: **a green run is not automatically a correct run.** The harness has
23
23
  several ways to no-op a check while still producing a green run (skip an assertion on replay — now
24
24
  flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
25
- allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
- the highest-value part. Read it.
25
+ allowlist). This skill exists mostly to keep you out of those traps — the *Invariants* below and the
26
+ full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
+ Read them.
27
28
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.9.0` (baseline
29
- > `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.0.0` (baseline
30
+ > `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
32
 
32
33
  ## Preflight — make sure the harness can actually run
@@ -35,14 +36,14 @@ The 10-second inner loop, once the CLI is on PATH:
35
36
 
36
37
  ```bash
37
38
  cowork-harness doctor # prerequisites OK? (Docker, agent, token, baseline)
38
- cowork-harness skill ./my-skill "do X" # run the skill once against the staged agent
39
+ cowork-harness skill ./my-skill "do X" --model claude-sonnet-5 # run once (or set COWORK_HARNESS_MODEL)
39
40
  ```
40
41
 
41
42
  Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
42
43
 
43
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.9.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.9.0"`. **Pin `@^3.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.0.0"`. **Pin `@^4.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
47
 
47
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -50,20 +51,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
50
51
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
51
52
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
52
53
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
53
- - **`--dotenv` is a GLOBAL flag — put it BEFORE the subcommand.** `cowork-harness --dotenv .env record …`, never `cowork-harness record … --dotenv .env`. Every *other* flag is subcommand-level, so muscle memory fights this one; the harness rejects the misplaced form with an exact-fix error, but placing it first avoids the round-trip. **One exception: `critique` also accepts `--dotenv` per-command** (`critique <folder> --prompt "…" --dotenv <path>`) — available as of **1.6.0** (documented but unreachable before then); `--run-dir` stays global-only everywhere.
54
+ - **Point at another `.env`** with `--dotenv <path>`, and relocate run output with `--run-dir <path>`; both work before or after the subcommand.
54
55
 
55
56
  ## Orient — the three loops
56
57
 
57
- Everything you do with the harness is one of **three loops**, and the rest of this skill is organized
58
- into three Parts to match: **author** a scenario (Part I), **run / record / lock** it into a
59
- reproducible regression (Part II), and **debug** a run that misbehaved or greened when it shouldn't
60
- (Part III — reachable straight from here in one hop).
58
+ Everything you do with the harness is one of **three loops**, and the detail lives in one reference
59
+ file per loop: **author** a scenario ([`references/authoring.md`](references/authoring.md)), **run /
60
+ record / lock** it into a reproducible regression ([`references/run-record-replay.md`](references/run-record-replay.md)),
61
+ and **debug** a run that misbehaved or greened when it shouldn't ([`references/debugging.md`](references/debugging.md)).
61
62
 
62
63
  Pick the entry point you need. The first three are the everyday path — a quick liveness check, the
63
64
  CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that hang off them:
64
65
 
65
- - **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>"`. Fastest; no
66
- scenario file.
66
+ - **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>" --model <id>`. Fastest;
67
+ no scenario file.
67
68
  - **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
68
69
  This is the CI-grade path and most of this skill.
69
70
  - **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
@@ -72,7 +73,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
72
73
  `cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
73
74
  then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
74
75
  the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
75
- output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
76
+ output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
76
77
  [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
77
78
  **"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
78
79
  grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
@@ -87,9 +88,9 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
87
88
  the cost, and it answers that question directly. Report and evidence-package shapes:
88
89
  `references/critique.md`.
89
90
  - **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
90
- TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
91
+ TTY, **not** an asserted test — see *Debugging with `chat`* in `references/debugging.md`).
91
92
  **"Interactive" splits two ways — don't take the wrong branch.** Want to answer gates yourself *and*
92
- still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path* below),
93
+ still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path* in `references/authoring.md`),
93
94
  **not** `chat`. Reach for `chat` only when you are exploring by hand and do NOT want a verdict.
94
95
 
95
96
  > **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
@@ -102,1070 +103,58 @@ lint-skill · analyze-skill · probe-dispatch ·
102
103
  verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
103
104
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
104
105
 
105
- **Two different `scaffold` tools — don't confuse them.** The native `cowork-harness scaffold <run-id>`
106
- above turns an already-*recorded* run into a scenario (needs a run to exist first). The bundled
107
- `scripts/scenario.py scaffold --name … --skill …` — see *Scaffold a valid scenario, then lint before
108
- you push* in **Part I** — builds a scenario from flags alone, no run required. Passing that section's
109
- flag set to the native command fails with `unknown flag: --name` (exit 2).
110
-
111
- ## Part I — AUTHOR a scenario
112
-
113
- Everything below composes one deterministic, asserted `scenarios/*.yaml`: the session/scenario split,
114
- how the skill mounts, the fidelity tier, the answer path, the two assertion axes, `web_fetch`
115
- provenance, and the scaffold/lint tools that keep the YAML honest.
116
-
117
- ### Two files: session vs scenario
118
-
119
- - **`sessions/*.yaml`** — pre-prompt setup: `model`, mounts (`folders`), and discovery
120
- (marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario that
121
- omits `session:` gets an all-defaults **inline** session (not a file on disk).
122
- - **`scenarios/*.yaml`** — the test: `prompt`, scripted `answers:`, and `assert:`.
123
-
124
- This split matters: release ground truth (`baseline:` / `baselines/`, produced by `sync`) is
125
- **separate** from authored setup (`session:` / `sessions/`). "profile" is retired vocabulary — do
126
- not use it. See `references/scenario-schema.md` for every field.
127
-
128
- ### Discovery: how the skill-under-test gets mounted
129
-
130
- The skill is **copied fresh into the sandbox each run**. Wire it via `plugins.local_plugins` +
131
- `plugins.enabled: [<plugin>@local]` in the session (or `--marketplace` / `--plugin` flags on
132
- `skill`). A missing mount source is now a **hard error** (`mount source(s) not found …`); set
133
- `COWORK_HARNESS_SOFT_MISSING=1` to fall back to warn-and-exclude. Mount names are always derived from
134
- the folder basename (collision-resolved); there is no `to:` override. See `references/scenario-schema.md`.
135
-
136
- > **`git add` a brand-new skill before testing it.** Inside a git repo the harness stages the
137
- > **git-tracked** files (the fidelity boundary — real Cowork installs from a repo and sees only committed
138
- > files). *Tracked* means **in the git index** (committed **or** `git add`-staged); the **content** staged
139
- > is your **working tree**, so an uncommitted edit to an already-tracked file *is* tested — you needn't
140
- > commit to iterate. Only brand-new (untracked) files must be `git add`-ed to appear. Commit before you
141
- > record the **locking cassette**, though: real Cowork ships the *committed* tree, so a green on
142
- > uncommitted edits isn't yet a green on what installs. An **all-untracked** skill folder mounts *empty* and the agent reports "the skill isn't
143
- > installed" then did the work itself — a green-looking run where the skill never loaded. That now
144
- > **hard-fails** (`BoundaryError`, exit 3) naming the dir, and a partially-tracked folder emits a loud
145
- > `::notice:: [stage]` listing the excluded files. Fix: `git add` the skill, or `COWORK_HARNESS_GITSET=0`
146
- > to copy untracked files (won't reflect what ships). A folder **outside** any repo is copied raw (no guard).
147
-
148
- ### Choose a fidelity tier
149
-
150
- | Tier | What it gives you | Use when |
151
- |---|---|---|
152
- | `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
153
- | `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
154
- | `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
155
- | `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
156
-
157
- Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
158
- (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
159
- `references/fidelity-and-answers.md`.
160
-
161
- **Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
162
- `/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
163
- **remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
164
- different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
165
- reproduces it and none can — that container is not something a local tool can stand up. Which lane a
166
- real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
167
- So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
168
- asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
169
- `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
170
- rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
171
- above).
172
-
173
- ### Choose an answer path (gates: AskUserQuestion + tool-permission)
174
-
175
- Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
176
- live model into answering flags the run `nonDeterministic` — keep those out of deterministic
177
- regressions.
178
-
179
- <!-- answer-channels:begin -->
180
- **Pick by asking one question about your situation**, not by scanning a table — the channels are not
181
- interchangeable and the wrong one either masks a gate or can't run at all:
182
-
183
- ```
184
- Will this run be re-executed UNATTENDED? (CI, a committed cassette, --repeat, --matrix)
185
- │
186
- ├─ YES ──► scripted `answers:` / `--answer` / `--answer-policy` + `on_unanswered: fail`
187
- │ The ONLY reproducible channel. Non-negotiable for CI and committed cassettes.
188
- │ Labels reworded every run? STAY HERE: pin a stable leading SUBSTRING
189
- │ (uniqueness-guarded, fails loud) or a positional `choose`. Both keep determinism.
190
- │
191
- └─ NO — a discovery / validation run. Who holds the context to answer?
192
- │
193
- ├─ a model, steered by one line of intent
194
- │ ──► `--decider-llm --intent "<…>"` [skill · record]
195
- │ NOT on `run` — there the spelling is the scenario-YAML `on_unanswered: llm`.
196
- │ Can false-green an oracle-less semantic gate.
197
- │
198
- ├─ deterministic logic you can write down
199
- │ ──► `--decider-cmd '<helper>'` [skill · run]
200
- │ Determinism is your helper's, not the harness's. NOT on `record`.
201
- │
202
- ├─ YOU, the driving agent, holding the task context
203
- │ ──► `--decider-dir <FRESH, EMPTY dir>` [skill · run · record]
204
- │ + `cowork-harness gates <dir> --follow` (arm a Monitor here)
205
- │ + `cowork-harness answer <dir> --gate N --choose "<label>"`
206
- │ Its ONE unique property: it needs no advance knowledge of the option SET.
207
- │ (Label *text* drift alone does not need this — substring anchors handle that.)
208
- │
209
- └─ a human at a keyboard, and you are NOT producing a test
210
- ──► `cowork-harness chat` (TTY; no pass/fail verdict — see below)
211
- ```
212
-
213
- | Channel | Deterministic? | Don't use it when |
214
- |---|---|---|
215
- | Scripted | ✅ the CI/agent default | you cannot know the option set in advance |
216
- | `--decider-llm` / `on_unanswered: llm` | ❌ nonDeterministic | the gate has no oracle a model could judge |
217
- | `--decider-cmd` | delegated to your helper | the logic needs task context code doesn't have |
218
- | `--decider-dir` | ❌ nonDeterministic | nobody is present to drive it — it BLOCKS per gate |
219
- | `on_unanswered: first` | ❌ nonDeterministic | the answer matters — it *masks* the gate |
220
-
221
- **Cost of `--decider-dir`, stated plainly:** flags the run `nonDeterministic`; needs a live driver + a
222
- Monitor, so it is **unusable unattended**; blocks at each gate, strictly serial; needs a fresh empty dir
223
- per run (a dirty one is refused); rejected with `--repeat`, `--on-unanswered`, `--decider-cmd`, and with
224
- `--matrix --concurrency > 1`; and a cassette recorded this way carries a **re-record cost** — regenerating
225
- it needs the driver present again.
226
-
227
- **Rehearse it in ~2s before wiring it into a real run** — `cowork-harness decide --decider-dir <dir>` fires
228
- one sample gate through the same channel, then blocks (10-min backstop) until you answer it with the two
229
- commands above. It is the cheapest way to see the protocol work. Full recipe, including the multiSelect
230
- wire shape and the `gates --follow` Monitor loop:
231
- [`docs/decider-dir.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/decider-dir.md)
232
- (repo-only — an npm install ships it at `node_modules/cowork-harness/docs/decider-dir.md`). <!-- npm-only-ok -->
233
-
234
- **It is a FEEDER for the scripted default, not a rival.** `record --decider-dir` is a first-class way to
235
- *produce* a cassette: the non-reproducibility is spent once at authoring time and the cassette replays
236
- deterministically forever. The loop is **discover → transcribe → script** — answer live, then paste the
237
- run's echoed `--answer "<q>=<choice>"` footer lines into the scenario's `answers:` so re-records go back to
238
- being unattended. Skip the transcribe step only for one-off/exploratory runs.
239
- <!-- answer-channels:end -->
240
-
241
- **For a QUESTION gate, never hand-write the `req-N.json`/`resp-N.json` files.** `gates` and `answer` wrap
242
- the protocol — the atomic temp+rename, the `{id, answers}` envelope, the multiSelect array shape.
243
- Hand-rolling a Monitor over the raw files is the single most common mistake on this channel.
244
-
245
- `answer` writes `{id, answers}` and nothing else, so it covers **question gates only**. The channel also
246
- carries **permission**, **dialog** and **elicit** gates, whose replies need `{behavior}` / `{action}` — for
247
- those, write `resp-N.json` yourself, following the `reply_with` template the gate's own `req-N.json`
248
- advertises (it spells out the exact shape, e.g. `{"id":"…","behavior":"allow|deny"}`).
249
-
250
- Exact accepted values (teach precisely): `--on-unanswered` takes `fail|prompt|first` on `skill`,
251
- only `fail|first` on `run`. **`llm` is NOT an `--on-unanswered` value** — the bare flag
252
- `--on-unanswered llm` is rejected (use `--decider-llm`); the YAML spelling is `on_unanswered: llm`.
253
- The word `agent` is **retired** — do not write `on_unanswered: agent` (the schema rejects it).
254
- `--on-unanswered` also conflicts with `--decider-dir`/`--decider-cmd`/`--decider-llm` (the channel or
255
- model IS the terminal, so a policy alongside it never applies) — pass one, not both. On `record`, a
256
- scenario setting `on_unanswered: prompt` is rejected too: the YAML field outranks the flag, and a TTY
257
- wait can't produce a deterministic committed fixture.
258
- `--on-unanswered first` is itself flagged `nonDeterministic` — it is *not* a deterministic stand-in
259
- for scripted answers. See `references/fidelity-and-answers.md`.
260
-
261
- **Which gates to anchor (re-record robustness).** The model rewords option labels (and sometimes the
262
- question) every run, so a brittle exact-label `choose:` is itself a re-record-fragility source — it drifts and
263
- forces a re-record. The practical rule: **label-anchor only the gates whose choice drives an `assert:`** (or
264
- materially changes behavior); for gates whose answer is immaterial to your assertions, `on_unanswered: first`
265
- is the more re-record-robust choice — accept the `nonDeterministic` flag rather than trade it for a flaky
266
- anchor. (When label *order* is stable but the text drifts, a positional `choose` is the middle option — the
267
- linter flags positional `choose` as order-dependent, so use it deliberately.) The caution stands: `first`
268
- *masks* an unanswered gate, so don't use it for a gate you actually need answered a specific way.
269
-
270
- **Drifting label TEXT and an unknowable option SET are different problems — don't reach past the cheap
271
- fix.** Text that rewords while the choices stay the same is a *scripted* problem with a deterministic
272
- answer: a uniqueness-guarded leading substring, or a positional `choose`. Only when you cannot know what
273
- the options will *be* — they're generated per input document, so no anchor can be written in advance — does
274
- the answer move to a live channel (`--decider-dir` if you're driving, `--decider-llm` if nobody is).
275
-
276
- #### External deciders and the "first" shorthand
277
-
278
- When using `--decider-cmd` or `--decider-dir`, the helper's output is passed through
279
- `coerceLabel` **with the "first" shorthand disabled**. This means a helper that returns the literal
280
- string `"first"` must match an actual label named `"first"` — it is **not** coerced to option 1.
281
- This prevents a helper bug (accidentally emitting `"first"`) from silently green-ing option 1.
282
-
283
- The `"first"` shorthand remains active only for the built-in `--on-unanswered first` path. If you
284
- write an external helper, return a label name or option index — never the bare word `"first"` unless
285
- your gate actually has a label called `"first"`.
286
-
287
- ### Assertions: two orthogonal axes
288
-
289
- Conflating these is the **biggest landmine**. An assertion key has two independent properties:
290
-
291
- - **Axis A — robust to LLM phrasing drift?** Structural/boundary keys (`subagent_dispatched`,
292
- `egress_*`, `file_exists`, `user_visible_artifact`, `result`) are robust. Free-text content is
293
- not: match prose with `transcript_matches` / `transcript_contains` (stable lexical markers only —
294
- not semantic content the model paraphrases, which re-records red); check structured JSON with YAML
295
- `artifact_json` (or the [pytest lane](https://github.com/yaniv-golan/cowork-harness/blob/main/python/README.md) for complex predicates), not via a transcript substring.
296
- - **Axis B — survives `replay`?** *Independent of Axis A.* On the token-free `replay` lane, only
297
- **content keys** evaluate; filesystem / egress keys are skipped (live-only) — loudly, via an
298
- `::warning::` annotation, not a silent no-op. A key
299
- being "robust" says nothing about whether it runs on your replay gate.
300
-
301
- Getting Axis B wrong means a check that **does nothing in CI** — the harness warns loudly when it skips
302
- (an `::warning::` annotation, not a silent no-op — see the Axis B bullet above), and the bundled linter
303
- catches it before you push — run it (see *Scaffold a valid scenario, then lint before you push* below).
304
-
305
- See `references/scenario-schema.md` for the full assertion catalog with each key's replay class.
306
-
307
- #### Which assertion for which question (goal → key)
308
-
309
- Beyond the outcome/content keys most scenarios reach for first (`result`, `transcript_*`,
310
- `file_exists`/`user_visible_artifact`, `artifact_json`), the harness surfaces the agent's *behavior*
311
- — tool health, sub-agent work, panels, skill attribution, resources — as assertable keys. Reach for
312
- them by what you're trying to prove:
313
-
314
- | You want to check that… | Reach for |
106
+ ## Invariants — how a green run lies
107
+
108
+ Each of these has produced a green run that tested nothing. The full catalog, with the reasoning
109
+ behind each, is [`references/gotchas.md`](references/gotchas.md).
110
+
111
+ 1. **`result: success` is not "the task completed".** It means the agent didn't error. Assert the
112
+ deliverable (`file_exists` / `artifact_json` / `transcript_matches`). A `skill`-lane `PASS` only means
113
+ no guard fired: read `skillsInvoked`, `models` and `ablated` before concluding anything from it.
114
+ 2. **`replay` skips live-only keys.** Filesystem and egress keys are skipped on replay (loudly), so a
115
+ mixed item like `{result, egress_denied}` greens on its content half. Keep one concern per `assert:`
116
+ item, put live-only checks on a live gate, and run `cowork-harness lint`.
117
+ 3. **Only scripted answers reproduce.** Scripted `answers:` + `on_unanswered: fail` is the CI channel.
118
+ `first`, an LLM decider and `--decider-dir` all flag the run `nonDeterministic`, and `first` masks
119
+ the gate it answers.
120
+ 4. **`replay` evaluates the FROZEN scenario.** Editing `scenarios/*.yaml` changes nothing on a plain
121
+ `replay`: re-check an `assert:` edit with `replay --assert-from <file>`, and re-record for any other key.
122
+ 5. **Some keys pass on absence.** `gate_answers_delivered` passes when no gate fired — pair it with
123
+ `gate_answer_count_min: 1`. `tool_called` proves a tool ran, not that it was attempted.
124
+ 6. **An untracked skill mounts empty.** `git add` a new skill before testing it, and commit before
125
+ recording the cassette that locks it.
126
+ 7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
127
+ (`container` serves `mcp__workspace__web_fetch`, not `WebFetch`), and every tier models Cowork's
128
+ desktop-local lane only.
129
+ 8. **A WARN signal never blocks a green.** Read the verdict signals after every run
130
+ (`prompt_asset_missing`, `undelivered_deliverables`, `model_fallback`, …).
131
+
132
+ ## Short workflows
133
+
134
+ - **Author, then lock:** `cowork-harness scaffold --name … --prompt …` → `cowork-harness lint scenarios/` →
135
+ `cowork-harness record <file.yaml> --dry-run` (free) → `record` once, with `--out` at a tracked path
136
+ (a cassette cannot be moved) → `replay` on the PR gate.
137
+ - **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
138
+ `answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
139
+ - **Debug:** the triage table in `references/debugging.md` → `inspect`, `trace --view …`,
140
+ `verify-run`, `diff`. For a green you don't trust: `replay --explain`, then the gotchas.
141
+ - **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
142
+ arm yourself, with the model pinned and a recoverable source frozen.
143
+
144
+ ## References — where the detail lives
145
+
146
+ | File | Read it for |
315
147
  |---|---|
316
- | the skill didn't error out of a tool | `tool_no_error: <regex>`, `max_tool_errors: <N>` |
317
- | it didn't waste repeated identical calls | `max_redundant_tool_calls: <N>` |
318
- | a deliverable reached the user | `user_visible_artifact: <path>` (+ `no_scratchpad_leak: true` if it delivers via `present_files` — **`container` only**) |
319
- | an internal name/path did **not** leak into a delivered file | `artifact_text: {artifact, not_contains}` — `artifact_json`'s companion for non-JSON bodies; literal path, no glob, so one entry per delivered surface |
320
- | a named path must **not** exist after the run | `file_absent: <path>` (**live/verify-run only**) — do NOT invert `no_unexpected_files`: that is an allowlist over *newly created* files and needs a pre-run manifest |
321
- | a to-do workflow finished | `all_tasks_completed: true`, `task_status: {match, status}` |
322
- | a skill / connector / tool was **offered** | `skill_available`, `connector_available`, `tool_available` (all `<regex>`) |
323
- | a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
324
- | a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
325
- | a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
326
- | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
327
- | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
328
- | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
329
- | the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
330
- | the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
331
- | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
332
- | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
333
- | a context compaction happened | `compaction_occurred: true` |
334
-
335
- Every one of these still obeys the two axes above — several are live-only or need a `controlOut`
336
- cassette on replay, so check the catalog's replay class before putting one on a PR gate.
337
- `cowork-harness assertions --list` prints the full, always-current key set with one-line semantics
338
- straight from the schema — treat it (and the catalog) as the source of truth; this map is a
339
- goal-oriented index into it, not a second catalog.
340
-
341
- ### web_fetch (fail-closed, two-path)
342
-
343
- `web_fetch` behaves unlike `curl`. A URL is gated by **provenance**, not the egress allowlist:
344
-
345
- - A URL is *provenanced* iff it appeared in the **prompt** or a **prior `web_fetch` result**. To
346
- make a fetch succeed, put the URL in the prompt.
347
- - **Provenanced** → fetches (still SSRF-guarded per redirect hop); the egress hostname allowlist is
348
- **not consulted**.
349
- - **Not provenanced** → raises a per-domain approval gate (`webfetch:<domain>`) that is
350
- **fail-closed** (it is *not* auto-allowed; `--on-unanswered first` won't allow it). Answer it with
351
- a scripted rule (`when_tool: "webfetch:<domain>"` + `grant: domain|once`), a session
352
- `web_fetch.approved_domains`, or a live decider.
353
-
354
- Surprise to remember: adding a host to `egress.extra_allow` is a **no-op** for a provenanced fetch.
355
- Full model in `references/scenario-schema.md`.
356
-
357
- ### Scaffold a valid scenario, then lint before you push
358
-
359
- Don't hand-write the YAML from memory — that's how invented keys (`assertions:` vs `assert:`,
360
- `json_file`, `answer_policy`) creep in. Start from the bundled generator, which emits the
361
- known-good skeleton (right tier, scripted `answers:` + `on_unanswered: fail`, content assertions
362
- separated from live-only ones, one concern per item) and **self-lints its own output**. The
363
- generator is the bundled `scripts/scenario.py` — installed as a plugin, point `S` at
364
- `${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`; from a repo checkout, use the literal path below:
365
-
366
- ```bash
367
- S=".claude/skills/cowork-harness/scripts/scenario.py"
368
- python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
369
- --prompt "Generate the weekly report to outputs/report.md." \
370
- --content 'weekly report' --artifact outputs/report.md \
371
- --egress-allowed api.weather.example.com --out scenarios/report-check.yaml
372
- ```
373
-
374
- Then lint every scenario — it encodes the no-silent-false-green invariants. Use the CLI wrapper
375
- `cowork-harness lint` (it runs the same bundled `scenario.py lint`):
376
-
377
- ```bash
378
- cowork-harness lint scenarios/*.yaml
379
- ```
380
-
381
- `lint` flags: filesystem/egress-only assertions on a `replay` gate (silent no-op), bad regex
382
- quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `hostloop`/`protocol`
383
- (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
384
- baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
385
- `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
386
- unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
387
- `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
388
- to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
389
- `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
390
- and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
391
- (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
392
- emits a scenario `lint` would reject.
393
-
394
- **`lint` is the LENIENT check — the loader is the strict one.** An unknown top-level key is a ⚠ WARN in
395
- `lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
396
- that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
397
- `cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
398
- `✗ broken:` file and exits 1). **Read the exit code, not just its sign:** `record <file>` — with or without
399
- `--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
400
- pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
401
- as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
402
- but **`replay` does not**: a frozen top-level key it doesn't recognize (e.g. `lane:` recorded pre-1.16.0) is
403
- silently ignored and can flip a lane-sensitive verdict green; only frozen **assertion** keys stay
404
- hard-rejected there. Full split + the v11 version-regime:
405
- [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md#unknown-keys-the-loader-is-strict-lint-is-lenient).
406
-
407
- ## Part II — RUN, RECORD & LOCK
408
-
409
- You have an authored scenario. This Part runs it, reads the verdict, locks it into a
410
- byte-deterministic cassette, checks a background run's liveness, and places the assertions in the
411
- right CI lane.
412
-
413
- ### Run, then lock determinism
414
-
415
- Read the verdict and the inline failing transcript. To pin a flaky-because-stochastic gate, paste
416
- the echoed `--answer "<q>=<choice>"` footer lines back into the scenario's `answers:` for a
417
- deterministic re-run. Use `cowork-harness trace <id>` to digest a run. If only an *assertion* is wrong (the
418
- run itself was fine), `cowork-harness verify-run <run-dir> <scenario.yaml>` re-checks the `assert:` block against
419
- a **kept** run dir (`--keep`, or a `--session-id` run) with no live re-record — tokens-free, ~1s per iteration.
420
- When the scenario declares `answers:`, verify-run **also** checks they still match the run's actual gates (a
421
- reworded gate or a `choose:` the run never offered fails here in ~1s instead of on a paid re-record). Or skip
422
- the discovery/encode/record dance entirely and answer gates **live during the recording** with
423
- `record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
424
- `run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
425
- `cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
426
- refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
427
- cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
428
- guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
429
- one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
430
- portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
431
- the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
432
- takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
433
- contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
434
- real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
435
- reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
436
-
437
- **Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
438
- concluded the free pre-flight was unavailable:
439
- - **"Does my whole corpus still load?"** → the **directory** arm (`record scenarios/ --dry-run --quiet`,
440
- the CI shape in `references/ci-recipe.md`). It reports every offender in one pass, and the
441
- destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
442
- are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
443
- knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
444
- with no `prompt:` key reports as `· skipped:` rather than broken (so a renamed or mis-indented
445
- `prompt:` reads as "not a scenario" and the batch still exits 0), exit 1 covers **refused**
446
- (path-independent only — prompt policy, assert contradiction, duplicate target) as well as broken, and
447
- `--quiet` suppresses the advisory notes entirely — it keeps `✗ broken:` / `✗ refused:` / `· skipped:`,
448
- which is what you want in CI but means the notes are not a thing you will see there.
449
- - **"Would THIS record be refused?"** → the **single-file** arm **with the flags and `--out` the real
450
- record will get**. The destination it evaluates is `--out` if given, else `cassettes/<slug>.cassette.json`
451
- *relative to your cwd* — so previewing from the repo root a record that really runs from a subdirectory
452
- asks about a path that may not even exist, and a scenario whose cassette IS committed can come back
453
- refused. Point it at the real destination and the answer is binding.
454
-
455
- Neither question is answered by passing `--allow-host-inventory-fixture` to get past the refusal: that
456
- flag is consent for a recording you intend to make, and reaching for it as a load-check habit is how it
457
- stops meaning anything.
458
-
459
- **Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
460
- Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
461
- default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
462
- That choice is permanent: the cassette rewrites `scenario.session` and `scenarioSource` **relative to
463
- its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
464
- into another repo — leaves those unresolvable and
465
- `verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
466
- re-record at the new location — or point `replay`/`verify-cassettes` at the session with `--session <file>`, which resolves it without a re-record. Since 2.0.0 a bare `replay` FAILS on this class rather than warning. **`record` now says so BEFORE it spends:** a pre-flight — at the same
467
- pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
468
- warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
469
- outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
470
- cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
471
- missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
472
- (`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25 below).
473
- The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
474
- redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
475
- for a cassette that cannot verify staleness from its own location — recoverable only by passing
476
- `--session <file>` on every invocation thereafter.
477
-
478
- **Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
479
- scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
480
- (and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
481
- `events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
482
- `description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
483
- only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
484
- was never delivered — the views are a digest, and the record is wider (`jq` recipes in
485
- [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
486
- your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
487
- skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
488
- current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
489
- deliberately. (Same fail-closed family: corrupt gate evidence — unparseable `events.jsonl` lines, or fewer
490
- gates than `trace.json` recorded questions — a structurally invalid `result.json`, a `command:"replay"`
491
- result (a replay is a re-check of a recorded cassette, not run evidence — verify the original live run dir),
492
- and a `mode:"chat"` result (chat carries no assertions or verdict by contract) also refuse rather than
493
- certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
494
-
495
- Run artifacts are written to `~/.cowork-harness/runs/…` by default — **outside any working tree**, so a run
496
- launched from a repo root never drops sensitive skill inputs/outputs into it. Pass `--run-dir <path>` (or set
497
- `COWORK_HARNESS_RUNS_DIR`) to relocate; in CI point it at a workspace path so an artifact-upload step can
498
- collect the runs.
499
-
500
- #### Validate a skill against real documents (not a cassette)
501
-
502
- The loops above build **deterministic regressions**. A different job — drive a skill against *real* input
503
- documents to judge whether it actually does the work (extraction, analysis), with no intent to record a
504
- cassette — has its own recipe:
505
-
506
- 1. **Explore with the LLM decider.** `cowork-harness skill <dir> --decider-llm --intent "<one line of what
507
- this run is testing>"` lets a model (Sonnet default) answer each gate steered by your intent. The model replies with
508
- the option **number** and the harness maps it to the exact label (so it can't whiff by mis-typing the
509
- label text); an out-of-set answer fails loud. This is exploration, **not** a deterministic regression —
510
- the run is flagged non-deterministic and a green here is not a scripted pass. The answering model
511
- defaults to a Sonnet id (a weaker model tends to prose-decline an ambiguous judgment gate → fail-loud);
512
- override it with `--decider-model <id>` — a cheaper model (e.g. Haiku) for simple gates to cut cost,
513
- or Opus for the hardest judgment gates; it won't make an under-specified gate deterministic. A live
514
- decider can false-green a semantic assertion on an oracle-less gate — see `references/fidelity-and-answers.md`.
515
- 2. **Script the load-bearing gates — especially binary confirm gates.** Once you know which gates fire
516
- (`trace <run-dir> --view questions`), pin the ones whose choice drives the outcome with
517
- `--answer "<q>=<label>"` / `--answer-policy <yaml>`. When a skill **re-words its option labels run-to-run**
518
- (LLM-authored gates), pin a **stable leading substring** instead of the full label — `--answer
519
- "<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
520
- **fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
521
- CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
522
- 3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
523
- **salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
524
- missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
525
- 4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
526
- produced — the artifacts plus a shallow field preview of each JSON artifact (e.g. the extracted figures).
527
- It works on a salvaged partial run too. (A partial run is marked `PARTIAL`; `verify-run` and `scaffold`
528
- refuse to treat its half-finished output as a passing result.)
529
- 5. **For image-only / scanned PDFs, use the full-parity image.** The default agent image omits OCR and
530
- PDF-table tooling; if a **scenario** sets `requires_capabilities` (a scenario field — not skill
531
- frontmatter) and the image provably omits one, the harness **aborts before the paid run (exit 3)** —
532
- unless the scenario asserts `allow_missing_capability: true`, which downgrades it to a notice and
533
- proceeds. Rebuild with `--build-arg COWORK_FULL_PARITY=1` and point `COWORK_AGENT_IMAGE` at it for those
534
- skills.
535
- 6. **Iterate across fixes — verify before you trust, and don't cross-pair generations.** A green run is
536
- not a correct run, and a skill's self-reported finding is not real until its cited evidence is found in
537
- the run's own output. Ground each finding against `result.json` (`finalMessage` = the skill's own
538
- answer/critique; `toolResults` = tool outputs) and the tool-call stream via
539
- `cowork-harness trace <run-dir> --output-format json` — add `--full-results` so a successful call's full
540
- input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
541
- pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
542
- produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
543
- kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
544
- `run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
545
- version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
546
- generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
547
- `--label <tag>`); an un-split window spanning more than one generation now warns.
548
- **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
549
- skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
550
- critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
551
- pairing step there silently groups on an absent key instead of erroring — check the field is present, or
552
- require ≥ 1.5.0. See [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)
553
- (repo-only) for the full loop.
554
-
555
- #### Interpreting verdict signals
556
-
557
- The run verdict may include `WARN`-severity signals in addition to pass/fail. One to watch for:
558
-
559
- - **`prompt_asset_missing`** — the run proceeded but a prompt asset referenced by the scenario was
560
- not found. The model ran against an incomplete prompt. This is a `WARN`, not a hard failure, so
561
- the run can still green. If you see it, fix the asset path — a green with a missing asset is
562
- not a valid pass.
563
-
564
- **False negatives — signals that are tier/image artifacts, not skill defects.** Some fail-severity
565
- signals read like a skill gap but are really a property of the reduced test image or the fidelity tier.
566
- Recognize these before "fixing" a non-bug:
567
-
568
- - **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
569
- rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
570
- `markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
571
- (`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
572
- Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
573
- behind that sentence). The message says so ("likely a FALSE
574
- NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
575
- `COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
576
- `allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
577
- or a declared `requires_capabilities` the tier can't provide, both lanes — an unknown family name
578
- hard-fails rather than silently passing.) **On an open-ended `skill` run** (no `assert:` block to carry
579
- the modifier), pass **`--allow-missing-capability`** — the CLI equivalent of the assertion.
580
- - **`ended_with_question`** (`WARN`, live lane) — a heuristic: the agent's final answer contains a
581
- question and the run wrote **no deliverable to `outputs/`** — it may have ended on a request for input
582
- instead of finishing. Warn-only; the fix is scripting/steering the answer (`answer:` / `--answer` / a
583
- decider, or `--decider-llm --intent`), not editing the skill's prose. The strict, fail-severity sibling
584
- `stalled` already catches a *trailing*-`?` final turn that did no tool work after the last gate; this
585
- covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
586
- the final message before acting — a legitimate question-posing answer that wrote a file never fires.
587
- Assert `allow_stall: true` if ending on a question is the intended terminal state.
588
- - **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
589
- and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
590
- they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
591
- does not get them. It fires with no assertion written — `present_files_called` covers the positive case
592
- only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
593
- **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
594
- scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
595
- **The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
596
- can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
597
- (what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
598
- `outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
599
- `fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
600
- bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
601
- a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
602
- decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
603
- [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
604
- nothing is delivered by location there, so only an explicit delivery counts. Assert
605
- **`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
606
- downloaded inputs) rather than a delivery gap.
607
- - **`delivery_unobservable`** (`WARN`, `lane: remote` only) — the run produced file(s) but the harness
608
- serves **no delivery tool on that lane**, so whether they reached the user is unanswerable. This is the
609
- honest cannot-verify companion to `undelivered_deliverables`: reporting every remote file as undelivered
610
- would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
611
- with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
612
- a harness coverage gap (see the *File delivery* section of fidelity-gaps).
613
- - **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
614
- `model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
615
- this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
616
- transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
617
- by a different model than the scenario names, so treat a green as evidence about the fallback model.
618
-
619
- - **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
620
- connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
621
- approval, not just outputs — a connected folder shows the identical default — so this run diverged from
622
- what production would have allowed. `WARN` rather than `FAIL` because the harness **detects** post-hoc
623
- what production **enforces**: by the time the scan sees it, the agent already proceeded where it would
624
- have hit `EPERM`, so failing the run would overstate what a post-hoc scan knows. Author
625
- `no_delete_in_mounts: true` to hard-fail on it, or `allow_delete_in: ["<mount>"]` to waive that mount
626
- (detection still runs and the hit is still recorded — the waiver is a verdict decision).
627
- - **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
628
- paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
629
- `container`/`microvm`, but only *fires* on an actual scanned leak with no authored
630
- `transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
631
- run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
632
- it's valid.
633
- - **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
634
- infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
635
- rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
636
- the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
637
- a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
638
- `Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
639
- failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
640
- suspiciously empty.
641
- - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
642
- `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
643
- pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
644
-
645
- The full 20-code signal table (severity + per-signal opt-out) is in
646
- [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
647
- the fuller narrative.
648
-
649
- ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
650
-
651
- A single green proves the run passed **once**. Two questions need more than that, and both have a
652
- discipline that is cheap to follow and expensive to skip.
653
-
654
- **"Did it pass, or pass once?"** → `--repeat N` (2-100, on `skill` AND `run`) samples the same
655
- skill+prompt N times and prints a variance rollup instead of a single verdict. `--min-pass-rate` sets
656
- the batch threshold, `--stop-on-diverge` stops the moment flakiness is proven, `--max-budget-usd` caps
657
- spend.
658
-
659
- **"Does the skill actually help?"** → `--ablate-skill` runs the prompt with every skill/plugin
660
- discovery source removed, so the agent answers from its own priors. **It is ONE arm, not a paired
661
- experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
662
- the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
663
- and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
664
- The rollup says so on its verdict line: `repeat "<skill>": PASS [ABLATED — control arm] — 5/5 passed`.
665
- Every ablated run is stamped `ablated: true` in `result.json` and carries `ablated=true` on its
666
- `[provenance]` footer line; a run that isn't stamped is a real run.
667
- What the harness gives you here is the run execution and the control arm — designing the comparison
668
- (scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
669
-
670
- **Measurement hygiene — four things that silently invalidate a batch:**
671
-
672
- 1. **Pin the model.** With no `model:` in the session (or `--model` on the `skill` lane) the run uses
673
- whatever the staged agent binary defaults to — not a harness constant, and it can move under a
674
- baseline bump. Read `result.json`'s `models` back before believing any cross-run comparison — and when
675
- you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
676
- fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
677
- array purely by whether such a turn occurred.
678
- 2. **Commit the skill first.** `fingerprint.skillHash` is content-exact, so an edit mid-batch silently
679
- splits your dataset into two generations — and a hash whose source was never committed identifies a
680
- generation that is unrecoverable. `stats --group-by skill-hash` separates them after the fact;
681
- nothing recovers the source.
682
- 3. **Check which arm you actually ran** before analysing anything: `ablated` and
683
- `context.availableSkills` in each `result.json`.
684
- 4. **Read `skillsInvoked`.** A rep where the skill never triggered is a measurement of the model, not
685
- of your skill — discard or re-run it.
686
-
687
- ### Checking whether a background run is alive
688
-
689
- Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
690
- processes in your OWN PID namespace, which is frequently NOT the harness process's namespace (e.g. when
691
- you're a sandboxed subagent). An empty `ps aux` match tells you nothing about whether the run is still
692
- going.
693
-
694
- Use **`cowork-harness status <dir> [--follow]`** instead — reads `<outDir>/status.json`, a file the
695
- harness writes/updates throughout the run's lifecycle (including a crash-safety net for a thrown
696
- error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
697
- either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
698
- checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
699
- run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
700
- which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
701
- modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
702
- `<dir>` also accepts the run-dir root
703
- passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
704
- newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
705
- rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
706
- payload; `cowork-harness status --help` has the flags.)
707
-
708
- **Poll with `--follow`, not with a shell loop over `status`'s stdout.** The one-shot text form prints to
709
- **stderr** and writes nothing to stdout; `--output-format json` (one envelope) and `--follow` (one JSON
710
- line per status change) are the **stdout** forms. A poll that greps `status`'s stdout therefore matches
711
- nothing, exits 1, and returns instantly against a run with minutes left to go — a silent false "done":
712
-
713
- ```bash
714
- # WRONG — stdout is empty, so grep exits 1, `!` inverts it, and the loop never sleeps.
715
- until ! cowork-harness status "$D" | grep -q '● running'; do sleep 30; done
716
-
717
- # RIGHT — the harness owns the poll loop and exits when the run reaches a terminal state.
718
- cowork-harness status "$D" --follow
719
- ```
720
-
721
- **A multi-minute `record`/`run` outlives a short-lived wrapper.** Don't launch a long record from a
722
- subagent that returns before it finishes — the returning agent tears down its process tree and kills the
723
- in-flight run mid-artifact-write. Run it foreground, or detached from any process that will exit first.
724
- (The `status.json` liveness above is exactly what surfaces such a teardown as `"error"`/`stale` rather
725
- than a stuck `"running"`.)
726
-
727
- ### Place assertions in the right CI lane
728
-
729
- CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
730
- (filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
731
- PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
732
- the four-stage pipeline.
733
-
734
- ## Part III — Debug
735
-
736
- A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
737
- afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
738
- re-run and hope. Start at the triage below, then use the observability output and, when you need to
739
- reproduce interactively, `chat`.
740
-
741
- > **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
742
- > `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
743
- > narrower record — `critique-evidence-package.txt`, what a grade was actually computed against — none of
744
- > the five tools above surface it; see `references/critique.md`.
745
-
746
- ### Triage — a run misbehaved, or a green looks wrong
747
-
748
- <!-- BEGIN triage-canonical -->
749
- Two situations need different tools — figure out which one you're in first, then reach for the tool
750
- instead of re-running and hoping. The run already wrote its evidence to a kept run dir (`--keep` prints
751
- the path; `trace <run-id>` finds it), and every tool below reads that evidence **token-free** — no
752
- Docker, no re-record.
753
-
754
- | Situation | Symptom | Reach for (in order) |
755
- |---|---|---|
756
- | **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
757
- | **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `replay --mutate` — perturbs a CAPPED SAMPLE of recorded JSON values (10/file, 50 total) and reports which perturbations NOTHING caught; the report names the sample size, so read it as a sample not a total (reporting only; never moves the verdict/exit code) · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
758
-
759
- A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
760
- agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
761
- already does.
762
- <!-- END triage-canonical -->
763
-
764
- **Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
765
- **deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
766
- so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
767
- `references/fidelity-and-answers.md` (shipped); the specific deltas vs. real Cowork and the sandbox
768
- boundary model live in [`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md) / [`docs/boundary.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/boundary.md) (repo-only, not in the installed
769
- payload). If the behavior is on that gap list, it's expected — stop debugging your skill.
770
-
771
- ### Inspecting a run's observability output
772
-
773
- A verdict is only the top of what a run records, and the run dir persists after the verdict
774
- (`~/.cowork-harness/runs/…`). Beyond pass/fail, every `run`/`skill`/`chat` writes a `result.json` and a
775
- trace you read back without a re-record — the debugging loop is *localize the failure from that
776
- already-written evidence*, not re-run-and-hope. Use them to diagnose a failure (and, secondarily, to
777
- decide which assertions from *Assertions: two orthogonal axes* are worth adding):
778
-
779
- - **`cowork-harness trace <run-dir> --view <view>`** — focuses one of the run's rollups (the per-tool
780
- call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
781
- bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
782
- the current list rather than relying on a fixed enumeration here.
783
- - **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
784
- is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
785
- cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
786
- location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
787
- session end), `present_files` is NOT served, and `user_visible_artifact` /
788
- `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
789
- to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
790
- `lane: remote` scenario still runs locally.
791
- - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
792
- `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
793
- plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
794
- `--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
795
- tier — instead of aggregating across them (a window spanning >1 generation warns, so an un-split A/B
796
- average announces itself rather than passing as one number;
797
- >1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
798
- individual runs behind each summary with their `skillHash`/`runLabel`, so
799
- you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
800
- - **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
801
- contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
802
- walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
803
- `total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
804
- `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
805
- `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
806
- `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
807
- `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
808
- `context` (tools/mcpServers/availableSkills), `tasks`,
809
- `workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
810
- (`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
811
- evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
812
- `gateProvenance` (per-gate `scripted`/`decided(llm|external)`/`first-option`/`prompt` with a
813
- `bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
814
- `egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
815
- cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
816
- `outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
817
- `jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
818
- `{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
819
- semantics: [`docs/cli.md` → What you get out](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md#what-you-get-out-inspectable-output) (repo-only); [`schema/run-result.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/run-result.json) is the
820
- machine source.)
821
- - **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
822
- **`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
823
- a re-record rarely tells you more than the captured stderr already does. Also check
824
- **`resultErrorKind`** (`"transport" | "agent" | "usage_limit"`) before spending another paid run: a
825
- `"usage_limit"` failure is a quota exhaustion, not a skill bug — retry after the limit resets rather
826
- than debugging; `"transport"`/`"agent"` means something actually broke, worth localizing before
827
- re-running.
828
- - **Attributing cost to sub-agent work.** `subagents[]` gives the dispatch tree — each sub-agent's
829
- `dispatchModel`/`resolvedModel`, `toolsUsed`, `prompt`/`output`, and `attributedSkillId` — but **not** its own token/cost;
830
- aggregate cost is per-**model** in `modelUsage` (and `trace --view usage`), not per-sub-agent. So a
831
- cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
832
- per-model usage — the harness doesn't line-item each sub-agent's tokens.
833
- - **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
834
- `tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
835
- `trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
836
- **Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
837
- **absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
838
- collection) is evidence **UNAVAILABLE**, not an empty run — `trace --view files` reports a loud
839
- UNAVAILABLE marker (`workspaceFilesRecorded: false` in JSON, and no phantom "removed" diff rows) and
840
- `inspect` prints `artifacts: UNAVAILABLE` (`artifactsRecorded: false`) instead of `artifacts (0):`.
841
-
842
- ### Debugging with `chat`
843
-
844
- `cowork-harness chat` opens an interactive multi-turn REPL against a live Cowork session. It is
845
- **not** an asserted test — no `assert:` block, no cassette. Use it to explore behavior, reproduce a
846
- bug interactively, or test a prompt before committing it to a scenario.
847
-
848
- Each session still writes an informational `result.json` (`mode: "chat"`, no `assertions`) plus a
849
- trace and index row under its run dir — the same telemetry (tool durations, model usage, resources,
850
- etc.) that `run`/`skill` produce — so `cowork-harness trace <chat-run-dir>` / `stats` work on a chat
851
- session too, even though it never yields a verdict.
852
-
853
- **`--plugin <dir>` flag (repeatable).** Load additional skill folders alongside the primary session
854
- plugin. Each `--plugin <dir>` appends the folder to `local_plugins`. Useful when the skill-under-test
855
- depends on a sibling plugin:
856
-
857
- ```bash
858
- cowork-harness chat ./skills/report-gen --plugin ./skills/shared-utils
859
- ```
860
-
861
- **Note:** `--raw` mode (native `docker run -it`) can't honor the harness-managed flags, so `--upload`,
862
- `--folder`, `--plugin`, and `--fidelity` are **rejected** with a usage error if combined with `--raw`;
863
- only `--model` is carried through.
864
-
865
- **`/help` in the REPL.** Type `/help` at the prompt to see available commands:
866
-
867
- ```
868
- Commands: /exit /quit /help
869
- ```
870
-
871
- The startup banner now reads `type your message (/help for commands)` as a reminder. `/exit` and
872
- `/quit` both terminate the session.
873
-
874
- ## Gotchas — the "✓ passed ≠ correct" landmines
875
-
876
- Stated as *symptom → why → fix*. This is the **workflow/record/answer-path** view — the broader of the
877
- two lists, but **not a strict superset**: `references/scenario-schema.md`'s *Authoring gotcha list*
878
- carries a few assertion-level landmines this one omits (`transcript_no_host_path`'s scan width,
879
- `egress.extra_allow`'s no-op on the provenanced `web_fetch` path, `replay_protocol_fidelity` not being
880
- authorable). Reach for this list when debugging a run's behavior, that one while authoring `assert:`.
881
- **The two lists are numbered independently** — a bare "gotcha N" means the list you are reading.
882
-
883
- 1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
884
- `replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
885
- `artifact_text`, `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
886
- *mixed* item like
887
- `{result, egress_denied}` greens on `result` while its `egress_denied` half is dropped. (`record`
888
- snapshots an `artifacts` manifest, which makes
889
- `file_exists`/`user_visible_artifact`/`artifact_json`/`artifact_text`/`computer_links_resolve`
890
- replay-checkable — but the live-only egress keys stay skipped, and `file_absent` is never
891
- replay-checkable at all: proving absence needs an exhaustive, healthy walk a manifest does not record.) *Fix:* put egress/live-only checks on
892
- a live gate; keep one concern per `assert:` item; run the linter. The harness warns loudly on skip.
893
-
894
- 2. **A steered gate answer never reached the model.** *Why:* `serializeDecision` must emit
895
- `updatedInput: { questions, answers }`; a header-only gate (empty `question`) can never be keyed.
896
- *Fix:* give every gate a non-empty `question`. (multiSelect gates ARE supported on **every** answer
897
- channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
898
- reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
899
- Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
900
- string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
901
- `questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
902
- old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
903
- on unobserved delivery (absence of evidence is failure, not neutral).
904
-
905
- 3. **A multi-key `assert:` item is an AND.** A single list item with more than one key passes iff
906
- **every** key passes. *Fix:* one concern per item unless you genuinely mean conjunction (and a
907
- mixed-class conjunction still loses its filesystem half on replay — see gotcha 1 below).
908
-
909
- 4. **`tool_called` doesn't mean "attempted".** Tool counts are authoritative and de-duped: a tool
910
- that was *requested then denied* does **not** register as called. *Fix:* don't assert `tool_called`
911
- to prove an attempt; it proves the tool actually ran.
912
-
913
- 5. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
914
- sub-agent used other tools. `subagent_dispatched` / `subagent_output_contains` match on dispatch
915
- type (`dispatchAgentType`), the binary-*resolved* type (`resolvedAgentType`), *or* the dispatch
916
- **description** — so a type-less dispatch that resolved to e.g. `general-purpose` is still
917
- selectable, by either the resolved type or the description. A `Task` dispatch that carries NO
918
- `subagent_type` at all falls back to the built-in `general-purpose` agent with a **wildcard tool
919
- surface** (`tools:["*"]`, including workspace bash) — faithful production behavior, and it fires
920
- routinely. The harness warns loudly on this fallback and records `subagents[].dispatchTypeOmitted`;
921
- an *explicit* `subagent_type: "general-purpose"` is a deliberate author choice and does not warn.
922
- Implication: `subagent_tool_absent` on a type-less dispatch is weaker evidence (wildcard surface) —
923
- pin `subagent_type` explicitly when you need a tight tool-absence guarantee.
924
-
925
- **Cross-tier "no shell" caveat.** On `hostloop`, native `Bash` calls route through the
926
- `mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
927
- `mcp__workspace__*` to hold across every tier.
928
-
929
- 6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
930
- reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
931
- run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
932
- per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
933
- (gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
934
- — repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
935
- `dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
936
- a fan-out you don't want.
937
-
938
- 7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
939
- assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
940
-
941
- 8. **Read-only mounts are enforced; delete-deny is a HARNESS gap — production DOES enforce it.**
942
- `mode:r` mounts get a real `:ro` bind (a write fails in-guest). But `rw` vs `rwd`
943
- (write-but-no-delete on `outputs/` / connected folders) is *not* mount-enforced **in the harness** —
944
- `rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`. **Real Cowork enforces it live:**
945
- outputs is a FUSE mount, and `unlink`/`rmdir` fail `Operation not permitted`; a skill must request
946
- approval via `allow_cowork_file_delete` (which re-mounts the folder `rwd` mid-session) to delete.
947
- **Only unlinking is denied.** Emptying a file in place — `truncate -s 0`, `> file`, `shred` without
948
- `-u` — and renaming *within* outputs both SUCCEED in production, so the harness does not flag them
949
- either. Renaming a file OUT of outputs fails (`EXDEV`, then `EPERM` on the copy-then-unlink
950
- fallback), so that stays a delete. Two consequences: a skill should not stage disposable scratch
951
- under `outputs/` (in production, cleanup there costs an approval prompt), and a skill's
952
- "catch-EPERM-then-request-approval" branch cannot be exercised at any harness tier (the `rm` just
953
- succeeds here). Do not read this gotcha as "delete-deny may not be real in production" — it is real.
954
- If a scenario's deletion IS intended, assert `allow_outputs_delete: true` rather than dropping
955
- `no_delete_in_outputs` — omitting it does not permit anything.
956
-
957
- 9. **Keep `.env` out of any mounted folder** — it is copied into the sandbox and the token could
958
- leak. Put it at a working-dir or install root (token resolution: env > `--dotenv` > `./.env` >
959
- install `.env`). **Inverse footgun — running from a git worktree:** a worktree's `./.env` is gitignored, so
960
- it's **absent** there and you'll get "no model credentials." *Fix:* pass `--dotenv <main-checkout>/.env`
961
- (or set the env var) — that's exactly what `--dotenv` is for.
962
-
963
- 10. **A base64 artifact that was scrubbed at record time will fail artifact assertions at replay.**
964
- When `record` detects a secret embedded in a base64 artifact, it replaces the entire artifact
965
- body with `[REDACTED:base64]` and emits a `::warning::`. Any `artifact_json` or content
966
- assertion targeting that artifact will fail at replay because the body no longer matches. *Fix:*
967
- do not let secrets flow into artifacts; if the artifact is intentionally opaque, drop the
968
- content assertion and gate on `file_exists` on the live lane instead.
969
-
970
- 11. **An external decider returning `"first"` does not select option 1.** The `"first"` keyword
971
- shorthand is disabled for `--decider-cmd` / `--decider-dir` helpers (see *Choose an answer path*
972
- → External deciders). If your helper
973
- accidentally emits `"first"` and no label named `"first"` exists, the gate fails — it does
974
- **not** silently pick the first option. This is intentional: a helper bug should fail loud, not
975
- green wrong. *Fix:* have helpers return a label name or numeric index.
976
-
977
- 12. **`prompt_asset_missing` is a WARN, not a hard failure — greens can hide it.** The
978
- `prompt_asset_missing` verdict signal (see *Interpreting verdict signals*) does not block a green verdict. Scan the verdict
979
- signals section after every run; a run that greened with this signal ran against an incomplete
980
- prompt. *Fix:* treat `prompt_asset_missing` as a blocking error in CI by checking the signals
981
- array.
982
- 13. **`result: success` means the agent didn't error, NOT that the task completed — always assert on
983
- artifacts/content.**
984
- - A turn that ends on a plain-text re-ask ("which file did you mean?") still reports
985
- `result: success`.
986
- - The harness catches this with a **`stalled`** verdict signal: a run that ends on a question and
987
- did **no productive work after its last gate** — both the no-gate case ("which file?" with no
988
- tool calls) AND the *answered-gate-then-re-ask* case (the agent answers an `AskUserQuestion`,
989
- then asks again in plain text and stops). Suppress with `allow_stall: true` if ending on a
990
- question is intended.
991
- - The signal is a **tool-position heuristic**, not deliverable detection, so it is imprecise both
992
- ways:
993
- - **False negative:** a post-gate tool *call* clears the flag whether it **succeeded or
994
- errored** — an agent that ran a tool after the gate and still stalled is not caught.
995
- - **False positive:** a deliverable written *before* a final confirmation gate does **not**
996
- clear it, so a write-then-confirm-then-question run is flagged — use `allow_stall: true` for a
997
- deliberate confirm-terminal skill.
998
- - The broad guard is therefore YOUR assertions — assert the deliverable (`file_exists` /
999
- `artifact_json` / `transcript_matches`), never just `result: success`.
1000
- - `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
1001
- stalling *after* one is answered — two different failure modes.
1002
- - **Free-text aside:** the scripted key for a "type-it-in-notes" option is **`answer:`** — an
1003
- arbitrary string delivered verbatim, bypassing label validation by author intent (Cowork
1004
- auto-provides an "Other" free-text path on every gate). Mutually exclusive with `choose:`; setting
1005
- both fails loud. What has no scripted equivalent is the `OTHER:` *directive*
1006
- (it works only on the LLM-decider path, not scripted `choose:`, and only on
1007
- **single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
1008
- options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
1009
- LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
1010
- `OTHER:` is marked `[via Other free-text]` in its `gateProvenance` rationale.
1011
- 14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
1012
- but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
1013
- index lands on a different option (a silent re-record flake). Prefer an exact label when order is
1014
- stable; `lint` flags positional `choose` with an advisory. Unstable option order is also what the
1015
- **user** sees — a reordered gate puts a different choice in the default slot — so pin what was shown
1016
- with `question_options`, rather than only hardening the answer rule against it.
1017
- 15. **A scripted `choose:` matching no offered option HARD-fails the run — `on_unanswered: first` does NOT
1018
- backstop it.** This is distinct from an *unanswered* gate (no rule matched → falls to `on_unanswered`): a
1019
- rule that DID match the gate but whose `choose:` names a label the gate never offered (the model reworded
1020
- it) is treated as an authoring bug and fails loud — `first`/`llm` won't absorb it. The error now prints the
1021
- **offered options** (and a closest-match suggestion), so fix the anchor from the error alone — no need to
1022
- dig through `events.jsonl`. (This is exactly the drift `verify-run` answer-coverage catches in ~1s; use it
1023
- before a paid record.)
1024
- 16. **Batch record keeps going — you don't need a one-at-a-time wrapper.** `record <dir>` and `record <dir>
1025
- --rerecord-stale` run **every** scenario, collect failures, and report them at the end (non-zero exit on
1026
- any failure) — a failing scenario does NOT abort the batch. So a single `cowork-harness record cassettes/
1027
- --rerecord-stale` surfaces ALL stale anchors in one pass (add `--concurrency <N>` to parallelize); a shell
1028
- wrapper that loops one cassette at a time with `set -e` defeats this and rediscovers stale anchors serially.
1029
- Two durability properties make the batch safe to trust: each cassette is written **atomically** (a
1030
- same-directory temp file + rename), so an interrupted or OOM-killed batch never leaves a partial/corrupt
1031
- cassette — a failed scenario simply produces none; and under `--concurrency <N>` each scenario runs **fully
1032
- isolated** (its own egress sidecar network + proxy, its own per-session run dir), so parallel records don't
1033
- cross-talk — the concurrency bound exists only for the Docker address pool + API rate limits, not correctness.
1034
-
1035
- 17. **Editing `scenarios/*.yaml` does NOT change a plain `replay` — the WHOLE scenario is frozen, not just
1036
- `assert:`.** *Why:* a cassette captures every key (`lane:`, `fidelity:`, `baseline:`, `prompt:`, `skills:` …)
1037
- and `replay` evaluates all of them from that frozen copy — byte-deterministic, ignoring the working tree (so
1038
- a committed cassette can't silently re-interpret against an uncommitted YAML). **Only `assert:`
1039
- (+`expect_denied:`) can be opted back to disk; any other edited key reaches a replay only by re-recording.**
1040
- This is *loud* rather than a *silent* no-op: plain `replay` prints a `::notice::` when a sibling's
1041
- `assert:`/`prompt:` differs, and when the sibling **fails to load** at all (a typo'd or too-new key) — and
1042
- points you at the fix.
1043
- *Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
1044
- `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
1045
- `prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
1046
- fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
1047
- are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the `session`
1048
- (model / data mounts / discovery) is NOT drift-checked or fingerprinted, so a **model change** between record
1049
- and re-assert is undetected — the notice flags this; re-record if the session changed. `verify-run` reads
1050
- on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
1051
-
1052
- 18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
1053
- bundle several sub-questions into a single gate; the assertion counts each sub-question, so a
1054
- 3-sub-question bundle counts as 3, not 1. `trace --view questions` shows the same per-gate
1055
- sub-question count and a matching footer total — read that off instead of the tool-call count when
1056
- sizing the budget.
1057
-
1058
- 19. **`gate_answers_delivered` passes vacuously when no gate fires — pair it, or drop it.** Whether a
1059
- gate fires is model-dependent, so `gate_answers_delivered: true` alone can't catch "the gate never
1060
- fired at all". If the scenario is meant to gate, pair it with `gate_answer_count_min: 1` (a floor of
1061
- `0` witnesses nothing). If the scenario is gate-clean by design, drop the key — it asserts nothing
1062
- there — and declare `questions_count_max: 0`, which fails loudly if a gate ever appears. Asserting
1063
- `questions_count_max: 0` alongside a gate-presence key is unsatisfiable: `run`/`skill`/`record`
1064
- refuse it before spending.
1065
-
1066
- 20. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
1067
- read-only folder's files as path + hash only (`truncated: true`, no `body`) — it's an input the agent
1068
- read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass against it on replay
1069
- (the hash-only entry still materializes a placeholder); `artifact_json` reports a clear
1070
- evidence-unavailable on every lane (live/verify-run/replay agree — no green-record/red-replay). This is
1071
- also why a `mode: r` input never trips the `binary` privacy finding or needs `--allow` — only a
1072
- *committed* body is scanned. `scaffold` won't emit `file_exists` for one either (it's not in
1073
- `RunResult.artifacts`). A `mode: rw`/`rwd` folder's contents are captured with a full body, same as
1074
- `outputs/`.
1075
-
1076
- 21. **A `fidelity: cowork` cassette can go stale in a way `skill`/`format` drift won't catch.** Its recorded
1077
- `effectiveFidelity` field pins which concrete tier (`hostloop` or `container`) the baseline resolved to
1078
- AT RECORD TIME. If a later Desktop baseline flips that resolution, `verify-cassettes` reports it as a
1079
- `resolved-tier` finding (re-record — the recording now exercises the wrong tier); a cassette with no
1080
- `effectiveFidelity` at all, or an unloadable pinned `baseline:`, reports `unverifiable-tier` instead
1081
- (couldn't check — also re-record). Both are `fidelity: cowork`-only; an explicit-tier scenario never
1082
- produces them. (Details: [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § tier staleness — repo-only.)
1083
-
1084
- 22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
1085
- `manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
1086
- assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
1087
- yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
1088
- fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
1089
- suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
1090
- rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
1091
- about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
1092
- `lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
1093
- `--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
1094
- 23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
1095
- *Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
1096
- servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
1097
- fail your gate. *Fix:* nothing, unless the scenario asserts `tool_available` on
1098
- `mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
1099
- re-recording would never produce those tools anyway.
1100
-
1101
- 24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
1102
- lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
1103
- emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
1104
- Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
1105
- optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
1106
- other — and probing a remote session makes this harness look like it emulates the wrong tool under the
1107
- wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
1108
- file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
1109
- `present_files_called` assertion keys are harness-side names and stay valid either way.
1110
- ([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
1111
- → "File delivery" has the binary-verified detail; repo-only.)
1112
-
1113
- 25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
1114
- --allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
1115
- (`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
1116
- `record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
1117
- fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
1118
- scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
1119
- pass it. Writing a recording the scan DID flag is the separate `record
1120
- --allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
1121
- exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
1122
- destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
1123
- before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
1124
- **quarantined** to `<runs-root>/quarantine/` with a `.findings.txt` naming what leaked, and the command
1125
- fails without writing the path you asked for (the recording is not discarded — you paid for it).
1126
- `verify-cassettes --allow-host-inventory <regex>` is unrelated: a per-finding suppressor for the
1127
- scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
1128
- wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
1129
-
1130
- 26. **A `skill`-lane `PASS` does not mean the skill ran, or that the run was the one you wanted.** *Why:*
1131
- an open-ended `skill` run has no `assert:` block, so its verdict reports only that **no guard fired**
1132
- (no error, stall, host-path leak, `outputs/` delete, permissive auto-allow or capability gap). On
1133
- `run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
1134
- means N runs cleared the guards — it says nothing about which model served them, whether the skill
1135
- was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
1136
- carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
1137
- `models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
1138
- exactly like skill output is not evidence: the skill's own source is mounted where the model can
1139
- read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
1140
- directly, with `skillActivity` empty.
1141
-
1142
- 27. **`allow_stall: true` is a scenario assertion, so the `skill` lane cannot use it.** *Why:* the
1143
- `stalled` guard fires when a run's final message ends in `?` with no productive tool call after the
1144
- last gate — which includes a complete answer that closes by *offering* a follow-up ("want me to run
1145
- this through a structured pass?"). The documented opt-out lives in an `assert:` block, and an
1146
- open-ended `skill` run has none, so the failure message names a remedy that lane can't perform.
1147
- *Fix:* on `skill`, read the final message before believing `stalled`, or move the check to a
1148
- `run` scenario where `allow_stall: true` is authorable.
1149
-
1150
- For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
1151
- files in `references/` (the gotchas above are the full list; the references repeat only the
1152
- assertion/replay-relevant ones).
1153
-
1154
- ## References
1155
-
1156
- - `references/task-recipes.md` — end-to-end recipes for the four jobs fleet owners actually hit:
1157
- evolve a cassette's `assert:` (usually no re-record), audit a fleet for tier drift, set up
1158
- redaction before the first hostloop/protocol record, derive budget assertions without a
1159
- two-pass record. Start here when the question is "how do I do X", not "what does flag Y mean".
1160
- - `references/scenario-schema.md` — scenario/session YAML schema, full assertion catalog (with each
1161
- key's replay class), the web_fetch model, and an assertion/replay-scoped gotcha subset (the full
1162
- landmine catalog lives in this SKILL's Gotchas section above).
1163
- - `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
1164
- - `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
1165
- GitHub Actions pipeline.
1166
- - `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
1167
- no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
1168
- (validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
1169
- - Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
1170
- alive* (Part II) above; the fuller recipe is in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) (repo-only, not shipped with the
1171
- installed skill).
148
+ | [`references/authoring.md`](references/authoring.md) | session vs scenario, discovery, fidelity tier, the answer-channel decision tree, `web_fetch`, scaffold + lint |
149
+ | [`references/assertions-guide.md`](references/assertions-guide.md) | the two assertion axes, the goal → key map |
150
+ | [`references/run-record-replay.md`](references/run-record-replay.md) | run / `verify-run`, recording and cassette placement, real-document validation, verdict signals (incl. where a relative or `outputs/` path lands per tier), background-run liveness, CI lanes |
151
+ | [`references/measurement.md`](references/measurement.md) | `--repeat`, `--ablate-skill`, measurement hygiene |
152
+ | [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
153
+ | [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
154
+ | [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality |
155
+ | [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
156
+ | [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
157
+ | [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
158
+ | [`references/ci-recipe.md`](references/ci-recipe.md) | the GitHub Action, replay-vs-live lanes, the four-stage pipeline |
159
+ | [`references/critique.md`](references/critique.md) | `critique` report and evidence-package shapes |
160
+ | `scripts/scenario.py` | `scaffold`, `lint`, `lint-skill`, `resolve-agent-types <plugin-dir>` (validates a pinned `subagent_type` against `plugin.json` + `agents/*.md`) |