cowork-harness 3.9.0 → 4.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +74 -1085
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +159 -0
- package/.claude/skills/cowork-harness/references/assertions-guide.md +60 -0
- package/.claude/skills/cowork-harness/references/authoring.md +258 -0
- package/.claude/skills/cowork-harness/references/ci-recipe.md +50 -39
- package/.claude/skills/cowork-harness/references/critique.md +13 -7
- package/.claude/skills/cowork-harness/references/debugging.md +145 -0
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +12 -7
- package/.claude/skills/cowork-harness/references/gotchas.md +285 -0
- package/.claude/skills/cowork-harness/references/measurement.md +56 -0
- package/.claude/skills/cowork-harness/references/run-record-replay.md +339 -0
- package/.claude/skills/cowork-harness/references/scenario-schema.md +41 -174
- package/.claude/skills/cowork-harness/references/task-recipes.md +38 -24
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +10 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +530 -32
- package/.cowork-redact.json +18 -7
- package/.env.example +4 -0
- package/AGENTS.md +7 -1
- package/CHANGELOG.md +792 -0
- package/CONTRIBUTING.md +7 -5
- package/DESIGN.md +2 -2
- package/README.md +20 -9
- package/RELEASING.md +2 -4
- package/SPEC.md +203 -34
- package/baselines/desktop-2.9939.4.json +1116 -0
- package/dist/agent/session.js +7 -0
- package/dist/assert.js +61 -18
- package/dist/baseline.js +6 -0
- package/dist/cli.js +365 -151
- package/dist/critique/command.js +223 -149
- package/dist/critique/evaluator.js +9 -4
- package/dist/critique/limitations.js +1 -1
- package/dist/critique/mount-check.js +52 -0
- package/dist/critique/package-evidence.js +33 -17
- package/dist/decide/decider.js +4 -8
- package/dist/decide/external-channel.js +57 -28
- package/dist/decide/llm-transport.js +2 -0
- package/dist/decide/semantic-judge.js +7 -2
- package/dist/egress/sidecar.js +18 -24
- package/dist/errors.js +38 -5
- package/dist/hostloop/pretooluse-path-hook.js +64 -0
- package/dist/hostloop/process-cwd.js +124 -0
- package/dist/redact.js +10 -2
- package/dist/redactable-literal.js +43 -0
- package/dist/run/analyze-skill.js +5 -3
- package/dist/run/budget.js +8 -2
- package/dist/run/cassette.js +567 -102
- package/dist/run/chat-result.js +4 -2
- package/dist/run/chat.js +40 -22
- package/dist/run/command-globals.js +136 -0
- package/dist/run/doctor.js +8 -6
- package/dist/run/envelope.js +36 -79
- package/dist/run/execute.js +773 -145
- package/dist/run/lint-load.js +235 -0
- package/dist/run/migrate-run-dir.js +3 -1
- package/dist/run/model-provenance.js +30 -17
- package/dist/run/outputs-delete-tier.js +59 -0
- package/dist/run/pre-run-manifest.js +57 -4
- package/dist/run/run.js +103 -11
- package/dist/run/runs-gc.js +4 -2
- package/dist/run/scaffold.js +2 -0
- package/dist/run/scenario-tool.js +193 -30
- package/dist/run/skill-flag-surface.js +9 -1
- package/dist/run/tier-vacuous-tools.js +12 -0
- package/dist/run/timeline-fold.js +48 -19
- package/dist/run/trace-view.js +109 -18
- package/dist/run/verdict.js +55 -10
- package/dist/runtime/agent-image.js +20 -1
- package/dist/runtime/argv.js +1 -0
- package/dist/runtime/hostloop-prompt.js +1 -1
- package/dist/runtime/hostloop.js +61 -16
- package/dist/runtime/microvm.js +111 -4
- package/dist/scan.js +69 -8
- package/dist/secrets.js +1 -1
- package/dist/session.js +87 -68
- package/dist/spawn-guard.js +13 -0
- package/dist/staging/resolve.js +11 -9
- package/dist/termination.js +143 -0
- package/dist/tool-call-assert.js +237 -0
- package/dist/types.js +66 -8
- package/docs/README.md +4 -4
- package/docs/boundary.md +7 -6
- package/docs/cassette.md +119 -19
- package/docs/chat.md +6 -2
- package/docs/ci.md +18 -17
- package/docs/cli.md +120 -35
- package/docs/companion-skill.md +5 -3
- package/docs/critique.md +48 -26
- package/docs/debugging.md +7 -4
- package/docs/decider-dir.md +24 -5
- package/docs/fidelity-gaps.md +84 -72
- package/docs/gotchas.md +3 -3
- package/docs/invariants.md +6 -5
- package/docs/maintenance.md +4 -1
- package/docs/plugin-root.md +32 -6
- package/docs/run-status.md +13 -4
- package/docs/scenario.md +123 -69
- package/docs/session.md +14 -13
- package/docs/stats.md +5 -0
- package/docs/subagents.md +14 -11
- package/examples/README.md +2 -1
- package/examples/replays/README.md +5 -5
- package/examples/replays/example-multiselect-gate.cassette.json +48 -48
- package/examples/replays/example-pdf-skill.cassette.json +92 -95
- package/examples/replays/hostloop-computer-links.cassette.json +58 -58
- package/examples/sessions/default.yaml +1 -1
- package/llms.txt +4 -4
- package/package.json +1 -1
- package/python/test_scenario_lint.py +532 -21
- package/schema/cassette.v13.json +330 -0
- package/schema/critique-report.json +1 -1
- package/schema/run-result.json +71 -4
- package/schema/scenario.schema.json +193 -12
- package/schema/verify-cassettes.json +2 -2
- package/scripts/bump-version.ts +99 -93
- package/scripts/check-claims.ts +32 -19
- package/scripts/check-surface.ts +58 -10
- package/scripts/check-versions.ts +68 -19
- package/scripts/gen-schema.ts +15 -6
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: cowork-harness
|
|
3
|
-
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers,
|
|
3
|
+
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version:
|
|
7
|
-
tracks-harness: cowork-harness
|
|
6
|
+
version: 4.0.0
|
|
7
|
+
tracks-harness: cowork-harness 4.0.0 (baseline desktop-2.9939.4)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,11 +22,12 @@ Anthropic. Say so if a user asks what it is.
|
|
|
22
22
|
The single most important idea: **a green run is not automatically a correct run.** The harness has
|
|
23
23
|
several ways to no-op a check while still producing a green run (skip an assertion on replay — now
|
|
24
24
|
flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an empty egress
|
|
25
|
-
allowlist). This skill exists mostly to keep you out of those traps — the
|
|
26
|
-
the highest-value part.
|
|
25
|
+
allowlist). This skill exists mostly to keep you out of those traps — the *Invariants* below and the
|
|
26
|
+
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
|
+
Read them.
|
|
27
28
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness
|
|
29
|
-
> `desktop-2.9939.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.0.0` (baseline
|
|
30
|
+
> `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
32
|
|
|
32
33
|
## Preflight — make sure the harness can actually run
|
|
@@ -35,14 +36,14 @@ The 10-second inner loop, once the CLI is on PATH:
|
|
|
35
36
|
|
|
36
37
|
```bash
|
|
37
38
|
cowork-harness doctor # prerequisites OK? (Docker, agent, token, baseline)
|
|
38
|
-
cowork-harness skill ./my-skill "do X"
|
|
39
|
+
cowork-harness skill ./my-skill "do X" --model claude-sonnet-5 # run once (or set COWORK_HARNESS_MODEL)
|
|
39
40
|
```
|
|
40
41
|
|
|
41
42
|
Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
|
|
42
43
|
|
|
43
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.0.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.0.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.0.0"`. **Pin `@^4.0.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
47
|
|
|
47
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -50,20 +51,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
50
51
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
51
52
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass. At `protocol` the plugin under test **is** delivered (`--plugin-dir`), but its hooks then run as native host processes — so a plugin declaring hooks is refused there until you pass `--allow-host-hooks` / `allow_host_hooks: true`.
|
|
52
53
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred), or `ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
53
|
-
-
|
|
54
|
+
- **Point at another `.env`** with `--dotenv <path>`, and relocate run output with `--run-dir <path>`; both work before or after the subcommand.
|
|
54
55
|
|
|
55
56
|
## Orient — the three loops
|
|
56
57
|
|
|
57
|
-
Everything you do with the harness is one of **three loops**, and the
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
58
|
+
Everything you do with the harness is one of **three loops**, and the detail lives in one reference
|
|
59
|
+
file per loop: **author** a scenario ([`references/authoring.md`](references/authoring.md)), **run /
|
|
60
|
+
record / lock** it into a reproducible regression ([`references/run-record-replay.md`](references/run-record-replay.md)),
|
|
61
|
+
and **debug** a run that misbehaved or greened when it shouldn't ([`references/debugging.md`](references/debugging.md)).
|
|
61
62
|
|
|
62
63
|
Pick the entry point you need. The first three are the everyday path — a quick liveness check, the
|
|
63
64
|
CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that hang off them:
|
|
64
65
|
|
|
65
|
-
- **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>"
|
|
66
|
-
scenario file.
|
|
66
|
+
- **"Is it even alive?"** (inner loop) → `cowork-harness skill <folder> "<prompt>" --model <id>`. Fastest;
|
|
67
|
+
no scenario file.
|
|
67
68
|
- **Repeatable, asserted regression** → author a `scenarios/*.yaml` and run `cowork-harness run`.
|
|
68
69
|
This is the CI-grade path and most of this skill.
|
|
69
70
|
- **A run failed — or greened and you don't trust it** (the debugging loop) → don't re-run and hope.
|
|
@@ -72,7 +73,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
72
73
|
`cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
|
|
73
74
|
then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
|
|
74
75
|
the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
|
|
75
|
-
output* sections in
|
|
76
|
+
output* sections in [`references/debugging.md`](references/debugging.md) are the detail (the fuller human-facing map lives in
|
|
76
77
|
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md) — repo-only, not shipped with the installed skill).
|
|
77
78
|
**"Evidence" here means the RUN's own record** — events, trace, transcript. `critique`'s evaluator
|
|
78
79
|
grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
|
|
@@ -87,9 +88,9 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
|
|
|
87
88
|
the cost, and it answers that question directly. Report and evidence-package shapes:
|
|
88
89
|
`references/critique.md`.
|
|
89
90
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
90
|
-
TTY, **not** an asserted test — see *Debugging with `chat`* in
|
|
91
|
+
TTY, **not** an asserted test — see *Debugging with `chat`* in `references/debugging.md`).
|
|
91
92
|
**"Interactive" splits two ways — don't take the wrong branch.** Want to answer gates yourself *and*
|
|
92
|
-
still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path*
|
|
93
|
+
still get an asserted, `assert:`-checked run? That is `--decider-dir` (*Choose an answer path* in `references/authoring.md`),
|
|
93
94
|
**not** `chat`. Reach for `chat` only when you are exploring by hand and do NOT want a verdict.
|
|
94
95
|
|
|
95
96
|
> **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
|
|
@@ -102,1070 +103,58 @@ lint-skill · analyze-skill · probe-dispatch ·
|
|
|
102
103
|
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
103
104
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
104
105
|
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
**
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
`
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
> to copy untracked files (won't reflect what ships). A folder **outside** any repo is copied raw (no guard).
|
|
147
|
-
|
|
148
|
-
### Choose a fidelity tier
|
|
149
|
-
|
|
150
|
-
| Tier | What it gives you | Use when |
|
|
151
|
-
|---|---|---|
|
|
152
|
-
| `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
|
|
153
|
-
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
|
|
154
|
-
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
|
|
155
|
-
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
|
|
156
|
-
|
|
157
|
-
Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
|
|
158
|
-
(it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
|
|
159
|
-
`references/fidelity-and-answers.md`.
|
|
160
|
-
|
|
161
|
-
**Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
|
|
162
|
-
`/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
|
|
163
|
-
**remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
|
|
164
|
-
different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
|
|
165
|
-
reproduces it and none can — that container is not something a local tool can stand up. Which lane a
|
|
166
|
-
real session gets is a Cowork setting ("Only on this computer"), observed **off** on a current install.
|
|
167
|
-
So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
|
|
168
|
-
asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
|
|
169
|
-
`lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
|
|
170
|
-
rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
|
|
171
|
-
above).
|
|
172
|
-
|
|
173
|
-
### Choose an answer path (gates: AskUserQuestion + tool-permission)
|
|
174
|
-
|
|
175
|
-
Default to **deterministic**: scripted `answers:` + `on_unanswered: fail`. Anything that brings a
|
|
176
|
-
live model into answering flags the run `nonDeterministic` — keep those out of deterministic
|
|
177
|
-
regressions.
|
|
178
|
-
|
|
179
|
-
<!-- answer-channels:begin -->
|
|
180
|
-
**Pick by asking one question about your situation**, not by scanning a table — the channels are not
|
|
181
|
-
interchangeable and the wrong one either masks a gate or can't run at all:
|
|
182
|
-
|
|
183
|
-
```
|
|
184
|
-
Will this run be re-executed UNATTENDED? (CI, a committed cassette, --repeat, --matrix)
|
|
185
|
-
│
|
|
186
|
-
├─ YES ──► scripted `answers:` / `--answer` / `--answer-policy` + `on_unanswered: fail`
|
|
187
|
-
│ The ONLY reproducible channel. Non-negotiable for CI and committed cassettes.
|
|
188
|
-
│ Labels reworded every run? STAY HERE: pin a stable leading SUBSTRING
|
|
189
|
-
│ (uniqueness-guarded, fails loud) or a positional `choose`. Both keep determinism.
|
|
190
|
-
│
|
|
191
|
-
└─ NO — a discovery / validation run. Who holds the context to answer?
|
|
192
|
-
│
|
|
193
|
-
├─ a model, steered by one line of intent
|
|
194
|
-
│ ──► `--decider-llm --intent "<…>"` [skill · record]
|
|
195
|
-
│ NOT on `run` — there the spelling is the scenario-YAML `on_unanswered: llm`.
|
|
196
|
-
│ Can false-green an oracle-less semantic gate.
|
|
197
|
-
│
|
|
198
|
-
├─ deterministic logic you can write down
|
|
199
|
-
│ ──► `--decider-cmd '<helper>'` [skill · run]
|
|
200
|
-
│ Determinism is your helper's, not the harness's. NOT on `record`.
|
|
201
|
-
│
|
|
202
|
-
├─ YOU, the driving agent, holding the task context
|
|
203
|
-
│ ──► `--decider-dir <FRESH, EMPTY dir>` [skill · run · record]
|
|
204
|
-
│ + `cowork-harness gates <dir> --follow` (arm a Monitor here)
|
|
205
|
-
│ + `cowork-harness answer <dir> --gate N --choose "<label>"`
|
|
206
|
-
│ Its ONE unique property: it needs no advance knowledge of the option SET.
|
|
207
|
-
│ (Label *text* drift alone does not need this — substring anchors handle that.)
|
|
208
|
-
│
|
|
209
|
-
└─ a human at a keyboard, and you are NOT producing a test
|
|
210
|
-
──► `cowork-harness chat` (TTY; no pass/fail verdict — see below)
|
|
211
|
-
```
|
|
212
|
-
|
|
213
|
-
| Channel | Deterministic? | Don't use it when |
|
|
214
|
-
|---|---|---|
|
|
215
|
-
| Scripted | ✅ the CI/agent default | you cannot know the option set in advance |
|
|
216
|
-
| `--decider-llm` / `on_unanswered: llm` | ❌ nonDeterministic | the gate has no oracle a model could judge |
|
|
217
|
-
| `--decider-cmd` | delegated to your helper | the logic needs task context code doesn't have |
|
|
218
|
-
| `--decider-dir` | ❌ nonDeterministic | nobody is present to drive it — it BLOCKS per gate |
|
|
219
|
-
| `on_unanswered: first` | ❌ nonDeterministic | the answer matters — it *masks* the gate |
|
|
220
|
-
|
|
221
|
-
**Cost of `--decider-dir`, stated plainly:** flags the run `nonDeterministic`; needs a live driver + a
|
|
222
|
-
Monitor, so it is **unusable unattended**; blocks at each gate, strictly serial; needs a fresh empty dir
|
|
223
|
-
per run (a dirty one is refused); rejected with `--repeat`, `--on-unanswered`, `--decider-cmd`, and with
|
|
224
|
-
`--matrix --concurrency > 1`; and a cassette recorded this way carries a **re-record cost** — regenerating
|
|
225
|
-
it needs the driver present again.
|
|
226
|
-
|
|
227
|
-
**Rehearse it in ~2s before wiring it into a real run** — `cowork-harness decide --decider-dir <dir>` fires
|
|
228
|
-
one sample gate through the same channel, then blocks (10-min backstop) until you answer it with the two
|
|
229
|
-
commands above. It is the cheapest way to see the protocol work. Full recipe, including the multiSelect
|
|
230
|
-
wire shape and the `gates --follow` Monitor loop:
|
|
231
|
-
[`docs/decider-dir.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/decider-dir.md)
|
|
232
|
-
(repo-only — an npm install ships it at `node_modules/cowork-harness/docs/decider-dir.md`). <!-- npm-only-ok -->
|
|
233
|
-
|
|
234
|
-
**It is a FEEDER for the scripted default, not a rival.** `record --decider-dir` is a first-class way to
|
|
235
|
-
*produce* a cassette: the non-reproducibility is spent once at authoring time and the cassette replays
|
|
236
|
-
deterministically forever. The loop is **discover → transcribe → script** — answer live, then paste the
|
|
237
|
-
run's echoed `--answer "<q>=<choice>"` footer lines into the scenario's `answers:` so re-records go back to
|
|
238
|
-
being unattended. Skip the transcribe step only for one-off/exploratory runs.
|
|
239
|
-
<!-- answer-channels:end -->
|
|
240
|
-
|
|
241
|
-
**For a QUESTION gate, never hand-write the `req-N.json`/`resp-N.json` files.** `gates` and `answer` wrap
|
|
242
|
-
the protocol — the atomic temp+rename, the `{id, answers}` envelope, the multiSelect array shape.
|
|
243
|
-
Hand-rolling a Monitor over the raw files is the single most common mistake on this channel.
|
|
244
|
-
|
|
245
|
-
`answer` writes `{id, answers}` and nothing else, so it covers **question gates only**. The channel also
|
|
246
|
-
carries **permission**, **dialog** and **elicit** gates, whose replies need `{behavior}` / `{action}` — for
|
|
247
|
-
those, write `resp-N.json` yourself, following the `reply_with` template the gate's own `req-N.json`
|
|
248
|
-
advertises (it spells out the exact shape, e.g. `{"id":"…","behavior":"allow|deny"}`).
|
|
249
|
-
|
|
250
|
-
Exact accepted values (teach precisely): `--on-unanswered` takes `fail|prompt|first` on `skill`,
|
|
251
|
-
only `fail|first` on `run`. **`llm` is NOT an `--on-unanswered` value** — the bare flag
|
|
252
|
-
`--on-unanswered llm` is rejected (use `--decider-llm`); the YAML spelling is `on_unanswered: llm`.
|
|
253
|
-
The word `agent` is **retired** — do not write `on_unanswered: agent` (the schema rejects it).
|
|
254
|
-
`--on-unanswered` also conflicts with `--decider-dir`/`--decider-cmd`/`--decider-llm` (the channel or
|
|
255
|
-
model IS the terminal, so a policy alongside it never applies) — pass one, not both. On `record`, a
|
|
256
|
-
scenario setting `on_unanswered: prompt` is rejected too: the YAML field outranks the flag, and a TTY
|
|
257
|
-
wait can't produce a deterministic committed fixture.
|
|
258
|
-
`--on-unanswered first` is itself flagged `nonDeterministic` — it is *not* a deterministic stand-in
|
|
259
|
-
for scripted answers. See `references/fidelity-and-answers.md`.
|
|
260
|
-
|
|
261
|
-
**Which gates to anchor (re-record robustness).** The model rewords option labels (and sometimes the
|
|
262
|
-
question) every run, so a brittle exact-label `choose:` is itself a re-record-fragility source — it drifts and
|
|
263
|
-
forces a re-record. The practical rule: **label-anchor only the gates whose choice drives an `assert:`** (or
|
|
264
|
-
materially changes behavior); for gates whose answer is immaterial to your assertions, `on_unanswered: first`
|
|
265
|
-
is the more re-record-robust choice — accept the `nonDeterministic` flag rather than trade it for a flaky
|
|
266
|
-
anchor. (When label *order* is stable but the text drifts, a positional `choose` is the middle option — the
|
|
267
|
-
linter flags positional `choose` as order-dependent, so use it deliberately.) The caution stands: `first`
|
|
268
|
-
*masks* an unanswered gate, so don't use it for a gate you actually need answered a specific way.
|
|
269
|
-
|
|
270
|
-
**Drifting label TEXT and an unknowable option SET are different problems — don't reach past the cheap
|
|
271
|
-
fix.** Text that rewords while the choices stay the same is a *scripted* problem with a deterministic
|
|
272
|
-
answer: a uniqueness-guarded leading substring, or a positional `choose`. Only when you cannot know what
|
|
273
|
-
the options will *be* — they're generated per input document, so no anchor can be written in advance — does
|
|
274
|
-
the answer move to a live channel (`--decider-dir` if you're driving, `--decider-llm` if nobody is).
|
|
275
|
-
|
|
276
|
-
#### External deciders and the "first" shorthand
|
|
277
|
-
|
|
278
|
-
When using `--decider-cmd` or `--decider-dir`, the helper's output is passed through
|
|
279
|
-
`coerceLabel` **with the "first" shorthand disabled**. This means a helper that returns the literal
|
|
280
|
-
string `"first"` must match an actual label named `"first"` — it is **not** coerced to option 1.
|
|
281
|
-
This prevents a helper bug (accidentally emitting `"first"`) from silently green-ing option 1.
|
|
282
|
-
|
|
283
|
-
The `"first"` shorthand remains active only for the built-in `--on-unanswered first` path. If you
|
|
284
|
-
write an external helper, return a label name or option index — never the bare word `"first"` unless
|
|
285
|
-
your gate actually has a label called `"first"`.
|
|
286
|
-
|
|
287
|
-
### Assertions: two orthogonal axes
|
|
288
|
-
|
|
289
|
-
Conflating these is the **biggest landmine**. An assertion key has two independent properties:
|
|
290
|
-
|
|
291
|
-
- **Axis A — robust to LLM phrasing drift?** Structural/boundary keys (`subagent_dispatched`,
|
|
292
|
-
`egress_*`, `file_exists`, `user_visible_artifact`, `result`) are robust. Free-text content is
|
|
293
|
-
not: match prose with `transcript_matches` / `transcript_contains` (stable lexical markers only —
|
|
294
|
-
not semantic content the model paraphrases, which re-records red); check structured JSON with YAML
|
|
295
|
-
`artifact_json` (or the [pytest lane](https://github.com/yaniv-golan/cowork-harness/blob/main/python/README.md) for complex predicates), not via a transcript substring.
|
|
296
|
-
- **Axis B — survives `replay`?** *Independent of Axis A.* On the token-free `replay` lane, only
|
|
297
|
-
**content keys** evaluate; filesystem / egress keys are skipped (live-only) — loudly, via an
|
|
298
|
-
`::warning::` annotation, not a silent no-op. A key
|
|
299
|
-
being "robust" says nothing about whether it runs on your replay gate.
|
|
300
|
-
|
|
301
|
-
Getting Axis B wrong means a check that **does nothing in CI** — the harness warns loudly when it skips
|
|
302
|
-
(an `::warning::` annotation, not a silent no-op — see the Axis B bullet above), and the bundled linter
|
|
303
|
-
catches it before you push — run it (see *Scaffold a valid scenario, then lint before you push* below).
|
|
304
|
-
|
|
305
|
-
See `references/scenario-schema.md` for the full assertion catalog with each key's replay class.
|
|
306
|
-
|
|
307
|
-
#### Which assertion for which question (goal → key)
|
|
308
|
-
|
|
309
|
-
Beyond the outcome/content keys most scenarios reach for first (`result`, `transcript_*`,
|
|
310
|
-
`file_exists`/`user_visible_artifact`, `artifact_json`), the harness surfaces the agent's *behavior*
|
|
311
|
-
— tool health, sub-agent work, panels, skill attribution, resources — as assertable keys. Reach for
|
|
312
|
-
them by what you're trying to prove:
|
|
313
|
-
|
|
314
|
-
| You want to check that… | Reach for |
|
|
106
|
+
## Invariants — how a green run lies
|
|
107
|
+
|
|
108
|
+
Each of these has produced a green run that tested nothing. The full catalog, with the reasoning
|
|
109
|
+
behind each, is [`references/gotchas.md`](references/gotchas.md).
|
|
110
|
+
|
|
111
|
+
1. **`result: success` is not "the task completed".** It means the agent didn't error. Assert the
|
|
112
|
+
deliverable (`file_exists` / `artifact_json` / `transcript_matches`). A `skill`-lane `PASS` only means
|
|
113
|
+
no guard fired: read `skillsInvoked`, `models` and `ablated` before concluding anything from it.
|
|
114
|
+
2. **`replay` skips live-only keys.** Filesystem and egress keys are skipped on replay (loudly), so a
|
|
115
|
+
mixed item like `{result, egress_denied}` greens on its content half. Keep one concern per `assert:`
|
|
116
|
+
item, put live-only checks on a live gate, and run `cowork-harness lint`.
|
|
117
|
+
3. **Only scripted answers reproduce.** Scripted `answers:` + `on_unanswered: fail` is the CI channel.
|
|
118
|
+
`first`, an LLM decider and `--decider-dir` all flag the run `nonDeterministic`, and `first` masks
|
|
119
|
+
the gate it answers.
|
|
120
|
+
4. **`replay` evaluates the FROZEN scenario.** Editing `scenarios/*.yaml` changes nothing on a plain
|
|
121
|
+
`replay`: re-check an `assert:` edit with `replay --assert-from <file>`, and re-record for any other key.
|
|
122
|
+
5. **Some keys pass on absence.** `gate_answers_delivered` passes when no gate fired — pair it with
|
|
123
|
+
`gate_answer_count_min: 1`. `tool_called` proves a tool ran, not that it was attempted.
|
|
124
|
+
6. **An untracked skill mounts empty.** `git add` a new skill before testing it, and commit before
|
|
125
|
+
recording the cassette that locks it.
|
|
126
|
+
7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
|
|
127
|
+
(`container` serves `mcp__workspace__web_fetch`, not `WebFetch`), and every tier models Cowork's
|
|
128
|
+
desktop-local lane only.
|
|
129
|
+
8. **A WARN signal never blocks a green.** Read the verdict signals after every run
|
|
130
|
+
(`prompt_asset_missing`, `undelivered_deliverables`, `model_fallback`, …).
|
|
131
|
+
|
|
132
|
+
## Short workflows
|
|
133
|
+
|
|
134
|
+
- **Author, then lock:** `cowork-harness scaffold --name … --prompt …` → `cowork-harness lint scenarios/` →
|
|
135
|
+
`cowork-harness record <file.yaml> --dry-run` (free) → `record` once, with `--out` at a tracked path
|
|
136
|
+
(a cassette cannot be moved) → `replay` on the PR gate.
|
|
137
|
+
- **Fix answers without paying:** `--keep` one run → `trace <run-dir> --view questions` → edit
|
|
138
|
+
`answers:` → `verify-run <run-dir> <scenario.yaml>` → record once.
|
|
139
|
+
- **Debug:** the triage table in `references/debugging.md` → `inspect`, `trace --view …`,
|
|
140
|
+
`verify-run`, `diff`. For a green you don't trust: `replay --explain`, then the gotchas.
|
|
141
|
+
- **Measure:** `--repeat N` for flakiness. `--ablate-skill` runs the control arm only; run the treatment
|
|
142
|
+
arm yourself, with the model pinned and a recoverable source frozen.
|
|
143
|
+
|
|
144
|
+
## References — where the detail lives
|
|
145
|
+
|
|
146
|
+
| File | Read it for |
|
|
315
147
|
|---|---|
|
|
316
|
-
|
|
|
317
|
-
|
|
|
318
|
-
|
|
|
319
|
-
|
|
|
320
|
-
|
|
|
321
|
-
|
|
|
322
|
-
|
|
|
323
|
-
|
|
|
324
|
-
|
|
|
325
|
-
|
|
|
326
|
-
|
|
|
327
|
-
|
|
|
328
|
-
|
|
|
329
|
-
| the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
|
|
330
|
-
| the user was **told something specific** at a gate | `question_context: {when_question, matches}` — a regex over the question label + option labels + option **descriptions**. Reach for this when the wording may land in an option's `description`, which `question_asked` and `question_options` cannot see |
|
|
331
|
-
| a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
|
|
332
|
-
| every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
|
|
333
|
-
| a context compaction happened | `compaction_occurred: true` |
|
|
334
|
-
|
|
335
|
-
Every one of these still obeys the two axes above — several are live-only or need a `controlOut`
|
|
336
|
-
cassette on replay, so check the catalog's replay class before putting one on a PR gate.
|
|
337
|
-
`cowork-harness assertions --list` prints the full, always-current key set with one-line semantics
|
|
338
|
-
straight from the schema — treat it (and the catalog) as the source of truth; this map is a
|
|
339
|
-
goal-oriented index into it, not a second catalog.
|
|
340
|
-
|
|
341
|
-
### web_fetch (fail-closed, two-path)
|
|
342
|
-
|
|
343
|
-
`web_fetch` behaves unlike `curl`. A URL is gated by **provenance**, not the egress allowlist:
|
|
344
|
-
|
|
345
|
-
- A URL is *provenanced* iff it appeared in the **prompt** or a **prior `web_fetch` result**. To
|
|
346
|
-
make a fetch succeed, put the URL in the prompt.
|
|
347
|
-
- **Provenanced** → fetches (still SSRF-guarded per redirect hop); the egress hostname allowlist is
|
|
348
|
-
**not consulted**.
|
|
349
|
-
- **Not provenanced** → raises a per-domain approval gate (`webfetch:<domain>`) that is
|
|
350
|
-
**fail-closed** (it is *not* auto-allowed; `--on-unanswered first` won't allow it). Answer it with
|
|
351
|
-
a scripted rule (`when_tool: "webfetch:<domain>"` + `grant: domain|once`), a session
|
|
352
|
-
`web_fetch.approved_domains`, or a live decider.
|
|
353
|
-
|
|
354
|
-
Surprise to remember: adding a host to `egress.extra_allow` is a **no-op** for a provenanced fetch.
|
|
355
|
-
Full model in `references/scenario-schema.md`.
|
|
356
|
-
|
|
357
|
-
### Scaffold a valid scenario, then lint before you push
|
|
358
|
-
|
|
359
|
-
Don't hand-write the YAML from memory — that's how invented keys (`assertions:` vs `assert:`,
|
|
360
|
-
`json_file`, `answer_policy`) creep in. Start from the bundled generator, which emits the
|
|
361
|
-
known-good skeleton (right tier, scripted `answers:` + `on_unanswered: fail`, content assertions
|
|
362
|
-
separated from live-only ones, one concern per item) and **self-lints its own output**. The
|
|
363
|
-
generator is the bundled `scripts/scenario.py` — installed as a plugin, point `S` at
|
|
364
|
-
`${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`; from a repo checkout, use the literal path below:
|
|
365
|
-
|
|
366
|
-
```bash
|
|
367
|
-
S=".claude/skills/cowork-harness/scripts/scenario.py"
|
|
368
|
-
python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
|
|
369
|
-
--prompt "Generate the weekly report to outputs/report.md." \
|
|
370
|
-
--content 'weekly report' --artifact outputs/report.md \
|
|
371
|
-
--egress-allowed api.weather.example.com --out scenarios/report-check.yaml
|
|
372
|
-
```
|
|
373
|
-
|
|
374
|
-
Then lint every scenario — it encodes the no-silent-false-green invariants. Use the CLI wrapper
|
|
375
|
-
`cowork-harness lint` (it runs the same bundled `scenario.py lint`):
|
|
376
|
-
|
|
377
|
-
```bash
|
|
378
|
-
cowork-harness lint scenarios/*.yaml
|
|
379
|
-
```
|
|
380
|
-
|
|
381
|
-
`lint` flags: filesystem/egress-only assertions on a `replay` gate (silent no-op), bad regex
|
|
382
|
-
quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `hostloop`/`protocol`
|
|
383
|
-
(ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
|
|
384
|
-
baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
|
|
385
|
-
`allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
|
|
386
|
-
unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
|
|
387
|
-
`present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
|
|
388
|
-
to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
|
|
389
|
-
`protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
|
|
390
|
-
and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
|
|
391
|
-
(CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
|
|
392
|
-
emits a scenario `lint` would reject.
|
|
393
|
-
|
|
394
|
-
**`lint` is the LENIENT check — the loader is the strict one.** An unknown top-level key is a ⚠ WARN in
|
|
395
|
-
`lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
|
|
396
|
-
that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
|
|
397
|
-
`cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
|
|
398
|
-
`✗ broken:` file and exits 1). **Read the exit code, not just its sign:** `record <file>` — with or without
|
|
399
|
-
`--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
|
|
400
|
-
pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
|
|
401
|
-
as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
|
|
402
|
-
but **`replay` does not**: a frozen top-level key it doesn't recognize (e.g. `lane:` recorded pre-1.16.0) is
|
|
403
|
-
silently ignored and can flip a lane-sensitive verdict green; only frozen **assertion** keys stay
|
|
404
|
-
hard-rejected there. Full split + the v11 version-regime:
|
|
405
|
-
[docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md#unknown-keys-the-loader-is-strict-lint-is-lenient).
|
|
406
|
-
|
|
407
|
-
## Part II — RUN, RECORD & LOCK
|
|
408
|
-
|
|
409
|
-
You have an authored scenario. This Part runs it, reads the verdict, locks it into a
|
|
410
|
-
byte-deterministic cassette, checks a background run's liveness, and places the assertions in the
|
|
411
|
-
right CI lane.
|
|
412
|
-
|
|
413
|
-
### Run, then lock determinism
|
|
414
|
-
|
|
415
|
-
Read the verdict and the inline failing transcript. To pin a flaky-because-stochastic gate, paste
|
|
416
|
-
the echoed `--answer "<q>=<choice>"` footer lines back into the scenario's `answers:` for a
|
|
417
|
-
deterministic re-run. Use `cowork-harness trace <id>` to digest a run. If only an *assertion* is wrong (the
|
|
418
|
-
run itself was fine), `cowork-harness verify-run <run-dir> <scenario.yaml>` re-checks the `assert:` block against
|
|
419
|
-
a **kept** run dir (`--keep`, or a `--session-id` run) with no live re-record — tokens-free, ~1s per iteration.
|
|
420
|
-
When the scenario declares `answers:`, verify-run **also** checks they still match the run's actual gates (a
|
|
421
|
-
reworded gate or a `choose:` the run never offered fails here in ~1s instead of on a paid re-record). Or skip
|
|
422
|
-
the discovery/encode/record dance entirely and answer gates **live during the recording** with
|
|
423
|
-
`record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
|
|
424
|
-
`run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
|
|
425
|
-
`cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
|
|
426
|
-
refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
|
|
427
|
-
cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
|
|
428
|
-
guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
|
|
429
|
-
one a paid run would give. On a **directory** the path-dependent verdicts (host-inventory, cassette
|
|
430
|
-
portability) are reported as `⚠ would-refuse (advisory)` / `⚠ would-warn (advisory)` notes — the label follows
|
|
431
|
-
the verdict kind, and portability can only ever warn — that do NOT affect the exit code — a dir target
|
|
432
|
-
takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
|
|
433
|
-
contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
|
|
434
|
-
real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
|
|
435
|
-
reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
|
|
436
|
-
|
|
437
|
-
**Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
|
|
438
|
-
concluded the free pre-flight was unavailable:
|
|
439
|
-
- **"Does my whole corpus still load?"** → the **directory** arm (`record scenarios/ --dry-run --quiet`,
|
|
440
|
-
the CI shape in `references/ci-recipe.md`). It reports every offender in one pass, and the
|
|
441
|
-
destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
|
|
442
|
-
are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
|
|
443
|
-
knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
|
|
444
|
-
with no `prompt:` key reports as `· skipped:` rather than broken (so a renamed or mis-indented
|
|
445
|
-
`prompt:` reads as "not a scenario" and the batch still exits 0), exit 1 covers **refused**
|
|
446
|
-
(path-independent only — prompt policy, assert contradiction, duplicate target) as well as broken, and
|
|
447
|
-
`--quiet` suppresses the advisory notes entirely — it keeps `✗ broken:` / `✗ refused:` / `· skipped:`,
|
|
448
|
-
which is what you want in CI but means the notes are not a thing you will see there.
|
|
449
|
-
- **"Would THIS record be refused?"** → the **single-file** arm **with the flags and `--out` the real
|
|
450
|
-
record will get**. The destination it evaluates is `--out` if given, else `cassettes/<slug>.cassette.json`
|
|
451
|
-
*relative to your cwd* — so previewing from the repo root a record that really runs from a subdirectory
|
|
452
|
-
asks about a path that may not even exist, and a scenario whose cassette IS committed can come back
|
|
453
|
-
refused. Point it at the real destination and the answer is binding.
|
|
454
|
-
|
|
455
|
-
Neither question is answered by passing `--allow-host-inventory-fixture` to get past the refusal: that
|
|
456
|
-
flag is consent for a recording you intend to make, and reaching for it as a load-check habit is how it
|
|
457
|
-
stops meaning anything.
|
|
458
|
-
|
|
459
|
-
**Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
|
|
460
|
-
Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
|
|
461
|
-
default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
|
|
462
|
-
That choice is permanent: the cassette rewrites `scenario.session` and `scenarioSource` **relative to
|
|
463
|
-
its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
|
|
464
|
-
into another repo — leaves those unresolvable and
|
|
465
|
-
`verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
|
|
466
|
-
re-record at the new location — or point `replay`/`verify-cassettes` at the session with `--session <file>`, which resolves it without a re-record. Since 2.0.0 a bare `replay` FAILS on this class rather than warning. **`record` now says so BEFORE it spends:** a pre-flight — at the same
|
|
467
|
-
pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
|
|
468
|
-
warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
|
|
469
|
-
outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
|
|
470
|
-
cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
|
|
471
|
-
missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
|
|
472
|
-
(`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25 below).
|
|
473
|
-
The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
|
|
474
|
-
redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
|
|
475
|
-
for a cassette that cannot verify staleness from its own location — recoverable only by passing
|
|
476
|
-
`--session <file>` on every invocation thereafter.
|
|
477
|
-
|
|
478
|
-
**Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
|
|
479
|
-
scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
|
|
480
|
-
(and `verify-run`) read the gates + every offered option's **label and `description`** out of that run's
|
|
481
|
-
`events.jsonl` for free — a skill routinely puts the sentence the user is actually deciding on in a
|
|
482
|
-
`description`, and `question_context:` is the key that gates on it (`question_options:` compares labels
|
|
483
|
-
only). When a view renders no field you need, read `events.jsonl` directly rather than concluding the text
|
|
484
|
-
was never delivered — the views are a digest, and the record is wider (`jq` recipes in
|
|
485
|
-
[`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)). Iterate
|
|
486
|
-
your `answers:` against that kept run, then record once. **But the kept run is a snapshot:** if you change the
|
|
487
|
-
skill's gate phrasing afterward, re-`--keep` — verify-run's answer-coverage *refuses* (exit 2, "predates the
|
|
488
|
-
current skill") rather than vouch against stale labels, but the trace/inspect path can't warn you, so re-keep
|
|
489
|
-
deliberately. (Same fail-closed family: corrupt gate evidence — unparseable `events.jsonl` lines, or fewer
|
|
490
|
-
gates than `trace.json` recorded questions — a structurally invalid `result.json`, a `command:"replay"`
|
|
491
|
-
result (a replay is a re-check of a recorded cassette, not run evidence — verify the original live run dir),
|
|
492
|
-
and a `mode:"chat"` result (chat carries no assertions or verdict by contract) also refuse rather than
|
|
493
|
-
certify.) (A token-free probe of "which gates fire" isn't possible — gates are model-decided per run.)
|
|
494
|
-
|
|
495
|
-
Run artifacts are written to `~/.cowork-harness/runs/…` by default — **outside any working tree**, so a run
|
|
496
|
-
launched from a repo root never drops sensitive skill inputs/outputs into it. Pass `--run-dir <path>` (or set
|
|
497
|
-
`COWORK_HARNESS_RUNS_DIR`) to relocate; in CI point it at a workspace path so an artifact-upload step can
|
|
498
|
-
collect the runs.
|
|
499
|
-
|
|
500
|
-
#### Validate a skill against real documents (not a cassette)
|
|
501
|
-
|
|
502
|
-
The loops above build **deterministic regressions**. A different job — drive a skill against *real* input
|
|
503
|
-
documents to judge whether it actually does the work (extraction, analysis), with no intent to record a
|
|
504
|
-
cassette — has its own recipe:
|
|
505
|
-
|
|
506
|
-
1. **Explore with the LLM decider.** `cowork-harness skill <dir> --decider-llm --intent "<one line of what
|
|
507
|
-
this run is testing>"` lets a model (Sonnet default) answer each gate steered by your intent. The model replies with
|
|
508
|
-
the option **number** and the harness maps it to the exact label (so it can't whiff by mis-typing the
|
|
509
|
-
label text); an out-of-set answer fails loud. This is exploration, **not** a deterministic regression —
|
|
510
|
-
the run is flagged non-deterministic and a green here is not a scripted pass. The answering model
|
|
511
|
-
defaults to a Sonnet id (a weaker model tends to prose-decline an ambiguous judgment gate → fail-loud);
|
|
512
|
-
override it with `--decider-model <id>` — a cheaper model (e.g. Haiku) for simple gates to cut cost,
|
|
513
|
-
or Opus for the hardest judgment gates; it won't make an under-specified gate deterministic. A live
|
|
514
|
-
decider can false-green a semantic assertion on an oracle-less gate — see `references/fidelity-and-answers.md`.
|
|
515
|
-
2. **Script the load-bearing gates — especially binary confirm gates.** Once you know which gates fire
|
|
516
|
-
(`trace <run-dir> --view questions`), pin the ones whose choice drives the outcome with
|
|
517
|
-
`--answer "<q>=<label>"` / `--answer-policy <yaml>`. When a skill **re-words its option labels run-to-run**
|
|
518
|
-
(LLM-authored gates), pin a **stable leading substring** instead of the full label — `--answer
|
|
519
|
-
"<q>=Israeli company"` binds whichever option starts with `Israeli company`. It is uniqueness-guarded and
|
|
520
|
-
**fails loud** if the anchor ever matches two options (the documented trade: drift-tolerance, not strict
|
|
521
|
-
CI reproducibility — for that, pin a full exact label or a free-text `answer:`).
|
|
522
|
-
3. **Budget ~1 re-run per file.** If a gate whiffs, the run does not vanish — it exits non-zero but
|
|
523
|
-
**salvages a PARTIAL run** (the extraction the agent already did is written to disk). So the cost of a
|
|
524
|
-
missed gate is one re-run with a better `--intent` or a scripted answer, not a lost paid run.
|
|
525
|
-
4. **Inspect the outputs to judge correctness.** `cowork-harness inspect <run-dir>` shows what the run
|
|
526
|
-
produced — the artifacts plus a shallow field preview of each JSON artifact (e.g. the extracted figures).
|
|
527
|
-
It works on a salvaged partial run too. (A partial run is marked `PARTIAL`; `verify-run` and `scaffold`
|
|
528
|
-
refuse to treat its half-finished output as a passing result.)
|
|
529
|
-
5. **For image-only / scanned PDFs, use the full-parity image.** The default agent image omits OCR and
|
|
530
|
-
PDF-table tooling; if a **scenario** sets `requires_capabilities` (a scenario field — not skill
|
|
531
|
-
frontmatter) and the image provably omits one, the harness **aborts before the paid run (exit 3)** —
|
|
532
|
-
unless the scenario asserts `allow_missing_capability: true`, which downgrades it to a notice and
|
|
533
|
-
proceeds. Rebuild with `--build-arg COWORK_FULL_PARITY=1` and point `COWORK_AGENT_IMAGE` at it for those
|
|
534
|
-
skills.
|
|
535
|
-
6. **Iterate across fixes — verify before you trust, and don't cross-pair generations.** A green run is
|
|
536
|
-
not a correct run, and a skill's self-reported finding is not real until its cited evidence is found in
|
|
537
|
-
the run's own output. Ground each finding against `result.json` (`finalMessage` = the skill's own
|
|
538
|
-
answer/critique; `toolResults` = tool outputs) and the tool-call stream via
|
|
539
|
-
`cowork-harness trace <run-dir> --output-format json` — add `--full-results` so a successful call's full
|
|
540
|
-
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
541
|
-
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
542
|
-
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
543
|
-
kept run predates the current skill). **The hazard is general, not critique-specific:** repeated
|
|
544
|
-
`run`/`skill` invocations of one scenario accumulate in the SAME scenario directory regardless of skill
|
|
545
|
-
version, so a plain `stats <scenario>` silently averages pre-fix and post-fix runs together. Compare
|
|
546
|
-
generations with **`stats <scenario> --group-by skill-hash`** (or narrow with `--skill-hash <prefix>` /
|
|
547
|
-
`--label <tag>`); an un-split window spanning more than one generation now warns.
|
|
548
|
-
**Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
|
|
549
|
-
skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
|
|
550
|
-
critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
|
|
551
|
-
pairing step there silently groups on an absent key instead of erroring — check the field is present, or
|
|
552
|
-
require ≥ 1.5.0. See [`docs/debugging.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/debugging.md)
|
|
553
|
-
(repo-only) for the full loop.
|
|
554
|
-
|
|
555
|
-
#### Interpreting verdict signals
|
|
556
|
-
|
|
557
|
-
The run verdict may include `WARN`-severity signals in addition to pass/fail. One to watch for:
|
|
558
|
-
|
|
559
|
-
- **`prompt_asset_missing`** — the run proceeded but a prompt asset referenced by the scenario was
|
|
560
|
-
not found. The model ran against an incomplete prompt. This is a `WARN`, not a hard failure, so
|
|
561
|
-
the run can still green. If you see it, fix the asset path — a green with a missing asset is
|
|
562
|
-
not a valid pass.
|
|
563
|
-
|
|
564
|
-
**False negatives — signals that are tier/image artifacts, not skill defects.** Some fail-severity
|
|
565
|
-
signals read like a skill gap but are really a property of the reduced test image or the fidelity tier.
|
|
566
|
-
Recognize these before "fixing" a non-bug:
|
|
567
|
-
|
|
568
|
-
- **`missing_capability`** — the lean `core` agent image is a deliberate partial mirror of real Cowork's
|
|
569
|
-
rootfs, so a skill that used `soffice`/LibreOffice (`office_convert`), `tesseract` (`ocr`),
|
|
570
|
-
`markitdown`/`magika` (`ml_extract`), `cv2` (`cv`), `camelot`/`tabula` (`pdf_tables`), or `wand`
|
|
571
|
-
(`magick`) can trip this even though real Cowork **ships** those (per the rootfs manifest captured at
|
|
572
|
-
Desktop `2.9939.2` — `baselines/provisioning/rootfs-provisioning.json`, which is the dated evidence
|
|
573
|
-
behind that sentence). The message says so ("likely a FALSE
|
|
574
|
-
NEGATIVE (real Cowork ships them)"). Fix: rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`, point
|
|
575
|
-
`COWORK_AGENT_IMAGE` at it), or — if the skill's fallback is genuinely equivalent — assert
|
|
576
|
-
`allow_missing_capability: true`. (Two sources: a skill *observed using* an omitted family, live lane;
|
|
577
|
-
or a declared `requires_capabilities` the tier can't provide, both lanes — an unknown family name
|
|
578
|
-
hard-fails rather than silently passing.) **On an open-ended `skill` run** (no `assert:` block to carry
|
|
579
|
-
the modifier), pass **`--allow-missing-capability`** — the CLI equivalent of the assertion.
|
|
580
|
-
- **`ended_with_question`** (`WARN`, live lane) — a heuristic: the agent's final answer contains a
|
|
581
|
-
question and the run wrote **no deliverable to `outputs/`** — it may have ended on a request for input
|
|
582
|
-
instead of finishing. Warn-only; the fix is scripting/steering the answer (`answer:` / `--answer` / a
|
|
583
|
-
decider, or `--decider-llm --intent`), not editing the skill's prose. The strict, fail-severity sibling
|
|
584
|
-
`stalled` already catches a *trailing*-`?` final turn that did no tool work after the last gate; this
|
|
585
|
-
covers the residual (a mid-message `?`, or tool work after the last gate that still ended asking). Read
|
|
586
|
-
the final message before acting — a legitimate question-posing answer that wrote a file never fires.
|
|
587
|
-
Assert `allow_stall: true` if ending on a question is the intended terminal state.
|
|
588
|
-
- **`undelivered_deliverables`** (`WARN`) — the skill produced file(s) **outside every user-visible root**
|
|
589
|
-
and never delivered them. On a **remote** Cowork session the workspace is reclaimed at session end, so
|
|
590
|
-
they are destroyed; on a **local** one they persist but stay invisible to the user. Either way the user
|
|
591
|
-
does not get them. It fires with no assertion written — `present_files_called` covers the positive case
|
|
592
|
-
only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
|
|
593
|
-
**Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
|
|
594
|
-
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
595
|
-
**The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
|
|
596
|
-
can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
|
|
597
|
-
(what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
|
|
598
|
-
`outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
|
|
599
|
-
`fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
|
|
600
|
-
bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
|
|
601
|
-
a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
|
|
602
|
-
decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
|
|
603
|
-
[docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
|
|
604
|
-
nothing is delivered by location there, so only an explicit delivery counts. Assert
|
|
605
|
-
**`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
|
|
606
|
-
downloaded inputs) rather than a delivery gap.
|
|
607
|
-
- **`delivery_unobservable`** (`WARN`, `lane: remote` only) — the run produced file(s) but the harness
|
|
608
|
-
serves **no delivery tool on that lane**, so whether they reached the user is unanswerable. This is the
|
|
609
|
-
honest cannot-verify companion to `undelivered_deliverables`: reporting every remote file as undelivered
|
|
610
|
-
would claim more than the evidence supports, and staying silent would read as clean. Mutually exclusive
|
|
611
|
-
with `undelivered_deliverables`, and quiet on a run that produced nothing to deliver. Not a skill defect —
|
|
612
|
-
a harness coverage gap (see the *File delivery* section of fidelity-gaps).
|
|
613
|
-
- **`model_fallback`** (`WARN`) — the agent switched off the requested model mid-run. Read the `trigger`:
|
|
614
|
-
`model_not_found` / `model_blocked` / `permission_denied` are properties of the **pin**, so every run of
|
|
615
|
-
this scenario falls back the same way until you change the pinned id; `overloaded` / `server_error` are
|
|
616
|
-
transient and a re-run may hold. The run's assertions still mean what they say — but they were produced
|
|
617
|
-
by a different model than the scenario names, so treat a green as evidence about the fallback model.
|
|
618
|
-
|
|
619
|
-
- **`mount_delete`** (`WARN`) — a delete touched a **delete-denied mount other than `outputs`**: a `rw`
|
|
620
|
-
connected folder. Production denies `unlink`/`rmdir` on *every* Cowork FUSE mount until per-mount
|
|
621
|
-
approval, not just outputs — a connected folder shows the identical default — so this run diverged from
|
|
622
|
-
what production would have allowed. `WARN` rather than `FAIL` because the harness **detects** post-hoc
|
|
623
|
-
what production **enforces**: by the time the scan sees it, the agent already proceeded where it would
|
|
624
|
-
have hit `EPERM`, so failing the run would overstate what a post-hoc scan knows. Author
|
|
625
|
-
`no_delete_in_mounts: true` to hard-fail on it, or `allow_delete_in: ["<mount>"]` to waive that mount
|
|
626
|
-
(detection still runs and the hit is still recorded — the waiver is a verdict decision).
|
|
627
|
-
- **`host_path_leak`** — skipped at **`hostloop` and `protocol`** fidelity (the agent runs on real host
|
|
628
|
-
paths there, so a host path in model-visible text is expected, not a leak); it is *armed* at
|
|
629
|
-
`container`/`microvm`, but only *fires* on an actual scanned leak with no authored
|
|
630
|
-
`transcript_no_host_path`. At `fidelity: cowork` the skip follows the **resolved** tier, so a `cowork`
|
|
631
|
-
run that lands on `container` is armed. Author `transcript_no_host_path` to enforce cleanliness where
|
|
632
|
-
it's valid.
|
|
633
|
-
- **`exec_infra_error`** (`WARN`, host-loop) — one or more container `exec` calls failed for
|
|
634
|
-
infrastructure reasons (daemon/container-level), so those tool calls returned an error to the agent
|
|
635
|
-
rather than the command's own output. Warn-severity because the run's other evidence is intact — unlike
|
|
636
|
-
the fail-severity `infra_error`, where a **supervising process** died and contaminated everything. Note
|
|
637
|
-
a model-requested `timeout_ms` expiry is *not* this: it returns the command's partial output with
|
|
638
|
-
`Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
|
|
639
|
-
failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
|
|
640
|
-
suspiciously empty.
|
|
641
|
-
- **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
|
|
642
|
-
`RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
|
|
643
|
-
pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
|
|
644
|
-
|
|
645
|
-
The full 20-code signal table (severity + per-signal opt-out) is in
|
|
646
|
-
[`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
|
|
647
|
-
the fuller narrative.
|
|
648
|
-
|
|
649
|
-
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
650
|
-
|
|
651
|
-
A single green proves the run passed **once**. Two questions need more than that, and both have a
|
|
652
|
-
discipline that is cheap to follow and expensive to skip.
|
|
653
|
-
|
|
654
|
-
**"Did it pass, or pass once?"** → `--repeat N` (2-100, on `skill` AND `run`) samples the same
|
|
655
|
-
skill+prompt N times and prints a variance rollup instead of a single verdict. `--min-pass-rate` sets
|
|
656
|
-
the batch threshold, `--stop-on-diverge` stops the moment flakiness is proven, `--max-budget-usd` caps
|
|
657
|
-
spend.
|
|
658
|
-
|
|
659
|
-
**"Does the skill actually help?"** → `--ablate-skill` runs the prompt with every skill/plugin
|
|
660
|
-
discovery source removed, so the agent answers from its own priors. **It is ONE arm, not a paired
|
|
661
|
-
experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
|
|
662
|
-
the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
|
|
663
|
-
and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
|
|
664
|
-
The rollup says so on its verdict line: `repeat "<skill>": PASS [ABLATED — control arm] — 5/5 passed`.
|
|
665
|
-
Every ablated run is stamped `ablated: true` in `result.json` and carries `ablated=true` on its
|
|
666
|
-
`[provenance]` footer line; a run that isn't stamped is a real run.
|
|
667
|
-
What the harness gives you here is the run execution and the control arm — designing the comparison
|
|
668
|
-
(scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
|
|
669
|
-
|
|
670
|
-
**Measurement hygiene — four things that silently invalidate a batch:**
|
|
671
|
-
|
|
672
|
-
1. **Pin the model.** With no `model:` in the session (or `--model` on the `skill` lane) the run uses
|
|
673
|
-
whatever the staged agent binary defaults to — not a harness constant, and it can move under a
|
|
674
|
-
baseline bump. Read `result.json`'s `models` back before believing any cross-run comparison — and when
|
|
675
|
-
you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
|
|
676
|
-
fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
|
|
677
|
-
array purely by whether such a turn occurred.
|
|
678
|
-
2. **Commit the skill first.** `fingerprint.skillHash` is content-exact, so an edit mid-batch silently
|
|
679
|
-
splits your dataset into two generations — and a hash whose source was never committed identifies a
|
|
680
|
-
generation that is unrecoverable. `stats --group-by skill-hash` separates them after the fact;
|
|
681
|
-
nothing recovers the source.
|
|
682
|
-
3. **Check which arm you actually ran** before analysing anything: `ablated` and
|
|
683
|
-
`context.availableSkills` in each `result.json`.
|
|
684
|
-
4. **Read `skillsInvoked`.** A rep where the skill never triggered is a measurement of the model, not
|
|
685
|
-
of your skill — discard or re-run it.
|
|
686
|
-
|
|
687
|
-
### Checking whether a background run is alive
|
|
688
|
-
|
|
689
|
-
Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
|
|
690
|
-
processes in your OWN PID namespace, which is frequently NOT the harness process's namespace (e.g. when
|
|
691
|
-
you're a sandboxed subagent). An empty `ps aux` match tells you nothing about whether the run is still
|
|
692
|
-
going.
|
|
693
|
-
|
|
694
|
-
Use **`cowork-harness status <dir> [--follow]`** instead — reads `<outDir>/status.json`, a file the
|
|
695
|
-
harness writes/updates throughout the run's lifecycle (including a crash-safety net for a thrown
|
|
696
|
-
error/`SIGTERM`, AND staleness detection for a hard `SIGKILL`/OOM-kill that no exit handler can catch —
|
|
697
|
-
either way you get `"error"`/`stale` instead of a permanently-trusted `"running"`), so liveness is
|
|
698
|
-
checkable regardless of PID namespace. The harness prints `[status] <outDir>` to stderr as soon as the
|
|
699
|
-
run starts, so capture stderr to get the exact directory — **unless you passed `--compact` (or `--demo`,
|
|
700
|
-
which implies it), which suppress that line** (it is a raw, un-tildeified host path, exactly what those shareable-output
|
|
701
|
-
modes exist to withhold; `status.json` is still written either way, so `status` still works) — but
|
|
702
|
-
`<dir>` also accepts the run-dir root
|
|
703
|
-
passed to `--run-dir` (a directory without its own `status.json`): it scans up to two levels down for the
|
|
704
|
-
newest session's `status.json` and reads that. `--follow` fails loud on a timeout/staleness
|
|
705
|
-
rather than hanging forever. (Fuller recipe in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) — repo-only, not in the installed
|
|
706
|
-
payload; `cowork-harness status --help` has the flags.)
|
|
707
|
-
|
|
708
|
-
**Poll with `--follow`, not with a shell loop over `status`'s stdout.** The one-shot text form prints to
|
|
709
|
-
**stderr** and writes nothing to stdout; `--output-format json` (one envelope) and `--follow` (one JSON
|
|
710
|
-
line per status change) are the **stdout** forms. A poll that greps `status`'s stdout therefore matches
|
|
711
|
-
nothing, exits 1, and returns instantly against a run with minutes left to go — a silent false "done":
|
|
712
|
-
|
|
713
|
-
```bash
|
|
714
|
-
# WRONG — stdout is empty, so grep exits 1, `!` inverts it, and the loop never sleeps.
|
|
715
|
-
until ! cowork-harness status "$D" | grep -q '● running'; do sleep 30; done
|
|
716
|
-
|
|
717
|
-
# RIGHT — the harness owns the poll loop and exits when the run reaches a terminal state.
|
|
718
|
-
cowork-harness status "$D" --follow
|
|
719
|
-
```
|
|
720
|
-
|
|
721
|
-
**A multi-minute `record`/`run` outlives a short-lived wrapper.** Don't launch a long record from a
|
|
722
|
-
subagent that returns before it finishes — the returning agent tears down its process tree and kills the
|
|
723
|
-
in-flight run mid-artifact-write. Run it foreground, or detached from any process that will exit first.
|
|
724
|
-
(The `status.json` liveness above is exactly what surfaces such a teardown as `"error"`/`stale` rather
|
|
725
|
-
than a stuck `"running"`.)
|
|
726
|
-
|
|
727
|
-
### Place assertions in the right CI lane
|
|
728
|
-
|
|
729
|
-
CI placement: a **token-free `replay` PR gate** (content/structure only) + a **nightly live `run`**
|
|
730
|
-
(filesystem/egress). Fastest setup: `uses: yaniv-golan/cowork-harness@v3` (a packaged GitHub Action with a
|
|
731
|
-
PR job-summary reporter). See `references/ci-recipe.md` for the Action, the manual step-by-step form, and
|
|
732
|
-
the four-stage pipeline.
|
|
733
|
-
|
|
734
|
-
## Part III — Debug
|
|
735
|
-
|
|
736
|
-
A run misbehaved, or greened when you don't trust it. Debugging is a first-class loop, not an
|
|
737
|
-
afterthought: the run already wrote its evidence, so you **localize the failure post-hoc** rather than
|
|
738
|
-
re-run and hope. Start at the triage below, then use the observability output and, when you need to
|
|
739
|
-
reproduce interactively, `chat`.
|
|
740
|
-
|
|
741
|
-
> **"Evidence" below means the run's own record** — events, trace, transcript; what `trace` / `inspect` /
|
|
742
|
-
> `diff` / `verify-run` / `replay --explain` read. `critique`'s **evaluator** grades against a separate,
|
|
743
|
-
> narrower record — `critique-evidence-package.txt`, what a grade was actually computed against — none of
|
|
744
|
-
> the five tools above surface it; see `references/critique.md`.
|
|
745
|
-
|
|
746
|
-
### Triage — a run misbehaved, or a green looks wrong
|
|
747
|
-
|
|
748
|
-
<!-- BEGIN triage-canonical -->
|
|
749
|
-
Two situations need different tools — figure out which one you're in first, then reach for the tool
|
|
750
|
-
instead of re-running and hoping. The run already wrote its evidence to a kept run dir (`--keep` prints
|
|
751
|
-
the path; `trace <run-id>` finds it), and every tool below reads that evidence **token-free** — no
|
|
752
|
-
Docker, no re-record.
|
|
753
|
-
|
|
754
|
-
| Situation | Symptom | Reach for (in order) |
|
|
755
|
-
|---|---|---|
|
|
756
|
-
| **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
|
|
757
|
-
| **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `replay --mutate` — perturbs a CAPPED SAMPLE of recorded JSON values (10/file, 50 total) and reports which perturbations NOTHING caught; the report names the sample size, so read it as a sample not a total (reporting only; never moves the verdict/exit code) · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
|
|
758
|
-
|
|
759
|
-
A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
|
|
760
|
-
agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
|
|
761
|
-
already does.
|
|
762
|
-
<!-- END triage-canonical -->
|
|
763
|
-
|
|
764
|
-
**Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
|
|
765
|
-
**deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
|
|
766
|
-
so a "bug" you see here that real Cowork also has isn't yours to fix. The tier semantics are in
|
|
767
|
-
`references/fidelity-and-answers.md` (shipped); the specific deltas vs. real Cowork and the sandbox
|
|
768
|
-
boundary model live in [`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md) / [`docs/boundary.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/boundary.md) (repo-only, not in the installed
|
|
769
|
-
payload). If the behavior is on that gap list, it's expected — stop debugging your skill.
|
|
770
|
-
|
|
771
|
-
### Inspecting a run's observability output
|
|
772
|
-
|
|
773
|
-
A verdict is only the top of what a run records, and the run dir persists after the verdict
|
|
774
|
-
(`~/.cowork-harness/runs/…`). Beyond pass/fail, every `run`/`skill`/`chat` writes a `result.json` and a
|
|
775
|
-
trace you read back without a re-record — the debugging loop is *localize the failure from that
|
|
776
|
-
already-written evidence*, not re-run-and-hope. Use them to diagnose a failure (and, secondarily, to
|
|
777
|
-
decide which assertions from *Assertions: two orthogonal axes* are worth adding):
|
|
778
|
-
|
|
779
|
-
- **`cowork-harness trace <run-dir> --view <view>`** — focuses one of the run's rollups (the per-tool
|
|
780
|
-
call-count/timing table, the sub-agent dispatch tree, the gate lifecycle, the tool/error rollups, …);
|
|
781
|
-
bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
|
|
782
|
-
the current list rather than relying on a fixed enumeration here.
|
|
783
|
-
- **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
|
|
784
|
-
is held to. Cowork picks the lane per session ("Run this task: In the cloud / On your computer") and
|
|
785
|
-
cloud is the default for new sessions; the lanes disagree about what *delivered* means. On `remote`,
|
|
786
|
-
location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
|
|
787
|
-
session end), `present_files` is NOT served, and `user_visible_artifact` /
|
|
788
|
-
`present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
|
|
789
|
-
to check a skill's delivery survives the lane most new sessions get. Orthogonal to `fidelity` — a
|
|
790
|
-
`lane: remote` scenario still runs locally.
|
|
791
|
-
- **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
|
|
792
|
-
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
|
|
793
|
-
plus `--skill-hash <prefix>`/`--label <tag>` to narrow to ONE skill generation and
|
|
794
|
-
`--group-by scenario|skill-hash|label|fidelity` to split per generation — or per effective fidelity
|
|
795
|
-
tier — instead of aggregating across them (a window spanning >1 generation warns, so an un-split A/B
|
|
796
|
-
average announces itself rather than passing as one number;
|
|
797
|
-
>1 tier warns too, independently, with `--group-by fidelity` as its own remedy). `--runs` lists the
|
|
798
|
-
individual runs behind each summary with their `skillHash`/`runLabel`, so
|
|
799
|
-
you can tell which arm a run belonged to without opening its `result.json`. `--last <n>` windows per group.
|
|
800
|
-
- **`result.json` carries the raw fields** the assertions read: `verdict`, `lane` (which Cowork delivery
|
|
801
|
-
contract the run was held to — see Gotcha 24), `scratchpadEvidenceComplete` (did a COMPLETE scratchpad
|
|
802
|
-
walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
|
|
803
|
-
`total_cost_usd` for the run — the authoritative single-run spend; NOT the same source as summing
|
|
804
|
-
`modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
|
|
805
|
-
`usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations`, `models`, `toolErrors`,
|
|
806
|
-
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
807
|
-
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
808
|
-
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
809
|
-
`workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
|
|
810
|
-
(`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
|
|
811
|
-
evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
|
|
812
|
-
`gateProvenance` (per-gate `scripted`/`decided(llm|external)`/`first-option`/`prompt` with a
|
|
813
|
-
`bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
|
|
814
|
-
`egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
|
|
815
|
-
cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
|
|
816
|
-
`outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
|
|
817
|
-
`jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
|
|
818
|
-
`{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
|
|
819
|
-
semantics: [`docs/cli.md` → What you get out](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md#what-you-get-out-inspectable-output) (repo-only); [`schema/run-result.json`](https://github.com/yaniv-golan/cowork-harness/blob/main/schema/run-result.json) is the
|
|
820
|
-
machine source.)
|
|
821
|
-
- **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
|
|
822
|
-
**`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
|
|
823
|
-
a re-record rarely tells you more than the captured stderr already does. Also check
|
|
824
|
-
**`resultErrorKind`** (`"transport" | "agent" | "usage_limit"`) before spending another paid run: a
|
|
825
|
-
`"usage_limit"` failure is a quota exhaustion, not a skill bug — retry after the limit resets rather
|
|
826
|
-
than debugging; `"transport"`/`"agent"` means something actually broke, worth localizing before
|
|
827
|
-
re-running.
|
|
828
|
-
- **Attributing cost to sub-agent work.** `subagents[]` gives the dispatch tree — each sub-agent's
|
|
829
|
-
`dispatchModel`/`resolvedModel`, `toolsUsed`, `prompt`/`output`, and `attributedSkillId` — but **not** its own token/cost;
|
|
830
|
-
aggregate cost is per-**model** in `modelUsage` (and `trace --view usage`), not per-sub-agent. So a
|
|
831
|
-
cost spike from fan-out reads as `trace --view dispatches` (how many, which agent) against that model's
|
|
832
|
-
per-model usage — the harness doesn't line-item each sub-agent's tokens.
|
|
833
|
-
- **Debugging a wrong Cowork UI panel.** Each panel is reconstructed in `result.json`: **Progress** =
|
|
834
|
-
`tasks[]`, **Working folder** = `workspaceFiles[]` (classified output/mount/input/scratchpad — the last being the agent's working area outside every user-visible root, with a
|
|
835
|
-
`trace --view files` diff), **Context / Connectors** = `context` (tools / mcpServers / availableSkills),
|
|
836
|
-
**Scratch-pad → outputs** = `presentedFiles[]`. If a panel looks wrong in a run, read its field. An
|
|
837
|
-
**absent** `workspaceFiles`/`artifacts` (a replay result, or a run whose workspace root was missing at
|
|
838
|
-
collection) is evidence **UNAVAILABLE**, not an empty run — `trace --view files` reports a loud
|
|
839
|
-
UNAVAILABLE marker (`workspaceFilesRecorded: false` in JSON, and no phantom "removed" diff rows) and
|
|
840
|
-
`inspect` prints `artifacts: UNAVAILABLE` (`artifactsRecorded: false`) instead of `artifacts (0):`.
|
|
841
|
-
|
|
842
|
-
### Debugging with `chat`
|
|
843
|
-
|
|
844
|
-
`cowork-harness chat` opens an interactive multi-turn REPL against a live Cowork session. It is
|
|
845
|
-
**not** an asserted test — no `assert:` block, no cassette. Use it to explore behavior, reproduce a
|
|
846
|
-
bug interactively, or test a prompt before committing it to a scenario.
|
|
847
|
-
|
|
848
|
-
Each session still writes an informational `result.json` (`mode: "chat"`, no `assertions`) plus a
|
|
849
|
-
trace and index row under its run dir — the same telemetry (tool durations, model usage, resources,
|
|
850
|
-
etc.) that `run`/`skill` produce — so `cowork-harness trace <chat-run-dir>` / `stats` work on a chat
|
|
851
|
-
session too, even though it never yields a verdict.
|
|
852
|
-
|
|
853
|
-
**`--plugin <dir>` flag (repeatable).** Load additional skill folders alongside the primary session
|
|
854
|
-
plugin. Each `--plugin <dir>` appends the folder to `local_plugins`. Useful when the skill-under-test
|
|
855
|
-
depends on a sibling plugin:
|
|
856
|
-
|
|
857
|
-
```bash
|
|
858
|
-
cowork-harness chat ./skills/report-gen --plugin ./skills/shared-utils
|
|
859
|
-
```
|
|
860
|
-
|
|
861
|
-
**Note:** `--raw` mode (native `docker run -it`) can't honor the harness-managed flags, so `--upload`,
|
|
862
|
-
`--folder`, `--plugin`, and `--fidelity` are **rejected** with a usage error if combined with `--raw`;
|
|
863
|
-
only `--model` is carried through.
|
|
864
|
-
|
|
865
|
-
**`/help` in the REPL.** Type `/help` at the prompt to see available commands:
|
|
866
|
-
|
|
867
|
-
```
|
|
868
|
-
Commands: /exit /quit /help
|
|
869
|
-
```
|
|
870
|
-
|
|
871
|
-
The startup banner now reads `type your message (/help for commands)` as a reminder. `/exit` and
|
|
872
|
-
`/quit` both terminate the session.
|
|
873
|
-
|
|
874
|
-
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
875
|
-
|
|
876
|
-
Stated as *symptom → why → fix*. This is the **workflow/record/answer-path** view — the broader of the
|
|
877
|
-
two lists, but **not a strict superset**: `references/scenario-schema.md`'s *Authoring gotcha list*
|
|
878
|
-
carries a few assertion-level landmines this one omits (`transcript_no_host_path`'s scan width,
|
|
879
|
-
`egress.extra_allow`'s no-op on the provenanced `web_fetch` path, `replay_protocol_fidelity` not being
|
|
880
|
-
authorable). Reach for this list when debugging a run's behavior, that one while authoring `assert:`.
|
|
881
|
-
**The two lists are numbered independently** — a bare "gotcha N" means the list you are reading.
|
|
882
|
-
|
|
883
|
-
1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
|
|
884
|
-
`replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
|
|
885
|
-
`artifact_text`, `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
|
|
886
|
-
*mixed* item like
|
|
887
|
-
`{result, egress_denied}` greens on `result` while its `egress_denied` half is dropped. (`record`
|
|
888
|
-
snapshots an `artifacts` manifest, which makes
|
|
889
|
-
`file_exists`/`user_visible_artifact`/`artifact_json`/`artifact_text`/`computer_links_resolve`
|
|
890
|
-
replay-checkable — but the live-only egress keys stay skipped, and `file_absent` is never
|
|
891
|
-
replay-checkable at all: proving absence needs an exhaustive, healthy walk a manifest does not record.) *Fix:* put egress/live-only checks on
|
|
892
|
-
a live gate; keep one concern per `assert:` item; run the linter. The harness warns loudly on skip.
|
|
893
|
-
|
|
894
|
-
2. **A steered gate answer never reached the model.** *Why:* `serializeDecision` must emit
|
|
895
|
-
`updatedInput: { questions, answers }`; a header-only gate (empty `question`) can never be keyed.
|
|
896
|
-
*Fix:* give every gate a non-empty `question`. (multiSelect gates ARE supported on **every** answer
|
|
897
|
-
channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
|
|
898
|
-
reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
|
|
899
|
-
Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
|
|
900
|
-
string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` / `question_context` /
|
|
901
|
-
`questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
|
|
902
|
-
old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
|
|
903
|
-
on unobserved delivery (absence of evidence is failure, not neutral).
|
|
904
|
-
|
|
905
|
-
3. **A multi-key `assert:` item is an AND.** A single list item with more than one key passes iff
|
|
906
|
-
**every** key passes. *Fix:* one concern per item unless you genuinely mean conjunction (and a
|
|
907
|
-
mixed-class conjunction still loses its filesystem half on replay — see gotcha 1 below).
|
|
908
|
-
|
|
909
|
-
4. **`tool_called` doesn't mean "attempted".** Tool counts are authoritative and de-duped: a tool
|
|
910
|
-
that was *requested then denied* does **not** register as called. *Fix:* don't assert `tool_called`
|
|
911
|
-
to prove an attempt; it proves the tool actually ran.
|
|
912
|
-
|
|
913
|
-
5. **`subagent_declared_but_unused` fires on declared-but-didn't-use-THAT-tool**, even if the
|
|
914
|
-
sub-agent used other tools. `subagent_dispatched` / `subagent_output_contains` match on dispatch
|
|
915
|
-
type (`dispatchAgentType`), the binary-*resolved* type (`resolvedAgentType`), *or* the dispatch
|
|
916
|
-
**description** — so a type-less dispatch that resolved to e.g. `general-purpose` is still
|
|
917
|
-
selectable, by either the resolved type or the description. A `Task` dispatch that carries NO
|
|
918
|
-
`subagent_type` at all falls back to the built-in `general-purpose` agent with a **wildcard tool
|
|
919
|
-
surface** (`tools:["*"]`, including workspace bash) — faithful production behavior, and it fires
|
|
920
|
-
routinely. The harness warns loudly on this fallback and records `subagents[].dispatchTypeOmitted`;
|
|
921
|
-
an *explicit* `subagent_type: "general-purpose"` is a deliberate author choice and does not warn.
|
|
922
|
-
Implication: `subagent_tool_absent` on a type-less dispatch is weaker evidence (wildcard surface) —
|
|
923
|
-
pin `subagent_type` explicitly when you need a tight tool-absence guarantee.
|
|
924
|
-
|
|
925
|
-
**Cross-tier "no shell" caveat.** On `hostloop`, native `Bash` calls route through the
|
|
926
|
-
`mcp__workspace__bash` alias, so a "sub-agent used no shell" check must glob **both** `Bash` and
|
|
927
|
-
`mcp__workspace__*` to hold across every tier.
|
|
928
|
-
|
|
929
|
-
6. **`dispatch_count_max` is your author-chosen budget UNDER Cowork's production cap, not a
|
|
930
|
-
reproduction of it.** It's a post-hoc count assertion: passing means "happened to dispatch ≤N this
|
|
931
|
-
run." Cowork DOES cap `Task` fan-out **agent-side** (`taskRegistry`: concurrent **20** /
|
|
932
|
-
per-session **200**, landed 2.1.212/2.1.217) — SEPARATE from the scheduled-task session limiter
|
|
933
|
-
(gate `1648655587`'s `{perTask:1, global:3}`, a different mechanism; binary-verified, `SPEC.md` §10
|
|
934
|
-
— repo-only). The harness **inherits** the production cap by spawning the real agent binary, so a
|
|
935
|
-
`dispatch_count_max` pass means "your tighter budget held," not "near a real limit"; use it to catch
|
|
936
|
-
a fan-out you don't want.
|
|
937
|
-
|
|
938
|
-
7. **`protocol` is rejected (not silently passed) if the scenario asserts egress** — boundary
|
|
939
|
-
assertions need a sandboxed tier (`container`+). Good: this one fails loud by design.
|
|
940
|
-
|
|
941
|
-
8. **Read-only mounts are enforced; delete-deny is a HARNESS gap — production DOES enforce it.**
|
|
942
|
-
`mode:r` mounts get a real `:ro` bind (a write fails in-guest). But `rw` vs `rwd`
|
|
943
|
-
(write-but-no-delete on `outputs/` / connected folders) is *not* mount-enforced **in the harness** —
|
|
944
|
-
`rm` succeeds and is only caught post-hoc by `no_delete_in_outputs`. **Real Cowork enforces it live:**
|
|
945
|
-
outputs is a FUSE mount, and `unlink`/`rmdir` fail `Operation not permitted`; a skill must request
|
|
946
|
-
approval via `allow_cowork_file_delete` (which re-mounts the folder `rwd` mid-session) to delete.
|
|
947
|
-
**Only unlinking is denied.** Emptying a file in place — `truncate -s 0`, `> file`, `shred` without
|
|
948
|
-
`-u` — and renaming *within* outputs both SUCCEED in production, so the harness does not flag them
|
|
949
|
-
either. Renaming a file OUT of outputs fails (`EXDEV`, then `EPERM` on the copy-then-unlink
|
|
950
|
-
fallback), so that stays a delete. Two consequences: a skill should not stage disposable scratch
|
|
951
|
-
under `outputs/` (in production, cleanup there costs an approval prompt), and a skill's
|
|
952
|
-
"catch-EPERM-then-request-approval" branch cannot be exercised at any harness tier (the `rm` just
|
|
953
|
-
succeeds here). Do not read this gotcha as "delete-deny may not be real in production" — it is real.
|
|
954
|
-
If a scenario's deletion IS intended, assert `allow_outputs_delete: true` rather than dropping
|
|
955
|
-
`no_delete_in_outputs` — omitting it does not permit anything.
|
|
956
|
-
|
|
957
|
-
9. **Keep `.env` out of any mounted folder** — it is copied into the sandbox and the token could
|
|
958
|
-
leak. Put it at a working-dir or install root (token resolution: env > `--dotenv` > `./.env` >
|
|
959
|
-
install `.env`). **Inverse footgun — running from a git worktree:** a worktree's `./.env` is gitignored, so
|
|
960
|
-
it's **absent** there and you'll get "no model credentials." *Fix:* pass `--dotenv <main-checkout>/.env`
|
|
961
|
-
(or set the env var) — that's exactly what `--dotenv` is for.
|
|
962
|
-
|
|
963
|
-
10. **A base64 artifact that was scrubbed at record time will fail artifact assertions at replay.**
|
|
964
|
-
When `record` detects a secret embedded in a base64 artifact, it replaces the entire artifact
|
|
965
|
-
body with `[REDACTED:base64]` and emits a `::warning::`. Any `artifact_json` or content
|
|
966
|
-
assertion targeting that artifact will fail at replay because the body no longer matches. *Fix:*
|
|
967
|
-
do not let secrets flow into artifacts; if the artifact is intentionally opaque, drop the
|
|
968
|
-
content assertion and gate on `file_exists` on the live lane instead.
|
|
969
|
-
|
|
970
|
-
11. **An external decider returning `"first"` does not select option 1.** The `"first"` keyword
|
|
971
|
-
shorthand is disabled for `--decider-cmd` / `--decider-dir` helpers (see *Choose an answer path*
|
|
972
|
-
→ External deciders). If your helper
|
|
973
|
-
accidentally emits `"first"` and no label named `"first"` exists, the gate fails — it does
|
|
974
|
-
**not** silently pick the first option. This is intentional: a helper bug should fail loud, not
|
|
975
|
-
green wrong. *Fix:* have helpers return a label name or numeric index.
|
|
976
|
-
|
|
977
|
-
12. **`prompt_asset_missing` is a WARN, not a hard failure — greens can hide it.** The
|
|
978
|
-
`prompt_asset_missing` verdict signal (see *Interpreting verdict signals*) does not block a green verdict. Scan the verdict
|
|
979
|
-
signals section after every run; a run that greened with this signal ran against an incomplete
|
|
980
|
-
prompt. *Fix:* treat `prompt_asset_missing` as a blocking error in CI by checking the signals
|
|
981
|
-
array.
|
|
982
|
-
13. **`result: success` means the agent didn't error, NOT that the task completed — always assert on
|
|
983
|
-
artifacts/content.**
|
|
984
|
-
- A turn that ends on a plain-text re-ask ("which file did you mean?") still reports
|
|
985
|
-
`result: success`.
|
|
986
|
-
- The harness catches this with a **`stalled`** verdict signal: a run that ends on a question and
|
|
987
|
-
did **no productive work after its last gate** — both the no-gate case ("which file?" with no
|
|
988
|
-
tool calls) AND the *answered-gate-then-re-ask* case (the agent answers an `AskUserQuestion`,
|
|
989
|
-
then asks again in plain text and stops). Suppress with `allow_stall: true` if ending on a
|
|
990
|
-
question is intended.
|
|
991
|
-
- The signal is a **tool-position heuristic**, not deliverable detection, so it is imprecise both
|
|
992
|
-
ways:
|
|
993
|
-
- **False negative:** a post-gate tool *call* clears the flag whether it **succeeded or
|
|
994
|
-
errored** — an agent that ran a tool after the gate and still stalled is not caught.
|
|
995
|
-
- **False positive:** a deliverable written *before* a final confirmation gate does **not**
|
|
996
|
-
clear it, so a write-then-confirm-then-question run is flagged — use `allow_stall: true` for a
|
|
997
|
-
deliberate confirm-terminal skill.
|
|
998
|
-
- The broad guard is therefore YOUR assertions — assert the deliverable (`file_exists` /
|
|
999
|
-
`artifact_json` / `transcript_matches`), never just `result: success`.
|
|
1000
|
-
- `on_unanswered` governs **unanswered** `AskUserQuestion` gates; the `stalled` signal covers
|
|
1001
|
-
stalling *after* one is answered — two different failure modes.
|
|
1002
|
-
- **Free-text aside:** the scripted key for a "type-it-in-notes" option is **`answer:`** — an
|
|
1003
|
-
arbitrary string delivered verbatim, bypassing label validation by author intent (Cowork
|
|
1004
|
-
auto-provides an "Other" free-text path on every gate). Mutually exclusive with `choose:`; setting
|
|
1005
|
-
both fails loud. What has no scripted equivalent is the `OTHER:` *directive*
|
|
1006
|
-
(it works only on the LLM-decider path, not scripted `choose:`, and only on
|
|
1007
|
-
**single-select** gates — a **multi-select** gate is index-only, so `OTHER:` fails loud there; on an
|
|
1008
|
-
options-bearing single-select gate a bare out-of-set LLM answer also fails loud (exit 2) — see the
|
|
1009
|
-
LLM-decider free-text note in `references/fidelity-and-answers.md`). An LLM decision answered via
|
|
1010
|
-
`OTHER:` is marked `[via Other free-text]` in its `gateProvenance` rationale.
|
|
1011
|
-
14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
|
|
1012
|
-
but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
|
|
1013
|
-
index lands on a different option (a silent re-record flake). Prefer an exact label when order is
|
|
1014
|
-
stable; `lint` flags positional `choose` with an advisory. Unstable option order is also what the
|
|
1015
|
-
**user** sees — a reordered gate puts a different choice in the default slot — so pin what was shown
|
|
1016
|
-
with `question_options`, rather than only hardening the answer rule against it.
|
|
1017
|
-
15. **A scripted `choose:` matching no offered option HARD-fails the run — `on_unanswered: first` does NOT
|
|
1018
|
-
backstop it.** This is distinct from an *unanswered* gate (no rule matched → falls to `on_unanswered`): a
|
|
1019
|
-
rule that DID match the gate but whose `choose:` names a label the gate never offered (the model reworded
|
|
1020
|
-
it) is treated as an authoring bug and fails loud — `first`/`llm` won't absorb it. The error now prints the
|
|
1021
|
-
**offered options** (and a closest-match suggestion), so fix the anchor from the error alone — no need to
|
|
1022
|
-
dig through `events.jsonl`. (This is exactly the drift `verify-run` answer-coverage catches in ~1s; use it
|
|
1023
|
-
before a paid record.)
|
|
1024
|
-
16. **Batch record keeps going — you don't need a one-at-a-time wrapper.** `record <dir>` and `record <dir>
|
|
1025
|
-
--rerecord-stale` run **every** scenario, collect failures, and report them at the end (non-zero exit on
|
|
1026
|
-
any failure) — a failing scenario does NOT abort the batch. So a single `cowork-harness record cassettes/
|
|
1027
|
-
--rerecord-stale` surfaces ALL stale anchors in one pass (add `--concurrency <N>` to parallelize); a shell
|
|
1028
|
-
wrapper that loops one cassette at a time with `set -e` defeats this and rediscovers stale anchors serially.
|
|
1029
|
-
Two durability properties make the batch safe to trust: each cassette is written **atomically** (a
|
|
1030
|
-
same-directory temp file + rename), so an interrupted or OOM-killed batch never leaves a partial/corrupt
|
|
1031
|
-
cassette — a failed scenario simply produces none; and under `--concurrency <N>` each scenario runs **fully
|
|
1032
|
-
isolated** (its own egress sidecar network + proxy, its own per-session run dir), so parallel records don't
|
|
1033
|
-
cross-talk — the concurrency bound exists only for the Docker address pool + API rate limits, not correctness.
|
|
1034
|
-
|
|
1035
|
-
17. **Editing `scenarios/*.yaml` does NOT change a plain `replay` — the WHOLE scenario is frozen, not just
|
|
1036
|
-
`assert:`.** *Why:* a cassette captures every key (`lane:`, `fidelity:`, `baseline:`, `prompt:`, `skills:` …)
|
|
1037
|
-
and `replay` evaluates all of them from that frozen copy — byte-deterministic, ignoring the working tree (so
|
|
1038
|
-
a committed cassette can't silently re-interpret against an uncommitted YAML). **Only `assert:`
|
|
1039
|
-
(+`expect_denied:`) can be opted back to disk; any other edited key reaches a replay only by re-recording.**
|
|
1040
|
-
This is *loud* rather than a *silent* no-op: plain `replay` prints a `::notice::` when a sibling's
|
|
1041
|
-
`assert:`/`prompt:` differs, and when the sibling **fails to load** at all (a typo'd or too-new key) — and
|
|
1042
|
-
points you at the fix.
|
|
1043
|
-
*Fix:* to re-check token-free against the edited block, `replay --assert-from <scenario.yaml>` (or
|
|
1044
|
-
`--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
|
|
1045
|
-
`prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
|
|
1046
|
-
fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
|
|
1047
|
-
are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the `session`
|
|
1048
|
-
(model / data mounts / discovery) is NOT drift-checked or fingerprinted, so a **model change** between record
|
|
1049
|
-
and re-assert is undetected — the notice flags this; re-record if the session changed. `verify-run` reads
|
|
1050
|
-
on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
|
|
1051
|
-
|
|
1052
|
-
18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
|
|
1053
|
-
bundle several sub-questions into a single gate; the assertion counts each sub-question, so a
|
|
1054
|
-
3-sub-question bundle counts as 3, not 1. `trace --view questions` shows the same per-gate
|
|
1055
|
-
sub-question count and a matching footer total — read that off instead of the tool-call count when
|
|
1056
|
-
sizing the budget.
|
|
1057
|
-
|
|
1058
|
-
19. **`gate_answers_delivered` passes vacuously when no gate fires — pair it, or drop it.** Whether a
|
|
1059
|
-
gate fires is model-dependent, so `gate_answers_delivered: true` alone can't catch "the gate never
|
|
1060
|
-
fired at all". If the scenario is meant to gate, pair it with `gate_answer_count_min: 1` (a floor of
|
|
1061
|
-
`0` witnesses nothing). If the scenario is gate-clean by design, drop the key — it asserts nothing
|
|
1062
|
-
there — and declare `questions_count_max: 0`, which fails loudly if a gate ever appears. Asserting
|
|
1063
|
-
`questions_count_max: 0` alongside a gate-presence key is unsatisfiable: `run`/`skill`/`record`
|
|
1064
|
-
refuse it before spending.
|
|
1065
|
-
|
|
1066
|
-
20. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
|
|
1067
|
-
read-only folder's files as path + hash only (`truncated: true`, no `body`) — it's an input the agent
|
|
1068
|
-
read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass against it on replay
|
|
1069
|
-
(the hash-only entry still materializes a placeholder); `artifact_json` reports a clear
|
|
1070
|
-
evidence-unavailable on every lane (live/verify-run/replay agree — no green-record/red-replay). This is
|
|
1071
|
-
also why a `mode: r` input never trips the `binary` privacy finding or needs `--allow` — only a
|
|
1072
|
-
*committed* body is scanned. `scaffold` won't emit `file_exists` for one either (it's not in
|
|
1073
|
-
`RunResult.artifacts`). A `mode: rw`/`rwd` folder's contents are captured with a full body, same as
|
|
1074
|
-
`outputs/`.
|
|
1075
|
-
|
|
1076
|
-
21. **A `fidelity: cowork` cassette can go stale in a way `skill`/`format` drift won't catch.** Its recorded
|
|
1077
|
-
`effectiveFidelity` field pins which concrete tier (`hostloop` or `container`) the baseline resolved to
|
|
1078
|
-
AT RECORD TIME. If a later Desktop baseline flips that resolution, `verify-cassettes` reports it as a
|
|
1079
|
-
`resolved-tier` finding (re-record — the recording now exercises the wrong tier); a cassette with no
|
|
1080
|
-
`effectiveFidelity` at all, or an unloadable pinned `baseline:`, reports `unverifiable-tier` instead
|
|
1081
|
-
(couldn't check — also re-record). Both are `fidelity: cowork`-only; an explicit-tier scenario never
|
|
1082
|
-
produces them. (Details: [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § tier staleness — repo-only.)
|
|
1083
|
-
|
|
1084
|
-
22. **`lint` floods CI with INFO advisories that don't apply to you.** *Why:* two rules —
|
|
1085
|
-
`manifest-needs-snapshot` and `gate-needs-controlout` — fire on the mere presence of manifest/gate
|
|
1086
|
-
assertion keys. The linter is **static**: it never reads your cassettes, so it cannot know whether
|
|
1087
|
-
yours already carry an `artifacts` manifest and `controlOut` (a current cassette does). On a healthy
|
|
1088
|
-
fleet every one of those lines is a false alarm. One exception: `manifest-needs-snapshot` is
|
|
1089
|
-
suppressed for `user_visible_artifact` on `lane: remote` — the only manifest-backed key that lane also
|
|
1090
|
-
rejects outright (`lane-remote-incompatible-key`, an ERROR), so the INFO would be redundant advice
|
|
1091
|
-
about a key the scenario can never even load with. `gate-needs-controlout` has no such exception. *Fix:*
|
|
1092
|
-
`lint --min-severity WARN` in CI (≥1.11.0) — the INFO advisories stay one flag away for interactive use.
|
|
1093
|
-
`--strict --min-severity ERROR` behaves as a plain lint, not a contradiction.
|
|
1094
|
-
23. **`verify-cassettes`/`replay` report a `discovery-surface` note on cassettes you just recorded fine.**
|
|
1095
|
-
*Why:* the cassette froze its `system/init` tool inventory from before the skills/plugins discovery
|
|
1096
|
-
servers existed at that tier (added 1.10.0). It is a non-gating **note**, never a finding — it cannot
|
|
1097
|
-
fail your gate. *Fix:* nothing, unless the scenario asserts `tool_available` on
|
|
1098
|
-
`mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
|
|
1099
|
-
re-recording would never produce those tools anyway.
|
|
1100
|
-
|
|
1101
|
-
24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
|
|
1102
|
-
lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
|
|
1103
|
-
emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
|
|
1104
|
-
Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
|
|
1105
|
-
optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
|
|
1106
|
-
other — and probing a remote session makes this harness look like it emulates the wrong tool under the
|
|
1107
|
-
wrong schema. It doesn't; the lanes genuinely disagree. *Fix:* describe the **outcome** ("deliver the
|
|
1108
|
-
file to the user") and let the model pick its surface's tool. The `no_scratchpad_leak` /
|
|
1109
|
-
`present_files_called` assertion keys are harness-side names and stay valid either way.
|
|
1110
|
-
([`docs/fidelity-gaps.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)
|
|
1111
|
-
→ "File delivery" has the binary-verified detail; repo-only.)
|
|
1112
|
-
|
|
1113
|
-
25. **Three host-inventory flags — two on `record`, one on `verify-cassettes`.** `record
|
|
1114
|
-
--allow-host-inventory-fixture` proceeds past the PRE-FLIGHT refusal when recording a host-inheriting
|
|
1115
|
-
(`protocol`/`hostloop`/`cowork`-resolving-to-hostloop) cassette into a repo-visible path — otherwise
|
|
1116
|
-
`record` refuses before it spends (freezing this machine's MCP servers/agents/account into a committed
|
|
1117
|
-
fixture is the risk). It bypasses that check and **nothing else**: the finished recording is still
|
|
1118
|
-
scanned, and a real finding still quarantines it, so you never have to audit the session by hand to
|
|
1119
|
-
pass it. Writing a recording the scan DID flag is the separate `record
|
|
1120
|
-
--allow-host-inventory-findings`. That pre-spend check **warns rather than refuses when the cassette already
|
|
1121
|
-
exists** — refusing would fire on every `--rerecord-stale` pass — and it reads the tier and the
|
|
1122
|
-
destination path, never the bytes. So `record` also scans the FINISHED recording, after redaction and
|
|
1123
|
-
before the write: a `host-inventory`/`machine-inventory` finding on a repo-visible path is
|
|
1124
|
-
**quarantined** to `<runs-root>/quarantine/` with a `.findings.txt` naming what leaked, and the command
|
|
1125
|
-
fails without writing the path you asked for (the recording is not discarded — you paid for it).
|
|
1126
|
-
`verify-cassettes --allow-host-inventory <regex>` is unrelated: a per-finding suppressor for the
|
|
1127
|
-
scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
|
|
1128
|
-
wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
|
|
1129
|
-
|
|
1130
|
-
26. **A `skill`-lane `PASS` does not mean the skill ran, or that the run was the one you wanted.** *Why:*
|
|
1131
|
-
an open-ended `skill` run has no `assert:` block, so its verdict reports only that **no guard fired**
|
|
1132
|
-
(no error, stall, host-path leak, `outputs/` delete, permissive auto-allow or capability gap). On
|
|
1133
|
-
`run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
|
|
1134
|
-
means N runs cleared the guards — it says nothing about which model served them, whether the skill
|
|
1135
|
-
was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
|
|
1136
|
-
carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
|
|
1137
|
-
`models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
|
|
1138
|
-
exactly like skill output is not evidence: the skill's own source is mounted where the model can
|
|
1139
|
-
read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
|
|
1140
|
-
directly, with `skillActivity` empty.
|
|
1141
|
-
|
|
1142
|
-
27. **`allow_stall: true` is a scenario assertion, so the `skill` lane cannot use it.** *Why:* the
|
|
1143
|
-
`stalled` guard fires when a run's final message ends in `?` with no productive tool call after the
|
|
1144
|
-
last gate — which includes a complete answer that closes by *offering* a follow-up ("want me to run
|
|
1145
|
-
this through a structured pass?"). The documented opt-out lives in an `assert:` block, and an
|
|
1146
|
-
open-ended `skill` run has none, so the failure message names a remedy that lane can't perform.
|
|
1147
|
-
*Fix:* on `skill`, read the final message before believing `stalled`, or move the check to a
|
|
1148
|
-
`run` scenario where `allow_stall: true` is authorable.
|
|
1149
|
-
|
|
1150
|
-
For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
|
|
1151
|
-
files in `references/` (the gotchas above are the full list; the references repeat only the
|
|
1152
|
-
assertion/replay-relevant ones).
|
|
1153
|
-
|
|
1154
|
-
## References
|
|
1155
|
-
|
|
1156
|
-
- `references/task-recipes.md` — end-to-end recipes for the four jobs fleet owners actually hit:
|
|
1157
|
-
evolve a cassette's `assert:` (usually no re-record), audit a fleet for tier drift, set up
|
|
1158
|
-
redaction before the first hostloop/protocol record, derive budget assertions without a
|
|
1159
|
-
two-pass record. Start here when the question is "how do I do X", not "what does flag Y mean".
|
|
1160
|
-
- `references/scenario-schema.md` — scenario/session YAML schema, full assertion catalog (with each
|
|
1161
|
-
key's replay class), the web_fetch model, and an assertion/replay-scoped gotcha subset (the full
|
|
1162
|
-
landmine catalog lives in this SKILL's Gotchas section above).
|
|
1163
|
-
- `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
|
|
1164
|
-
- `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
|
|
1165
|
-
GitHub Actions pipeline.
|
|
1166
|
-
- `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
|
|
1167
|
-
no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
|
|
1168
|
-
(validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
|
|
1169
|
-
- Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
|
|
1170
|
-
alive* (Part II) above; the fuller recipe is in [`docs/run-status.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/run-status.md) (repo-only, not shipped with the
|
|
1171
|
-
installed skill).
|
|
148
|
+
| [`references/authoring.md`](references/authoring.md) | session vs scenario, discovery, fidelity tier, the answer-channel decision tree, `web_fetch`, scaffold + lint |
|
|
149
|
+
| [`references/assertions-guide.md`](references/assertions-guide.md) | the two assertion axes, the goal → key map |
|
|
150
|
+
| [`references/run-record-replay.md`](references/run-record-replay.md) | run / `verify-run`, recording and cassette placement, real-document validation, verdict signals (incl. where a relative or `outputs/` path lands per tier), background-run liveness, CI lanes |
|
|
151
|
+
| [`references/measurement.md`](references/measurement.md) | `--repeat`, `--ablate-skill`, measurement hygiene |
|
|
152
|
+
| [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
|
|
153
|
+
| [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
|
|
154
|
+
| [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality |
|
|
155
|
+
| [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
|
|
156
|
+
| [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
|
|
157
|
+
| [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
|
|
158
|
+
| [`references/ci-recipe.md`](references/ci-recipe.md) | the GitHub Action, replay-vs-live lanes, the four-stage pipeline |
|
|
159
|
+
| [`references/critique.md`](references/critique.md) | `critique` report and evidence-package shapes |
|
|
160
|
+
| `scripts/scenario.py` | `scaffold`, `lint`, `lint-skill`, `resolve-agent-types <plugin-dir>` (validates a pinned `subagent_type` against `plugin.json` + `agents/*.md`) |
|