cowork-harness 3.9.0 → 3.10.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +62 -23
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +16 -14
  3. package/.claude/skills/cowork-harness/references/critique.md +10 -5
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +5 -3
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +17 -15
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +3 -2
  7. package/.claude/skills/cowork-harness/scripts/scenario.py +59 -0
  8. package/.cowork-redact.json +18 -7
  9. package/AGENTS.md +1 -1
  10. package/CHANGELOG.md +370 -0
  11. package/CONTRIBUTING.md +7 -5
  12. package/DESIGN.md +1 -1
  13. package/README.md +16 -5
  14. package/SPEC.md +26 -2
  15. package/dist/agent/session.js +7 -0
  16. package/dist/assert.js +24 -8
  17. package/dist/baseline.js +6 -0
  18. package/dist/cli.js +11 -4
  19. package/dist/critique/command.js +167 -127
  20. package/dist/critique/limitations.js +1 -1
  21. package/dist/critique/mount-check.js +52 -0
  22. package/dist/critique/package-evidence.js +33 -17
  23. package/dist/decide/decider.js +4 -8
  24. package/dist/decide/external-channel.js +17 -17
  25. package/dist/egress/sidecar.js +18 -24
  26. package/dist/errors.js +38 -5
  27. package/dist/hostloop/pretooluse-path-hook.js +64 -0
  28. package/dist/hostloop/process-cwd.js +124 -0
  29. package/dist/redact.js +10 -2
  30. package/dist/run/cassette.js +7 -0
  31. package/dist/run/chat-result.js +1 -0
  32. package/dist/run/chat.js +3 -1
  33. package/dist/run/doctor.js +2 -2
  34. package/dist/run/execute.js +621 -84
  35. package/dist/run/lint-load.js +233 -0
  36. package/dist/run/outputs-delete-tier.js +59 -0
  37. package/dist/run/pre-run-manifest.js +57 -4
  38. package/dist/run/run.js +53 -6
  39. package/dist/run/scenario-tool.js +74 -10
  40. package/dist/run/verdict.js +47 -6
  41. package/dist/runtime/agent-image.js +20 -1
  42. package/dist/runtime/argv.js +1 -0
  43. package/dist/runtime/hostloop-prompt.js +1 -1
  44. package/dist/runtime/hostloop.js +61 -16
  45. package/dist/runtime/microvm.js +111 -4
  46. package/dist/scan.js +69 -8
  47. package/dist/secrets.js +1 -1
  48. package/dist/termination.js +143 -0
  49. package/dist/types.js +2 -2
  50. package/docs/README.md +4 -4
  51. package/docs/boundary.md +7 -6
  52. package/docs/cassette.md +73 -11
  53. package/docs/ci.md +2 -2
  54. package/docs/cli.md +105 -25
  55. package/docs/companion-skill.md +2 -2
  56. package/docs/critique.md +37 -18
  57. package/docs/decider-dir.md +3 -1
  58. package/docs/fidelity-gaps.md +43 -55
  59. package/docs/gotchas.md +1 -1
  60. package/docs/invariants.md +5 -5
  61. package/docs/maintenance.md +3 -0
  62. package/docs/run-status.md +12 -4
  63. package/docs/scenario.md +83 -47
  64. package/docs/session.md +6 -5
  65. package/docs/subagents.md +14 -11
  66. package/examples/README.md +2 -1
  67. package/examples/replays/README.md +1 -1
  68. package/examples/replays/hostloop-computer-links.cassette.json +59 -60
  69. package/llms.txt +3 -3
  70. package/package.json +1 -1
  71. package/python/test_scenario_lint.py +98 -0
  72. package/schema/run-result.json +35 -3
  73. package/schema/scenario.schema.json +3 -3
  74. package/scripts/check-versions.ts +46 -4
  75. package/scripts/gen-schema.ts +6 -5
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.9.0
7
- tracks-harness: cowork-harness 3.9.0 (baseline desktop-2.9939.2)
6
+ version: 3.10.0
7
+ tracks-harness: cowork-harness 3.10.0 (baseline desktop-2.9939.2)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.9.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.10.0` (baseline
29
29
  > `desktop-2.9939.2`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.9.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.9.0"`. **Pin `@^3.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.10.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.10.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.10.0"`. **Pin `@^3.10.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -372,7 +372,8 @@ python3 "$S" scaffold --name report-check --skill ./skills/report-gen \
372
372
  ```
373
373
 
374
374
  Then lint every scenario — it encodes the no-silent-false-green invariants. Use the CLI wrapper
375
- `cowork-harness lint` (it runs the same bundled `scenario.py lint`):
375
+ `cowork-harness lint`: it runs the bundled `scenario.py lint` **and** the harness's own scenario loader,
376
+ so a file `run`/`record` would refuse fails lint too (running `scenario.py lint` directly skips the loader):
376
377
 
377
378
  ```bash
378
379
  cowork-harness lint scenarios/*.yaml
@@ -391,11 +392,17 @@ and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is
391
392
  (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
392
393
  emits a scenario `lint` would reject.
393
394
 
394
- **`lint` is the LENIENT check — the loader is the strict one.** An unknown top-level key is a ⚠ WARN in
395
- `lint` (exit 0) but a **hard error** in the runtime (`Unrecognized key: "<k>"`, exit 2) — so a scenario
396
- that lints with warnings may still not run. To check whether a scenario actually loads, without spending:
397
- `cowork-harness record <file.yaml> --dry-run` (exit 2 on a schema error; a directory reports each
398
- `✗ broken:` file and exits 1). **Read the exit code, not just its sign:** `record <file>` — with or without
395
+ **`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
396
+ the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
397
+ a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
398
+ naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
399
+ does not check what depends on the machine the run happens on (the session file and its mounts, an
400
+ absolute `baseline:` path, environment variables). A session or matrix YAML in a linted directory is not
401
+ a scenario and is reported as one that does not load — keep those out of the linted set. `python3
402
+ scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
403
+ `cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
404
+ on a schema error; a directory reports each `✗ broken:` file and exits 1). **Read the exit code, not just
405
+ its sign:** `record <file>` — with or without
399
406
  `--dry-run` — answers `2` for "did not load" and `1` for "loaded fine, but this record is refused" (a
400
407
  pre-spend policy refusal; `--max-budget-usd` is the one refusal that keeps exit 2). Treating any non-zero
401
408
  as "scenario broken" mis-reports every refused-but-valid scenario. Corollary: **the loader** fails LOUD on an unknown key (never silently) —
@@ -421,8 +428,9 @@ When the scenario declares `answers:`, verify-run **also** checks they still mat
421
428
  reworded gate or a `choose:` the run never offered fails here in ~1s instead of on a paid re-record). Or skip
422
429
  the discovery/encode/record dance entirely and answer gates **live during the recording** with
423
430
  `record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
424
- `run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
425
- `cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
431
+ `run` takes no `--dry-run`: to check that a scenario **loads** without spending, `cowork-harness lint
432
+ <file.yaml>` runs the real loader (and resolves a named `baseline:`); `cowork-harness record <file.yaml>
433
+ --dry-run` runs the real loader too AND the same scenario-level
426
434
  refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
427
435
  cassette-portability pre-flight below**, so it cannot green something a paid run would reject. **That binding
428
436
  guarantee is the SINGLE-FILE form only** — it takes the real `--out` and the real flags, so its verdict is the
@@ -432,12 +440,14 @@ the verdict kind, and portability can only ever warn — that do NOT affect the
432
440
  takes no `--out`, so the destination is a guess — and only the path-independent ones (prompt policy, assert
433
441
  contradiction, duplicate cassette target) gate the batch. So a directory dry-run CAN exit 0 on a scenario the
434
442
  real `record` would refuse; re-run that one file with its real flags for a binding answer. A directory also
435
- reports every offender and the batch cost estimate. `lint` checks the assertion invariants (both above).
443
+ reports every offender and the batch cost estimate. `lint` checks the assertion invariants AND that each file loads (the same loader, plus a named `baseline:`), but not the pre-spend refusals.
436
444
 
437
445
  **Which arm to reach for.** They answer different questions, and picking the wrong one is why a consumer
438
446
  concluded the free pre-flight was unavailable:
439
- - **"Does my whole corpus still load?"** → the **directory** arm (`record scenarios/ --dry-run --quiet`,
440
- the CI shape in `references/ci-recipe.md`). It reports every offender in one pass, and the
447
+ - **"Does my whole corpus still load?"** → `cowork-harness lint scenarios/` answers it (every file the
448
+ loader rejects is an ERROR, and so is a `baseline:` naming no shipped baseline), or the **directory** arm
449
+ (`record scenarios/ --dry-run --quiet`, the CI shape in `references/ci-recipe.md`) when you also want the
450
+ pre-spend refusals. The directory arm reports every offender in one pass, and the
441
451
  destination-policy verdict cannot red it: that arm knows no `--out`, so host-inventory and portability
442
452
  are advisory `notes[]` at exit 0 while a file that cannot load is `✗ broken:` at exit 1. Limits worth
443
453
  knowing: it is **non-recursive** (`readdirSync` — scenarios in subdirectories are never opened), a file
@@ -593,11 +603,15 @@ Recognize these before "fixing" a non-bug:
593
603
  **Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
594
604
  scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
595
605
  **The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
596
- can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
597
- (what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
598
- `outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
599
- `fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
600
- bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
606
+ can see them, and give the file tools an **absolute** path under the outputs directory the agent's prompt
607
+ names. On the desktop-local host-loop lane (what production runs), against Desktop **2.7032.0 and later**,
608
+ the agent process runs outside the session (`/var/empty`), so a relative `Read`/`Write`/`Edit` — a bare
609
+ filename or `outputs/x.md` alike — is **refused** ("File is in a directory that is denied by your
610
+ permission settings."); only a pathless or relative `Grep`/`Glob` is redirected to outputs. (Before
611
+ 2.7032.0 the file tools were rooted at `outputs/`, so a bare filename landed there and `outputs/x.md`
612
+ doubled to `outputs/outputs/x.md`.) At `fidelity: container`/`microvm` (VM-loop, the harness default) the
613
+ base is the session root, so a bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an
614
+ explicit delivery. Addressing
601
615
  a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
602
616
  decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
603
617
  [docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
@@ -638,11 +652,36 @@ Recognize these before "fixing" a non-bug:
638
652
  `Command timed out after <duration>` in stderr, matching production. Known gap: if **every** exec
639
653
  failed, the agent ran nothing yet the run still only warns — read `result.infraErrors` when a run looks
640
654
  suspiciously empty.
655
+ - **`outputs_delete_unconfirmed`** (`WARN`) — a delete-shaped command near `mnt/outputs` that nothing
656
+ confirms: the per-turn filesystem diff shows no output present at turn start was deleted, and no flagged
657
+ delete has an `outputs/` path as its own operand. The classic case is a Python variable named `rm` in a
658
+ `python3 -c` body that also reads a report from outputs: `rm = json.load(open(".../outputs/r.json"))`.
659
+ A file the turn created and then deleted is invisible to the diff, so real deletes land here too:
660
+ - a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`);
661
+ - a `cd` then a relative path;
662
+ - chained variables (`A=…; B=$A/x; rm "$B"`);
663
+ - a Python path held in a variable set on another line (`p = …` then `os.remove(p)`);
664
+ - wrapper flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`);
665
+ - calls outside the modelled set, such as Node's `fs.promises.rm(…)`.
666
+
667
+ Read the command before dismissing it. A literal-path delete (`rm -f mnt/outputs/x`,
668
+ `os.remove(".../outputs/x")`) still fails `outputs_delete`. So do two non-deletes: quoted text where a
669
+ delete command with an outputs operand follows a separator, subshell or keyword
670
+ (`echo 'note; rm mnt/outputs/x'` — the classifier does not track quotes), and a heredoc that *writes* a
671
+ script instead of running it. A statement over 4 KiB or a command over 16 KiB is judged by the stricter
672
+ original rule, so a huge one-line body with a variable named `rm` fails again, and a command whose variable
673
+ expansion would exceed the scanner's work budget (about a hundred distinct variables in one 10 KB line) is
674
+ not expanded — every mount it names literally counts as deleted in. Waive any of these with
675
+ `allow_outputs_delete`. The warn is raised even when `no_delete_in_outputs` is authored (the assertion
676
+ passes; this warn is how the hit stays visible in text output).
677
+ - **`outputs_diff_unavailable`** (`WARN`) — the outputs filesystem diff could not verify this turn and the
678
+ text scan saw nothing, so a delete by a script file or a non-bash tool would have gone unseen.
641
679
  - **`scan_unavailable`** (`WARN`) — emitted only on the live lane: `events.jsonl` was missing/corrupt, so
642
- `RunResult.scan` is undefined and the host-path + outputs-delete guards **did not run this run**. Not a
680
+ `RunResult.scan` is undefined and the host-path guard and the outputs-delete **text scan did not run this
681
+ run** (the outputs filesystem diff still did, and a delete it proves still fails). Not a
643
682
  pass or a defect — assert `no_delete_in_outputs` / `transcript_no_host_path` to hard-fail on it instead.
644
683
 
645
- The full 20-code signal table (severity + per-signal opt-out) is in
684
+ The full 22-code signal table (severity + per-signal opt-out) is in
646
685
  [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
647
686
  the fuller narrative.
648
687
 
@@ -1066,7 +1105,7 @@ authorable). Reach for this list when debugging a run's behavior, that one while
1066
1105
  20. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
1067
1106
  read-only folder's files as path + hash only (`truncated: true`, no `body`) — it's an input the agent
1068
1107
  read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass against it on replay
1069
- (the hash-only entry still materializes a placeholder); `artifact_json` reports a clear
1108
+ (the hash-only entry still materializes a placeholder); `artifact_json`/`artifact_text` report a clear
1070
1109
  evidence-unavailable on every lane (live/verify-run/replay agree — no green-record/red-replay). This is
1071
1110
  also why a `mode: r` input never trips the `binary` privacy finding or needs `--allow` — only a
1072
1111
  *committed* body is scanned. `scaffold` won't emit `file_exists` for one either (it's not in
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
3
+ Self-contained reference. Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.9.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.10.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -77,7 +77,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
77
77
  GitHub-hosted runners, no token/Docker/agent:
78
78
 
79
79
  ```yaml
80
- - run: npm i -g "cowork-harness@^3.9.0"
80
+ - run: npm i -g "cowork-harness@^3.10.0"
81
81
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
82
82
  # no silent false-greens. WITHOUT --strict this
83
83
  # step cannot fail on a WARN-class rule (e.g.
@@ -208,7 +208,7 @@ no live filesystem and no network — and it verifies an artifact's *content* on
208
208
  carries an `artifacts` manifest (recorded `outputs/` + connected folders; then `file_exists` /
209
209
  `user_visible_artifact` / `artifact_json` evaluate on replay). On a manifest-less cassette the
210
210
  deliverable is invisible to the gate. Don't let a green replay gate convince you the deliverable is correct.
211
- Run `cowork-harness lint` (the bundled `scenario.py lint`) in CI to catch a scenario that put a
211
+ Run `cowork-harness lint` (the bundled `scenario.py lint` plus the harness's own loader) in CI to catch a scenario that put a
212
212
  filesystem/egress-only check on the replay lane (a silent no-op). Author new scenarios with
213
213
  `scenario.py scaffold` so they start from a valid, self-linted skeleton.
214
214
 
@@ -238,8 +238,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
238
238
  `.cowork-redact.json` next to your scenarios (or set `COWORK_HARNESS_REDACT_PATTERNS` /
239
239
  `_KEYS`); empty by default. The policy is searched in **cwd → the scenario's dir → the cassette's
240
240
  dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
241
- packaged reference template (local-path prefixes + a generic email regex) into the cwd as a starting
242
- point — review and tailor it. Redaction is **verdict-preserving** — `record` refuses to write if
241
+ packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
242
+ like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
243
+ it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` or add them. Redaction is **verdict-preserving** — `record` refuses to write if
243
244
  redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
244
245
  inputs.
245
246
  - **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
@@ -307,14 +308,15 @@ cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness onl
307
308
  A typical skill repo runs four stages, fastest/cheapest first:
308
309
 
309
310
  1. **Unit** — your skill's own tests (pytest/vitest of its scripts). Not the harness's job.
310
- 2. **Boundary / lint** — `cowork-harness lint scenarios/*.yaml` (no-silent-false-green invariants; needs
311
- python3 — PyYAML is bundled) + `cowork-harness verify-cassettes cassettes/` (privacy scan + staleness) +
311
+ 2. **Boundary / lint** — `cowork-harness lint scenarios/*.yaml` (no-silent-false-green invariants, and
312
+ every scenario the loader would refuse is an ERROR; needs python3 — PyYAML is bundled) + `cowork-harness verify-cassettes cassettes/` (privacy scan + staleness) +
312
313
  `cowork-harness boundary-check` where relevant. Token-free, agent-free. **Don't `|| true` the lint
313
314
  step** — a missing python3 (exit 127) or a lint error makes `scenario.py` exit non-zero, and swallowing
314
315
  that turns the false-green guard itself into a silent no-op.
315
316
 
316
- **Add a load gate next to `lint`** — they answer different questions, and `lint` is the more permissive
317
- of the two (it *warns* on an unknown key; the runtime *refuses* to load one):
317
+ **Optionally add a load gate next to `lint`** — `lint` already runs the loader (an unknown key, a wrong
318
+ value type or an unknown `baseline:` name is an ERROR), and `record --dry-run` adds the pre-spend
319
+ refusals a real record would apply:
318
320
 
319
321
  ```bash
320
322
  cowork-harness record scenarios/ --dry-run --quiet # does every scenario LOAD? no tokens, writes nothing
@@ -324,8 +326,8 @@ A typical skill repo runs four stages, fastest/cheapest first:
324
326
  failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
325
327
  offending file *and* the rejected key, one line per file, and the step still exits 1. (Silent is
326
328
  literal only when your scenarios name a `fidelity:`; one still in the deprecation window prints one
327
- defaulted-fidelity notice per scenario.) A scenario that lints with only
328
- warnings can still be unloadable, so a green `lint` is not evidence the suite runs.
329
+ defaulted-fidelity notice per scenario.) Point `lint` at scenarios only: a session or matrix YAML in
330
+ the linted set is reported as a file that does not load.
329
331
 
330
332
  **If the repo pays for `critique`, gate the evidence corpus here first, for free:**
331
333
 
@@ -364,7 +366,7 @@ jobs:
364
366
  with: { node-version: '24' }
365
367
  - uses: actions/setup-python@v5
366
368
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
367
- - run: npm i -g "cowork-harness@^3.9.0"
369
+ - run: npm i -g "cowork-harness@^3.10.0"
368
370
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
369
371
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
370
372
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -393,7 +395,7 @@ jobs:
393
395
  echo "live=true" >> "$GITHUB_OUTPUT"
394
396
  fi
395
397
  - if: steps.guard.outputs.live == 'true'
396
- run: npm i -g "cowork-harness@^3.9.0"
398
+ run: npm i -g "cowork-harness@^3.10.0"
397
399
  - if: steps.guard.outputs.live == 'true'
398
400
  run: cowork-harness run scenarios/ --output-format json
399
401
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -144,10 +144,15 @@ non-git folder is measured raw, as staging copies it). Unlike `lint-skill`'s sta
144
144
  packager: untracked files and symlinks outside the plugin are excluded by the same filter and containment
145
145
  rule a critique applies, bytes are the same UTF-8-decoded measurement, and the one clause no static
146
146
  instrument can see (a plugin-root reference read at run time) is stated as the floor rather than guessed
147
- at. Known gap: a skill that is a git submodule of its plugin (or any `--skill` subdirectory with nothing
148
- tracked under it) is REFUSED by `--corpus-only` in staging's terms, but a live critique's packager still
149
- accepts it from the directory's own index and grades a skill the mount never delivers — pre-existing,
150
- rare, not fixed here.
147
+ at. A skill that is a git submodule of its plugin (or any `--skill` subdirectory with nothing tracked
148
+ under it) is refused in staging's terms by `--corpus-only` and by a paid critique alike, before any spend;
149
+ the corpus only ever holds files the mount carries.
150
+
151
+ `critique <plugin>/skills/<name>` is the same run as `critique <plugin> --skill <name>`: critique mounts
152
+ the plugin, as Cowork does, and grades `<name>` (same corpus, `skillHash`, `gradedSkill`), with a notice.
153
+ It mounts the skill folder alone — and says why — only when `--skill` cannot reach it (not at exactly
154
+ `skills/<name>`, a submodule, or a case mismatch); that run and its corpus lack the plugin's agents and
155
+ shared references.
151
156
 
152
157
  The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
153
158
  from `dist/` source:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`).
3
+ Self-contained reference. Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -309,8 +309,10 @@ up often enough to spell out:
309
309
 
310
310
  - `COWORK_HARNESS_RUNS_DIR` (or `--run-dir <path>`) — override the default run-output root `~/.cowork-harness/runs` (out of any working tree). flag > env > default.
311
311
  - `COWORK_HARNESS_DECIDER_CMD_TIMEOUT_MS` / `COWORK_HARNESS_LLM_TIMEOUT_MS` — decider backstops
312
- (default 600 s; **fail loud** on timeout).
313
- - `COWORK_HARNESS_DECIDER_DIR_POLL_MS` / `_TIMEOUT_MS` — the `--decider-dir` rendezvous (poll defaults: 300 ms for the run-side rendezvous, 500 ms for `gates --follow`).
312
+ (default 600 s; **fail loud** on timeout). A `--decider-cmd` timeout, or a helper that exits before
313
+ answering, ends the run as an unanswered-gate partial (`result.json` written, exit 2); a timeout is
314
+ recorded as `errorSource: "decider_timeout"`.
315
+ - `COWORK_HARNESS_DECIDER_DIR_POLL_MS` / `_TIMEOUT_MS` — the `--decider-dir` rendezvous (poll defaults: 300 ms for the run-side rendezvous, 500 ms for `gates --follow`). A backstop timeout ends the run as an unanswered-gate partial recorded as `errorSource: "decider_timeout"`.
314
316
  - `COWORK_HARNESS_DIALOG_TIMEOUT_MS` — dialog auto-cancel (default 6 s).
315
317
  - `COWORK_HARNESS_LLM_MAX_BYTES` — stdout bound on `--decider-llm` (default 8 MiB).
316
318
  - `COWORK_HARNESS_LLM_RETRIES` — bounded retries for a transient non-zero `claude -p` exit on the
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.9.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.10.0`
4
4
  (baseline `desktop-2.9939.2`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -307,7 +307,7 @@ same set live from the schema.
307
307
  | `transcript_not_matches: <regex>` | it does not match (e.g. no leaked stack trace) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so text emitted inside a tool call (a gate question, an option label or description) can never match; use `question_context`/`question_asked` or `tool_result_contains` for those. |
308
308
  | `file_exists: <path>` | the path exists under the run's `work/` (anchored at `mnt/`, e.g. `outputs/x.md`). For a user-facing deliverable prefer `user_visible_artifact` — with a connected folder the file lands in `mnt/<folder>` (= `{{workspaceFolder}}`), not `mnt/outputs`, so `file_exists: outputs/x.md` misses it |
309
309
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
310
- | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan, not mount enforcement, so a green means none was *detected* |
310
+ | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected. **Omitting it does NOT allow deletes** — a detected delete still fails via the `outputs_delete` verdict signal, which fires *because* the key was not authored; use `allow_outputs_delete: true` to accept an intended one. Detection is a post-run bash-command scan plus a per-turn filesystem diff of `outputs/`, not mount enforcement, so a green means none was *detected*. Fails when the diff proves the delete, when a delete in command/call position has an `outputs/` path as its own operand, or when the diff could not verify; a hit resting only on the detector's inference (e.g. a Python variable named `rm`) with a clean diff passes, and the `outputs_delete_unconfirmed` warn is still raised in the run output — see that code for the classes of real delete that land there |
311
311
  | `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
312
312
  | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
313
313
  | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
@@ -315,7 +315,7 @@ same set live from the schema.
315
315
  | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
316
316
  | `artifact_text: {artifact, contains?, not_contains?, matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read |
317
317
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
318
- | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
318
+ | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. `cowork-harness lint` reports it too (ERROR `scenario-invalid` — the wrapper runs the same loader); the bundled `scenario.py lint` run directly does NOT |
319
319
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
320
320
  | **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
321
321
  | **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
@@ -358,7 +358,7 @@ same set live from the schema.
358
358
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
359
359
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
360
360
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
361
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs`, inside it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
361
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, `/private/var/empty` from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
362
362
  | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
363
363
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
364
364
  | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
@@ -367,7 +367,7 @@ same set live from the schema.
367
367
  | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
368
368
  | `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
369
369
  | `gate_answer_count_min: <N>` | at least N AskUserQuestion gates fired AND were delivered non-error — presence companion to `gate_answers_delivered`'s vacuous-pass. **`: 0` asserts nothing** and does not satisfy that pairing; `>= 1` is **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
370
- | `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream) |
370
+ | `hook_blocked: <regex>` | a PreToolUse hook blocked a tool whose name matches the regex (`RunResult.hookEvents`) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette (a custom hook's decision lives only there, not the recorded stream). **Does not see the agent's own refusal**: on `hostloop` against Desktop 2.7032.0+ a relative `Read`/`Write`/`Edit` is denied by the agent's permission rules before any hook runs, so neither this key nor `path_denied` records it — assert it with `tool_result_contains: "denied by your permission settings"` |
371
371
  | `no_hook_blocked: true` | no tool was hook-blocked during the run (distinguishes a real tool crash from an intentional hook block) — evidence-unavailable if hook telemetry is absent. Replay: needs a `controlOut` cassette. **Only `true` is valid** |
372
372
  | `hook_event_fired: <HookEvent>` | a **command hook** for this event (a plugin's `hooks/hooks.json` or manifest hook — `Stop`, `SessionStart`, `PostToolUse`, …) ran: a `hook_response` system frame with that `hook_event` was recorded (`RunResult.contextEvents`). Any outcome counts. The harness passes `--include-hook-events` whenever a staged plugin declares hooks — that is what puts events other than SessionStart/Setup on the stream — so a recording made without it reports "never fired". Content-class, grades on replay. Recorded end-to-end for `Stop` ([stop-hook-probe.scenario.yaml](https://github.com/yaniv-golan/cowork-harness/blob/main/examples/probes/stop-hook-probe.scenario.yaml)); the other names match the same frame but have not each been recorded |
373
373
  | `hook_event_blocked: <HookEvent>` | that command hook **blocked** at least once — a `hook_response` frame for the event carried `exit_code: 2`. Fails naming the exit codes seen when it fired without blocking (a frame with no `exit_code` is reported as such, never counted); fails "never fired" otherwise; cannot-verify when the run has no context events. Content-class |
@@ -379,9 +379,9 @@ same set live from the schema.
379
379
  | `allow_l0_host_config_contamination: true` | verdict modifier — accept a contaminated L0 environment: suppresses the default-fail when `protocol` runs against your REAL config dir, where your installed plugins/skills/auto-memory/MCP servers are visible and may answer instead of the thing under test. Live tiers only |
380
380
  | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
381
381
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
382
- | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
382
+ | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Silences `outputs_delete`, `outputs_delete_unconfirmed` and `outputs_diff_unavailable`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
383
383
  | `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
384
- | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
384
+ | `transcript_no_host_path: true` | no host path (`/Users/`, `/opt/cowork/`, `/home/`, `/root/`, and the macOS `/private/var/`, `/private/tmp/`, `/var/folders/`, `/Volumes/` roots — also inside a `file://` or `computer://` link) leaked into model-visible text — **incompatible with `hostloop` AND `protocol`**: hostloop's native file tools legitimately expose real host paths, and protocol (L0) runs the agent's file tools on the real host cwd with no sealed filesystem, so this fails BY DESIGN on both (the harness warns at run start if asserted anyway); use `container`/`microvm` for this check |
385
385
  | `egress_denied: <host>` | the host was blocked by the egress proxy |
386
386
  | `egress_allowed: <host>` | the host was allowed through |
387
387
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
@@ -405,7 +405,7 @@ dotted path.
405
405
 
406
406
  **VerdictSignals in `result.verdict.signals`:** `computeVerdict` pushes signals into `result.verdict.signals`; eleven
407
407
  are **fail**-severity (they flip the run's pass/exit code even though `result.result` itself stays
408
- `"success"`) and nine are **warn**-severity (informational, never flip pass/fail). All twenty signal
408
+ `"success"`) and eleven are **warn**-severity (informational, never flip pass/fail). All twenty-two signal
409
409
  codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
410
410
 
411
411
  | Code | Severity | Meaning |
@@ -415,7 +415,9 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
415
415
  | `usage_limit` | fail | Usage/quota limit hit (not a skill failure) — retry after the limit resets. Emitted when `RunResult.resultErrorKind === "usage_limit"` |
416
416
  | `transport_error` | fail | The connection dropped mid/after-run |
417
417
  | `permissive_auto_allow` | fail | A cowork-parity auto-allow real Cowork would block (opt out: `allow_permissive_auto_allow`) |
418
- | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs` (opt out: author `no_delete_in_outputs`) |
418
+ | `outputs_delete` | fail | An unauthorized delete touched `mnt/outputs`, confirmed: the per-turn filesystem diff proves it, a delete in command/call position has an `outputs/` path as its own operand, or the diff could not verify the turn. Authoring `no_delete_in_outputs` moves it into that assertion; `allow_outputs_delete` waives it |
419
+ | `outputs_delete_unconfirmed` | warn | A delete-shaped command near `mnt/outputs` nothing confirms: no output present at turn start was deleted and no flagged delete has an outputs path as its own operand (e.g. a Python variable named `rm` next to an outputs path, quoted prose, a `sed`/`grep` pattern). A *statement* is a fragment split on newline/`;`/`&&`/`\|\|`, quote-blind, so some real deletes also land here — a loop body whose operand is the loop variable (`for f in …; do rm "$f"; done`), a `cd` then a relative path, chained variables (`A=…; B=$A/x; rm "$B"`), a Python path held in a variable set on another line (`p = …` then `os.remove(p)`, or `for p in …:` then `p.unlink()`), wrappers with flag combinations the classifier does not model (`sudo -Hu user rm`, `git -C dir rm`), and calls outside the modelled set such as Node's `fs.promises.rm(…)` — read the command. Raised even when `no_delete_in_outputs` is authored (that assertion passes on it). Kept false positives that still fail `outputs_delete`: quoted text in which a delete command with an outputs operand follows a shell separator, subshell or keyword — the classifier does not track quotes (`echo 'note; rm mnt/outputs/x'`, `echo "a & rm …/outputs/x"`), and a heredoc that *writes* a script rather than running it (`cat <<EOF > clean.sh` with an `rm …/outputs/x` line). Waive: `allow_outputs_delete` |
420
+ | `outputs_diff_unavailable` | warn | The per-turn filesystem diff of `outputs/` could not verify this turn and the text scan flagged nothing — a delete made without a bash command would have gone undetected. With a text hit the turn fails `outputs_delete` instead |
419
421
  | `mount_delete` | warn | A delete touched a delete-denied mount other than `outputs` (a `rw` connected folder). Production denies `unlink`/`rmdir` there until per-mount approval, so the run diverged. Warn, not fail: the harness detects post-hoc what production enforces. Author `no_delete_in_mounts` to hard-fail, or `allow_delete_in` to waive |
420
422
  | `host_path_leak` | fail | A host path leaked into model-visible text (opt out: author `transcript_no_host_path`) |
421
423
  | `l0_host_config_contamination` | fail | `protocol` ran against the operator's REAL config dir, so host-installed plugins/skills/memory/MCP may have answered instead of the thing under test — the run did not necessarily measure it (opt out: `allow_l0_host_config_contamination`). Name predates the meaning: delivery at L0 works, contamination is what this reports |
@@ -425,7 +427,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
425
427
  | `non_deterministic` | warn | The run was LLM/external/human-decided — not reproducible |
426
428
  | `model_fallback` | warn | The agent fell back off the requested model mid-run (SDK `model_fallback` event); a `model_not_found`/`model_blocked` trigger repeats every run until the pin changes |
427
429
  | `prompt_asset_missing` | warn | The run proceeded with a missing prompt asset (`COWORK_HARNESS_ALLOW_MISSING_PROMPT=1`); fidelity is degraded |
428
- | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path and outputs-delete guards did not run this run |
430
+ | `scan_unavailable` | warn | Post-run scan evidence unavailable (`RunResult.scan` undefined) — the host-path guard and the outputs-delete text scan did not run this run; the outputs filesystem diff still ran |
429
431
  | `ended_with_question` | warn | Live-lane heuristic: the final answer contains a question and the run wrote no deliverable to `outputs/` — the lenient sibling of `stalled` (covers a mid-message `?`, or tool work after the last gate that still ended asking). Opt out: `allow_stall` |
430
432
  | `undelivered_deliverables` | warn | The skill produced file(s) OUTSIDE every user-visible root and never delivered them, so they stay invisible to the user. Fires on every run without opting in, because the scenarios that most need it are the ones whose author never considered delivery. Silent when the evidence cannot answer the question (no workspace walk, a tier that runs no scratchpad walk, absent delivery telemetry, a resumed turn, or a lane where delivery is unobservable — see `delivery_unobservable`) — never a vacuous clean. **`lane: local` only**: on remote, delivery cannot be measured at all, so that lane reports `delivery_unobservable` instead of guessing. Opt out: `allow_undelivered_deliverables` |
431
433
  | `delivery_unobservable` | warn | `lane: remote` only — the run produced file(s), but whether any reached the user CANNOT be verified: nothing is delivered by location on that lane and the harness models no remote delivery tool (production uses the agent-native `SendUserFile`). The honest counterpart to `undelivered_deliverables`, which would otherwise fire on every remote run that writes anything — a signal that always fires carries no information. Mutually exclusive with it; quiet when the run produced nothing to deliver. A harness coverage gap, not a skill defect. Opt out: `allow_undelivered_deliverables` |
@@ -433,7 +435,7 @@ codes (`VerdictSignal["code"]` in `src/run/verdict.ts`):
433
435
 
434
436
  A **fail**-severity signal does not change `result.result` (still `"success"`), but it DOES fail the
435
437
  overall run verdict and exit code — `assert result: success` alone won't catch it; check
436
- `result.verdict.signals[].severity` or the run's exit code. Only the nine **warn** codes are truly benign.
438
+ `result.verdict.signals[].severity` or the run's exit code. Only the eleven **warn** codes are truly benign.
437
439
 
438
440
  ## Replay class
439
441
 
@@ -580,8 +582,8 @@ debugging a run's behavior. The two are **numbered independently**: a bare "gotc
580
582
  the stochastic path flags the run `nonDeterministic`. The LLM decider is one mechanism, two
581
583
  spellings: `on_unanswered: llm` (YAML) and `--decider-llm` (CLI). The bare `--on-unanswered llm`
582
584
  is rejected (use `--decider-llm`). `agent` is **retired** — `on_unanswered: agent` is rejected by
583
- the schema. (`src/types.ts` — the `on_unanswered` enum; `src/cli.ts:899` — the CLI-side
584
- `--on-unanswered` value check.)
585
+ the schema. (`src/types.ts` — the `on_unanswered` enum; `src/cli.ts` — the CLI-side
586
+ `--on-unanswered` value check; grep `--on-unanswered llm is not a user flag`.)
585
587
 
586
588
  4. **`--on-unanswered first` is non-deterministic too** — it picks option 1 and is flagged
587
589
  `nonDeterministic`; not a deterministic substitute for scripted answers.
@@ -662,8 +664,8 @@ debugging a run's behavior. The two are **numbered independently**: a bare "gotc
662
664
  21. **A `mode: r` connected folder's contents are recorded body-less, not excluded.** `record` captures a
663
665
  read-only folder's files as `path` + `bytes` + `sha256` only (`truncated: true`, no `body`) — it's an
664
666
  input the agent read, not a deliverable it wrote. `file_exists`/`computer_links_resolve` still pass
665
- against it on replay (the hash-only entry still materializes a 0-byte placeholder); `artifact_json`
666
- reports a clear evidence-unavailable on every lane (live/verify-run/replay agree). This is also why a
667
+ against it on replay (the hash-only entry still materializes a 0-byte placeholder); `artifact_json`/`artifact_text`
668
+ report a clear evidence-unavailable on every lane (live/verify-run/replay agree). This is also why a
667
669
  `mode: r` input never trips the `binary` privacy finding
668
670
  or needs `--allow` in `verify-cassettes` — only a *committed* body is scanned. `scaffold` won't emit
669
671
  `file_exists` for one either, since it isn't in `RunResult.artifacts`. A `mode: rw`/`rwd` folder's
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.9.0` (baseline `desktop-2.9939.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.10.0` (baseline `desktop-2.9939.2`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -84,7 +84,8 @@ the transcript and would be committed inside the cassette. Set the policy up **b
84
84
  record — retrofitting means re-recording:
85
85
 
86
86
  1. `cowork-harness init-redact` — copies the reference `.cowork-redact.json` (host-path +
87
- email patterns) into the current directory. Version-control it.
87
+ email patterns) into the current directory. Version-control it. Re-run with `--force` after an
88
+ upgrade to pick up new rules (save your tailoring first).
88
89
  2. **Search set:** record looks for `.cowork-redact.json` in the **cwd, the scenario's directory,
89
90
  and the cassette's output directory** — every distinct file found is MERGED (plus
90
91
  `COWORK_HARNESS_REDACT_PATTERNS`/`_KEYS` from the env). Repo root (= cwd in CI) is the
@@ -50,6 +50,15 @@ lint-skill flags (skill bodies + any sibling hooks.json):
50
50
  live-verified) but has no assertion key, so a scenario can't gate on it
51
51
  W `${CLAUDE_PLUGIN_ROOT}` in a VM bash step / host-side hook seeding (host-loop footguns)
52
52
 
53
+ Through the `cowork-harness lint` CLI wrapper (not when this script is run directly), lint ALSO reports
54
+ every file the harness's own scenario loader rejects -- the check `run`/`record` apply before anything
55
+ runs:
56
+ E `scenario-invalid` the loader refuses the file (schema, value types, unknown keys, bad regex,
57
+ reserved values) -- including a non-scenario YAML in a linted directory
58
+ E `baseline-unknown` `baseline:` names no baseline this installed cowork-harness ships
59
+ Run directly, this script stays offline and never parses with the loader, so prefer the CLI wrapper
60
+ when it is installed.
61
+
53
62
  Designed for agents and CI: non-interactive, --help, --json, meaningful exit codes,
54
63
  idempotent. `lint` exits 1 on any ERROR (or any finding with --strict); else 0.
55
64
 
@@ -1322,6 +1331,55 @@ def _print_findings(findings, n_files, kind="scenario", clean_suffix=" — no si
1322
1331
  print(f"\n{n_err} error(s), {n_warn} warning(s), {n_info} info across {n_files} file(s).")
1323
1332
 
1324
1333
 
1334
+ _EXTRA_FINDINGS_ENV = "COWORK_HARNESS_LINT_EXTRA_FINDINGS"
1335
+
1336
+
1337
+ def _wrapper_loader_findings():
1338
+ """Findings from the harness's own scenario loader, handed over by the `cowork-harness lint` wrapper.
1339
+
1340
+ The wrapper runs the loader `run`/`record` use before spawning this script and writes what it rejects
1341
+ (`scenario-invalid`, `baseline-unknown`) to a JSON file named by COWORK_HARNESS_LINT_EXTRA_FINDINGS.
1342
+ Merged here so ONE renderer, ONE --min-severity filter and ONE exit rule apply to both sets.
1343
+
1344
+ Honoured only together with COWORK_HARNESS_PROG, which the wrapper always sets: a direct
1345
+ `python3 scenario.py lint` never reads the variable, even if the shell happens to carry it. A blank
1346
+ value is unset. A file that can't be read or doesn't have the expected shape is an ERROR, never a
1347
+ silent drop -- a dropped loader finding would be a false green."""
1348
+ path = (os.environ.get(_EXTRA_FINDINGS_ENV) or "").strip()
1349
+ if not path or not os.environ.get("COWORK_HARNESS_PROG"):
1350
+ return []
1351
+
1352
+ def bad(why):
1353
+ return [
1354
+ Finding(
1355
+ "ERROR",
1356
+ "linter-extra-findings-invalid",
1357
+ f"could not read the scenario-loader findings handed over by cowork-harness: {why}",
1358
+ "This is a harness bug, not a scenario problem -- please report it. "
1359
+ "`cowork-harness record <file> --dry-run` checks that a scenario loads.",
1360
+ "(scenario.py)",
1361
+ )
1362
+ ]
1363
+
1364
+ try:
1365
+ entries = json.loads(Path(path).read_text(encoding="utf-8"))
1366
+ except (OSError, ValueError) as e:
1367
+ return bad(str(e))
1368
+ if not isinstance(entries, list):
1369
+ return bad("expected a JSON array")
1370
+ out = []
1371
+ for i, x in enumerate(entries):
1372
+ if (
1373
+ not isinstance(x, dict)
1374
+ or x.get("severity") not in SEV_ORDER
1375
+ or not all(isinstance(x.get(k), str) for k in ("rule", "message", "fix", "file"))
1376
+ or not (x.get("line") is None or (isinstance(x.get("line"), int) and not isinstance(x.get("line"), bool)))
1377
+ ):
1378
+ return bad(f"entry {i} is not a finding")
1379
+ out.append(Finding(x["severity"], x["rule"], x["message"], x["fix"], x["file"], x.get("line")))
1380
+ return out
1381
+
1382
+
1325
1383
  def cmd_lint(args):
1326
1384
  all_findings = []
1327
1385
  # Expand directory args to their scenario files — mirrors src/run/inputs.ts `resolveInputs`: a SINGLE
@@ -1363,6 +1421,7 @@ def cmd_lint(args):
1363
1421
  )
1364
1422
  for f in args.files:
1365
1423
  all_findings.extend(lint_file(f))
1424
+ all_findings.extend(_wrapper_loader_findings())
1366
1425
  # Filter BEFORE rendering AND before the exit computation — deliberately, so --min-severity narrows
1367
1426
  # what the run actually cares about. Filtering at render only would make `--strict --min-severity ERROR`
1368
1427
  # print "0 findings" and still exit 1 (because --strict keys off the unfiltered set), which is