cowork-harness 1.22.0 → 1.24.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +94 -13
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +46 -11
  3. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +2 -2
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +34 -11
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +24 -7
  7. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +3 -0
  8. package/.claude/skills/cowork-harness/scripts/scenario.py +57 -10
  9. package/CHANGELOG.md +516 -0
  10. package/DESIGN.md +3 -3
  11. package/README.md +40 -15
  12. package/SPEC.md +1 -1
  13. package/baselines/desktop-1.28929.0.json +13 -3
  14. package/baselines/desktop-1.30096.1.json +577 -0
  15. package/baselines/desktop-1.32352.0.json +580 -0
  16. package/baselines/prompts/cowork-system-prompt-fingerprints.json +241 -7
  17. package/dist/assert.js +259 -2
  18. package/dist/cli.js +36 -2
  19. package/dist/critique/command.js +53 -6
  20. package/dist/run/artifacts.js +11 -0
  21. package/dist/run/cassette.js +263 -29
  22. package/dist/run/execute.js +4 -0
  23. package/dist/run/run.js +14 -1
  24. package/dist/run/verdict.js +5 -3
  25. package/dist/runtime/image-capabilities.js +33 -0
  26. package/dist/sync/cowork-sync.js +667 -42
  27. package/dist/types.js +47 -1
  28. package/docs/README.md +9 -2
  29. package/docs/cassette.md +39 -4
  30. package/docs/critique.md +25 -5
  31. package/docs/debugging.md +22 -0
  32. package/docs/discovery.md +9 -0
  33. package/docs/fidelity-gaps.md +88 -9
  34. package/docs/gotchas.md +6 -0
  35. package/docs/maintenance.md +5 -3
  36. package/docs/scenario.md +27 -8
  37. package/docs/session.md +10 -1
  38. package/examples/replays/README.md +1 -1
  39. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  40. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  41. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  42. package/llms.txt +2 -0
  43. package/package.json +1 -1
  44. package/schema/cassette.v10.json +18 -0
  45. package/schema/cassette.v11.json +18 -0
  46. package/schema/run-result.json +15 -2
  47. package/schema/scenario.schema.json +82 -1
  48. package/scripts/check-versions.ts +198 -0
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.22.0
7
- tracks-harness: cowork-harness 1.22.0 (baseline desktop-1.28929.0)
6
+ version: 1.24.0
7
+ tracks-harness: cowork-harness 1.24.0 (baseline desktop-1.32352.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.22.0` (baseline
26
- > `desktop-1.28929.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.24.0` (baseline
26
+ > `desktop-1.32352.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -39,7 +39,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.22.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.22.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.22.0"`. **Pin `@>=1.22.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.24.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.24.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@>=1.24.0"`. **Pin `@>=1.24.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
44
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
45
45
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -293,6 +293,8 @@ them by what you're trying to prove:
293
293
  | the skill didn't error out of a tool | `tool_no_error: <regex>`, `max_tool_errors: <N>` |
294
294
  | it didn't waste repeated identical calls | `max_redundant_tool_calls: <N>` |
295
295
  | a deliverable reached the user | `user_visible_artifact: <path>` (+ `no_scratchpad_leak: true` if it delivers via `present_files` — **`container` only**) |
296
+ | an internal name/path did **not** leak into a delivered file | `artifact_text: {artifact, not_contains}` — `artifact_json`'s companion for non-JSON bodies; literal path, no glob, so one entry per delivered surface |
297
+ | a named path must **not** exist after the run | `file_absent: <path>` (**live/verify-run only**) — do NOT invert `no_unexpected_files`: that is an allowlist over *newly created* files and needs a pre-run manifest |
296
298
  | a to-do workflow finished | `all_tasks_completed: true`, `task_status: {match, status}` |
297
299
  | a skill / connector / tool was **offered** | `skill_available`, `connector_available`, `tool_available` (all `<regex>`) |
298
300
  | a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
@@ -301,6 +303,7 @@ them by what you're trying to prove:
301
303
  | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
302
304
  | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
303
305
  | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
306
+ | the user was **shown** the right choices, in order | `question_options: {when_question, equals}` — the option SET/ORDER a gate offered (`question_asked` matches the text only); order is compared by default |
304
307
  | a hook blocked / didn't block a tool | `hook_blocked: <regex>`, `no_hook_blocked: true` (replay needs a `controlOut` cassette) |
305
308
  | every MCP round-trip succeeded | `no_mcp_error: true` (**live-only**) |
306
309
  | a context compaction happened | `compaction_occurred: true` |
@@ -393,10 +396,28 @@ the discovery/encode/record dance entirely and answer gates **live during the re
393
396
  `record --decider-dir`/`--decider-llm` (the cassette is flagged non-deterministic but replays deterministically).
394
397
  `run` takes no `--dry-run`: to check that a scenario **loads** without spending, use
395
398
  `cowork-harness record <file.yaml> --dry-run` — it runs the real loader AND the same scenario-level
396
- refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing), so it
397
- cannot green something a paid run would reject. On a directory it reports every offender and the batch
399
+ refusals the real `record` applies (`on_unanswered: prompt`, and an unsatisfiable assert pairing) **plus the
400
+ cassette-portability pre-flight below**, so it cannot green something a paid run would reject. On a directory it reports every offender and the batch
398
401
  cost estimate. `lint` checks the assertion invariants (both above).
399
402
 
403
+ **Decide WHERE the cassette lives before you record it — a cassette cannot be moved afterwards.**
404
+ Without `--out`, `record` writes `cassettes/<scenario-name-slug>.cassette.json` (gitignored by
405
+ default); pass `--out <path>` to put it somewhere tracked, e.g. `examples/replays/<name>.cassette.json`.
406
+ That choice is permanent: the cassette rewrites `scenario.session` and `scenarioSource` **relative to
407
+ its own directory** at record time, so moving the file later — a different `--out`, a `git mv`, a copy
408
+ into another repo — leaves those unresolvable and
409
+ `verify-cassettes` reports `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you
410
+ re-record at the new location. **`record` now says so BEFORE it spends:** a pre-flight — at the same
411
+ pre-spend point as the host-inventory refusal, and in `record --dry-run`, so the rehearsal is free —
412
+ warns when the cassette would be written outside the scenario's tree, or when `session:` itself lives
413
+ outside it (an absolute or `~` path: the mirror case, invisible to a check that only looks at where the
414
+ cassette lands). A warning, not a refusal — an out-of-tree throwaway cassette is legitimate; what was
415
+ missing was anything saying so while you could still act. Related: recording at a **host-inheriting** tier
416
+ (`protocol`/`hostloop`/`cowork`→hostloop) into a repo-visible path is refused outright (gotcha 25).
417
+ The clean answer there is `fidelity: container` (sealed, `HOME=/tmp`, nothing to leak) — **not**
418
+ redirecting `--out` outside the repo and moving the file in afterwards, which trades a loud refusal
419
+ for a permanently unverifiable cassette.
420
+
400
421
  **Author answers WITHOUT re-paying — the cheap loop.** You don't need a fresh paid record to discover a
401
422
  scenario's gates or their labels: `--keep` ONE run, then `cowork-harness trace <run-dir> --view questions`
402
423
  (and `verify-run`) read the gates + offered option labels out of that run's `events.jsonl` for free. Iterate
@@ -548,6 +569,42 @@ The full 17-code signal table (severity + per-signal opt-out) is in
548
569
  [`references/scenario-schema.md`](./references/scenario-schema.md); [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) (repo-only) carries
549
570
  the fuller narrative.
550
571
 
572
+ ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
573
+
574
+ A single green proves the run passed **once**. Two questions need more than that, and both have a
575
+ discipline that is cheap to follow and expensive to skip.
576
+
577
+ **"Did it pass, or pass once?"** → `--repeat N` (2-100, on `skill` AND `run`) samples the same
578
+ skill+prompt N times and prints a variance rollup instead of a single verdict. `--min-pass-rate` sets
579
+ the batch threshold, `--stop-on-diverge` stops the moment flakiness is proven, `--max-budget-usd` caps
580
+ spend.
581
+
582
+ **"Does the skill actually help?"** → `--ablate-skill` runs the prompt with every skill/plugin
583
+ discovery source removed, so the agent answers from its own priors. **It is ONE arm, not a paired
584
+ experiment**: this invocation is the control. Run the same prompt a second time *without* the flag for
585
+ the treatment arm and compare them yourself. Composed with `--repeat 5` it produces **5 ablated runs
586
+ and 0 treatment runs** — N samples of the control, which is the intended reading and is not an A/B.
587
+ Every ablated run is stamped `ablated: true` in `result.json`; a run that isn't stamped is a real run.
588
+ What the harness gives you here is the run execution and the control arm — designing the comparison
589
+ (scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
590
+
591
+ **Measurement hygiene — four things that silently invalidate a batch:**
592
+
593
+ 1. **Pin the model.** With no `model:` in the session (or `--model` on the `skill` lane) the run uses
594
+ whatever the staged agent binary defaults to — not a harness constant, and it can move under a
595
+ baseline bump. Read `result.json`'s `models` back before believing any cross-run comparison — and when
596
+ you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
597
+ fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
598
+ array purely by whether such a turn occurred.
599
+ 2. **Commit the skill first.** `fingerprint.skillHash` is content-exact, so an edit mid-batch silently
600
+ splits your dataset into two generations — and a hash whose source was never committed identifies a
601
+ generation that is unrecoverable. `stats --group-by skill-hash` separates them after the fact;
602
+ nothing recovers the source.
603
+ 3. **Check which arm you actually ran** before analysing anything: `ablated` and
604
+ `context.availableSkills` in each `result.json`.
605
+ 4. **Read `skillsInvoked`.** A rep where the skill never triggered is a measurement of the model, not
606
+ of your skill — discard or re-run it.
607
+
551
608
  ### Checking whether a background run is alive
552
609
 
553
610
  Never use `ps aux` to check on a `cowork-harness` run you launched in the background — it only sees
@@ -741,11 +798,13 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
741
798
 
742
799
  1. **An assertion passed but tested nothing on the PR gate.** *Why:* on a manifest-less cassette
743
800
  `replay` skips filesystem/egress keys (`file_exists`, `user_visible_artifact`, `artifact_json`,
744
- `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a *mixed* item like
801
+ `artifact_text`, `egress_*`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`); a
802
+ *mixed* item like
745
803
  `{result, egress_denied}` greens on `result` while its `egress_denied` half is dropped. (`record`
746
804
  snapshots an `artifacts` manifest, which makes
747
- `file_exists`/`user_visible_artifact`/`artifact_json`/`computer_links_resolve`
748
- replay-checkable — but the live-only egress keys stay skipped.) *Fix:* put egress/live-only checks on
805
+ `file_exists`/`user_visible_artifact`/`artifact_json`/`artifact_text`/`computer_links_resolve`
806
+ replay-checkable — but the live-only egress keys stay skipped, and `file_absent` is never
807
+ replay-checkable at all: proving absence needs an exhaustive, healthy walk a manifest does not record.) *Fix:* put egress/live-only checks on
749
808
  a live gate; keep one concern per `assert:` item; run the linter. The harness warns loudly on skip.
750
809
 
751
810
  2. **A steered gate answer never reached the model.** *Why:* `serializeDecision` must emit
@@ -754,8 +813,8 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
754
813
  channel: scripted `choose:` list, in-band `--decider-dir` via a repeated `--choose` / a JSON-array
755
814
  reply, and `--decider-cmd` via a JSON-array reply — all deliver the same `", "`-joined wire shape.
756
815
  Free-text "Other" via `answer:`. Do NOT hand-write a multiSelect reply as a bare comma-joined
757
- string — send an array; a scalar is treated as one selection.) `question_asked` / `questions_count_max` /
758
- `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
816
+ string — send an array; a scalar is treated as one selection.) `question_asked` / `question_options` /
817
+ `questions_count_max` / `gate_answers_delivered` only evaluate on replay **with a `controlOut` cassette** — re-record an
759
818
  old cassette or they're excluded (loudly), not vacuously passed. `gate_answers_delivered` *fails*
760
819
  on unobserved delivery (absence of evidence is failure, not neutral).
761
820
 
@@ -865,7 +924,9 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
865
924
  14. **A positional `choose` (`first` / index) is order-dependent.** `choose: "2"` survives label drift
866
925
  but NOT option *re-ordering* — if the gate presents its options in a different order run-to-run, the
867
926
  index lands on a different option (a silent re-record flake). Prefer an exact label when order is
868
- stable; `lint` flags positional `choose` with an advisory.
927
+ stable; `lint` flags positional `choose` with an advisory. Unstable option order is also what the
928
+ **user** sees — a reordered gate puts a different choice in the default slot — so pin what was shown
929
+ with `question_options`, rather than only hardening the answer rule against it.
869
930
  15. **A scripted `choose:` matching no offered option HARD-fails the run — `on_unanswered: first` does NOT
870
931
  backstop it.** This is distinct from an *unanswered* gate (no rule matched → falls to `on_unanswered`): a
871
932
  rule that DID match the gate but whose `choose:` names a label the gate never offered (the model reworded
@@ -970,6 +1031,26 @@ repeats the assertion/replay-relevant ones alongside the schema (a scoped subset
970
1031
  scanner's `host-inventory` class on an already-committed cassette. Passing one where the other command
971
1032
  wants it fails as an unrecognized flag — they don't interchange. Depth: `references/ci-recipe.md`.
972
1033
 
1034
+ 26. **A `skill`-lane `PASS` does not mean the skill ran, or that the run was the one you wanted.** *Why:*
1035
+ an open-ended `skill` run has no `assert:` block, so its verdict reports only that **no guard fired**
1036
+ (no error, stall, host-path leak, `outputs/` delete, permissive auto-allow or capability gap). On
1037
+ `run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
1038
+ means N runs cleared the guards — it says nothing about which model served them, whether the skill
1039
+ was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
1040
+ carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
1041
+ `models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
1042
+ exactly like skill output is not evidence: the skill's own source is mounted where the model can
1043
+ read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
1044
+ directly, with `skillActivity` empty.
1045
+
1046
+ 27. **`allow_stall: true` is a scenario assertion, so the `skill` lane cannot use it.** *Why:* the
1047
+ `stalled` guard fires when a run's final message ends in `?` with no productive tool call after the
1048
+ last gate — which includes a complete answer that closes by *offering* a follow-up ("want me to run
1049
+ this through a structured pass?"). The documented opt-out lives in an `assert:` block, and an
1050
+ open-ended `skill` run has none, so the failure message names a remedy that lane can't perform.
1051
+ *Fix:* on `skill`, read the final message before believing `stalled`, or move the check to a
1052
+ `run` scenario where `allow_stall: true` is authorable.
1053
+
973
1054
  For the assertion catalog, the YAML schema, the fidelity/answer tables, and the CI recipe, read the
974
1055
  files in `references/` (the gotchas above are the full list; the references repeat only the
975
1056
  assertion/replay-relevant ones).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`).
3
+ Self-contained reference. Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.22.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.24.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -32,7 +32,7 @@ jobs:
32
32
  - uses: actions/checkout@v4
33
33
  - name: Stage the agent binary (official channel, sha256-verified — see https://github.com/yaniv-golan/cowork-harness/blob/main/docs/maintenance.md)
34
34
  run: |
35
- V=2.1.227 # match your scenario's pinned baseline's agentVersion
35
+ V=2.1.229 # match your scenario's pinned baseline's agentVersion
36
36
  curl -fSL "https://downloads.claude.ai/claude-code-releases/$V/linux-arm64/claude" -o "$RUNNER_TEMP/claude-$V"
37
37
  chmod +x "$RUNNER_TEMP/claude-$V"
38
38
  # verify against the committed baseline's sha256 (baselines/desktop-*.json → agentBinary.sha256)
@@ -58,13 +58,17 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
58
58
  GitHub-hosted runners, no token/Docker/agent:
59
59
 
60
60
  ```yaml
61
- - run: npm i -g "cowork-harness@>=1.22.0"
61
+ - run: npm i -g "cowork-harness@>=1.24.0"
62
62
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
63
63
  # no silent false-greens. WITHOUT --strict this
64
64
  # step cannot fail on a WARN-class rule (e.g.
65
65
  # vacuous-gate-assert) — it would print the
66
66
  # finding and still exit 0. --min-severity WARN
67
- # keeps the advisory INFO class advisory.
67
+ # keeps the advisory INFO class advisory — pair them.
68
+ # Bare `lint --strict` fails on INFO too, which reds
69
+ # on scenarios that are perfectly fine. (`lint-skill
70
+ # --strict` never fails on INFO: same flag name, a
71
+ # different rule. Do not carry one over to the other.)
68
72
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness — FAILS on a stale recording
69
73
  # ALSO fails on a leaked host inventory: recording at
70
74
  # protocol/hostloop freezes YOUR machine's MCP servers,
@@ -114,16 +118,22 @@ The split is not just about tokens — it decides **where each lane can run**:
114
118
  `transcript_*`, `tool_*`, `subagent_*`, `dispatch_count_max`, `skill_triggered`, `no_skill_triggered`,
115
119
  `max_cost_usd`, `max_tokens`, `tool_calls_max`, `result`, and the verdict modifiers
116
120
  `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
117
- `allow_stall` (no-op passes); plus the gate keys `question_asked` / `questions_count_max` /
118
- `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
119
- (`file_exists` / `user_visible_artifact` / `artifact_json`) **if** it carries an artifact manifest.
121
+ `allow_stall` (no-op passes); plus the gate keys `question_asked` / `question_options` /
122
+ `questions_count_max` / `gate_answers_delivered` **if** the cassette has `controlOut`, and the manifest keys
123
+ (`file_exists` / `user_visible_artifact` / `artifact_json` / `artifact_text`) **if** it carries an artifact
124
+ manifest. `file_absent` is in neither class — it is live/verify-run only.
120
125
  **That list is illustrative, not the authoritative set** — more keys are replay-checkable than fit a
121
126
  paragraph, and a hand-typed enumeration is exactly what goes stale. For the current set, ask the CLI:
122
127
 
123
128
  ```bash
124
- cowork-harness assertions --list --output-format json # every key, with its replay class
129
+ cowork-harness assertions --list --output-format json # every key + its one-line semantics
125
130
  ```
126
131
 
132
+ It emits `{key, description}` — there is no structured replay-class field to filter on; the
133
+ live-only / manifest / `controlOut` preconditions are stated in each key's `description` prose and,
134
+ in full, in the catalog's replay-class tables in
135
+ [`references/scenario-schema.md`](./scenario-schema.md).
136
+
127
137
  **`replay --mutate`** is a distinct, reporting-only diagnostic on this same lane: it perturbs each
128
138
  recorded JSON artifact value one at a time, re-runs the assertions against the perturbed cassette,
129
139
  and reports which perturbations NOTHING caught — those are the fields your `assert:` block leaves
@@ -279,7 +289,7 @@ jobs:
279
289
  with: { node-version: '24' }
280
290
  - uses: actions/setup-python@v5
281
291
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
282
- - run: npm i -g "cowork-harness@>=1.22.0"
292
+ - run: npm i -g "cowork-harness@>=1.24.0"
283
293
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
284
294
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
285
295
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -308,7 +318,7 @@ jobs:
308
318
  echo "live=true" >> "$GITHUB_OUTPUT"
309
319
  fi
310
320
  - if: steps.guard.outputs.live == 'true'
311
- run: npm i -g "cowork-harness@>=1.22.0"
321
+ run: npm i -g "cowork-harness@>=1.24.0"
312
322
  - if: steps.guard.outputs.live == 'true'
313
323
  run: cowork-harness run scenarios/ --output-format json
314
324
  env:
@@ -335,6 +345,31 @@ scenario = `result === "success" && assertions.every(pass)`. Exit code is non-ze
335
345
  fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-ready without parsing
336
346
  JSON.
337
347
 
348
+ **Telling *why* a run failed, without scraping stderr.** Each result carries a `verdict` whose
349
+ `failures[]` is the one place every failure reason is enumerated, in one shape (the same object lands
350
+ in `result.json` and in the stdout envelope, by construction). Each entry is
351
+ `{kind, assertion?: "<key>", message}`, and **`kind` is the discriminator**:
352
+
353
+ - **`assertion`** — one of *your* `assert:` items failed; `assertion` names its key (`file_exists`,
354
+ `semantic_matches`, …).
355
+ - **`guard`** — a signal you didn't author: `stalled`, `permissive_auto_allow`, `missing_capability`, a
356
+ host-path leak, an `outputs/` delete, an infra or transport error, an unanswered gate.
357
+ - **`staleness`** — on `replay --strict` / `--assert-from` / `--reassert`, skill or baseline drift,
358
+ which those modes escalate to a hard failure on purpose (frozen events must not green an edited
359
+ assert against a skill whose current source produces something else).
360
+ - **`cassette-format`** — the cassette is too new for this build to interpret.
361
+ - **`coverage`** — a `verify-run` answer-coverage miss: a gate the run fired that your `answers:` block
362
+ does not cover.
363
+
364
+ So `jq '[.verdict.failures[] | select(.kind=="assertion")]'` answers "did MY assertions pass?" and
365
+ `select(.kind=="staleness")` answers "is the cassette stale?", from the envelope alone. Do not infer
366
+ either from the exit code: every kind lands on exit 1.
367
+
368
+ > **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
369
+ > existed and it never worked in both directions: `coverage` entries carry a key too (an internal
370
+ > `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
371
+ > `cassette-format` all arrive key-less and indistinguishable from one another.
372
+
338
373
  **For the commands in this recipe, stdout carries the machine envelope and nothing else.** Without
339
374
  `--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
340
375
  human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`).
3
+ Self-contained reference. Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -210,7 +210,7 @@ up often enough to spell out:
210
210
  - **A green `replay` proves "same as when recorded," not "correct today."** `replay` never touches
211
211
  a filesystem or network — it re-evaluates assertions from the frozen cassette. A fixed set of
212
212
  keys is live-only and **skipped outright** on replay (absent from `assertions[]`, not vacuously
213
- passed): `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
213
+ passed): `file_absent`, `no_delete_in_outputs`, `self_heal_ran`, `transcript_no_host_path`, `egress_denied`,
214
214
  `egress_allowed`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back`, and `expect_denied`.
215
215
  Everything else that *is* evaluated is checked against the **recording**, not fresh behavior — a
216
216
  green replay says the skill produced these events when it was recorded, not that it still does
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.22.0`
4
- (baseline `desktop-1.28929.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.24.0`
4
+ (baseline `desktop-1.32352.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -103,6 +103,13 @@ allow_host_writes: true # OPTIONAL — required to run `hostloop` fi
103
103
  Relative paths resolve from the file's own directory, so a scenario + session + referenced files
104
104
  form a relocatable bundle. `~` expands to home.
105
105
 
106
+ > **A recorded cassette is NOT part of that bundle and is NOT relocatable.** It rewrites its own
107
+ > references relative to **its own** directory at record time (`scenario.session` and
108
+ > `scenarioSource`), so moving it afterwards — a different `--out`, a
109
+ > `git mv`, a copy into another repo — leaves them unresolvable and `verify-cassettes` reports
110
+ > `unverifiable-skill` ("can't verify ⇒ not green", exit 3) until you re-record at the new location.
111
+ > Decide where a cassette will live *before* you record it.
112
+
106
113
  ## Session YAML
107
114
 
108
115
  A session (`sessions/*.yaml`) captures everything you'd configure in Cowork **before the first
@@ -111,7 +118,15 @@ session = your setup.**
111
118
 
112
119
  ```yaml
113
120
  # model & reasoning
114
- model: claude-opus-4-8 # omit for the agent default
121
+ model: claude-opus-4-8 # omit ONLY to track the agent default: with no `model:` the harness emits
122
+ # no --model flag at all, so the run uses whatever the STAGED AGENT BINARY
123
+ # defaults to — not a harness constant, and it can move under a baseline
124
+ # bump. PIN IT for anything you will compare (a --repeat batch, a
125
+ # before/after, a with/without); read result.json's `models` back to confirm,
126
+ # IGNORING any `<…>`-wrapped entry (`<synthetic>` = a turn the agent
127
+ # fabricated locally, not a model id).
128
+ # On the ad-hoc `skill` lane there is no session file, so `--model <id>`
129
+ # (or COWORK_HARNESS_MODEL) is the ONLY way to pin it.
115
130
  account_name: my-account # OPTIONAL — display name rendered into {{accountName}} / the prompt's
116
131
  # "User name:" line; NOT a credential/identity selector (see src/prompt.ts,
117
132
  # https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md)
@@ -282,9 +297,11 @@ same set live from the schema.
282
297
  | `user_visible_artifact: <path>` | exists **and** under a user-visible root (`outputs/` + each connected folder's mount name) — the right primitive for a workspace deliverable when a folder is connected |
283
298
  | `no_delete_in_outputs: true` | no delete op touched `mnt/outputs` — **only `true` is valid**; `false` is rejected (omit to allow deletes) |
284
299
  | `no_delete_in_mounts: true` | no delete op touched ANY delete-denied mount — `outputs` plus every `rw` connected folder — except those waived by `allow_delete_in`. Production denies `unlink`/`rmdir` on every such mount, so `no_delete_in_outputs` covers only part of the real rule. **only `true` is valid** |
285
- | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning |
300
+ | `no_unexpected_files: [<glob>, …]` | every **newly created** file under a user-visible root matches ≥1 glob (workRoot-relative paths; `**` = whole path segment for any depth — use `outputs/handoff/**` for per-run subdirs); `[]` = no new files; **new-files-only** — overwriting a pre-existing file in place is invisible (use content-level producer stamping); live/verify-run without a pre-run manifest ⇒ evidence-unavailable hard-fail (live runs capture the baseline only when this key is asserted; recordings always capture); an **incomplete post-run filesystem walk** (an unreadable subtree — a permission/I-O error) also fails evidence-unavailable rather than reporting "no strays" over a partial tree — distinct from the missing-manifest case (a `--resume` run), which fails for a different reason; captured on every live sandbox tier including microvm (its outputs are snapshotted from the VM into the run dir); replay needs `cassette.preRunPaths` (≥0.24 recordings) — cassettes without it **exclude** the key with a loud warning. ⚠️ **Not a stand-in for "file X must not exist":** there is no negative-existence key, and inverting this one is a different claim — it is new-files-only (a pre-existing file is invisible to it) and needs a pre-run manifest. Assert the content-level consequence instead (`tool_not_called`, or the absent `artifact_json`) rather than widening an allowlist for every incidental lock/temp file |
286
301
  | `input_unmodified: <glob>` or `[<glob>, …]` | a single glob or a list; every **pre-existing** file (incl. uploaded files under `uploads/**`) whose workRoot-relative path matches ≥1 glob keeps an unchanged content hash after the run — the in-place-mutation companion to `no_unexpected_files`'s new-files check (`[]` is rejected by the schema — list at least one glob); a glob that matches **no** pre-run path fails loud (a typo or renamed mount would otherwise verify zero files and pass vacuously); a matched file that was deleted counts as a content change (fails); live/verify-run without a pre-run hash manifest ⇒ evidence-unavailable hard-fail (a `--resume` run); captured on every live sandbox tier including microvm; replay needs `cassette.preRunHashes` — cassettes without it **exclude** the key with a loud warning; on replay it compares against the manifest's recorded `sha256`, never a re-hash of the materialized tree |
287
302
  | `self_heal_ran: <bool>` | a plugin-root self-heal script was (not) invoked |
303
+ | `file_absent: <path>` | the named path does **not** exist under the work root after the run — the direct negative-existence key. Do NOT invert `no_unexpected_files` for this: that is an allowlist over NEW files only, and it needs a pre-run manifest. **LIVE/verify-run only** (a cassette records no walk health, so absence is unprovable on replay); evidence-unavailable on `lane: remote` / `preRunOrigin: remote-unavailable` |
304
+ | `artifact_text: {artifact, contains?, not_contains?, matches?, not_matches?}` | assert over a delivered artifact's TEXT body — `artifact_json`'s companion for non-JSON files, and how you prove an internal name did not leak into a file the user receives. Literal path (no glob), so one entry per delivered surface. Manifest-class; body-less / symlinked / over-cap targets fail evidence-unavailable, and a non-UTF-8 body fails the NEGATIVE matchers rather than passing against bytes it never read |
288
305
  | `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
289
306
  | `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
290
307
  | `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
@@ -312,7 +329,7 @@ same set live from the schema.
312
329
  `skill_triggered`/`no_skill_triggered`, `skill_available`, `connector_available`, `skill_tool_used`,
313
330
  `subagent_type` are **regex** (unanchored, case-insensitive). So `tool_called: mcp__workspace__*` (glob) but
314
331
  `tool_available: mcp__workspace__.*` (regex) — a `.*` in a `tool_called` glob is a load-time schema error, not silently-matches-nothing.
315
- | `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary) |
332
+ | `skill_tool_used: {skill, tool}` | a tool whose name matches `tool` ran inside a skill-activation window whose `skillId` matches `skill` (`RunResult.skillActivity`) — evidence-unavailable if skill-activity telemetry is absent; heuristic for inline skills (a sticky, sequential window matching the agent's `activeSkill` scope, not an exact per-tool boundary). **Scope:** the window's tool counts **include sub-agent calls** made during it, so this key can't say which agent called (`subagent_tool_used` is the sub-agent-only claim), and it matches tool **names** only — never the path/args, so "did it read *this* file" isn't expressible (per-sub-agent reads are recorded at `subagents[].referencesRead`, readable but not assertable) |
316
333
  | `max_cost_usd: <N>` | the run's SDK-reported cost is ≤ N USD — evidence-unavailable if cost telemetry is absent. **Replay asserts the frozen recording's cost, not fresh spend** — a real regression needs a live `run` |
317
334
  | `max_tokens: <N>` | `usage.input_tokens + usage.output_tokens` ≤ N (cache tokens excluded) — same replay caveat as `max_cost_usd` |
318
335
  | `tool_calls_max: <N>` | total top-level tool calls (sum of `toolCounts`) ≤ N — meaningfully replay-checkable (re-drive recomputes `toolCounts` deterministically) |
@@ -327,7 +344,8 @@ same set live from the schema.
327
344
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
328
345
  | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives identically on replay (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: the agent's cwd there already *is* the outputs dir, so a write-to-cwd delivery is under a user-visible root from the start and `user_visible_artifact` passes. **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
329
346
  | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentedFiles` is non-empty) — the presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
330
- | `question_asked: <regex>` | the agent asked an AskUserQuestion whose text matches |
347
+ | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` |
348
+ | `question_options: {when_question?, equals?, contains?, order?}` | the option SET and ORDER a gate offered — what the user was actually SHOWN. `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable |
331
349
  | `questions_count_max: <N>` | at most N **sub-questions** asked — a bundled `AskUserQuestion` with K sub-questions counts as K, not 1; `trace --view questions`'s footer total uses the same definition |
332
350
  | `gate_answers_delivered: true` | every answered gate's answer reached the model (observed `tool_result`; unobserved = fail); **zero gates fired passes vacuously** — pair with `gate_answer_count_min: >= 1` to also require a gate, or drop this key and declare `questions_count_max: 0` in a scenario that expects none |
333
351
  | `gate_answers_delivered: false` | asserts at least one answered gate's answer was **confirmed not delivered** (an observed delivery failure); an unobserved/null delivery does **not** satisfy this — for negative-path delivery tests. Requires a gate, so **mutually exclusive** with `questions_count_max: 0` (refused by `run`/`skill`/`record`) |
@@ -340,7 +358,7 @@ same set live from the schema.
340
358
  | `allow_permissive_auto_allow: true` | verdict modifier — suppresses the default-fail when the run recorded a cowork-parity permissive auto-allow; for tests that deliberately assert Cowork's permissive behavior |
341
359
  | `allow_missing_capability: true` | verdict modifier — suppresses the default-fail when the (partial "core") agent image omits a capability the skill used but real Cowork ships (OCR/LibreOffice/markitdown/opencv/PDF-tables). Assert only when the skill's fallback is genuinely equivalent; otherwise rebuild full parity (`--build-arg COWORK_FULL_PARITY=1`). Also opts out of the `requires_capabilities` declared-need check. Live tiers only |
342
360
  | `allow_l0_plugin_divergence: true` | verdict modifier — opt into L0/protocol plugin divergence: suppresses the default-fail when a plugin behaves differently at `protocol` (L0) fidelity than under a sandboxed tier. Live tiers only |
343
- | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider) |
361
+ | `allow_stall: true` | verdict modifier — suppresses the `stalled` default-fail when a run ends on a question having done no productive tool work after its last gate (the agent asked for input and stopped — incl. re-asking in plain text after answering an `AskUserQuestion`); assert only when ending on a question is intended, else script the answer (`answer:` / `--answer` / a decider). **Scenario-only:** an open-ended `skill` run has no `assert:` block, so it cannot opt out — a helpful closing offer ("want me to run this through a structured pass?") fails `stalled` there with no suppressor. Read the final message before believing it, or move the check to a `run` scenario |
344
362
  | `allow_undelivered_deliverables: true` | verdict modifier — suppresses the `undelivered_deliverables` WARN. Working in the scratchpad is Cowork's designed pattern, so a skill that legitimately leaves intermediates, caches or downloaded inputs behind can say so instead of carrying permanent noise. The signal is warn-only and never fails a run on its own; reach for this when the scratch activity is intentional, not to silence a real delivery gap |
345
363
  | `allow_outputs_delete: true` | verdict modifier — accepts a detected outputs delete instead of failing the run, for a skill whose deletion is intended. Omitting `no_delete_in_outputs` does **not** permit deletes (a detected delete fails via the `outputs_delete` signal precisely because the key was not authored), so this is the way to accept one. **Mutually exclusive** with `no_delete_in_outputs`. Waives the harness's post-hoc detection; it does not model Cowork's `allow_cowork_file_delete` approval handshake |
346
364
  | `allow_delete_in: [<mount>…]` | verdict modifier — accepts detected deletes in the named mounts (the per-mount analogue of `allow_outputs_delete`, mirroring production's `fileDeleteApprovedMounts`). Waives the verdict only; detection still runs and the hits stay in `result.json`. Listing `outputs` conflicts with `no_delete_in_outputs` |
@@ -423,7 +441,7 @@ sourcing ≠ evaluation (replay warns when you edit one). `verify-run` is the on
423
441
  modifiers `allow_permissive_auto_allow` / `allow_missing_capability` / `allow_l0_plugin_divergence` /
424
442
  `allow_stall` are also kept on replay, evaluated as no-op passes.
425
443
 
426
- **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `questions_count_max`,
444
+ **Gate keys — replay only with a `controlOut` cassette:** `question_asked`, `question_options`, `questions_count_max`,
427
445
  `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`,
428
446
  `path_denied`, `no_path_denied` (the latter three are also `fidelity: hostloop`-only — see the assertion
429
447
  table). With `controlOut` present they evaluate; on an old
@@ -435,13 +453,16 @@ the question keys: a custom hook's block/allow decision is an opaque async reply
435
453
  built-in Task hook's view and could vacuously pass `no_hook_blocked` even if a custom hook genuinely blocked.
436
454
 
437
455
  **Filesystem — replay-checkable WITH an artifact manifest:** `file_exists`, `user_visible_artifact`,
438
- `artifact_json`, `computer_links_resolve` run on replay when the cassette carries an `artifacts` snapshot
456
+ `artifact_json`, `artifact_text`, `computer_links_resolve` (+ `computer_links_resolve_if_present`) run on
457
+ replay when the cassette carries an `artifacts` snapshot
439
458
  (`record` captures `outputs/` + connected folders; `replay` materializes it). `artifact_json` needs the
440
459
  small-file JSON `body` inlined; a hash-only entry still satisfies `file_exists`. `computer_links_resolve`
441
460
  resolves a `/sessions/…/mnt/…`-shaped link directly against the manifest, and a host-shaped (hostloop) link
442
461
  by first normalizing it to a mount-relative path (recorded connected-folder prefixes + the outputs/uploads
443
462
  mounts) — replay has no live filesystem to check a host path against directly (that only happens on a live
444
- `run`/`verify-run`). Without a manifest (older cassettes) all five are skipped (five need the manifest; two
463
+ `run`/`verify-run`). `artifact_text` is manifest-class for the same reason `artifact_json` is — a
464
+ body-less, symlinked or over-cap entry fails evidence-unavailable rather than passing.
465
+ Without a manifest (older cassettes) all six are skipped (six need the manifest; two
445
466
  more — `no_unexpected_files` and `input_unmodified` — need the pre-run path/hash capture, below); `no_unexpected_files` also
446
467
  needs `preRunPaths` (≥0.24 recordings) — without it the key is excluded with a loud warning (live/verify-run
447
468
  hard-fails evidence-unavailable instead). `input_unmodified` is the same shape but needs `preRunHashes`
@@ -451,7 +472,9 @@ warning. A green replay re-confirms
451
472
  staleness `fingerprint` shows ANY skill/baseline drift, or `replay --fail-on-skill-drift` only on
452
473
  skill-source drift; every replay result also reports it class-tagged in `staleness[]` for a JSON gate.
453
474
 
454
- **Egress + other filesystem — still skipped on replay (live-only):** `no_delete_in_outputs`,
475
+ **Egress + other filesystem — still skipped on replay (live-only):** `file_absent` (proving a path is
476
+ ABSENT needs an exhaustive, healthy walk; a manifest records no walk health, so "not captured" and "not
477
+ there" are indistinguishable and the key would pass while proving nothing), `no_delete_in_outputs`,
455
478
  `self_heal_ran`, `transcript_no_host_path`, `egress_*` / `expect_denied`, `no_mcp_error`, `max_peak_rss_bytes`,
456
479
  `no_lost_write_back`. These run only on a live `run`/`record`.
457
480
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 1.22.0` (baseline `desktop-1.28929.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 1.24.0` (baseline `desktop-1.32352.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -17,11 +17,12 @@ a paid live re-record? Walk this tree — the answer is usually no:
17
17
  `cowork-harness replay <cassette> --assert-from <scenario.yaml>`. Token-free, no re-record.
18
18
  If the recording genuinely lacks the telemetry a key needs (very old cassettes), the key fails
19
19
  **loud** as `evidence-unavailable` — that is correct behavior, not a bug; only then re-record.
20
- 2. **Gate keys** (`question_asked`, `questions_count_max`, `gate_answers_delivered`) on a cassette
20
+ 2. **Gate keys** (`question_asked`, `question_options`, `questions_count_max`, `gate_answers_delivered`) on a cassette
21
21
  **with `controlOut`** (any modern recording) → same token-free `--assert-from` path.
22
22
  3. **Gate keys** on a **pre-`controlOut`** cassette → one re-record unlocks gate asserts for that
23
23
  cassette permanently.
24
- 4. **Filesystem / egress keys** (`file_exists` without an artifact manifest, `egress_allowed`,
24
+ 4. **Filesystem / egress keys** (`file_exists` without an artifact manifest, `file_absent` — live-only
25
+ whatever the cassette carries — `egress_allowed`,
25
26
  `egress_denied`, `no_delete_in_outputs`) → live lane **by design**; replay skips them with a
26
27
  loud `::warning::`. Keep them in the scenario, run them on the nightly live gate.
27
28
 
@@ -69,6 +70,7 @@ Top-level fields of a `*.cassette.json` (schema [`schema/cassette.v11.json`](htt
69
70
  | `folderPrefixMap` | Optional even on v9+: the record-time connected-folder host-path → mount-name map. Replay's `computer_links_resolve` uses THIS (never the current session file); absent → the link is treated as evidence-unavailable, never reconstructed from the current session |
70
71
  | `timeline`, `timelineHeader` | The recorded per-event timeline (harness-observation timestamps for tool_use/tool_result/subagent_dispatch/thinking/decision/result, in total order) plus its header (`startedAtWall`/`startedAtMono` anchors); informational only — never affects the replay verdict. Absent on a cassette recorded before this field existed |
71
72
  | `environment` | Recording provenance: `location` (`"local"` on every cassette this harness produces), the resolved `tier`, `agentBinaryFormat`, and `harnessVersion` — the CLI that RECORDED the cassette (≥1.11.0). A harness-code change can shift recorded behaviour at an UNCHANGED baseline, which no staleness class keys off, so this is the only provenance for that class. **Absent** on pre-1.11.0 cassettes, and never backfilled — the absence is itself the signal |
73
+ | `environment.agentImage` | The `agentImage` sub-block records the rootfs image the recording ran against, stamped only for the tiers whose capabilities come from it (`container`, `hostloop`; `microvm` probes the Lima guest instead). `ref` is the resolved image ref, `configId` the local config id (built **and** pulled images, NOT comparable across machines), `registryDigest` the registry manifest digest (pulled only, and the only cross-machine-comparable identity). The image decides `missingCapabilityUse`, which `computeVerdict` fails on, so replaying against a different rootfs can move a verdict with nothing in the cassette having changed — `replay` compares `registryDigest` first and prints an advisory note, never a failure. Re-inspected at cassette-WRITE time, so a rebuild between run and record records the later image. **Absent** on cassettes recorded before this field existed, and never backfilled |
72
74
 
73
75
  ## Recipe 3 — Set up redaction BEFORE your first hostloop/protocol record
74
76
 
@@ -175,9 +177,16 @@ degrade the advice. It is real work to calibrate; these steps are the traps that
175
177
  from `RunResult.assertions[].semanticClaims`. Do **not** chase single-run all-pass — set `min_pass` to
176
178
  the reliably-hit core for a green verdict, and treat the per-claim rates as the real signal. (N=1
177
179
  routinely mislabels a stable 0/3 as "intermittent" and vice-versa.)
178
- 5. **Check discrimination — does the skill actually help?** Run one rep with the skill NOT installed (or
179
- inspect a not-invoked rep). If the answer still scores high, that claim is answerable from priors and
180
- tests the model, not your skill — strengthen it (a skill-specific fact) or drop it.
180
+ 5. **Check discrimination — does the skill actually help?** Run one rep with the skill NOT installed and
181
+ compare. `--ablate-skill` is the flag for it: it empties every skill/plugin discovery source for
182
+ **that one invocation**, so the agent answers from its own priors, and stamps the result
183
+ `ablated: true`. It is the **control arm only** — run the same prompt again *without* the flag for
184
+ the treatment arm; `--ablate-skill --repeat 5` gives you 5 control runs and 0 treatment runs, not an
185
+ A/B. (Inspecting an organically not-invoked rep works too, but is not a substitute: that rep may
186
+ differ for other reasons.) If the answer still scores high without the skill, that claim is
187
+ answerable from priors and tests the model, not your skill — strengthen it (a skill-specific fact) or
188
+ drop it. Everything past "run both arms" — scrubbing giveaways, shuffling, judging blind, unblinding
189
+ after grading — is yours to build; the harness supplies the runs and the control.
181
190
  6. **Gate a change on the profile diff.** Capture the per-claim profile before your edit (the baseline),
182
191
  make the edit, re-capture, and compare per claim: a claim that DROPPED (e.g. 3/3 → 0/3) is a regression
183
192
  your edit caused; a claim already at 0/3 (a known gap) cannot regress. That turns "did my SKILL.md
@@ -228,7 +237,15 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
228
237
  - `cowork-harness inspect <run-dir>` → what the run produced, plus the run's `label` and `skillHash`.
229
238
  - In-run alternative: dispatch a checker **sub-agent** (maker/checker) whose result folds into the
230
239
  verdict.
231
- 2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
240
+ 2. **Commit the skill before a measurement batch, and pin the model.** Two cheap disciplines that are
241
+ unrecoverable if skipped. `skillHash` is content-exact, so an edit mid-batch silently splits the
242
+ dataset into two generations — `stats --group-by skill-hash` separates them afterwards, but a hash
243
+ whose source was never committed names a generation that is unrecoverable, which makes the
244
+ comparison uninterpretable rather than merely noisy. And with no `model:` in the session (or
245
+ `--model` on the `skill` lane) each run uses whatever the staged agent binary defaults to, so a
246
+ before/after can silently straddle two models; read `result.json`'s `models` back to confirm —
247
+ ignoring any `<…>`-wrapped entry (`<synthetic>` marks a turn the agent fabricated locally, not a model).
248
+ 3. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
232
249
  `result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
233
250
  content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
234
251
  no `skillHash` on the `skill` lane at all, so verify the field is present before pairing on it** (a run that mounts nothing
@@ -10,6 +10,7 @@
10
10
  "allow_stall",
11
11
  "allow_undelivered_deliverables",
12
12
  "artifact_json",
13
+ "artifact_text",
13
14
  "compaction_occurred",
14
15
  "computer_links_resolve",
15
16
  "computer_links_resolve_if_present",
@@ -17,6 +18,7 @@
17
18
  "dispatch_count_max",
18
19
  "egress_allowed",
19
20
  "egress_denied",
21
+ "file_absent",
20
22
  "file_exists",
21
23
  "gate_answer_count_min",
22
24
  "gate_answers_delivered",
@@ -41,6 +43,7 @@
41
43
  "path_denied",
42
44
  "present_files_called",
43
45
  "question_asked",
46
+ "question_options",
44
47
  "questions_count_max",
45
48
  "replay_protocol_fidelity",
46
49
  "result",