cowork-harness 4.1.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (93) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +14 -10
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +4 -4
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +2 -2
  4. package/.claude/skills/cowork-harness/references/authoring.md +65 -14
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
  6. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  7. package/.claude/skills/cowork-harness/references/debugging.md +9 -3
  8. package/.claude/skills/cowork-harness/references/eval.md +79 -0
  9. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  10. package/.claude/skills/cowork-harness/references/gotchas.md +19 -5
  11. package/.claude/skills/cowork-harness/references/measurement.md +10 -2
  12. package/.claude/skills/cowork-harness/references/run-record-replay.md +14 -7
  13. package/.claude/skills/cowork-harness/references/scenario-schema.md +16 -4
  14. package/.claude/skills/cowork-harness/references/task-recipes.md +32 -13
  15. package/.claude/skills/cowork-harness/scripts/scenario.py +56 -1
  16. package/CHANGELOG.md +381 -0
  17. package/DESIGN.md +2 -2
  18. package/README.md +10 -5
  19. package/SPEC.md +25 -13
  20. package/baselines/desktop-2.16120.0.json +1148 -0
  21. package/baselines/prompts/cowork-system-prompt-fingerprints.json +10 -1
  22. package/dist/agent/session.js +26 -0
  23. package/dist/assert.js +102 -13
  24. package/dist/baseline.js +39 -1
  25. package/dist/cli.js +156 -373
  26. package/dist/critique/command.js +19 -9
  27. package/dist/critique/skill-invocation.js +61 -1
  28. package/dist/decide/decider.js +25 -3
  29. package/dist/decide/semantic-judge.js +170 -37
  30. package/dist/decide/usage.js +52 -0
  31. package/dist/errors.js +41 -0
  32. package/dist/eval/classify.js +308 -0
  33. package/dist/eval/command.js +591 -0
  34. package/dist/eval/invocation.js +44 -0
  35. package/dist/eval/job-runner.js +50 -0
  36. package/dist/eval/manifest.js +11 -0
  37. package/dist/eval/pins.js +45 -0
  38. package/dist/eval/report.js +479 -0
  39. package/dist/eval/runs.js +127 -0
  40. package/dist/eval/schedule.js +31 -0
  41. package/dist/eval/snapshot.js +328 -0
  42. package/dist/eval/stats.js +240 -0
  43. package/dist/eval/usage.js +62 -0
  44. package/dist/hillclimb/schema-check.js +657 -0
  45. package/dist/run/api-retries.js +31 -0
  46. package/dist/run/artifacts.js +5 -4
  47. package/dist/run/authored-capture-opts.js +23 -0
  48. package/dist/run/cassette.js +211 -40
  49. package/dist/run/chat-result.js +9 -1
  50. package/dist/run/chat.js +125 -63
  51. package/dist/run/command-globals.js +15 -2
  52. package/dist/run/doctor.js +62 -45
  53. package/dist/run/execute.js +280 -80
  54. package/dist/run/lint-load.js +5 -2
  55. package/dist/run/model-provenance.js +50 -3
  56. package/dist/run/provenance.js +29 -11
  57. package/dist/run/renderer.js +15 -1
  58. package/dist/run/run-index.js +6 -0
  59. package/dist/run/run.js +8 -0
  60. package/dist/run/runs-gc.js +26 -1
  61. package/dist/run/verify-context.js +403 -0
  62. package/dist/runtime/agent-tree.js +480 -0
  63. package/dist/runtime/hostloop.js +6 -5
  64. package/dist/runtime/protocol.js +6 -1
  65. package/dist/scan.js +20 -0
  66. package/dist/session.js +5 -0
  67. package/dist/sync/cowork-sync.js +266 -5
  68. package/dist/termination.js +76 -10
  69. package/dist/types.js +2 -2
  70. package/docs/README.md +2 -1
  71. package/docs/boundary.md +7 -0
  72. package/docs/cassette.md +13 -6
  73. package/docs/chat.md +7 -1
  74. package/docs/ci.md +25 -1
  75. package/docs/cli.md +25 -14
  76. package/docs/companion-skill.md +2 -2
  77. package/docs/critique.md +2 -1
  78. package/docs/debugging.md +12 -4
  79. package/docs/eval.md +244 -0
  80. package/docs/fidelity-gaps.md +78 -9
  81. package/docs/run-status.md +7 -3
  82. package/docs/scenario.md +15 -12
  83. package/docs/session.md +1 -1
  84. package/docs/stats.md +15 -3
  85. package/examples/replays/README.md +1 -1
  86. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  87. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  88. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  89. package/llms.txt +3 -2
  90. package/package.json +1 -1
  91. package/python/test_scenario_lint.py +71 -0
  92. package/schema/run-result.json +77 -1
  93. package/schema/scenario.schema.json +2 -2
@@ -1,10 +1,10 @@
1
1
  ---
2
2
  name: cowork-harness
3
- description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
3
+ description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did the change make its answers worse? (`eval`: paired, interleaved A/B of two plugin versions, pinned models). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.1.0
7
- tracks-harness: cowork-harness 4.1.0 (baseline desktop-2.9939.4)
6
+ version: 4.2.0
7
+ tracks-harness: cowork-harness 4.2.0 (baseline desktop-2.16120.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.1.0` (baseline
30
- > `desktop-2.9939.4`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.2.0` (baseline
30
+ > `desktop-2.16120.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
33
33
  ## Preflight — make sure the harness can actually run
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.1.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.1.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.1.0"`. **Pin `@^4.1.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.2.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.2.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.2.0"`. **Pin `@^4.2.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -76,8 +76,10 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
76
76
  grades against a different artifact, `critique-evidence-package.txt`, which none of these tools
77
77
  surface; see `references/critique.md`.
78
78
  - **Regression-test your skill's ANSWER quality** (not just its behavior — does its guidance still lead to
79
- correct answers after you edit it?) → author `semantic_matches` scenarios and gate on the per-claim
80
- profile. See **Recipe 5** in `references/task-recipes.md` (validity, N≥3, discrimination — the traps).
79
+ correct answers after you edit it?) → author `semantic_matches` scenarios, then compare the version
80
+ before your edit with the one after using `cowork-harness eval` (EXPERIMENTAL, live: 10 runs per
81
+ scenario at the defaults). See **Recipe 5** step 6 in `references/task-recipes.md` (validity,
82
+ discrimination — the traps) and [`references/eval.md`](references/eval.md).
81
83
  - **"What is WRONG with this skill?"** (a graded critique, not a pass/fail) → `cowork-harness critique
82
84
  <folder> --prompt "<probe>"`. Up to four model workloads (zero with `--corpus-only`; pass 2 is skipped with no self-report) and 10–20 minutes; budget from
83
85
  `report.costUsd.totalUsd`. Reach for it when you want **findings**. **For "what does this skill
@@ -97,7 +99,7 @@ CI-grade scenario, and the post-hoc debug loop; the rest are narrower tools that
97
99
 
98
100
  Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · migrate-run-dir · lint ·
99
101
  lint-skill · analyze-skill · probe-dispatch ·
100
- verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
102
+ verify-run · trace · inspect · diff · critique · eval · eval report · stats · decide · gates · answer · scaffold · assertions --list · sync ·
101
103
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
102
104
 
103
105
  ## Invariants — how a green run lies
@@ -107,7 +109,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
107
109
 
108
110
  1. **`result: success` is not "the task completed".** It means the agent didn't error. Assert the
109
111
  deliverable (`file_exists` / `artifact_json` / `transcript_matches`). A `skill`-lane `PASS` only means
110
- no guard fired: read `skillsInvoked`, `models` and `ablated` before concluding anything from it.
112
+ no guard fired: read `skillsInvoked` (plus `slashInvokedSkills` — a `/<skill>` prompt runs the skill
113
+ with no `Skill` call), `models` and `ablated` before concluding anything from it.
111
114
  2. **`replay` skips live-only keys.** Filesystem and egress keys are skipped on replay (loudly), so a
112
115
  mixed item like `{result, egress_denied}` greens on its content half. Keep one concern per `assert:`
113
116
  item, put live-only checks on a live gate, and run `cowork-harness lint`.
@@ -154,4 +157,5 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
154
157
  | [`references/fidelity-and-answers.md`](references/fidelity-and-answers.md) | tier semantics, answer paths, the determinism contract |
155
158
  | [`references/ci-recipe.md`](references/ci-recipe.md) | the GitHub Action, replay-vs-live lanes, the four-stage pipeline |
156
159
  | [`references/critique.md`](references/critique.md) | `critique` report and evidence-package shapes |
160
+ | [`references/eval.md`](references/eval.md) | `eval`: paired before/after of two plugin versions — labels, refusals, exit codes, files |
157
161
  | `scripts/scenario.py` | `scaffold`, `lint`, `lint-skill`, `resolve-agent-types <plugin-dir>` (validates a pinned `subagent_type` against `plugin.json` + `agents/*.md`) |
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -55,8 +55,8 @@ same set live from the schema.
55
55
  | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
56
56
  | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut. **Covers what the run dispatches** (`Agent`/`Task`, including `Agent(subagent_type:"fork")`), **not a `context: fork` skill** invoked through the `Skill` tool: that skill's own answer is never a dispatch — it comes back as the `Skill` tool result, which agent 2.1.284 builds as `Skill "<name>" completed (forked execution).`, a `Result:` line, then the answer. Assert on it with `tool_result_matches` anchored on that prefix, e.g. `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` (`[\s\S]*` because `.` stops at a newline; the match is case-insensitive, has no multiline flag so `^` is the start of the result, and sees the first 10 KB of each result). This covers a **foreground** fork only: a backgrounded fork's result is the line `Skill "<name>" launched (forked execution, running in the background).`, which carries no answer |
57
57
  | `dispatch_count_max: <N>` | at most N sub-agents dispatched — your author-chosen budget under Cowork's agent-side fan-out cap (concurrent 20 / per-session 200, inherited by the harness); records only, does not itself enforce — see gotcha 12 in `scenario-schema.md` |
58
- | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked via the `Skill` tool — evidence-unavailable (not a normal fail) if the agent's init tools have no `Skill` tool |
59
- | `no_skill_triggered: <regex>` | no invoked skill id matched — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent or the `Skill` tool is unobservable |
58
+ | `skill_triggered: <regex>` | a skill matching the regex (invoked id, e.g. `"plugin:skill"`) was invoked — via the `Skill` tool, or by a prompt starting `/<skill> …` / `/<plugin>:<skill> …` (the agent expands that itself with no `Skill` call; recorded as `slashInvokedSkills`) — evidence-unavailable (not a normal fail) if neither matched and the agent's init tools have no `Skill` tool, or the leading `/name` can't be resolved (ambiguous bare name, no skill inventory) |
59
+ | `no_skill_triggered: <regex>` | no invoked skill id matched, counting a slash-command invocation as well as a `Skill` call — the negative-control / description-collision catcher; evidence-unavailable (never a vacuous pass) if invocation data is absent, the `Skill` tool is unobservable, or the prompt's leading `/name` can't be resolved |
60
60
  | `skill_available: <regex>` | a staged skill's id matched the regex (offered, not necessarily invoked — see `skill_triggered` for invocation) — content-class: the id list comes from the agent's init `skills` listing, so it replays from the frozen init event (id-only; the `whenToUse` enrichment is live-disk and thus absent on replay, but the id is what's matched); evidence-unavailable only if `RunResult.context.availableSkills` is absent entirely (an older cassette recorded before the available-skills listing was captured) |
61
61
  | `connector_available: <regex>` | an MCP server/connector's name matched the regex (available, not necessarily used) — evidence-unavailable if `RunResult.context.mcpServers` is absent |
62
62
  | `tool_available: <regex>` | a tool in the init manifest matched the regex (available, not necessarily called — see `tool_called` for invocation) — evidence-unavailable if `RunResult.context.tools` is absent. The `mcp__skills__*`/`mcp__plugins__*` discovery tools are modeled (as `alwaysLoad`) on `container`/`hostloop`/`cowork` — a miss there is a real absence; `microvm`/`protocol` declare no such server, so a miss on those two tiers means "not modeled at this tier", not "provably unavailable" |
@@ -107,7 +107,7 @@ same set live from the schema.
107
107
  | `egress_allowed: <host>` | the host was allowed through |
108
108
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
109
109
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
110
- | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, evidence_files?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). **A `context: fork` skill's own answer is not graded either**, even with `include_subagent_text: true`: it is not a dispatch (so it has no `subagents[]` entry to fold in), and it reaches the main agent as the `Skill` tool result, which the judged document excludes. The judge sees it only if the main agent restates it; to check the fork's answer directly, use `tool_result_matches` (see `subagent_output_contains`). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass}]`, so a consumer can diff the per-claim profile across runs); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message |
110
+ | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, evidence_files?}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer — the **union of the agent's final result text (`RunResult.finalMessage`), the transcript, and the final on-disk content of any files the agent authored during the run** — so a claim about content the skill led the agent to *write to a file* grades as reliably as one about inlined prose. **"The transcript" is narrower than it reads: top-level `assistant_text` ONLY.** It excludes every `tool_use`/`tool_result` and **all sub-agent-originated text** (including fork-scoped `Skill`/`Agent(fork)` dispatches — the harness attributes their *tool* calls to the main agent, but not their text). **A `context: fork` skill's own answer is not graded either**, even with `include_subagent_text: true`: it is not a dispatch (so it has no `subagents[]` entry to fold in), and it reaches the main agent as the `Skill` tool result, which the judged document excludes. The judge sees it only if the main agent restates it; to check the fork's answer directly, use `tool_result_matches` (see `subagent_output_contains`). ⚠️ **Consequence: a rubric claim about whether a tool was called can NEVER grade true** — the evidence is not in the judged document. Such a claim looks reasonable and silently caps your pass rate; assert tool use with `tool_called` / `present_files_called` / `subagent_dispatched` / `hook_blocked` instead. Sub-agent text is captured in `RunResult.subagents[].reasoning` and reaches the judge only via opt-in `include_subagent_text: true` (`kind:"text"` turns only — sub-agent *thinking* arrives empty with `redacted:true`, so it would pad the document with blanks) (authored-file evidence is captured on every live sandbox tier including **microvm** — its session tree is snapshotted from the VM into the run dir). When the authored-file evidence backing the judged document is **incomplete** — a file dropped at the capture-size cap, unreadable at read-back, or (on `--resume`) the scratchpad walk skipped — the assert fails evidence-unavailable rather than trusting a judge grade over a partial document; this is separate from the malformed-grade `judgeInvalid` path below. The assert passes iff ≥ `min_pass` claims pass (default: all — avoid for a gating scenario). Results align by claim index and are recorded per-claim in `RunResult.assertions[].semanticClaims` (`[{index, claim, pass, rationale?}]`, so a consumer can diff the per-claim profile across runs; `rationale` is the judge's one-sentence reason, printed under each failed claim in the failure footer: untrusted model text that can quote the judged document, whose content never affects `pass`, absent when the judge gave none (a reply whose shape is broken, such as unparseable JSON, a malformed `{"results": …}` group beside a valid grade, or a partial restatement that contradicts it, is retried once and then marked `judgeInvalid`), and comparable only between runs that share `judgePromptHash`); a rep whose grade can't be parsed (after one retry) is marked `RunResult.assertions[].judgeInvalid` and **never silently dropped** — it is excluded from the pass denominator, and the guard against a misleading score from that exclusion is the gate's minimum-valid-rep floor (`MIN_VALID` ≥ 4) plus this visibility, not a claim that denominator-shrinking inflation is impossible. Within a rep, a grade that's still unparseable after the retry **fails that assert outright** (evidence-unavailable, not a vacuous pass) — a persistently-flaky judge reds the run rather than silently passing. `judge_model` pins the grader (default when neither it nor `COWORK_HARNESS_JUDGE_MODEL` is set: `claude-opus-4-8`; a dated id keeps a before/after comparison reproducible — or let `eval` compare two skill versions per claim, with the judge pinned). Each graded assert records the judge's provenance: `RunResult.assertions[].judgeModel` (the resolved model), `judgeCostUsd` (judge spend over both attempts, reported beside `cost.usd` and never inside it; absent when unpriced) and `judgePromptHash` (the grading-prompt template — compare only runs that share it). Live-only: the judge is a live model call, so evidence-unavailable / skipped-loud on replay (never a vacuous pass) **`evidence_files: [globs]` scopes which authored files are graded** — reach for it the moment a run authors more than a couple of files. The capture budget (64 KiB total by default) is spent prefix-major then alphabetically, so a pipeline that stages intermediates (`outputs/_work/*.json`) exhausts it before reaching its own deliverable and the verdict is refused evidence-unavailable over files no rubric mentions. Scoping also makes the capture spend the budget on the named files FIRST and exempts them from the per-file cap. Paths are `<root>/<rel>` (`outputs/report.md`, never a bare `report.md`; session-root writes are `scratchpad/<rel>`); globs are `*`/`?`/`**`, not regex. A glob matching nothing FAILS and the message lists every authored path — read it rather than guessing. Still too big? Raise `$COWORK_HARNESS_AUTHORED_TOTAL_BYTES`. The typed reason is on `RunResult.assertions[].semanticEvidence` — check `.reason` instead of parsing the message |
111
111
  | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
112
112
  | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
113
113
  | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -43,7 +43,7 @@ them by what you're trying to prove:
43
43
  | a skill actually **ran** (or must NOT) | `skill_triggered: <regex>`, `no_skill_triggered: <regex>` |
44
44
  | a tool ran **inside** a skill's scope | `skill_tool_used: {skill, tool}` |
45
45
  | a sub-agent did the work | `subagent_output_contains: {contains}`, `subagent_dispatched: <regex>`, `dispatch_count_max: <N>` |
46
- | a `context: fork` skill answered correctly | `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` — its answer is the `Skill` tool result, not a sub-agent output, so `subagent_output_contains` and `semantic_matches` never see it (foreground fork only — a backgrounded fork's result carries no answer) |
46
+ | a `context: fork` skill answered correctly | `tool_result_matches: '^Skill "[^"]*" completed \(forked execution\)[\s\S]*<pattern>'` — its answer is the `Skill` tool result, not a sub-agent output, so `subagent_output_contains` and `semantic_matches` never see it (foreground fork only — a backgrounded fork's result carries no answer). Only when the MODEL invokes the skill: a `/<skill> …` prompt runs the fork with no `Skill` call and no such result — use `skill_triggered` + `transcript_matches` there |
47
47
  | a pre-existing input wasn't mutated (incl. `uploads/**`) | `input_unmodified: <glob>` or `[<glob>, …]` (live/verify-run) |
48
48
  | no authored interactive artifact silently loses its Submit under Cowork | `no_lost_write_back: true` (**live-only**; static Tier A over the run's authored `.html`/`.py`/`.js`; per-scenario gate for the same class `analyze-skill` scans) |
49
49
  | a resource ceiling held | `max_peak_rss_bytes: <N>` (**live-only**) |
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -220,25 +220,76 @@ so a file `run`/`record` would refuse fails lint too (running `scenario.py lint`
220
220
  cowork-harness lint scenarios/*.yaml
221
221
  ```
222
222
 
223
- `lint` flags: filesystem/egress-only assertions on a `replay` gate (silent no-op), bad regex
224
- quoting, an egress assert on `protocol` fidelity, `transcript_no_host_path` on `hostloop`/`protocol`
225
- (ERROR — fails by design at those tiers; WARN on `fidelity: cowork`, whose tier resolves per the
226
- baseline's host-loop gate), non-empty `requires_capabilities` on `protocol` without
227
- `allow_missing_capability` (ERROR — the capability probe can't run there, so the run hard-fails as
228
- unverifiable), `no_scratchpad_leak` off `container` (ERROR on `protocol`/`microvm`/`hostloop` — hostloop's
229
- `present_files` passes a validated path through without promoting, so there is no scratch→outputs copy
230
- to leak; WARN on `cowork`, whose tier resolves per the baseline gate) or `present_files_called` on
231
- `protocol`/`microvm` (ERROR — served only at `container`/`hostloop`), or `present_files_called`/`no_scratchpad_leak`/`user_visible_artifact` on `lane: remote` (ERROR — the runtime rejects those at scenario load time, so the tier rules are suppressed there), a `controlOut`-gated key on a non-`controlOut` replay, mixed-class assertion items,
232
- and hallucinated schema (`assertions:` vs `assert:`, unknown keys). Exit code is non-zero on errors
233
- (CI-friendly). `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never
234
- emits a scenario `lint` would reject.
223
+ `lint` exits non-zero on any ERROR (CI-friendly); `--strict` also fails on WARN. Every rule it reports:
224
+
225
+ | Rule | Severity | Fires on |
226
+ |---|---|---|
227
+ | `assert-contradiction` | ERROR | assert items no single run can satisfy together |
228
+ | `assertions-key` | ERROR | `assertions:` instead of `assert:` — none of the checks would run |
229
+ | `authored-replay-fidelity` | ERROR | an authored `replay_protocol_fidelity` (only the replay lane synthesizes it) |
230
+ | `capabilities-on-protocol` | ERROR | non-empty `requires_capabilities` on `protocol` without `allow_missing_capability` — the probe cannot run there, so the run fails as unverifiable |
231
+ | `cassette-evidence-skipped` | INFO | with `--cassette-dir`: a cassette (or the directory) could not be read, so it cannot quiet replay-evidence advice |
232
+ | `container-only-key-off-container` | ERROR | `no_scratchpad_leak` off `container` — hostloop's `present_files` never promotes, so there is nothing to leak (WARN on `cowork`, whose tier resolves per the baseline gate) |
233
+ | `egress-on-protocol` | ERROR | an egress assertion (`egress_*` / `expect_denied`) on `protocol`, which enforces no egress |
234
+ | `enum-value-invalid` | ERROR | a field value outside its allowed set |
235
+ | `fidelity-missing` | ERROR | no `fidelity:` (required since 4.0.0) |
236
+ | `file-absent-contradiction` | ERROR | one path under both `file_exists` and `file_absent` |
237
+ | `gate-needs-controlout` | INFO | gate assertions, which evaluate on replay only when the cassette has `controlOut` |
238
+ | `host-path-assert-cowork` | WARN | `transcript_no_host_path` on `cowork` — it fails by design if the tier resolves to hostloop |
239
+ | `host-path-assert-tier` | ERROR | `transcript_no_host_path` on `hostloop` / `protocol`, where it fails by design |
240
+ | `lane-remote-incompatible-key` | ERROR | `present_files_called` / `no_scratchpad_leak` / `user_visible_artifact` on `lane: remote` (the runtime rejects them at load, so the tier rules are suppressed there) |
241
+ | `linter-extra-findings-invalid` | ERROR | the loader findings `cowork-harness lint` hands the linter could not be read |
242
+ | `linter-unclassified-key` | ERROR | a valid assertion key this linter cannot classify (the linter is out of date) |
243
+ | `manifest-needs-snapshot` | INFO | manifest-backed keys, which evaluate on replay only when the cassette carries an `artifacts` manifest |
244
+ | `mixed-assert-item` | WARN | one assert item mixing replay-checkable and live-only keys (replay drops the live-only half) |
245
+ | `no-scenarios` | ERROR | a linted directory with no `*.yaml` / `*.yml` |
246
+ | `not-found` | ERROR | a named file that does not exist |
247
+ | `parse` | ERROR | a file that is not YAML, or not a mapping |
248
+ | `positional-choose-order` | INFO | an answer rule with a positional `choose` (first / index), which option re-ordering can move |
249
+ | `present-files-key-off-tier` | ERROR | `present_files_called` on `protocol` / `microvm` (served only at `container` / `hostloop`) |
250
+ | `prompt-slash-not-leading` | WARN | a `prompt:` that names `/<skill>` without starting with it, so it is never expanded |
251
+ | `reference-access-contradiction` | ERROR | one reference under both `reference_read` and `no_observed_reference_access` |
252
+ | `regex-double-quoted` | WARN | a double-quoted regex with an unescaped backslash (YAML strips it) |
253
+ | `replay-noop` | WARN | every assertion is live-only or a verdict modifier, so a replay gate verifies nothing |
254
+ | `slash-prompt-forked-result-anchor` | WARN | a `prompt:` starting with `/<skill>` plus a `tool_result_*` anchored on `forked execution` — a slash-invoked skill makes no `Skill` call, so that tool result never exists; assert `skill_triggered` instead |
255
+ | `tool-called-always-passes` | INFO | `tool_called` with `count: {min: 0}` and no `max` — it asserts nothing |
256
+ | `tool-input-regex-redactable` | WARN | a `tool_not_called` input literal the redaction policy rewrites in the committed cassette (or a policy pattern it cannot check offline) |
257
+ | `tool-input-shell-tier` | INFO | the object form with `tool: Bash` and a `command` on `hostloop` / `cowork`, where shell runs as `mcp__workspace__bash` — list both |
258
+ | `tool-not-called-tier-vacuous` | WARN | `tool_not_called` / `subagent_tool_absent` naming a tool the tier never serves |
259
+ | `transcript-command-shaped` | WARN | a `transcript_*` value shaped like a shell command — those keys read prose only, never a tool call |
260
+ | `unknown-assert-key` | WARN | an assertion key not in the catalog (the loader rejects it) |
261
+ | `unknown-top-key` | WARN | a scenario key not in the schema |
262
+ | `vacuous-gate-assert` | WARN | `gate_answers_delivered` with no presence companion (zero gates passes it), or inert beside `questions_count_max: 0` |
263
+ | `scenario-invalid` | ERROR | the harness's scenario loader refuses the file (via `cowork-harness lint` only — see below) |
264
+ | `baseline-unknown` | ERROR | `baseline:` names no baseline this CLI ships (via `cowork-harness lint` only) |
265
+ | `lint-loader-internal` | ERROR | the wrapper could not run its loader check on a file — a harness bug; it never falls back to a lint that skipped the loader |
266
+
267
+ `scaffold` auto-upgrades the tier if you ask for egress on `protocol`, so it never emits a scenario `lint`
268
+ would reject.
269
+
270
+ **Lint the skill itself: `cowork-harness lint-skill <skill-dir>`.** It checks the skill, not a scenario:
271
+ Cowork host-loop footguns (`${CLAUDE_PLUGIN_ROOT}` in a VM bash step, hook events, a misplaced
272
+ `hooks.json`, an unresolvable `subagent_type`), the evidence corpus a `critique` can package
273
+ (`references/critique.md`), and two size caps. `skill-body-over-reattach-cap` (WARN) fires when the
274
+ `SKILL.md` body, frontmatter excluded, passes 19,000 B — after a compaction the agent re-attaches only the
275
+ start of an invoked skill — and `skill-body-near-reattach-cap` (INFO) from 80% of that;
276
+ `skill-reference-over-read-cap` (WARN) fires on a `references/**.md` over 60,000 B, past which a
277
+ whole-file Read returns a partial view. `--strict` fails on WARN, never on INFO. To accept a reviewed
278
+ judgement-call finding, pass `--ignore-rule <rule>[=<glob>]` (repeatable; the glob matches the finding's
279
+ file) or fence the text in `SKILL.md` with `<!-- lint-skill: ignore-start <rule>[,<rule>…]: <reason> -->`
280
+ … `<!-- lint-skill: ignore-end -->` (outside any code fence). A suppressed finding is still printed; it
281
+ stops gating. A provable rule (an ERROR, a misplaced `hooks.json`, a missing pinned agent) cannot be
282
+ suppressed: naming it, or an unknown rule, in `--ignore-rule` is a usage error (exit 2); in a marker it is
283
+ WARN `lint-skill-ignore-invalid`, as is any other malformed marker. An unclosed marker is WARN
284
+ `lint-skill-ignore-unclosed`, and one that suppresses nothing is INFO `lint-skill-ignore-unused`.
235
285
 
236
286
  **`cowork-harness lint` runs the loader: a file it calls clean is one `run`/`record` will load.** Anything
237
287
  the loader refuses — an unknown key, a wrong value type (a scalar `semantic_matches.rubric`), a bad regex,
238
288
  a reserved value — is ✗ ERROR `scenario-invalid` (exit 1, with or without `--strict`), and a `baseline:`
239
289
  naming no baseline this installed CLI ships is ✗ ERROR `baseline-unknown` (`latest` always resolves). It
240
290
  does not check what depends on the machine the run happens on (the session file and its mounts, an
241
- absolute `baseline:` path, environment variables). A session or matrix YAML in a linted directory is not
291
+ absolute `baseline:` path that does not exist here — one that exists is checked (4.1.1 and later) —
292
+ environment variables). A session or matrix YAML in a linted directory is not
242
293
  a scenario and is reported as one that does not load — keep those out of the linted set. `python3
243
294
  scenario.py lint` run directly stays offline and lenient: there an unknown key is only a ⚠ WARN (exit 0).
244
295
  `cowork-harness record <file.yaml> --dry-run` also runs the loader and adds the pre-spend refusals (exit 2
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`).
3
+ Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "4.1.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "4.2.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.1.0"
85
+ - run: npm i -g "cowork-harness@^4.2.0"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -241,7 +241,9 @@ cowork-harness replay cassettes/ # replay every *.cassette.jso
241
241
  Re-record whenever the protocol or your scenario's expected content changes. An old cassette without
242
242
  `controlOut` excludes the gate keys (with a loud warning) — re-record to enable them. `record` **refuses
243
243
  to freeze a failing live run** into a cassette (pass `--allow-failing` to override) — a committed red
244
- cassette is a latent false-signal.
244
+ cassette is a latent false-signal. An `--allow-failing` recording of a red run exits 0 with `ok: true`:
245
+ `record`'s `ok` means a cassette was written, and the run's own verdict is `results[0].verdict.pass`
246
+ (`items[].verdict` on `record <dir/>`).
245
247
 
246
248
  ## Privacy: cassettes are committed fixtures → record only against SYNTHETIC inputs
247
249
 
@@ -256,7 +258,9 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
256
258
  dir** (first file found per dir; env vars merge on top). `cowork-harness init-redact` copies the
257
259
  packaged reference template (local-path prefixes, incl. macOS temp roots and slugged home segments
258
260
  like `-Users-<user>-…`, + a generic email regex) into the cwd as a starting point — review and tailor
259
- it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` or add them. Redaction is **verdict-preserving** — `record` refuses to write if
261
+ it; a copy from an earlier release lacks the temp-root rules, so re-copy with `--force` (without it
262
+ `init-redact` refuses to overwrite an existing policy; with it your tailoring is replaced — save and
263
+ re-apply it) or add them by hand. Redaction is **verdict-preserving** — `record` refuses to write if
260
264
  redaction would flip an assertion (a manufactured green). `--no-redact` skips it for known-synthetic
261
265
  inputs.
262
266
  - **Pre-spawn preflight**: `record` warns (`::warning::`, before the paid run starts — once per batch
@@ -342,8 +346,8 @@ A typical skill repo runs four stages, fastest/cheapest first:
342
346
  This is the shape a CI step wants: **silent on success (no output, exit 0), loud and specific on
343
347
  failure** — `--quiet` suppresses the readiness preview but never the `✗ broken:` lines, which name the
344
348
  offending file *and* the rejected key, one line per file, and the step still exits 1. It exits 0,
345
- though, on an input the real record would refuse (a missing path, an unknown baseline name, a
346
- tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
349
+ though, on an input the real record would refuse (a `session:` file that cannot be read — 4.1.1 and
350
+ later — a missing path, an unknown baseline name, a tier-vacuous `tool_not_called`): that prints a `⚠ input error:` line and lands in `inputErrors[]`. To
347
351
  gate on those too, use the JSON form (4.1.0 and later):
348
352
 
349
353
  ```bash
@@ -391,7 +395,7 @@ jobs:
391
395
  with: { node-version: '24' }
392
396
  - uses: actions/setup-python@v5
393
397
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
394
- - run: npm i -g "cowork-harness@^4.1.0"
398
+ - run: npm i -g "cowork-harness@^4.2.0"
395
399
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
396
400
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
397
401
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -420,7 +424,7 @@ jobs:
420
424
  echo "live=true" >> "$GITHUB_OUTPUT"
421
425
  fi
422
426
  - if: steps.guard.outputs.live == 'true'
423
- run: npm i -g "cowork-harness@^4.1.0"
427
+ run: npm i -g "cowork-harness@^4.2.0"
424
428
  - if: steps.guard.outputs.live == 'true'
425
429
  run: cowork-harness run scenarios/ --output-format json
426
430
  env:
@@ -443,7 +447,8 @@ sandbox).
443
447
 
444
448
  `--output-format json` emits a machine envelope on stdout (human output goes to stderr):
445
449
  `{tool, version, command, ok, results[], error}` — one `RunResult` per scenario. **Overall pass for a
446
- scenario is `verdict.pass`** (envelope-wide: `ok`), and it is strictly stronger than
450
+ scenario is `verdict.pass`** (envelope-wide: `ok`; on `record`, `ok` only means the command exited 0 —
451
+ see below), and it is strictly stronger than
447
452
  `result === "success" && assertions.every(pass)`: the verdict also carries ~20 signal codes that fail a run
448
453
  with no failing assertion at all — `stalled`, `outputs_delete`, `mount_delete`, `host_path_leak`,
449
454
  `undelivered_deliverables`, `missing_capability`, `permissive_auto_allow`, `ended_with_question`,
@@ -477,6 +482,11 @@ Keep the `.results[]?` hop and the `?` operators. `.results[0]` silently ignores
477
482
  the first when you pass a directory, and a bare `.verdict` does not exist at the envelope root at all —
478
483
  both read as "no failures" against a run that failed.
479
484
 
485
+ **`record` is shaped differently.** Its `ok` is the exit code's verdict (a cassette was written), not the
486
+ run's: `record --allow-failing` on a red run is `ok: true`. The verdict is in `results[0].verdict` on
487
+ `record <file>` and in `items[].verdict` on `record <dir/>` and `--rerecord-stale`, which carry no
488
+ `results[]` — so `.results[]?` reads nothing there. Use `[.items[]? | .verdict.failures[]?]` on a batch.
489
+
480
490
  > **Do not filter on whether `assertion` is present.** That was the only discriminator before `kind`
481
491
  > existed and it never worked in both directions: `coverage` entries carry a key too (an internal
482
492
  > `answer_coverage` marker), so they read as authored asserts, while `guard`, `staleness` and
@@ -486,7 +496,9 @@ both read as "no failures" against a run that failed.
486
496
  `--output-format json`, `run` / `record` / `replay` / `verify-cassettes` / `status` write their whole
487
497
  human rendering — warnings, verdict, `status`'s summary line — to **stderr**, and stdout stays empty.
488
498
  A wrapper that captures only stdout gets an empty log and, if it greps that for a state, a silent false
489
- negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout.
499
+ negative. Capture stderr for the human trail (`2> run.stderr.log`), or ask for JSON and parse stdout —
500
+ `COWORK_HARNESS_OUTPUT_FORMAT=json` makes JSON the default for every command that takes `--output-format`
501
+ (an explicit flag still wins).
490
502
  (Commands whose whole job is to print a value — `--version`, `assertions --list`, `scaffold`, `gates`,
491
503
  `skill --dry-run` — write it to stdout by design. Under `--output-format json` the value rides inside the
492
504
  envelope: `scaffold`'s YAML is `.scenario`, `skill --dry-run`'s preview is the envelope's own fields;
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
@@ -37,7 +37,11 @@ already does.
37
37
  **microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
38
38
  stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
39
39
  `cowork-harness vm status` — a `provisioning` other than `ready` confirms it — and if a run does not
40
- recover it on its own, `cowork-harness vm delete` and retry.
40
+ recover it on its own, `cowork-harness vm delete` and retry. A run on such a VM can also have cached an
41
+ empty toolchain for it in the capability probe's cache: `vm delete` drops that VM's entry (`vm prune`
42
+ forgets only the orphaned VMs it deletes, never the current one); otherwise delete `capability-cache.json`
43
+ from the runs root (`~/.cowork-harness/runs/` unless
44
+ `COWORK_HARNESS_RUNS_DIR` is set) so the next run probes again.
41
45
 
42
46
  **Is it your skill's bug, or a known harness gap?** Before deep-debugging a wrong behavior, rule out a
43
47
  **deliberate fidelity gap** — the harness intentionally does *not* reproduce a few real-Cowork behaviors,
@@ -80,7 +84,9 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
80
84
  walk observe this run — what distinguishes "nothing was left undelivered" from "cannot tell"), `cost` (`cost.usd` = the SDK's
81
85
  `total_cost_usd` for the run — the authoritative single-run spend of the agent session, which leaves out the `semantic_matches` judge and the LLM decider calls; NOT the same source as summing
82
86
  `modelUsage[].costUSD`, which is what `trace --view usage` reports, so the two can differ),
83
- `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`, `toolErrors`,
87
+ `usage` (`input_tokens`/`output_tokens`/`turns`), `toolDurations` (with `toolDurationsBasis`), `models`,
88
+ `toolCalls` (every tool call in stream order: `name`, top-level `input` fields each capped at 10 KB, and
89
+ `origin` `main`/`subagent`/`unknown` — what the object form of `tool_called`/`tool_not_called` reads), `toolErrors`,
84
90
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
85
91
  `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
86
92
  `context` (tools/mcpServers/availableSkills), `tasks`,
@@ -0,0 +1,79 @@
1
+ # `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
2
+
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full guide is
4
+ [docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
5
+ need while running it.
6
+
7
+ ```bash
8
+ cowork-harness eval <scenario.yaml | dir/> --arm before=git:HEAD:plugins/my-skill --arm after=./plugins/my-skill \
9
+ --model <concrete id> --judge-model <concrete id> [--holdout <scenario.yaml>]... [--fail-on possible|confirmed]
10
+ cowork-harness eval report <eval-dir> # rebuild the report from the eval dir, no spend
11
+ ```
12
+
13
+ ## What it does
14
+
15
+ - Two arms: a plugin directory, or `git:<ref>:<path>` (read from the commit, `<path>` relative to the repo
16
+ root). The FIRST arm is the baseline; a drop is the second arm passing less often.
17
+ - Only the session's single `plugins.local_plugins` entry is substituted. Each arm is copied once, before
18
+ the first run, and every rep mounts the copy.
19
+ - The plugin is mounted at the `local_plugins` path (`mnt/.local-plugins/marketplaces/<marketplace>/<plugin>`),
20
+ not the `remote_plugins` path a UI-installed plugin has (`mnt/.remote-plugins/plugin_<id>`). A skill that
21
+ locates its own files at runtime sees a different path under each, so the comparison holds for the
22
+ `local_plugins` layout only.
23
+ - scenarios × 2 × `--reps` live runs (10 per scenario at the default `--reps 5`), interleaved, plus one judge
24
+ call per `semantic_matches` assert per run. The start-up line prints the job count. No budget flag.
25
+ - In a scenario directory, YAML with no `prompt:` (a session file) is skipped.
26
+ - Run an A/A first (`--allow-identical-arms`, the same source twice) to see your scenarios' noise.
27
+
28
+ ## Reading the labels
29
+
30
+ | Label | Meaning |
31
+ |---|---|
32
+ | `confirmed drop/rise` | significant after the correction (bh q = 0.10, or holm) |
33
+ | `possible drop/rise` | p ≤ `--alpha` (0.05), not confirmed |
34
+ | `no detectable change` | with the smallest change this n could have detected (MDD) |
35
+ | `underpowered` | no outcome at these sizes could reach `--alpha` — NOT "no change" |
36
+ | `insufficient` | too few valid reps in an arm (4 of 5 needed by default) |
37
+
38
+ A drop is a signal to investigate, not proof: open the run dirs the report links for that row. In order,
39
+ first match wins: an infrastructure failure is excluded and reported — including a rep where no model
40
+ answered (the agent's `Not logged in` / `Authentication required` reply, rule `auth`; a usage or
41
+ spend limit as its final message, even on a nonzero exit after spend, rule `usage_limit`; or only
42
+ `<synthetic>` models at $0, rule `no_model_answered`). An agent-caused failure (timeout, max turns,
43
+ unanswered question, crash) then fails every row of its rep, even with its pin unknown. Only after that
44
+ are a pin the agent did not honour (false, or unknown on a rep that completed) and a snapshot that changed
45
+ excluded and reported. A loud UNCLASSIFIED count means a termination the classifier does not know — read
46
+ those runs. Per scenario: if EVERY rep of both arms errored, or every rep of one arm is infrastructure,
47
+ that scenario compared nothing — its rows are `insufficient` and the eval exits 1. If one arm's every rep
48
+ is the agent's own failure and the other arm ran, the reps are scored (a real drop) — exit 0 unless
49
+ `--fail-on`. Either way the header
50
+ names the arm, scenario, dominant error and a matching hint (`Every rep of arm <label> in <scenario> errored — …`).
51
+
52
+ ## Refused before any run (exit 2)
53
+
54
+ - a model or judge that is an alias (`opus`, `best`) rather than a concrete id;
55
+ - an eval dir inside any git work tree (the snapshots would mount empty);
56
+ - identical arms (unless `--allow-identical-arms`);
57
+ - an arm that contains the eval's own scenario or session files (by location, copy or symlink), an
58
+ `evals.json`, or a symlink resolving outside it;
59
+ - a scenario input a run would refuse (a missing path, a `tool_not_called` the tier can never violate);
60
+ - `--fail-on confirmed` when no row could reach `confirmed` at this `--reps`;
61
+ - no usable agent credential for a scenario's tier — the same check as `doctor --tier <tier>`'s `token` row,
62
+ with its fix. A Keychain login or a `.credentials.json` in the config dir, without an env/.env token,
63
+ passes only at `protocol`;
64
+ - a session whose plugin is declared only under `plugins.remote_plugins` (the session must declare exactly
65
+ one `local_plugins` entry). Workaround: eval a copy of the session that declares the same directory under
66
+ `local_plugins`, and check the `remote_plugins` path handling with an ordinary `run`.
67
+
68
+ ## Exit codes
69
+
70
+ `0` completed — no drop fails the eval unless you pass `--fail-on`. `1` a drop at the `--fail-on` level,
71
+ every row `insufficient`, a scenario that compared nothing, or the judge model differed across reps (an A/A run under `--fail-on possible`
72
+ can exit 1 on noise). `2` usage or a refusal. `3` an arm snapshot could not be copied or staged.
73
+
74
+ ## Files
75
+
76
+ `<eval-dir>/` (default `~/.cowork-harness/evals/<eval-id>/`): `manifest.json`, `arms/`, `runs.jsonl`,
77
+ `report.json` (every rep with its bucket), `report.md`. The runs are ordinary run dirs labelled
78
+ `eval:<eval-id>:<arm>`; a bare `prune` keeps 5 per scenario and says which evals it trimmed — `eval report`
79
+ still works, but the evidence links then dangle.
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`).
3
+ Self-contained reference. Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -175,9 +175,10 @@ authorable). Reach for this list when debugging a run's behavior, that one while
175
175
  `--reassert`). That opt-in path is safe by construction for the authored fields — it **hard-fails** if
176
176
  `prompt`/`answers`/`baseline`/`fidelity`/`lane`/`skills`/`requires_capabilities` or the skill content (when a
177
177
  fingerprint exists) drifted from the recording (re-record then), and `expect_denied`/filesystem/egress keys
178
- are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the `session`
179
- (model / data mounts / discovery) is NOT drift-checked or fingerprinted, so a **model change** between record
180
- and re-assert is undetected — the notice flags this; re-record if the session changed. `verify-run` reads
178
+ are sourced but stay **live-only** (it warns; they don't move the replay verdict). **Caveat:** the session's
179
+ `model:` IS in the cassette's `sessionFingerprint` (a `--model`/env model is not; `environment.model`
180
+ records what ran). `verify-cassettes` reports a changed session as staleness (exit 1); `replay` never
181
+ checks it — plain, `--strict` or `--assert-from` — so re-record if the session changed. `verify-run` reads
181
182
  on-disk `assert:` against a kept *run dir*; `replay --assert-from` is the equivalent for a *cassette*.
182
183
 
183
184
  18. **`questions_count_max` counts sub-questions, not gates.** One `AskUserQuestion` tool call can
@@ -234,6 +235,17 @@ authorable). Reach for this list when debugging a run's behavior, that one while
234
235
  fail your gate. *Fix:* nothing, unless the scenario asserts `tool_available` on
235
236
  `mcp__skills__*`/`mcp__plugins__*`; then re-record. It stays silent at `microvm`/`protocol`, where
236
237
  re-recording would never produce those tools anyway.
238
+ A sibling **`agent-version:` note** means the agent version the cassette's own `system/init` event
239
+ reports differs from the one the baseline its `fingerprint.baseline` names pins for that tier: the
240
+ `agentVersion` at `container`/`microvm`, the native agent in `agentBinary.nativeStagedPath` at
241
+ `hostloop`. It never appears at `protocol`, which runs the unpinned `claude` on your `PATH`. *Why:* one
242
+ of a fingerprint re-stamped by hand across an agent bump, a recording made under
243
+ `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1`, an explicit binary override (`COWORK_AGENT_BINARY`, or
244
+ `COWORK_HOST_AGENT_BINARY` at `hostloop`), or at `hostloop` the default patch-bump substitution of the
245
+ native agent; the note lists that tier's causes and does not pick one. It is non-gating too:
246
+ `verify-cassettes` puts it in the result's `notes[]`, and `replay` prints one
247
+ `::notice:: [replay] <file> — … [agent-version]` line per cassette on stderr (also under
248
+ `--output-format json`). *Fix:* re-record against the pinned agent.
237
249
 
238
250
  24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
239
251
  lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
@@ -270,7 +282,9 @@ authorable). Reach for this list when debugging a run's behavior, that one while
270
282
  `run` the same word additionally means *your assertions held*; on `skill --repeat N`, `PASS — N/N`
271
283
  means N runs cleared the guards — it says nothing about which model served them, whether the skill
272
284
  was invoked, or whether they were the ablated arm. *Fix:* read the three fields the record already
273
- carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all),
285
+ carries before drawing any conclusion — `skillsInvoked` / `skillActivity` (was it invoked at all;
286
+ a `/<skill> …` prompt runs the skill with NO `Skill` call, so read `slashInvokedSkills` too — and
287
+ `models` is then just `["<synthetic>"]`, with the real model only in `modelUsage`),
274
288
  `models` (which model), `ablated` + `context.availableSkills` (which arm). An answer that reads
275
289
  exactly like skill output is not evidence: the skill's own source is mounted where the model can
276
290
  read it — in production too — so on a self-referential prompt it may read `SKILL.md` and answer
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.1.0` (baseline `desktop-2.9939.4`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.2.0` (baseline `desktop-2.16120.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -25,6 +25,11 @@ Every ablated run is stamped `ablated: true` in `result.json` and carries `ablat
25
25
  What the harness gives you here is the run execution and the control arm — designing the comparison
26
26
  (scrubbing giveaways, shuffling, judging blind, unblinding only after grading) is still yours.
27
27
 
28
+ **"Did my edit change it?"** → `cowork-harness eval` (EXPERIMENTAL): the version before your edit and the
29
+ one after, interleaved, with the agent and judge models pinned, compared per assertion and per rubric
30
+ claim with an exact test. It is a regression signal to investigate, not proof — see
31
+ [`eval.md`](eval.md) and Recipe 5 step 6 in [`task-recipes.md`](task-recipes.md).
32
+
28
33
  ### Tool timing — what `toolDurations` measures
29
34
 
30
35
  `result.json`'s `toolDurations` and `trace <run> --view tool-durations` report, per tool, the **wall gap
@@ -40,7 +45,10 @@ Compare timings between runs of the same tier and model only.
40
45
 
41
46
  1. **Pin the model in the session.** A run that resolves no model is refused, but one pinned only by
42
47
  `COWORK_HARNESS_MODEL` takes its model from the machine, so two shells can run two models. Set
43
- `model:` in the session (or pass the same `--model` on every `skill` run). Read `result.json`'s `models` back before believing any cross-run comparison — and when
48
+ `model:` in the session (or pass the same `--model` on every `skill` run). Adding `model:` to a session
49
+ that already has cassettes re-stales them (`verify-cassettes` exits 1 — the model is in the session
50
+ fingerprint); to pin without re-recording now, use `--model` or `COWORK_HARNESS_MODEL` and move the
51
+ pin into the session at the next re-record. Read `result.json`'s `models` back before believing any cross-run comparison — and when
44
52
  you do, **ignore any entry wrapped in angle brackets**: `<synthetic>` is the agent marking a turn it
45
53
  fabricated locally (no API call), not a model, so two runs of the same pinned model can differ on this
46
54
  array purely by whether such a turn occurred.