cowork-harness 4.4.0 → 4.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +8 -7
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +3 -3
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
  4. package/.claude/skills/cowork-harness/references/authoring.md +14 -10
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +21 -7
  6. package/.claude/skills/cowork-harness/references/critique.md +4 -2
  7. package/.claude/skills/cowork-harness/references/debugging.md +29 -11
  8. package/.claude/skills/cowork-harness/references/eval.md +4 -1
  9. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +7 -3
  10. package/.claude/skills/cowork-harness/references/gotchas.md +3 -3
  11. package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +37 -6
  12. package/.claude/skills/cowork-harness/references/hillclimb.md +17 -8
  13. package/.claude/skills/cowork-harness/references/measurement.md +2 -2
  14. package/.claude/skills/cowork-harness/references/run-record-replay.md +38 -7
  15. package/.claude/skills/cowork-harness/references/scenario-schema.md +2 -2
  16. package/.claude/skills/cowork-harness/references/semantic-judging.md +1 -1
  17. package/.claude/skills/cowork-harness/references/task-recipes.md +123 -1
  18. package/.claude/skills/cowork-harness/scripts/scenario.py +1 -1
  19. package/CHANGELOG.md +158 -0
  20. package/DESIGN.md +6 -6
  21. package/README.md +10 -14
  22. package/RELEASING.md +17 -2
  23. package/baselines/desktop-2.19675.1.json +1458 -0
  24. package/baselines/desktop-2.26454.0.json +1492 -0
  25. package/dist/cli.js +10 -2
  26. package/dist/decide/llm-transport.js +32 -9
  27. package/dist/hostloop/workspace-handler.js +3 -1
  28. package/dist/run/chat.js +3 -0
  29. package/dist/run/doctor.js +52 -3
  30. package/dist/run/execute.js +11 -4
  31. package/dist/run/lane-notice.js +72 -0
  32. package/dist/run/verdict.js +1 -1
  33. package/dist/runtime/hostloop.js +2 -2
  34. package/dist/session.js +15 -2
  35. package/dist/sync/baseline-diff.js +62 -2
  36. package/dist/sync/cowork-sync.js +203 -6
  37. package/dist/sync/remote-devices.js +430 -0
  38. package/dist/types.js +22 -4
  39. package/docs/ci.md +2 -2
  40. package/docs/cli.md +6 -5
  41. package/docs/companion-skill.md +2 -2
  42. package/docs/fidelity-gaps.md +73 -32
  43. package/docs/invariants.md +1 -0
  44. package/docs/maintenance.md +32 -8
  45. package/docs/plugin-root.md +4 -0
  46. package/docs/scenario.md +137 -16
  47. package/docs/session.md +5 -2
  48. package/examples/replays/README.md +1 -1
  49. package/examples/replays/example-multiselect-gate.cassette.json +1 -1
  50. package/examples/replays/example-pdf-skill.cassette.json +1 -1
  51. package/examples/replays/hostloop-computer-links.cassette.json +1 -1
  52. package/package.json +1 -1
  53. package/schema/run-result.json +1 -1
  54. package/schema/scenario.schema.json +3 -3
  55. package/scripts/check-versions.ts +27 -2
  56. package/scripts/release-preflight.ts +5 -3
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.4.0
7
- tracks-harness: cowork-harness 4.4.0 (baseline desktop-2.19675.0)
6
+ version: 4.5.0
7
+ tracks-harness: cowork-harness 4.5.0 (baseline desktop-2.26454.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,8 +26,8 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.4.0` (baseline
30
- > `desktop-2.19675.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.5.0` (baseline
30
+ > `desktop-2.26454.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
33
33
  ## Preflight — make sure the harness can actually run
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.4.0"`. **Pin `@^4.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.5.0"`. **Pin `@^4.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -130,7 +130,8 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
130
130
  recording the cassette that locks it.
131
131
  7. **The tier decides what exists.** `protocol` has no sandbox and no egress, tool names differ per tier
132
132
  (`container` serves `mcp__workspace__web_fetch`, not `WebFetch`), and every tier models Cowork's
133
- desktop-local lane only.
133
+ local lane only; new Pro and Max tasks do not use it from 2026-10-06. Behaviour-shaped results transfer;
134
+ path, mount, delivery and egress results do not.
134
135
  8. **A WARN signal never blocks a green.** Read the verdict signals after every run
135
136
  (`prompt_asset_missing`, `undelivered_deliverables`, `model_fallback`, …).
136
137
 
@@ -156,7 +157,7 @@ behind each, is [`references/gotchas.md`](references/gotchas.md).
156
157
  | [`references/measurement.md`](references/measurement.md) | `--repeat`, `--ablate-skill`, measurement hygiene |
157
158
  | [`references/debugging.md`](references/debugging.md) | triage, `result.json` fields and `trace` views, `chat` |
158
159
  | [`references/gotchas.md`](references/gotchas.md) | the full "✓ passed ≠ correct" landmine catalog |
159
- | [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality |
160
+ | [`references/task-recipes.md`](references/task-recipes.md) | start here for "how do I X": evolve `assert:`, audit tier drift, redaction, budgets, answer quality, and goals with no flag (force a compaction, ablate a section, a form reply, resume in a new conversation, hook JSON decisions, schema checks, unattended runs) |
160
161
  | [`references/assertion-catalog.md`](references/assertion-catalog.md) | every `assert:` key's semantics, the verdict-signal table |
161
162
  | [`references/semantic-judging.md`](references/semantic-judging.md) | `semantic_matches` in full: what the judge reads, fork results, refusal reasons, provenance |
162
163
  | [`references/scenario-schema.md`](references/scenario-schema.md) | every YAML field, which keys survive `replay`, the `web_fetch` model |
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -79,7 +79,7 @@ same set live from the schema.
79
79
  | `all_tasks_completed: true` | every task in `RunResult.tasks[]` reached status `"completed"` — **requires ≥1 task** (a zero-task run fails; assert `task_count_min` for presence); evidence-unavailable if tasks telemetry is absent |
80
80
  | `task_count_min: <N>` | at least N tasks were created (`RunResult.tasks.length >= N`) — the presence companion for task assertions |
81
81
  | `task_status: {match, status}` | a task whose `subject` OR `id` matches the regex `match` reached `status` — evidence-unavailable if tasks telemetry is absent; also fails **malformed** when a TaskCreate result was unparseable (corrupt task telemetry), mirroring the guard `all_tasks_completed`/`task_count_min` already had |
82
- | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the desktop-local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
82
+ | `no_scratchpad_leak: true` | every file presented via `present_files` that was in the scratchpad was successfully promoted to `mnt/outputs` (none left behind) — vacuously passes if nothing was presented (pair with a presence check to require a delivery); content-class: both the `present_files` tool_use and its own tool_result live in the ordinary events stream, so `RunResult.presentedFiles` re-derives on replay at container, where the agent's cwd IS the session root the live lane measures from; at hostloop the live lane measures from the session root while a re-drive has only the recorded cwd (`mnt/outputs` before Desktop 2.7032.0, an empty macOS system dir from it), so the promoted/leaked booleans there are not equivalent — immaterial to this key, which evaluates at container only (meaningfully replay-checkable, same as `skill_triggered`); evidence-unavailable if `presentedFiles` telemetry is absent (an older run predating the feature). **Container-only on the merits**: hostloop serves `present_files` but never promotes (its handler passes a validated path through), so there is no scratch→outputs copy that could leak. On microvm/protocol a scratchpad-delivered file is neither promoted to `mnt/outputs` nor detected there (a skill that delivers via write-to-cwd→`present_files` will false-red `user_visible_artifact` on those tiers; use `container`, or write directly to `outputs/`). Hostloop is not one of them: at hostloop a delivered file under the outputs dir is visible there immediately, so `user_visible_artifact` passes (before Desktop 2.7032.0 a write-to-cwd landed there; from it the agent runs at `/var/empty` and a relative write is refused). **The tool name is lane-specific:** `present_files` is the local lane's tool (the one this harness emulates); remote Cowork delivers via the agent-native `SendUserFile` instead, so a skill should describe the delivery outcome rather than naming either tool — this key asserts the harness-side delivery record either way. **Only `true` is valid** |
83
83
  | `present_files_called: true` | at least one file was actually delivered via the `present_files` tool (`RunResult.presentFilesCalls > 0` — the count of invocations that carried a well-formed `file_path`). Presence is read at the **invocation**, deliberately not off the classified `presentedFiles` list: that list drops any path it cannot resolve, and a host-path redaction policy makes that certain at `hostloop`, where a real host path redacts to `[REDACTED:…]/mnt/outputs/f`. So this key is safe to assert in a scenario you intend to `record` with redaction active — `record` refuses a cassette whose verdict redaction changed, and an invocation count cannot be changed by it. A run that called the tool but whose every call carried an unusable path reports **cannot verify**, never a claim of non-delivery. The presence companion to `no_scratchpad_leak` (which passes vacuously when nothing was presented). Pair them to require a delivery **and** require it not to leak. Content-class (re-derives identically on replay). **`fidelity: container` or `hostloop`** — the harness serves `present_files` at both (hostloop via a handler mirroring production's own host-loop branch: validate the path, pass it through, no promotion). `protocol` and `microvm` report cannot-verify. See the `no_scratchpad_leak` row, which stays container-only for a different reason; and see the lane note there — the tool name differs on remote Cowork. **Only `true` is valid** |
84
84
  | `question_asked: <regex>` | the agent asked an AskUserQuestion whose **question text** matches (`question`, falling back to `header`). Text only — for the options it offered, use `question_options` ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
85
85
  | `question_options: {when_question?, equals?: [..], contains?: [..], order?}` | the option SET and ORDER a gate offered, **by label** — descriptions are not compared (see `question_context`). `when_question` selects the sub-question by the same label `question_asked` matches (omit only if exactly one fired; ambiguity FAILS). Exactly one of `equals` (complete set) / `contains` (subset); `order: exact` is the default because a re-ordered list is the defect this exists for. Captured at ask time, so a gate that was shown then denied/stalled still counts; unreadable evidence fails evidence-unavailable ⚠️ **Model-composed text, reworded run to run** — pin a producer-authored constant, not model prose. |
@@ -110,7 +110,7 @@ same set live from the schema.
110
110
  | `no_mcp_error: true` | no MCP round-trip failed (`RunResult.mcpErrors` is empty — no unhandled server, no handler throw) — live-only: MCP round-trips are harness-computed, not in the SDK stdout stream, so evidence-unavailable on replay (never a vacuous pass). **Only `true` is valid** |
111
111
  | `max_peak_rss_bytes: <N>` | peak sampled RSS of the agent sandbox ≤ N bytes (`RunResult.resources.peakRssBytes`) — live-only: replay never spawns a sandbox to sample, so evidence-unavailable on replay/protocol (never a vacuous pass); also evidence-unavailable when sampling captured no RSS value |
112
112
  | `semantic_matches: {rubric: [...], min_pass?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge grades each fixed `rubric` claim against the run's answer: the final result text, the transcript and the files the agent authored. The transcript is top-level `assistant_text` only — it excludes every `tool_use`/`tool_result` and sub-agent text unless opted in. Full semantics (evidence scope, fork results, refusals, judge provenance, cost): [semantic-judging.md](semantic-judging.md). |
113
- | `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline". A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
113
+ | `semantic_pairwise: {refs?: [..], rubric?: [..], pass_if?, order?, judge_model?, include_subagent_text?, include_fork_results?, evidence_files?: [..]}` | a pinned LLM judge compares this run's judged document with a **frozen reference** — the same document an earlier run (usually the baseline) produced, written once by `ref freeze` and never regenerated — and answers win, tie, loss or `both_bad`. **The judged document is the one `semantic_matches` builds** (final message + transcript + authored files, same `evidence_files` / `include_subagent_text` / `include_fork_results` options), so the transcript **excludes every `tool_use`/`tool_result`** and a criterion about whether a tool was called cannot be judged from it. `refs` are reference stores relative to the scenario file; the run is judged against every one (under `hillclimb run`, against the flow's own references instead — only the baseline's gates `pass_if`, a later variant's is a metric recorded with `gate: false`), and `pass_if` (default `not_worse`) must hold against each gating one: `win`, `not_worse` (win or tie), or `any` (graded at all — a metric only). `both_bad` fails `win` and `not_worse`. Which output the judge sees first is a seeded coin per run, assert and reference (`order: both` judges both orders: a win in one and a loss in the other is position bias and scores as a tie; any other disagreement keeps the worse outcome; each order's own outcome is recorded as `pairwise[].orders.candidate_first` / `.ref_first`, absent for a single-order grade, and a judge that favours whichever output it sees first shows across runs as `candidate_first` winning more often than `ref_first`). The judge never sees the words "reference" or "baseline", and the stored reference is scrubbed with the grading process's secret set before it reads it (`pairwise[].refRedactions`). A reference must have been frozen with the same evidence options and for the same prompt; a missing, damaged, differently-scoped or other-prompt one **refuses the run before it spends anything**, as does a reference store inside any mounted source (the agent could read its own answer key). Unavailable evidence refuses with `semantic_matches`' typed reasons. Per-reference outcomes land in `assertions[].pairwise`. **LIVE-ONLY** — skipped on replay. |
114
114
  | `artifact_json: {artifact, path, …}` | assert a JSON artifact's contents — `equals`/`gt`/`in`/`exists`/`absent`/`is_null` over a dotted `path` (`in` = membership in a list, for a stochastic/LLM value; `absent` ≠ `is_null`; an unresolved intermediate fails loud) |
115
115
  | `computer_links_resolve: true` | every `computer://` link in the model-visible transcript resolves to an artifact that exists in the run's collected outputs/mounts — a dangling link fails, naming which target was checked (a live host path, the collected work tree, or the replay manifest). **Requires ≥1 link** (zero links fails — use `computer_links_resolve_if_present` for the presence-free variant). **Only `true` is valid** (`false` is rejected by the schema) **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
116
116
  | `computer_links_resolve_if_present: true` | like `computer_links_resolve` but passes vacuously when the transcript has zero `computer://` links — the presence-free variant. **Only `true` is valid** **Sees top-level `assistant_text` only — it excludes every `tool_use`/`tool_result`**, so a `computer://` link that appeared only inside a tool call or its result is invisible to it. |
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -11,8 +11,10 @@ provenance, and the scaffold/lint tools that keep the YAML honest.
11
11
  ### Two files: session vs scenario
12
12
 
13
13
  - **`sessions/*.yaml`** — pre-prompt setup: `model`, mounts (`folders`), and discovery
14
- (marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario that
15
- omits `session:` gets an all-defaults **inline** session (not a file on disk).
14
+ (marketplaces / plugins / skills / mcp). One session is reused by many scenarios. A scenario's
15
+ `session:` is a **path** to such a file, never a nested block: `session: { plugins: … }` fails to load. A
16
+ scenario that omits `session:` gets an all-defaults session with no plugin declared, so a scenario that tests a
17
+ plugin needs a session file.
16
18
  - **`scenarios/*.yaml`** — the test: `prompt`, scripted `answers:`, and `assert:`.
17
19
 
18
20
  This split matters: release ground truth (`baseline:` / `baselines/`, produced by `sync`) is
@@ -52,18 +54,16 @@ Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejec
52
54
  (it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
53
55
  `references/fidelity-and-answers.md`.
54
56
 
55
- **Every tier models Cowork's DESKTOP-LOCAL lane** — agent on the user's machine, shell rooted at
57
+ **Every tier models Cowork's LOCAL lane** — agent on the user's machine, shell rooted at
56
58
  `/sessions/<id>`, folders at `/sessions/<id>/mnt/<name>`, delivery via `present_files`. Cowork's
57
59
  **remote** lane runs server-side in a cloud container with a different filesystem (`$HOME/mnt/`),
58
60
  different delivery (`/mnt/user-data/outputs/` + `SendUserFile`) and a server-authored prompt; no tier
59
- reproduces it and none can — that container is not something a local tool can stand up. No setting
60
- reliably decides which lane a real session gets: sessions ran in the cloud with "Only on this computer"
61
- **on** (observed 2026-10-02), and for Pro and Max plans Anthropic
62
- [announces](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile) that new
63
- tasks run in the cloud from 2026-10-06. Check the session's own lane
61
+ reproduces it and none can — that container is not something a local tool can stand up.
62
+ [From 2026-10-06 new Pro and Max tasks run in the cloud](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile);
63
+ before then no setting reliably decided the lane. Check the session's own lane
64
64
  ([how](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md#which-lane-a-session-actually-ran-on)).
65
65
  So: behaviour conclusions (triggering, tool sequencing, gate handling) travel between lanes; anything
66
- asserting a **path, mount or delivery mechanism** is a claim about the local lane only. Declare
66
+ asserting a **path, mount, delivery mechanism or egress rule** is a claim about the local lane only. Declare
67
67
  `lane: remote` when the scenario is about that lane — the affected assertions then refuse to grade
68
68
  rather than passing (see the `delivery_unobservable` WARN and the `lane: remote` load-time rejections
69
69
  in `run-record-replay.md`).
@@ -213,6 +213,10 @@ cowork-harness scaffold --name report-check --skill ./skills/report-gen \
213
213
  --egress-allowed api.weather.example.com --out scenarios/report-check.yaml
214
214
  ```
215
215
 
216
+ Each repeatable flag adds one item: `--content` a `transcript_matches`, `--tool` a `tool_called`, `--subagent` a
217
+ `subagent_dispatched`, `--file` a `file_exists`, `--artifact` a `user_visible_artifact`, `--egress-allowed` /
218
+ `--egress-denied` an `egress_allowed` / `egress_denied`, `--gate REGEX=CHOICE` a scripted answer and `--web-fetch`
219
+ an approval rule; `--no-delete` adds `no_delete_in_outputs: true`, and `--no-validate` skips the self-lint.
216
220
  The flag-built form runs the bundled `scripts/scenario.py scaffold`, which also runs directly with the
217
221
  same flags (installed as a plugin, `${CLAUDE_PLUGIN_ROOT}/scripts/scenario.py`).
218
222
 
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "4.4.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "4.5.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -36,7 +36,7 @@ jobs:
36
36
  - uses: actions/checkout@v4
37
37
  - name: Stage the agent binary (official channel, sha256-verified against the pinned baseline)
38
38
  run: |
39
- V=2.1.286 # match your scenario's pinned baseline's agentVersion
39
+ V=2.1.289 # match your scenario's pinned baseline's agentVersion
40
40
  # The release channel is NOT always the stable one. Desktop also stages release CANDIDATES,
41
41
  # served from .../claude-code-releases/rc/<commit>/. For some versions the stable path 404s
42
42
  # (2.1.255); for others it returns 200 and serves a DIFFERENT BUILD UNDER THE SAME VERSION
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.4.0"
85
+ - run: npm i -g "cowork-harness@^4.5.0"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -131,6 +131,16 @@ exits 0, so dropping the
131
131
  `verify-cassettes` step means a skill edit silently stops being tested. (One command instead of two:
132
132
  `replay --fail-on-skill-drift`.)
133
133
 
134
+ `lint --json` and `lint-skill --json` print their findings as JSON for a script to read. `analyze-skill --strict
135
+ <plugin-dir>/` is the static pre-flight for host-loop path fidelity and lost artifact write-backs (exit 3 when it
136
+ could not parse a candidate); `--runtime` adds a headless-DOM confirmation (needs `jsdom`) that never changes
137
+ the exit code.
138
+
139
+ Other `verify-cassettes` flags: `--skip-privacy` or `--skip-staleness` runs only one half of the gate;
140
+ `--skip-scenario-drift` drops the scenario-prompt drift check; `--allow-empty` lets an existing directory with no
141
+ cassettes exit 0 (a missing path still fails); `--margins` prints each count-bound assert's recorded value against
142
+ its budget (diagnostic, never the verdict).
143
+
134
144
  The rest of this doc explains the lane split, recording, privacy, and the full pipeline + live job.
135
145
 
136
146
  ## The core split: token-free PR gate + live nightly (self-hosted)
@@ -398,7 +408,7 @@ jobs:
398
408
  with: { node-version: '24' }
399
409
  - uses: actions/setup-python@v5
400
410
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
401
- - run: npm i -g "cowork-harness@^4.4.0"
411
+ - run: npm i -g "cowork-harness@^4.5.0"
402
412
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
403
413
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
404
414
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -427,7 +437,7 @@ jobs:
427
437
  echo "live=true" >> "$GITHUB_OUTPUT"
428
438
  fi
429
439
  - if: steps.guard.outputs.live == 'true'
430
- run: npm i -g "cowork-harness@^4.4.0"
440
+ run: npm i -g "cowork-harness@^4.5.0"
431
441
  - if: steps.guard.outputs.live == 'true'
432
442
  run: cowork-harness run scenarios/ --output-format json
433
443
  env:
@@ -520,7 +530,11 @@ set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path
520
530
  artifact-upload step can collect them. Each run dir holds `events.jsonl`, `control-out.jsonl` and
521
531
  `egress.log` at the root, plus each turn's `run.jsonl` / `trace.json` / `result.json` under `turns/<N>/`
522
532
  (a single-turn run has just `turns/1/`; there is no root compat copy of any of these). Digest one with `cowork-harness trace <run-id | dir>`.
523
- Secrets are scrubbed from every persisted log by value.
533
+ Secrets are scrubbed from every persisted log by value. The first live run under a runs root also creates its
534
+ scrub-set key, `scrubset.key`, beside the runs root (in the working directory for `--run-dir runs`): add it to
535
+ `.gitignore`, and never upload or commit it (an upload of `runs/` does not include it). A fresh runner creates a new
536
+ key, so `regrade` of a CI run on another machine cannot prove its scrub set covered: new or edited rubric text is
537
+ refused there; re-run the case instead.
524
538
 
525
539
  ## Don't assume a fixed assertion count across lanes
526
540
 
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -35,7 +35,9 @@ Two things a harvester needs: roll-ups are excluded from `stats` aggregation (th
35
35
  counting them adds a phantom run and drags `passRate` toward 1), so **filter them out of any pass-rate
36
36
  computed over raw rows** — that exclusion also governs `stats --group-by skill-hash`, so a per-generation
37
37
  **total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
38
- UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
38
+ UNDERCOUNT. The index is the only cost record that survives run-dir pruning. `--evaluator-model <id>` changes
39
+ only the evaluator passes' model (and the armor's injection-resistance check covers the default evaluator only);
40
+ when the task turn dominates cost, the levers are `--model`, `--timeout` and the probe's scope.
39
41
 
40
42
  ## An exit-2 report — which turn failed, and whether it was really infrastructure
41
43
 
@@ -1,6 +1,6 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
@@ -40,10 +40,30 @@ rubric changed (or you want another judge model) on a run you already paid for,
40
40
  kept run without re-running the agent — unlike the tools above it is not token-free (the judge call is its
41
41
  spend) — writes the grade beside the run, and says whether the judge read the same document the live judge did.
42
42
  Content the live judge never read (a widened `evidence_files` / `include_subagent_text` / `include_fork_results` scope, or a larger
43
- `--authored-total-bytes`) is refused unless you pass `--allow-unchecked`. A `task_unverifiable` /
43
+ `--authored-total-bytes`) is refused unless you pass `--allow-unchecked` (`error.code` `unchecked_content`); a
44
+ judged document that differs from the one the live judge read is refused as `doc_drift` unless you pass
45
+ `--allow-doc-drift`, and a scenario with no judged assert is `no_semantic_asserts`. A `task_unverifiable` /
44
46
  `rubric_unverifiable` / `evidence_unverifiable` / `reference_unverifiable` refusal means a part of the judge's input
45
- cannot be proven scrubbed with the run's scrub set (typically a pre-4.4 run with an edited rubric): re-run the
46
- case, or pass `--allow-scrub-change` after checking `COWORK_HARNESS_SCRUB_VALUES` / `_KEYS`.
47
+ cannot be proven scrubbed with the run's scrub set (typically an edited rubric on a run recorded before the scrub-set
48
+ fingerprint); `refusals[].scrubSet` says why (`legacy`, `unrecorded`, `mangled`, `key`, `smaller`). Re-run the
49
+ case, regrade with the run's `COWORK_HARNESS_SCRUB_VALUES` / `_KEYS`, or pass `--allow-scrub-change` after checking
50
+ them. Content accepted with `--allow-unchecked`, and the document of an `unknown` or `live_refused` assert, are
51
+ outside that proof (scrubbed with this process's set only); `regrade` names both in a `::warning::` before judging.
52
+
53
+ **`::warning:: [scrub-set] no scrub-set fingerprint is recorded: the installation key …`** means the run could not
54
+ use its installation key, `scrubset.key` beside the runs root (`~/.cowork-harness/scrubset.key` for the default
55
+ root; the current directory for `--run-dir runs`). The warning names the path and the defect: not a regular file
56
+ (a symlink or a directory), readable or writable by others (it must be mode 600), owned by another user, empty (an
57
+ interrupted create), unreadable, not a 64-hex-digit key, or could not be created. Such a run records
58
+ `scrubSetUnavailable` instead of a `scrubSet`, so a later `regrade` of new or edited judge input refuses with
59
+ `scrubSet: "unrecorded"` ("this run recorded no scrub set (its key …)"). The harness never removes or replaces an
60
+ existing key file: fix it as the warning says (`chmod 600` a key you own, or delete a bad file so the next run
61
+ creates a key), then re-run the case to record a fingerprint. Keep `scrubset.key` out of version control.
62
+
63
+ Two more read-only views: `trace <run> --translate-paths` rewrites VM paths to host paths in the text
64
+ `tools`/default views (an effective `hostloop` run with its `mounts.json`), and `diff <a> <b>` masks per-run noise
65
+ (ids, timestamps, host paths) unless you pass `--no-normalize`; on two baselines, `--changelog` renders the
66
+ known fields as prose.
47
67
 
48
68
  **microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
49
69
  stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
@@ -74,16 +94,14 @@ decide which assertions from *Assertions: two orthogonal axes* in `assertions-gu
74
94
  bare `trace` digests the whole run. The view set is actively being extended — run `trace --help` for
75
95
  the current list rather than relying on a fixed enumeration here.
76
96
  - **`lane: local|remote`** (scenario key, default `local`) — which Cowork lane's DELIVERY CONTRACT the run
77
- is held to. As of Desktop 2.19675.0 the composer offers no per-session lane
78
- picker and no setting reliably decides the lane (for Pro and Max plans, Anthropic
79
- [announces](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile) that
80
- new tasks run in the cloud from 2026-10-06); the lanes disagree about what *delivered* means. On `remote`,
97
+ is held to. `local` is the default and models the local lane, which new Pro and Max tasks do not use from
98
+ 2026-10-06 (they run in the cloud); the lanes disagree about what *delivered* means. A live `local` run with an
99
+ environment-shaped assertion (a path, mount, delivery or egress key) prints one `[lane]` line on stderr saying
100
+ so, once per process; `--compact`/`--demo`, `CI` and `COWORK_HARNESS_NO_LANE_NOTICE=1` silence it. On `remote`,
81
101
  location delivers nothing (a remote container has no auto-delivering outputs dir and is reclaimed at
82
102
  session end), `present_files` is NOT served, and `user_visible_artifact` /
83
103
  `present_files_called` / `no_scratchpad_leak` are rejected at LOAD time as unable to pass. Reach for it
84
- to check a skill's delivery survives the cloud lane (Anthropic
85
- [announces](https://support.claude.com/en/articles/15520349-use-claude-cowork-on-web-desktop-and-mobile) that
86
- from 2026-10-06 new Pro and Max tasks run in the cloud). Orthogonal to `fidelity` — a `lane: remote`
104
+ to check a skill's delivery survives the cloud lane, where new Pro and Max tasks run from 2026-10-06. Orthogonal to `fidelity` — a `lane: remote`
87
105
  scenario still runs locally.
88
106
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
89
107
  `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`. Filters: `--since`/`--baseline`/`--branch`,
@@ -1,6 +1,6 @@
1
1
  # `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full guide is
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). The full guide is
4
4
  [docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
5
5
  need while running it.
6
6
 
@@ -31,6 +31,9 @@ cowork-harness eval report <eval-dir> # rebuild the report from the eval dir
31
31
  worst observed runs × 2 × `--reps` exceed x — pre-flight only, agent cost only, never a mid-run stop.
32
32
  - In a scenario directory, YAML with no `prompt:` (a session file) is skipped.
33
33
  - Run an A/A first (`--allow-identical-arms`, the same source twice) to see your scenarios' noise.
34
+ - A directory arm's snapshot leaves out untracked files unless you pass `--include-untracked` (not with a
35
+ `git:` arm). `--correction bh|holm` picks the multiple-comparison correction behind `confirmed` (default
36
+ `bh` at q = 0.10; `holm` is stricter).
34
37
 
35
38
  ## Reading the labels
36
39
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -207,7 +207,11 @@ permissive behaviour is deliberately what the scenario is about.
207
207
 
208
208
  A Desktop update deletes the prior version's staged agent while often leaving an empty version dir, so
209
209
  a scenario pinning that agent version resolves to nothing. `doctor` validates the agent for its own
210
- current baseline, not what each scenario pins, so it can report ready seconds before the run fails.
210
+ current baseline, not what each scenario pins, so it can report ready seconds before the run fails. Its
211
+ staged-agent row ends with what Desktop has staged against that pin: staged and pinned; a newer agent staged
212
+ (upgrade cowork-harness, or run `cowork-harness sync` if you maintain the baseline); this Desktop older than the
213
+ pin; the pinned agent not staged (staging may be withheld, or no task has booted the VM
214
+ since the update); or none found.
211
215
 
212
216
  To keep the exact pin, **recover the pinned ELF**: re-download that version from the release channel,
213
217
  check its sha256 against the baseline's, and set `COWORK_AGENT_BINARY` to it
@@ -298,7 +302,7 @@ up often enough to spell out:
298
302
  behavior). See [`docs/cassette.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cassette.md) § "Still skipped on replay" and [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md) § "Which
299
303
  assertions survive replay."
300
304
  - **`present_files` assertions can't verify off the tiers that serve the tool.** `no_scratchpad_leak` and
301
- `present_files_called` check the `present_files` delivery path — the desktop-local lane's tool; remote
305
+ `present_files_called` check the `present_files` delivery path — the local lane's tool (not used by new Pro and Max tasks from 2026-10-06); remote
302
306
  Cowork delivers via the agent-native `SendUserFile` instead, so never hardcode a delivery tool name in
303
307
  a SKILL.md (Gotcha 24 in `gotchas.md`). **The harness** serves `present_files` on `container` **and `hostloop`**
304
308
  — not `microvm`/`protocol`. `present_files_called` works at both; `no_scratchpad_leak` stays
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -270,8 +270,8 @@ authorable). Reach for this list when debugging a run's behavior, that one while
270
270
  `--output-format json`). *Fix:* re-record against the pinned agent.
271
271
 
272
272
  24. **Never name the file-delivery tool in a `SKILL.md`.** *Why:* Cowork has **two**, one per product
273
- lane, and an agent only sees the one for the surface it is on. The desktop-local sandbox this harness
274
- emulates is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
273
+ lane, and an agent only sees the one for the surface it is on. The local-lane sandbox this harness
274
+ emulates (not used by new Pro and Max tasks from 2026-10-06) is served `mcp__cowork__present_files` (`{files:[{file_path}]}`); **remote** cloud-container
275
275
  Cowork instead gives the agent the native `SendUserFile` (`files: string[]`, required `status`,
276
276
  optional `caption`/`display`). A skill that hardcodes either name works on one lane and fails on the
277
277
  other — and probing a remote session makes this harness look like it emulates the wrong tool under the
@@ -1,6 +1,6 @@
1
1
  # Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
5
5
  `/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
6
6
  row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
@@ -59,6 +59,9 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
59
59
  - **Never pass `--approve-harness`.** It records the harness sha and is the user's to run. A refusal before
60
60
  spending prints `refusing to run: …` and exits 2: stop and show the user what it names. `freeze-ref` needs
61
61
  no approval.
62
+ - **Never pass `--allow-scrub-change` unattended,** and never put it in an allowed prefix. It sends the judge a
63
+ part of its input that cannot be proven scrubbed with the run's scrub set, so it can disclose a secret the run
64
+ scrubbed. Pass it only after the user has checked the scrub settings and said yes for that listing.
62
65
  - **`--skill` takes the bare skill name** (a `skills/<dir>` name or the registered name), never `plugin:name`
63
66
  or a path.
64
67
 
@@ -76,10 +79,27 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
76
79
  - **Recompute the headline** from `F/<variant>/results.jsonl`, never from `summary.json`.
77
80
  - **Spot-check grading.** Read the lowest-scoring baseline rows' `explanation` and traces. If a rubric is wrong,
78
81
  tell the user; after they edit it and approve the new sha, `cowork-harness hillclimb regrade T --flow F`
79
- re-evaluates every row in place from its kept run without running the agent: changed deterministic
82
+ re-evaluates the rows in place from their kept runs without running the agent: changed deterministic
80
83
  assertions and every metric without a judge call, a judged assertion re-judged because its rubric changed. An
81
84
  unchanged assertion whose evaluation a harness upgrade changed keeps its recorded outcome (noted); editing the
82
85
  assertion, or re-running the case, re-grades it.
86
+ - **Rows listed instead of re-graded (exit 1).** A new or edited rubric line (or another part of the judge's input)
87
+ is sent only when it can be proven scrubbed with the run's scrub set. When it cannot, the row is listed, untouched
88
+ and still carrying its old grade, with no judge call; stderr names the parts and why the run's set is not proven,
89
+ and `F/<variant>/regrade.md` lists the row. The usual causes: a run recorded before the scrub-set fingerprint (no
90
+ `scrubSet` in its `result.json`; one stderr summary line counts them), a run whose key was unusable ("this run
91
+ recorded no scrub set"), a run from another machine or under a replaced `scrubset.key`, or a token the run
92
+ scrubbed that has rotated since. What to do, in order: if the run recorded no scrub set, first have the user fix
93
+ `scrubset.key` as its `::warning:: [scrub-set]` line names (see the debugging reference); the harness never
94
+ replaces an existing bad key, so a new run before that fix records no scrub set again and is listed again. If this
95
+ process lacks a scrub value the run had (`COWORK_HARNESS_SCRUB_VALUES` / `COWORK_HARNESS_SCRUB_KEYS`, a rotated
96
+ token), ask the user to set the run's settings and regrade again; otherwise the listed cases need new runs,
97
+ which record a fresh fingerprint once the key is usable. A pass
98
+ runs nothing for a slot that already has a row, so tell the user and propose running the change as a new variant,
99
+ or a fresh flow dir from the baseline when the baseline's rows are listed. Only if the user has checked the scrub
100
+ settings and asks for it, re-run the same `regrade` with `--allow-scrub-change` (the regrade file then records
101
+ `scrubAcceptedBy`). Never pass it unattended; `--allow-doc-drift`, `--allow-unchecked` and `--rejudge` never
102
+ imply it. Don't compare variants while rows are listed: their grades are from the old rubric.
83
103
  - **Triage every zero.** An agent's own failure is a scored row (`meta.failure_class: "errored_agent"`, with
84
104
  `meta.termination_rule`). Infrastructure, timeouts, a wrong served model and invalid judge grades are
85
105
  `errors.jsonl` rows, never in the scored denominator. A variant that rewords or batches its questions can miss
@@ -182,7 +202,11 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
182
202
  - **Grader drift.** If a rubric is wrong, the user edits it and approves the sha; then
183
203
  `cowork-harness hillclimb regrade T --flow F` re-grades every variant's rows (default `--variant all`) and writes
184
204
  `F/<variant>/regrade.md` when a row moved or was listed. A row whose judge evidence itself changed since it was graded is
185
- listed instead: ask the user before re-running with `--rejudge`.
205
+ listed instead: ask the user before re-running with `--rejudge`. A row whose new rubric text cannot be proven
206
+ scrubbed with its run's scrub set is listed too, and so is one whose evidence would be less redacted than the
207
+ graded document: handle both as in Step 0.5 (*Rows listed instead of re-graded*), never with `--allow-doc-drift`.
208
+ For the less-redacted row the override, if the user asks for it, is `--rejudge --allow-scrub-change` together:
209
+ `--allow-scrub-change` alone does not re-judge changed evidence.
186
210
  - **A new metric.** Add `metrics:` to the scenario, have the user approve the sha, re-run
187
211
  `cowork-harness hillclimb state-template T --flow F` and merge only the new `metrics` entries into
188
212
  `_state.json`. Rows written before it lack the key (`check` notes them); `hillclimb regrade T --flow F` fills it
@@ -192,9 +216,16 @@ Each line matches one entry of the list in [the hillclimb guide](https://github.
192
216
  `cowork-harness hillclimb regrade T --flow F --fill-refs` (no `--case`: rows of cases with no pairwise
193
217
  assertion need the column too, at no judge cost), then `state-template T --flow F` and merge the new
194
218
  `win_vN` entries (it declares them only once no scored row lacks them). Only the baseline's reference
195
- decides `pass`.
196
- - **After a rubric re-grade, compare the ranking.** If the order of the variants flipped, tell the user and
197
- propose restarting the climb from the baseline.
219
+ decides `pass`. When an entry already frozen lacks a compose key (an assertion added or re-scoped since),
220
+ `freeze-ref` and a baseline pass add it only when this process's scrub set provably covers the run it was frozen
221
+ from; otherwise the case is refused (exit 1). The refusal says to re-run the variant, but a pass never re-runs a
222
+ filled slot and a baseline pass refuses again each time: restore the run's scrub settings and re-run `freeze-ref`,
223
+ or start a fresh flow dir. A `--fill-refs` comparison against a
224
+ reference the row's run never judged is proven only when the run's scrub set is covered; otherwise the row is
225
+ listed (exit 1) and keeps no `win_vN` column. Handle it as in Step 0.5: same scrub settings, or new runs; ask
226
+ the user before any `--allow-scrub-change`.
227
+ - **After a rubric re-grade, compare the ranking** once no row is listed. If the order of the variants flipped,
228
+ tell the user and propose restarting the climb from the baseline.
198
229
 
199
230
  ## Step 5 — report and hand back
200
231
 
@@ -1,6 +1,6 @@
1
1
  # `hillclimb` — the runner for a `/claude-api hillclimb` loop
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
5
5
  [docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
6
6
  while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
@@ -224,7 +224,9 @@ measured) whose run delivered an output, under the variant's lock: it refuses wh
224
224
  its `meta.run_dir`, else by its run id under the current runs root (`--run-dir` / `COWORK_HARNESS_RUNS_DIR`). An
225
225
  entry that is already complete is reported (`exists`), never rewritten. One that lacks a compose key (an assert
226
226
  added or re-scoped) gains it from the run it was frozen from, marked `unchecked`, and is refused when that run is
227
- gone: start a fresh flow dir then.
227
+ gone: start a fresh flow dir then. The new document is composed with this process's scrub set, so it is added
228
+ (by `freeze-ref` or a baseline pass) only when that set provably covers the one the run recorded (`scrubSet`);
229
+ otherwise the case is refused with the remedy: re-run the variant, or add it under the run's scrub settings.
228
230
 
229
231
  Flags: `--variant ID` (required), `--flow DIR`, `--case ID` (repeatable), `--output-format text|json` (json
230
232
  carries `{frozen, added, exists, refused}`), `--dotenv FILE`, `--run-dir DIR`.
@@ -280,10 +282,12 @@ every climb there is finished.
280
282
  the judge's input is covered. Otherwise each part must equal the run's own scrubbed record: the pairwise `## Task`
281
283
  line (the prompt), each rubric line, each evidence note, and each reference, by its bytes whatever it is called
282
284
  (its send must hash to what a live comparison of the run under the same compose key was sent, `refSentSha256`; a
283
- pre-4.4 grade's stored text must be one that comparison recorded). A reference re-frozen
284
- since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven neither way is listed, with no judge call. A run from before 4.4 has no fingerprint, so
285
- its new or edited rubric text is listed (one stderr line says so); so is a run from another machine, or one whose
286
- token has rotated since. Re-run the case, or pass `--allow-scrub-change` after checking the scrub settings: the
285
+ grade recorded before the reference-send hash, the stored text must be one that comparison recorded). A reference
286
+ re-frozen since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven
287
+ neither way is listed, with no judge call. A run recorded before the scrub-set fingerprint (no `scrubSet`) proves
288
+ no coverage, so its new or edited rubric text is listed (one stderr line counts such rows); so is a run whose key
289
+ was unusable (`scrubSetUnavailable`), a run from another machine or under a replaced key, or one whose token has
290
+ rotated since. Re-run the case, or pass `--allow-scrub-change` after checking the scrub settings: the
287
291
  regrade file then records `scrubAcceptedBy`. Neither `--rejudge` nor `--allow-doc-drift` implies it.
288
292
  - **`--rejudge`:** every judged assert of every selected row is re-judged, with the flow's references as they are
289
293
  now. Use it after a judge change the triggers above do not see. Not with `--fill-refs`.
@@ -370,7 +374,11 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
370
374
  assert does not line up with the scenario or whose deterministic outcome changed (run a default `regrade` first;
371
375
  an agent-failed row aside); one whose
372
376
  re-grade is judge-invalid; in a fill one whose kept outcome was judged against a reference that has changed
373
- since; and an open `judge_invalid` slot.
377
+ since; one with a part of the judge's input not proven scrubbed with its run's scrub set, one whose evidence
378
+ would be less redacted than the graded document, or one whose re-judge needs an assert whose scrubbed literal
379
+ this process cannot reproduce (each released only by `--allow-scrub-change`, with `--rejudge` for a less-redacted
380
+ row, never by `--allow-doc-drift`); one whose assert has an edit inside a scrubbed literal (only a re-run applies
381
+ it); and an open `judge_invalid` slot.
374
382
 
375
383
  ## Exit codes
376
384
 
@@ -384,7 +392,8 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
384
392
  - `state-template`: `0`, or `2` on usage or a refusal.
385
393
  - `freeze-ref`: `0` no case refused (an entry already complete is reported, not refused); `1` a case refused (no
386
394
  good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a
387
- different prompt, a missing compose key whose run is gone, a store write that failed, or a run whose judged
395
+ different prompt, a missing compose key whose run is gone or whose scrub set this process cannot prove it
396
+ covers, a store write that failed, or a run whose judged
388
397
  document cannot be composed, differs from the one its live judge read, or has no live fingerprint); `2` usage
389
398
  (a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, the variant's
390
399
  lock held by a live run).
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.5.0` (baseline `desktop-2.26454.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -56,7 +56,7 @@ Compare timings between runs of the same tier and model only.
56
56
  `fingerprint.skillHash` is content-exact but one-way, so an edit mid-batch silently splits your
57
57
  dataset into two generations — and a hash whose source was never frozen identifies a generation that
58
58
  is unrecoverable. `stats --group-by skill-hash` separates them after the fact; nothing recovers the
59
- source.
59
+ source. (`stats --reindex` rebuilds the runs index from the run dirs when it is lost or predates it.)
60
60
  3. **Check which arm you actually ran** before analysing anything: `ablated` and
61
61
  `context.availableSkills` in each `result.json`.
62
62
  4. **Classify each rep three ways**: invocation (`skillsInvoked`), observed source access (did it read