cowork-harness 4.3.0 → 4.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +4 -4
  2. package/.claude/skills/cowork-harness/references/assertion-catalog.md +1 -1
  3. package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
  4. package/.claude/skills/cowork-harness/references/authoring.md +1 -1
  5. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  6. package/.claude/skills/cowork-harness/references/critique.md +1 -1
  7. package/.claude/skills/cowork-harness/references/debugging.md +5 -2
  8. package/.claude/skills/cowork-harness/references/eval.md +1 -1
  9. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  10. package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
  11. package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +1 -1
  12. package/.claude/skills/cowork-harness/references/hillclimb.md +14 -4
  13. package/.claude/skills/cowork-harness/references/measurement.md +1 -1
  14. package/.claude/skills/cowork-harness/references/run-record-replay.md +5 -2
  15. package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
  16. package/.claude/skills/cowork-harness/references/semantic-judging.md +1 -1
  17. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  18. package/CHANGELOG.md +93 -0
  19. package/DESIGN.md +1 -1
  20. package/README.md +3 -3
  21. package/SPEC.md +23 -8
  22. package/dist/hillclimb/cli.js +1 -0
  23. package/dist/hillclimb/freeze-ref.js +21 -4
  24. package/dist/hillclimb/regrade.js +61 -10
  25. package/dist/hillclimb/usage.js +11 -5
  26. package/dist/refs/cli-usage.js +5 -3
  27. package/dist/refs/cli.js +33 -4
  28. package/dist/refs/compose.js +4 -1
  29. package/dist/refs/store.js +5 -3
  30. package/dist/run/cassette.js +4 -0
  31. package/dist/run/chat-result.js +2 -0
  32. package/dist/run/execute.js +13 -0
  33. package/dist/run/pairwise-prepass.js +37 -8
  34. package/dist/run/regrade-usage.js +4 -3
  35. package/dist/run/regrade.js +168 -4
  36. package/dist/scrub-set.js +196 -0
  37. package/docs/ci.md +1 -1
  38. package/docs/cli.md +73 -15
  39. package/docs/companion-skill.md +2 -2
  40. package/docs/hillclimb.md +13 -2
  41. package/examples/replays/README.md +1 -1
  42. package/package.json +1 -1
  43. package/schema/regrade.json +45 -6
  44. package/schema/run-result.json +15 -0
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 4.3.0
7
- tracks-harness: cowork-harness 4.3.0 (baseline desktop-2.19675.0)
6
+ version: 4.4.0
7
+ tracks-harness: cowork-harness 4.4.0 (baseline desktop-2.19675.0)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
26
26
  full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
27
27
  Read them.
28
28
 
29
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.3.0` (baseline
29
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.4.0` (baseline
30
30
  > `desktop-2.19675.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
31
31
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
32
32
 
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
43
43
 
44
44
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
45
45
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
46
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.3.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.3.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.3.0"`. **Pin `@^4.3.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.4.0"`. **Pin `@^4.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
47
47
 
48
48
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
49
49
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # Assertion catalog
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Every `assert:` key with its semantics, and the
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Every `assert:` key with its semantics, and the
4
4
  verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
5
5
  the scenario and session YAML fields are there too.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Assertions guide
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
4
4
 
5
5
  ### Assertions: two orthogonal axes
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Authoring a scenario
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
4
4
 
5
5
  ## Part I — AUTHOR a scenario
6
6
 
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "4.3.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "4.4.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
82
82
  GitHub-hosted runners, no token/Docker/agent:
83
83
 
84
84
  ```yaml
85
- - run: npm i -g "cowork-harness@^4.3.0"
85
+ - run: npm i -g "cowork-harness@^4.4.0"
86
86
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
87
87
  # no silent false-greens. WITHOUT --strict this
88
88
  # step cannot fail on a WARN-class rule (e.g.
@@ -398,7 +398,7 @@ jobs:
398
398
  with: { node-version: '24' }
399
399
  - uses: actions/setup-python@v5
400
400
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
401
- - run: npm i -g "cowork-harness@^4.3.0"
401
+ - run: npm i -g "cowork-harness@^4.4.0"
402
402
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
403
403
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
404
404
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -427,7 +427,7 @@ jobs:
427
427
  echo "live=true" >> "$GITHUB_OUTPUT"
428
428
  fi
429
429
  - if: steps.guard.outputs.live == 'true'
430
- run: npm i -g "cowork-harness@^4.3.0"
430
+ run: npm i -g "cowork-harness@^4.4.0"
431
431
  - if: steps.guard.outputs.live == 'true'
432
432
  run: cowork-harness run scenarios/ --output-format json
433
433
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -1,6 +1,6 @@
1
1
  # Debugging a run
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
4
4
 
5
5
  ## Part III — Debug
6
6
 
@@ -40,7 +40,10 @@ rubric changed (or you want another judge model) on a run you already paid for,
40
40
  kept run without re-running the agent — unlike the tools above it is not token-free (the judge call is its
41
41
  spend) — writes the grade beside the run, and says whether the judge read the same document the live judge did.
42
42
  Content the live judge never read (a widened `evidence_files` / `include_subagent_text` / `include_fork_results` scope, or a larger
43
- `--authored-total-bytes`) is refused unless you pass `--allow-unchecked`.
43
+ `--authored-total-bytes`) is refused unless you pass `--allow-unchecked`. A `task_unverifiable` /
44
+ `rubric_unverifiable` / `evidence_unverifiable` / `reference_unverifiable` refusal means a part of the judge's input
45
+ cannot be proven scrubbed with the run's scrub set (typically a pre-4.4 run with an edited rubric): re-run the
46
+ case, or pass `--allow-scrub-change` after checking `COWORK_HARNESS_SCRUB_VALUES` / `_KEYS`.
44
47
 
45
48
  **microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
46
49
  stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
@@ -1,6 +1,6 @@
1
1
  # `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). The full guide is
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full guide is
4
4
  [docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
5
5
  need while running it.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`).
3
+ Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
4
4
 
5
5
  > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
6
  > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
@@ -1,6 +1,6 @@
1
1
  # Gotchas
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). The full "✓ passed ≠ correct" landmine catalog.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full "✓ passed ≠ correct" landmine catalog.
4
4
 
5
5
  ## Gotchas — the "✓ passed ≠ correct" landmines
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
5
5
  `/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
6
6
  row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
@@ -1,6 +1,6 @@
1
1
  # `hillclimb` — the runner for a `/claude-api hillclimb` loop
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
4
4
  `hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
5
5
  [docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
6
6
  while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
@@ -274,7 +274,17 @@ every climb there is finished.
274
274
  markers than the graded one (a secret the run scrubbed that this process does not), or a changed authored file
275
275
  whose graded fingerprint has no marker count, is listed even under `--rejudge`, by path only, with no judge call.
276
276
  Set the run's `COWORK_HARNESS_SCRUB_VALUES` / `COWORK_HARNESS_SCRUB_KEYS` and regrade again, or, after checking
277
- them, `--rejudge --allow-doc-drift` grades it anyway.
277
+ them, `--rejudge --allow-scrub-change` grades it anyway (`--allow-doc-drift` does not).
278
+ - **Never a part the run scrubbed, unproven.** A run records a keyed fingerprint of its scrub set (`result.json`
279
+ `scrubSet`; the key is `scrubset.key` beside the runs root). When this process's set provably covers it, every part of
280
+ the judge's input is covered. Otherwise each part must equal the run's own scrubbed record: the pairwise `## Task`
281
+ line (the prompt), each rubric line, each evidence note, and each reference, by its bytes whatever it is called
282
+ (its send must hash to what a live comparison of the run under the same compose key was sent, `refSentSha256`; a
283
+ pre-4.4 grade's stored text must be one that comparison recorded). A reference re-frozen
284
+ since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven neither way is listed, with no judge call. A run from before 4.4 has no fingerprint, so
285
+ its new or edited rubric text is listed (one stderr line says so); so is a run from another machine, or one whose
286
+ token has rotated since. Re-run the case, or pass `--allow-scrub-change` after checking the scrub settings: the
287
+ regrade file then records `scrubAcceptedBy`. Neither `--rejudge` nor `--allow-doc-drift` implies it.
278
288
  - **`--rejudge`:** every judged assert of every selected row is re-judged, with the flow's references as they are
279
289
  now. Use it after a judge change the triggers above do not see. Not with `--fill-refs`.
280
290
  - **`--fill-refs`:** only the `semantic_pairwise` comparisons a row lacks are judged (a reference frozen after the
@@ -293,7 +303,7 @@ every climb there is finished.
293
303
 
294
304
  Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant with rows), `--case ID`
295
305
  (repeatable), `--judge-model ID`, `--fill-refs`, `--rejudge`, `--approve-harness`, `--allow-doc-drift`, `--allow-unchecked`,
296
- `--output-format text|json`, `--dotenv FILE`, `--run-dir DIR`.
306
+ `--allow-scrub-change`, `--output-format text|json`, `--dotenv FILE`, `--run-dir DIR`.
297
307
 
298
308
  - **Everything that can refuse does so before the first judge call**, and then writes nothing: a host `claude`
299
309
  that cannot run the judge isolated (asked only when a judge would be called — again under the locks when a
@@ -323,7 +333,7 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
323
333
  its literal (one scrubbed value for another) cannot be told from no edit: on a row whose `meta.assert_sig` is not
324
334
  the scenario's now, it lists the row, untouched, on every regrade until the case is re-run. One this process cannot
325
335
  reproduce is kept unchanged too — never re-evaluated or re-judged over scrubbed evidence; a re-judge it would need
326
- lists the row (same scrub settings, or `--allow-doc-drift` after checking — then the judge sees the RAW rubric
336
+ lists the row (same scrub settings, or `--allow-scrub-change` after checking — then the judge sees the RAW rubric
327
337
  against the scrubbed evidence, so its grade may not match the live run's) — and named on stderr (an edited one
328
338
  takes a re-run). A row no judge re-grades is re-evaluated too: a case with no judged assert, a row whose judged
329
339
  asserts all keep their entries, an agent-failed row (it gains the metric signature and `<id>_present: 0`, never a
@@ -1,6 +1,6 @@
1
1
  # Measurement
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
4
4
 
5
5
  ### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Run, record and lock
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
4
4
 
5
5
  ## Part II — RUN, RECORD & LOCK
6
6
 
@@ -20,7 +20,10 @@ reworded gate or a `choose:` the run never offered fails here in ~1s instead of
20
20
  never calls the semantic judge, so a `semantic_matches` or `semantic_pairwise` assert is not re-graded by it; after a rubric change,
21
21
  `cowork-harness regrade <run-dir> --scenario <scenario.yaml>` re-grades those against the kept run (the judge call
22
22
  is the only spend) and reports whether the judge read the same document the live judge did (widening the evidence
23
- scope needs `--allow-unchecked`: content the live judge never read is refused otherwise). A run dir moved or
23
+ scope needs `--allow-unchecked`: content the live judge never read is refused otherwise). A part of the judge's
24
+ input that cannot be proven scrubbed with the run's scrub set is refused too: on a run from before 4.4 (no
25
+ `scrubSet` in `result.json`), from another machine, or after a token rotated, unchanged rubric lines re-grade but a
26
+ new or edited one is refused until the case is re-run or `--allow-scrub-change` is passed. A run dir moved or
24
27
  downloaded from where it ran is read from where it is; a COPY beside its still-present original is refused (its
25
28
  `result.json` names the original's files), so grade the original, or re-run the scenario. Or skip
26
29
  the discovery/encode/record dance entirely and answer gates **live during the recording** with
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, replay class, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.3.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.4.0`
4
4
  (baseline `desktop-2.19675.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -1,6 +1,6 @@
1
1
  # `semantic_matches` — the full semantics
2
2
 
3
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`). The one-line summary is in
3
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The one-line summary is in
4
4
  [assertion-catalog.md](assertion-catalog.md); this is the whole contract, split out so the catalog stays within the
5
5
  agent's single-read size.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 4.3.0` (baseline `desktop-2.19675.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,99 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [4.4.0] — 2026-10-05
10
+
11
+ A security fix: `regrade` and `hillclimb regrade` no longer send a judge a value the run scrubbed. A run now records
12
+ a keyed fingerprint of its scrub set, and a re-grade sends only the parts of the judge's input it can prove are
13
+ scrubbed with a set covering the run's. **A run recorded before 4.4 has no fingerprint**, so re-grading it with new
14
+ or edited judge input (rubric text, an evidence note, a reference re-frozen since) is refused by `regrade` and
15
+ listed by `hillclimb regrade` until the case is re-run or `--allow-scrub-change` is passed.
16
+
17
+ ### Security
18
+
19
+ - **A re-grade no longer sends a judge a value the run scrubbed.** Before, `regrade` and `hillclimb regrade`
20
+ (`--rejudge` and `--fill-refs` included) rebuilt the judge's input with the scrub set of the process re-grading.
21
+ When that set was smaller than the run's, a value the run had scrubbed reached the host `claude` judge raw, and
22
+ nothing was listed or refused: the `semantic_pairwise` `## Task` line (the scenario prompt), a new or edited
23
+ rubric line (core `regrade` had no rubric check; `hillclimb regrade` checked only asserts matched to the run's),
24
+ and the frozen pairwise reference, which was sent as stored on every lane, live included. The judge's rationale,
25
+ which can quote its input, is then stored in the regrade file and in `results.jsonl`. Now:
26
+ - A run records a keyed fingerprint of its scrub set, `result.json` `scrubSet` (`{v, keyId, values}`): one
27
+ HMAC-SHA256 per scrubbed string under a random per-installation key, `scrubset.key` beside the runs root (mode
28
+ 0600; `~/.cowork-harness/scrubset.key` for the default root), so a runs root uploaded or shared carries no key.
29
+ It holds no value, and without the key there is nothing to brute-force.
30
+ - Before any judge call, a re-grade proves each part it sends: covered when the run's set is a subset of this
31
+ process's under the same key (a grown set included), else by equality with the run's own scrubbed record —
32
+ the task line (checked only when a pairwise comparison will be judged), each rubric line, each evidence-health
33
+ and scratch note, and each reference, by its bytes whatever it is called: its send now must hash to what a live
34
+ comparison of the run under the same compose key was sent (`pairwise[].refSentSha256`, new), or, for a grade
35
+ recorded before 4.4 (whose judge got the stored text raw), the stored text must be one such a comparison
36
+ recorded (`refDocSha256`). A reference re-frozen since, or one the run never judged (a `--fill-refs` column),
37
+ proves nothing. A part proven neither way is refused by `regrade` (exit 2, `error.code` `task_unverifiable`,
38
+ `rubric_unverifiable`, `evidence_unverifiable` or `reference_unverifiable`, in `refusals[]` beside `doc_drift`)
39
+ and listed by `hillclimb regrade` (exit 1),
40
+ naming the parts, never their text: each `refusals[]` entry carries `scrubSet` (why the set is not proven),
41
+ `scrubSetDetail` (the run's own reason, when it recorded one), and `rubricLines`, `evidenceSections` or
42
+ `references`.
43
+ - The new `--allow-scrub-change` flag (both commands) sends such a part anyway, recorded as `scrubAcceptedBy` on
44
+ the regrade file and `runs[]`. `--allow-doc-drift`, `--allow-unchecked` and `--rejudge` never imply it; an
45
+ accepted drift whose authored file lost scrub markers is part of what is proven.
46
+ - Every pairwise comparison scrubs the stored reference with the grading process's set before the judge reads it,
47
+ live runs included, and records how many markers that added (`pairwise[].refRedactions`). A reference document's
48
+ sidecar records `scrubCount`, a count only, and a reference frozen with fewer strings is noted
49
+ (`ref_scrub_weaker`).
50
+ - `hillclimb freeze-ref` and a baseline pass add a missing compose key to a reference only when this process's set
51
+ provably covers the run's; otherwise the case is refused with "re-run the variant". `ref freeze` freezes an
52
+ unchecked document (`--allow-unchecked`) under the same proof, or with `--allow-scrub-change` (which
53
+ `ref verify` refuses, like the other freeze-only flags).
54
+ - An unusable installation key (a symlink, a file readable by others, another owner, an empty file, not a key) is
55
+ warned about once, naming its path and the defect, and recorded on the result as `scrubSetUnavailable`, so a
56
+ re-grade names that cause instead of calling the run older than 4.4. An existing key file is never removed or
57
+ replaced. A new key is written whole before it is linked into place; on a filesystem with no hard links it is
58
+ created exclusively in place instead.
59
+ - An invalid pairwise judge reply that quotes its input is scrubbed before it is warned on stderr or stored, and
60
+ every other warning a pairwise comparison prints (a reference's name, a model id) is scrubbed too.
61
+
62
+ ### Upgrade notes
63
+
64
+ - **Cassettes: no re-record needed.** Nothing under `src/runtime`, `src/hostloop`, `src/staging`, `src/session.ts`,
65
+ `baselines/` or `docker/` changed, `latest` still resolves to `desktop-2.19675.0`, and `CASSETTE_VERSION` is still
66
+ 14. `src/run/cassette.ts` changed only to leave the new `scrubSet` / `scrubSetUnavailable` result fields unset on
67
+ a replay; a cassette's recorded content and fingerprints are unchanged.
68
+ - **Re-grading a run from before 4.4.** Such a run records no scrub-set fingerprint, so nothing can prove its set
69
+ is covered. Its unchanged task and rubric lines still re-grade, and so does a reference still byte-identical to the
70
+ one its live grade recorded. New or edited rubric text, a changed evidence note, or a reference re-frozen since (or
71
+ never judged by the run, as a `--fill-refs` column is) is refused by `regrade` and listed by
72
+ `hillclimb regrade`, with one summary line naming the remedy: re-run the case (a new run records its set), or
73
+ pass `--allow-scrub-change` after checking the scrub settings. The same applies to a run recorded on another
74
+ machine, and to a run whose auth token has rotated since, since a named key's value is part of the set.
75
+ - **Keep `scrubset.key` out of version control.** The first live run under a runs root creates the installation key
76
+ beside it: `~/.cowork-harness/scrubset.key` for the default root, but the current directory for `--run-dir runs`.
77
+ When the runs root sits in a repository, add `scrubset.key` to its `.gitignore`.
78
+
79
+ ### Changed
80
+
81
+ - **`--allow-doc-drift` no longer unlocks `hillclimb regrade`'s scrubbed-literal or less-redacted listings;
82
+ `--allow-scrub-change` replaces it for those.** A row listed because an assert's scrubbed literal cannot be
83
+ reproduced, or because its evidence would be less redacted than the graded document, was released by
84
+ `--allow-doc-drift` in 4.3.0 (with `--rejudge` for the latter). Passing `--allow-doc-drift` now leaves such a row
85
+ listed; pass `--allow-scrub-change` (with `--rejudge` for a less-redacted row) after checking the scrub settings.
86
+ 4.3.0's re-grade surfaces are new and experimental, so this change to the flag's meaning ships in a minor release.
87
+ - **What the scrub-set proof does not cover.** Content accepted with `--allow-unchecked`, and the documents of
88
+ `unknown` / `live_refused` asserts (no live fingerprint to compare with), are still protected only by the
89
+ re-grading process's scrub set, as before; `regrade` warns about both before any judge call. So is an authored
90
+ file a drift accepted with `--allow-doc-drift` changed on a run whose set is not proven covered, when its scrub
91
+ markers did not drop: it is sent scrubbed with the current set only.
92
+
93
+ ### Fixed
94
+
95
+ - **Corrections to the 4.3.0 notes.** 4.3.0 said a `hillclimb regrade` re-judge over a scrubbed literal "lists the
96
+ row, with `--allow-doc-drift` the only override", and the docs described the drift and less-redacted checks as
97
+ what keeps a value the run scrubbed from the judge. Those checks covered the judged documents only: the pairwise
98
+ task line, an edited or added rubric line and the frozen reference were not covered, and core `regrade` had no
99
+ rubric check. The `ref freeze` docs said the run under test "gets the same redaction" as the stored reference;
100
+ that held for host paths, not for the secret scrub. All of these are fixed above.
101
+
9
102
  ## [4.3.0] — 2026-10-04
10
103
 
11
104
  Full support for Claude Code's `/claude-api hillclimb` loop, with `cowork-harness hillclimb` as its runner,
package/DESIGN.md CHANGED
@@ -208,7 +208,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
208
208
  > Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
209
209
  > 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
210
210
 
211
- > **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **Final pass, on the release commit `8a5e479f`:** (1) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip (the auto-memory prerequisites check, which runs only under `COWORK_LIVE_REQUIRE=1`); before it ran, `vitest list --staticParse=false` listed all 23 with no `SKIPPED` warning, and 6 with four warnings when no token was resolvable, so a silent skip would have shown; (2) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (3) the graders' pinned effort: every judged assert that called the judge in the live hillclimb acceptance runs (45, Opus 4.8: 25 `semantic_matches`, 20 `semantic_pairwise`; host `claude` 2.1.288 and 2.1.289) records `judgeTransport.effort: "high"`; the five pairwise asserts without a `judgeTransport` are the reference variant's own rows, which call no judge. **Earlier in the same release, on the release candidate's tree:** (4) a `microvm` smoke on a freshly booted VM (agent 2.1.286); (5) `boundary-check`, 6/6 constraints enforced; (6) an `eval` A/A at `--reps 4` on `container`, every row `no detectable change`, its report rebuilt byte-identical; (7) the companion skill's router check, three prompts on `container`: routing was correct on all three; one answer was wrong because a reference showed `artifact_text`'s `contains` without its list shape (fixed in the references, with a test that keeps every list-typed field shown as a list), and the same prompt's verdict went red on a false `host_path_leak`, because a reference spelled out the host roots that check looks for and the agent echoed it (fixed in the skill's files and in the check, which no longer counts a literal that comes from the plugin's or a local skill's own staged files); re-run after the fixes, the prompt routed correctly, wrote a scenario that passes `lint` and loads, and carried no `host_path_leak`; (8) the `hostloop` detached-agent kill check: the first pass found the workspace sidecar container, its `docker run` client and its network surviving SIGINT, which was fixed, and the re-run passed on a normal run, SIGINT to the harness process and SIGINT to the whole process group (exit 130, nothing left behind); (9) `vm delete`'s usage refusal; (10) `chat` under a real terminal, on `protocol`: Ctrl-C mid-turn exits 130 in about 2 s with nothing left running and no `result.json`; `/exit` followed by Ctrl-C exits 130 after writing the turn's `result.json`; Ctrl-C at the idle prompt ends the session normally (exit 0, `result.json` written); SIGHUP mid-turn exits 129 with nothing left running; and closing the terminal window mid-turn (run by hand) leaves no harness or agent process and no `result.json`. Total live spend about **$11.20**, including the re-runs. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction, and these checks are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (b) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all four committed cassettes were **re-recorded** on 2026-10-03 against `2.19675.0` with auto-memory off, each on its original model: `example-pdf-skill` and `test/fixtures/tool-call-dispatch/dispatch-shell.cassette.json` (`container`, agent 2.1.286), `hostloop-computer-links` (`hostloop`, agent 2.1.286) and `example-multiselect-gate` (`protocol`, which runs the host CLI, here 2.1.287), about $0.75 in all. Their init frames carry no `memory_paths`. The recorder's own scrub (now applied to every recording) removed, across the four, the account's model menu, a subscription account's `rate_limit_info` (from `example-pdf-skill`, `dispatch-shell` and `hostloop-computer-links`) and the agent's sub-agent hand-back frame (from `dispatch-shell`). `verify-cassettes` exits 0 on all four, with one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
211
+ > **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **The 4.4.0 pass, 2026-10-05, on the release-prep commit `e9c516df`** (same baseline and staged agent; host `claude` 2.1.289 for `protocol`): (i) `vitest list --staticParse=false` listed all 23 live tests with no `SKIPPED` warning when the token was exported, and 6 with four warnings when none was resolvable; (ii) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip; (iii) the auto-memory check on all three tiers, `Tests 2 passed (2)` each; (iv) the companion skill's router check, three prompts on `container`: the authoring prompt read `authoring.md` and wrote a scenario that passes `lint`, the tool-timing prompt read `measurement.md` and answered with `trace … --view tool-durations`, and the debugging prompt read no reference and answered with correct clarifying questions, so `debugging.md` routing was not exercised in this pass; (v) the re-grade scrub-set check: a live run on the default runs root created `scrubset.key` beside the runs directory (mode 0600) and recorded `scrubSet` (`v: 1`) with the configured scrub literal absent from `result.json`; a re-grade of that run with an edited rubric line and a smaller scrub set was refused as `rubric_unverifiable` (exit 2) naming only the edited line, with no judge call and the run directory byte-identical; and a re-grade of a run recorded before 4.4 with an edited rubric line was refused with the `legacy` reason and the one-line remedy, again with no judge call and the run directory byte-identical. Live spend about **$2.83** recorded. **The 4.3.0 pass, on the release commit `8a5e479f`:** (1) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip (the auto-memory prerequisites check, which runs only under `COWORK_LIVE_REQUIRE=1`); before it ran, `vitest list --staticParse=false` listed all 23 with no `SKIPPED` warning, and 6 with four warnings when no token was resolvable, so a silent skip would have shown; (2) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (3) the graders' pinned effort: every judged assert that called the judge in the live hillclimb acceptance runs (45, Opus 4.8: 25 `semantic_matches`, 20 `semantic_pairwise`; host `claude` 2.1.288 and 2.1.289) records `judgeTransport.effort: "high"`; the five pairwise asserts without a `judgeTransport` are the reference variant's own rows, which call no judge. **Earlier in the same release, on the release candidate's tree:** (4) a `microvm` smoke on a freshly booted VM (agent 2.1.286); (5) `boundary-check`, 6/6 constraints enforced; (6) an `eval` A/A at `--reps 4` on `container`, every row `no detectable change`, its report rebuilt byte-identical; (7) the companion skill's router check, three prompts on `container`: routing was correct on all three; one answer was wrong because a reference showed `artifact_text`'s `contains` without its list shape (fixed in the references, with a test that keeps every list-typed field shown as a list), and the same prompt's verdict went red on a false `host_path_leak`, because a reference spelled out the host roots that check looks for and the agent echoed it (fixed in the skill's files and in the check, which no longer counts a literal that comes from the plugin's or a local skill's own staged files); re-run after the fixes, the prompt routed correctly, wrote a scenario that passes `lint` and loads, and carried no `host_path_leak`; (8) the `hostloop` detached-agent kill check: the first pass found the workspace sidecar container, its `docker run` client and its network surviving SIGINT, which was fixed, and the re-run passed on a normal run, SIGINT to the harness process and SIGINT to the whole process group (exit 130, nothing left behind); (9) `vm delete`'s usage refusal; (10) `chat` under a real terminal, on `protocol`: Ctrl-C mid-turn exits 130 in about 2 s with nothing left running and no `result.json`; `/exit` followed by Ctrl-C exits 130 after writing the turn's `result.json`; Ctrl-C at the idle prompt ends the session normally (exit 0, `result.json` written); SIGHUP mid-turn exits 129 with nothing left running; and closing the terminal window mid-turn (run by hand) leaves no harness or agent process and no `result.json`. Total live spend about **$11.20**, including the re-runs. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction, and these checks are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (b) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all four committed cassettes were **re-recorded** on 2026-10-03 against `2.19675.0` with auto-memory off, each on its original model: `example-pdf-skill` and `test/fixtures/tool-call-dispatch/dispatch-shell.cassette.json` (`container`, agent 2.1.286), `hostloop-computer-links` (`hostloop`, agent 2.1.286) and `example-multiselect-gate` (`protocol`, which runs the host CLI, here 2.1.287), about $0.75 in all. Their init frames carry no `memory_paths`. The recorder's own scrub (now applied to every recording) removed, across the four, the account's model menu, a subscription account's `rate_limit_info` (from `example-pdf-skill`, `dispatch-shell` and `hostloop-computer-links`) and the agent's sub-agent hand-back frame (from `dispatch-shell`). `verify-cassettes` exits 0 on all four, with one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
212
212
 
213
213
  > The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
214
214
 
package/README.md CHANGED
@@ -36,7 +36,7 @@ npm ci && npm run build
36
36
  node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
37
37
  ```
38
38
 
39
- (Installing globally — `npm install -g "cowork-harness@^4.3.0"` — gives you the `cowork-harness` CLI for your own
39
+ (Installing globally — `npm install -g "cowork-harness@^4.4.0"` — gives you the `cowork-harness` CLI for your own
40
40
  scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
41
41
 
42
42
  Full setup → [Quick start](./docs/cli.md#quick-start).
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
49
49
 
50
50
  | I want to… | Start here | Needs |
51
51
  |---|---|---|
52
- | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.3.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
- | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.3.0"` |
52
+ | **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.4.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
53
+ | **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.4.0"` |
54
54
  | **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v4`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
55
55
 
56
56
  Improving a skill round by round with Claude Code's `/claude-api hillclimb` loop? The harness is its runner, with
package/SPEC.md CHANGED
@@ -606,7 +606,7 @@ there are three families:
606
606
  `dryRun: true` with the estimate at `plan.cost`, where `eval --dry-run` puts it), `hillclimb check` (`reading`, `disclaimer`, `profile`, `findings`, `errors`, `notes`, `warnings`),
607
607
  `hillclimb state-template` (`state`, `metrics_md`, and with `--flow` `metrics_md_file`; `notes` when a column was left
608
608
  undeclared, experimental), `hillclimb freeze-ref` (payload experimental, §12: `frozen`, `added`, `exists`, `refused`,
609
- `exitCode`), `hillclimb regrade` (payload experimental, §12: `flow`, `variants` (each `{variant, rewritten, judged, rebuilt, ownRefOnly, listedAfterJudge, judgeUsd?, judgeUnpriced, judgeStopped, listed, reevaluated, remeasured, agentFailed, evidenceChanged, regradeFiles, backup?}`), `exitCode`; so are `<variant>/regrade.md`, the `regrade-<sha16>.bak.jsonl` backups and the rows' `meta.regrade_*` and `meta.assert_sig` keys; a row a judged entry of which the current harness would show its judge different evidence than it was graded on is in `evidenceChanged` and is never silently kept or re-graded — listed and kept without `--rejudge`, re-judged on the current evidence with it, recording `evidence_changed` and both document hashes in `meta.regrade_evidence` — except a row whose current document would be less redacted than the graded one, listed in both modes unless `--allow-doc-drift` is also given), `verify-cassettes` (§11.1), `doctor` (§11.2), `rehash`,
609
+ `exitCode`), `hillclimb regrade` (payload experimental, §12: `flow`, `variants` (each `{variant, rewritten, judged, rebuilt, ownRefOnly, listedAfterJudge, judgeUsd?, judgeUnpriced, judgeStopped, listed, reevaluated, remeasured, agentFailed, evidenceChanged, regradeFiles, backup?}`), `exitCode`; so are `<variant>/regrade.md`, the `regrade-<sha16>.bak.jsonl` backups and the rows' `meta.regrade_*` and `meta.assert_sig` keys; a row a judged entry of which the current harness would show its judge different evidence than it was graded on is in `evidenceChanged` and is never silently kept or re-graded — listed and kept without `--rejudge`, re-judged on the current evidence with it, recording `evidence_changed` and both document hashes in `meta.regrade_evidence` — except a row whose current document would be less redacted than the graded one, listed in both modes unless `--allow-scrub-change` is also given; and a row with a part of its judge input not proven scrubbed with a set covering its run's, listed unless `--allow-scrub-change`), `verify-cassettes` (§11.1), `doctor` (§11.2), `rehash`,
610
610
  `answer` (`gate`, `answers`), `fixture export` (exit `0` written, `2` usage or refusal; payload experimental,
611
611
  §12: `message`, `written`, `skipped`, `refused` (`[]`), `notes`, `bytes`, `outputsDir`, `partial`, `result`; a
612
612
  refusal is the error envelope, its message in `error.message`, carrying `refused[]` and whichever of the other
@@ -711,6 +711,19 @@ Two evidence refusals are decided before any judge call, for every run dir:
711
711
  secrets) are never unchecked content, so a smaller budget — which only drops or truncates file content, and may
712
712
  add a health note saying so — is never refused as unchecked; content it truncates makes that assert refuse its
713
713
  own evidence. An assert whose evidence will be refused sends nothing and is not measured.
714
+ - **Scrub coverage.** A run records `scrubSet` (`{v, keyId, values}`: one HMAC-SHA256 per string its scrub used,
715
+ under a per-installation key, `scrubset.key`, beside the runs root). When this process's set provably covers the
716
+ run's (every recorded HMAC recomputed under the same key), every part of the judge's input is covered. Otherwise
717
+ each part sent must equal the run's own scrubbed record: the `semantic_pairwise` task line (only when a pairwise
718
+ comparison is judged), each rubric line, each evidence-health or scratch note, an accepted drift's authored file
719
+ that carries fewer scrub markers, and each reference: its send must hash to the `refSentSha256` any live comparison of the
720
+ run recorded under the same compose key, whatever the reference is called (a pre-4.4 grade, which sent the stored
721
+ text raw: the stored text's sha256 must equal such a comparison's `refDocSha256`). An
722
+ unproven part refuses unless `--allow-scrub-change` is passed, with `task_unverifiable`, `rubric_unverifiable`,
723
+ `evidence_unverifiable` or `reference_unverifiable`; accepted, the run's entry and file record `scrubAcceptedBy`.
724
+ Neither `--allow-doc-drift` nor `--allow-unchecked` implies it. A run with no `scrubSet` (harness < 4.4), another
725
+ key, or a set lacking a string the run scrubbed (a rotated token included) proves nothing: only parts equal to its
726
+ record are sent.
714
727
 
715
728
  Not checked: a run in which no live assert recorded a `judgedDoc` has no rebuilt document to measure against — its
716
729
  asserts are `unknown` or `live_refused`, neither drift- nor secret-checked, and are warned about before the judge
@@ -720,11 +733,12 @@ so its extra content can refuse. The envelope is scrubbed with the same secret s
720
733
  **Exit codes:** `0` every re-graded assert passes · `1` any fails or is judge-invalid · `2`, with three meanings:
721
734
  a usage error; a refusal before any judge call (a multi-turn, partial, replay or chat run dir, a pruned work dir,
722
735
  a missing transcript sidecar, a run that did not record `authoredCapture` without `--authored-total-bytes`, an
723
- alias judge model, a scenario with neither a `semantic_matches` nor a `semantic_pairwise` assert (`error.code: "no_semantic_asserts"`), and the two evidence refusals above); or a failure
736
+ alias judge model, a scenario with neither a `semantic_matches` nor a `semantic_pairwise` assert (`error.code: "no_semantic_asserts"`), and the evidence and scrub refusals above); or a failure
724
737
  writing a regrade file after earlier run dirs were graded. Each is the shared error envelope. The evidence
725
738
  refusals are collected over every run dir and carry `error.code` — `doc_drift` when any dir drifted, else
726
- `unchecked_content` — and a top-level `refusals[]`, one entry per run dir and code: `{runDir, code,
727
- uncheckedCount?, uncheckedSections?, liveDocDrift?}`; every other refusal stops at the first run dir that
739
+ `unchecked_content`, else the first of `task_unverifiable`, `rubric_unverifiable`, `evidence_unverifiable`,
740
+ `reference_unverifiable` — and a top-level `refusals[]`, one entry per run dir and code: `{runDir, code,
741
+ uncheckedCount?, uncheckedSections?, liveDocDrift?, scrubSet?, scrubSetDetail?, rubricLines?, evidenceSections?, references?}`; every other refusal stops at the first run dir that
728
742
  fires it and carries no code. So `refusals[]` is complete only when no other refusal fires: a batch with a
729
743
  drifted dir and a later dir refused for another reason (a pruned work dir, say) reports only the latter, with no
730
744
  code and no `refusals[]`.
@@ -878,11 +892,11 @@ assertions (never user-authored themselves):
878
892
  "results":[], // [] except record's post-run refusal: the refused run, beside the non-null error
879
893
  "budget?": { /* §11 --max-budget-usd marker — present when a pre-flight ran */ },
880
894
  "error": { "category": "usage|unanswered|boundary|runtime|internal", "message": "string", "hint?": "string",
881
- "code?": "budget_exceeded|doc_drift|unchecked_content|no_semantic_asserts", "budget?": { /* §11 --max-budget-usd */ } } }
895
+ "code?": "budget_exceeded|doc_drift|unchecked_content|task_unverifiable|rubric_unverifiable|evidence_unverifiable|reference_unverifiable|no_semantic_asserts", "budget?": { /* §11 --max-budget-usd */ } } }
882
896
  ```
883
897
  `error.code` narrows a category, never replaces it: `budget_exceeded` is the `--max-budget-usd` refusal (§11);
884
- `doc_drift` and `unchecked_content` are `regrade`'s evidence refusals, whose error envelope also carries a
885
- top-level `refusals[]` (see `regrade` above); `no_semantic_asserts` is `regrade`'s `usage` refusal of a scenario
898
+ `doc_drift` and `unchecked_content` are `regrade`'s evidence refusals and the four `*_unverifiable` codes its scrub
899
+ refusals, whose error envelope also carries a top-level `refusals[]` (see `regrade` above); `no_semantic_asserts` is `regrade`'s `usage` refusal of a scenario
886
900
  with neither a `semantic_matches` nor a `semantic_pairwise` assert (nothing to re-grade).
887
901
  Categories come from TYPED errors (`UnansweredError`→`unanswered`, `BoundaryError`→`boundary`).
888
902
  `results` is `[]` with one exception: when `record` refuses to write a cassette after the agent finished,
@@ -936,7 +950,7 @@ prints the same exclusion warning the run prints. `verify-run` follows the same
936
950
  not exist or is a file, or a scenario file that does not load, is `usage`; a directory holding no completed
937
951
  run stays `runtime`. `answer` splits the same way: a directory or gate that is not there is `usage`; a gate
938
952
  request that exists but cannot be read or parsed, or an answer that cannot be written, is `runtime`.
939
- **Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `eval` (and `eval report`) exits `0` when the comparison completed — whatever drops the rows show, unless `--fail-on` was given; `1` for a drop at the `--fail-on` level (`possible` or `confirmed`; no gating without the flag — under `possible` a drop includes an `insufficient_refusals` row, one the candidate's excess `semantic_matches` refusals for unavailable evidence took below the rep threshold, which would otherwise hide a drop as `insufficient`), every row `insufficient`, or a judge model that differed across reps; `2` for a usage error or any refusal before the first run (an alias model, an eval dir inside a git work tree, identical arms, the answer-key guard, a scenario input a run would refuse, an unreachable `--fail-on confirmed`, a `--max-budget-usd` refusal — `error.code: "budget_exceeded"`); `3` when an arm snapshot could not be copied or failed its staging preflight. Its `--output-format json` envelope is `{tool, version, command:"eval", ok, evalDir, arms, pins, sections, summary, cost, stoppedEarly, error}`, with `ok` ⇔ exit `0`. `eval --dry-run` runs no agent and creates no eval dir: it exits `0` with the plan (`{tool, version, command:"eval", ok:true, dryRun:true, plan, budget?, error:null}`), `2` for any refusal the real eval would make before its first run (the budget refusal included), `3` as above, or when its temp dir is inside a git work tree or git cannot tell whether it is (set TMPDIR) — its arm snapshots go to a temp dir that is removed afterwards. `hillclimb run` exits `0` when every attempted (case, rep) was scored, `1` when any attempt failed (an `errors.jsonl` row, or a scored row whose trace or copies could not be written), the pass stopped mid-run (rows already written are kept), a baseline pass could not freeze a pairwise case's reference, or `summary.json` could not be written after the pass, and `2` for any refusal before spending: a usage error, the harness gate, no usable agent credential at a selected case's tier (`doctor`'s check, as on `eval`; at `protocol` a login only in the real config dir too, since hillclimb runs protocol with a managed config dir), an alias model, a session that does not declare exactly one `plugins.local_plugins` entry (or an inline session, or cases tuning different plugins), `on_unanswered: prompt`, the `protocol` tier without a managed config dir, a flow-dir or `_state.json` problem, split or duplicate case ids, an unknown `--case`, the answer-key guard, a `harness_paths` entry inside the tuned plugin, an unknown, untracked or ambiguous `--skill` (the message lists the plugin's skills), a tracked skill (or none) that differs from a `meta.skill_tracked` recorded by the variant's existing rows, a requested model or effort for a case that differs from the `meta.model_requested` / `meta.effort` the variant's rows for that case recorded (scored or error rows) (a row recording neither is held to its served `model`), an `--effort` level the pinned model does not offer (or any effort on a model with no effort selector, dated ids included), `xhigh` or `max` with `extended_thinking: false`, `extended_thinking: false` on a model whose baseline entry sets `disallowThinkingDisabled`, a missing or incomplete variant snapshot, a scenario input a run would refuse, a `semantic_pairwise` reference that is missing, damaged or exposed through a mount (the gate a run applies), a host `claude` that cannot run the judge or LLM decider isolated (checked when a scenario uses `semantic_matches` or `semantic_pairwise`, or `on_unanswered: llm` with no decider channel, as on `eval`), or a decider with `--concurrency` above 1, or one scenario metric id declared differently in two scenarios, or declared differently from the signature the flow's existing rows carry for it in `meta.metric_sigs` (adding or removing a metric is allowed; removing one warns); `--dry-run` exits `0` unless such a refusal fires; the harness gate does not refuse a dry run, which reports its status instead. Under `--case`, the per-case refusals (the session, its model pins, `on_unanswered: prompt`, the `protocol` tier, a scenario input, a `semantic_pairwise` reference, the isolation check, and the mounts the answer-key guard checks) cover the selected cases only; the flow-level ones cover every case: every scenario file must parse, every case's baseline must load, the harness gate hashes every case's scenario and session file (an unselected case's inputs it cannot read — missing, a directory, unreadable — or that its unparseable session cannot name, are left out of it; the gate hashes each path's name, so the sha still moves), the answer-key guard keeps every case's scenario and session file unreadable, and the one-plugin rule covers every case whose session parses (an unselected one whose session is inline or does not parse is skipped, with a note). `hillclimb regrade --case` follows the same rule, with the same harness sha. `--timeout-s` bounds the whole attempt: the agent phase through the scenario's `timeout_ms` (lowered to it), and the judge phase through a deadline no judge may start after; an attempt whose agent phase reaches its bound, or whose judge phase reaches the deadline, is an `errors.jsonl` row (`timeout`), never a scored one. So is an attempt whose agent did not send the requested effort: its own session transcript shows a main-loop assistant message sent with another effort, with a value that is not an effort level, or with none (even when the agent then failed; a model with no effort selector may send none), or, on an otherwise valid run whose main loop answered, records no main-loop assistant message or cannot be found (read from the run's config root, or the session's pinned `plugins.config_dir` on hostloop and protocol, by the run's session id): `serving_substitution` with `meta.failure_rule: "effort_not_sent"`. Its envelope's `ok` ⇔ exit `0`; a refusal or a mid-run stop is the shared error envelope, its reason in `error.message` (category `usage` for a refusal, `runtime` for a stop), with the payload keys (`flow`, `variant`, `scheduled`, `scored`, `failed`, `exitCode`, and a dry run's `dryRun`) riding on it; `failed` counts failed (case, rep) slots, so a failed reference freeze or `summary.json` write exits `1` without adding to it. Its `--dry-run` `plan.cost` is the same object `eval --dry-run` emits at `plan.cost` (covered keys `jobs`, `meanUsd`, `p50Usd`, `p95Usd`, `worstObservedUsd`, `lowerBound`, `unpriced`, `pricedRuns`, `thinnest`; `p95Usd` and `worstObservedUsd` are pessimistic figures, never bounds, and `worstObservedUsd` is not the `--max-budget-usd` gate's figure), priced on `eval --dry-run`'s basis exactly — the scenario's runs on this machine at its effective tier and baseline, turn 1, every `hillclimb:` run excluded — so each covered key means the same on both commands. `jobs` is the slots the pass would still run after resuming. A scored row records how its judge ran in `meta.judge_transport` (`assertions[].judgeTransport`'s shape, `{isolation, cliVersion?, strictMcp?}`), or `meta.judge_transports` listing each distinct one when its asserts were judged differently; neither is present when no host judge recorded one, so a round judged under a different isolation can be told apart from a regression. `skill_invoked` is `1` or `0` for whether the run invoked the tracked skill; a blank cell means not measured (no tracked skill, or a record that could not tell), never "not invoked". The tracked skill is matched by the id the agent registers for it, `<plugin>:<name>`, the name being a skill directory's name, or a `SKILL.md`'s frontmatter `name` where the agent reads one (else its directory's name), with every character outside `[a-zA-Z0-9_-]` replaced by `-` (`skills/my.skill/` registers `<plugin>:my-skill`), among the skills the agent loads: `skills/*/` whenever `skills/` exists, each path in the plugin manifest's `skills` field (a string or an array of paths relative to the plugin root, each `.` or starting with `./`, as the agent's manifest schema requires, and each a skill directory or a directory of them), and a root `SKILL.md` only when the manifest has no `skills` field at all (`"skills": []` is a field, naming nothing) and there is no `skills/` dir; a `SKILL.md` that is not a regular file or is over 1 MiB is skipped, as the agent skips it. A plugin with one such skill is tracked; with several, `--skill <name>` (a skill directory's name or its registered name) picks one, and without it the column is omitted. Each scored row records the tracked id in `meta.skill_tracked`, absent when the column is omitted. A `--skill` selection enters the harness sha as `skill:<registered name>` and `--approve-harness` records it as `_state.json` `harness_skill`, so a changed, added or dropped `--skill` is a harness-gate refusal that names the change; without `--skill` the sha is unchanged. `--skill` is not remembered between passes: the runner command must pass the same one every time. A scenario's declared `metrics` are grade keys on every scored row of the flow, over the union of the cases' declarations: `<id>` only when measured — an unmeasured metric is omitted, never 0 — and `<id>_present` (1 measured; 0 on a case that does not declare it, on an agent-caused failure, or when unavailable, with the reason in `meta.metrics_unavailable`). `hillclimb state-template` declares each `<id>_present` with the other `_present` companions after `pass`, and each metric last as `{id, kind: "float", label, better}`, plus `scale` and `min` when the scenario declares them; every scored row records each flow metric's declaration signature in `meta.metric_sigs`. `hillclimb check`'s headroom takes a float headline's good end from `scale` (higher is better) or `min`, 0 when absent (lower is better), and without a `scale` says so instead of computing one; a row that lacks a declared metric (its `meta.metric_sigs` lacks the id) is a note, not an error — it predates the metric when a row after it carries it, and no scenario declares the metric any more when it comes after the last row that does (variants in order, then file order) — a float outside `[min, scale]` is a warning, and so are a case's rows carrying more than one `meta.assert_sig` (graded under different assertion sets) and rows whose `meta.assert_sig` is not their scenario's now — the scenario target `hillclimb check [<scenario.yaml | dir/>]` was given, else the scenario files the last `--approve-harness` hashed (the `.yaml` keys of `_state.json` `harness_files` whose stem is a case's id), else those `harness_paths` lists (a recorded file is a case's scenario when it has a `prompt:` and its stem is the case's id; a note when nothing is recorded, or names a recorded scenario not found from the current directory or that does not parse, or a case the target holds no scenario for, and skips that case; the remedy's `regrade` target is one that hashes exactly the files the flow was approved over — the scenarios' directory when it holds exactly them, else the case's own file or its directory, else `<scenarios>`). `hillclimb check` exits `0` clean, `1` on any error finding (no warning or note changes it), `2` usage (a target that does not load included); `hillclimb state-template` exits `0`, or `2` on usage or a refusal (a metric id declared differently in two scenarios included); `hillclimb regrade` re-evaluates every selected row from its kept run before any judge call — its deterministic asserts and `expect_denied` hosts with `verify-run`'s evaluation (a changed, added or removed one applied in a default re-grade), its metrics re-measured there (a row no judge re-grades included: a case with no judged assert, an agent-failed row — signatures and `<id>_present: 0`, never a value — and a fill row that needs no comparison; every rewritten row of a case that declares a metric, a re-judged one included, is marked `meta.regrade_remeasured: true`, and each variant's payload entry counts the rows re-measured with no judge call in `remeasured` and those re-evaluated in `reevaluated`); a row whose re-evaluation changes nothing stays byte for byte. It exits `0` when every selected row was rewritten or there was nothing to do, `1` when a row was listed instead (no scenario file for its case in the target, without `--case`; its kept run gone or refused, its kept work dir gone while its case declares a metric or a filesystem assert needs it, an assert the recorded `workspace_fixture` satisfies on its own; an assert whose literal the run scrubbed, on a row graded under another assertion list, since an edit inside a scrubbed literal cannot be applied; a re-judge of an assert scrubbed in the run that this process's scrub does not reproduce, without `--allow-doc-drift`; judged evidence that cannot be recomposed; judged evidence that changed since its grade, without `--rejudge`; current evidence less redacted than the graded document, in both modes unless `--allow-doc-drift`; a scenario with no judged assert left to re-grade; in a fill an assert list, a rubric or a deterministic outcome that does not line up with the scenario as graded, a kept outcome judged against a reference that has changed since, an unreadable regrade file, or a fill that would move `pass`; its re-grade judge-invalid, misaligned or stopped; an open `judge_invalid` slot, which `hillclimb run <target> --flow <dir> --variant <v> --case <id> --reps <rep+1>` re-runs — it is never moved into `results.jsonl`; an assert unchanged since the run that re-evaluates differently from the kept run is never listed: the row keeps the run's own outcome and records `meta.regrade_kept_live`), `1` also for a failure after the first judge call (naming the variants already rewritten), and `2` on usage or a refusal before any judge call (a host `claude` that cannot run the judge isolated, the gate, a held lock, a scenario metric declared differently in two scenarios or from the flow's rows' `meta.metric_sigs`, as on `run`, an evidence refusal anywhere — nothing is written then); `hillclimb freeze-ref` exits `0` when no case was refused (an entry already complete is reported, not refused), `1` when a case was (no good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a different prompt, a missing compose key whose recorded run is gone, a store write that failed, or a run whose judged document cannot be composed, differs from the one its live judge read, or has no live fingerprint to check it against), and `2` on usage (a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, its variant's lock held by a live run). `lint` exits `127` when `python3` is missing (spawn error), and `1` — never `0` — when the scenario loader rejected a file but its findings could not be handed to the linter (an unwritable temp directory); `replay` exits
953
+ **Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `eval` (and `eval report`) exits `0` when the comparison completed — whatever drops the rows show, unless `--fail-on` was given; `1` for a drop at the `--fail-on` level (`possible` or `confirmed`; no gating without the flag — under `possible` a drop includes an `insufficient_refusals` row, one the candidate's excess `semantic_matches` refusals for unavailable evidence took below the rep threshold, which would otherwise hide a drop as `insufficient`), every row `insufficient`, or a judge model that differed across reps; `2` for a usage error or any refusal before the first run (an alias model, an eval dir inside a git work tree, identical arms, the answer-key guard, a scenario input a run would refuse, an unreachable `--fail-on confirmed`, a `--max-budget-usd` refusal — `error.code: "budget_exceeded"`); `3` when an arm snapshot could not be copied or failed its staging preflight. Its `--output-format json` envelope is `{tool, version, command:"eval", ok, evalDir, arms, pins, sections, summary, cost, stoppedEarly, error}`, with `ok` ⇔ exit `0`. `eval --dry-run` runs no agent and creates no eval dir: it exits `0` with the plan (`{tool, version, command:"eval", ok:true, dryRun:true, plan, budget?, error:null}`), `2` for any refusal the real eval would make before its first run (the budget refusal included), `3` as above, or when its temp dir is inside a git work tree or git cannot tell whether it is (set TMPDIR) — its arm snapshots go to a temp dir that is removed afterwards. `hillclimb run` exits `0` when every attempted (case, rep) was scored, `1` when any attempt failed (an `errors.jsonl` row, or a scored row whose trace or copies could not be written), the pass stopped mid-run (rows already written are kept), a baseline pass could not freeze a pairwise case's reference, or `summary.json` could not be written after the pass, and `2` for any refusal before spending: a usage error, the harness gate, no usable agent credential at a selected case's tier (`doctor`'s check, as on `eval`; at `protocol` a login only in the real config dir too, since hillclimb runs protocol with a managed config dir), an alias model, a session that does not declare exactly one `plugins.local_plugins` entry (or an inline session, or cases tuning different plugins), `on_unanswered: prompt`, the `protocol` tier without a managed config dir, a flow-dir or `_state.json` problem, split or duplicate case ids, an unknown `--case`, the answer-key guard, a `harness_paths` entry inside the tuned plugin, an unknown, untracked or ambiguous `--skill` (the message lists the plugin's skills), a tracked skill (or none) that differs from a `meta.skill_tracked` recorded by the variant's existing rows, a requested model or effort for a case that differs from the `meta.model_requested` / `meta.effort` the variant's rows for that case recorded (scored or error rows) (a row recording neither is held to its served `model`), an `--effort` level the pinned model does not offer (or any effort on a model with no effort selector, dated ids included), `xhigh` or `max` with `extended_thinking: false`, `extended_thinking: false` on a model whose baseline entry sets `disallowThinkingDisabled`, a missing or incomplete variant snapshot, a scenario input a run would refuse, a `semantic_pairwise` reference that is missing, damaged or exposed through a mount (the gate a run applies), a host `claude` that cannot run the judge or LLM decider isolated (checked when a scenario uses `semantic_matches` or `semantic_pairwise`, or `on_unanswered: llm` with no decider channel, as on `eval`), or a decider with `--concurrency` above 1, or one scenario metric id declared differently in two scenarios, or declared differently from the signature the flow's existing rows carry for it in `meta.metric_sigs` (adding or removing a metric is allowed; removing one warns); `--dry-run` exits `0` unless such a refusal fires; the harness gate does not refuse a dry run, which reports its status instead. Under `--case`, the per-case refusals (the session, its model pins, `on_unanswered: prompt`, the `protocol` tier, a scenario input, a `semantic_pairwise` reference, the isolation check, and the mounts the answer-key guard checks) cover the selected cases only; the flow-level ones cover every case: every scenario file must parse, every case's baseline must load, the harness gate hashes every case's scenario and session file (an unselected case's inputs it cannot read — missing, a directory, unreadable — or that its unparseable session cannot name, are left out of it; the gate hashes each path's name, so the sha still moves), the answer-key guard keeps every case's scenario and session file unreadable, and the one-plugin rule covers every case whose session parses (an unselected one whose session is inline or does not parse is skipped, with a note). `hillclimb regrade --case` follows the same rule, with the same harness sha. `--timeout-s` bounds the whole attempt: the agent phase through the scenario's `timeout_ms` (lowered to it), and the judge phase through a deadline no judge may start after; an attempt whose agent phase reaches its bound, or whose judge phase reaches the deadline, is an `errors.jsonl` row (`timeout`), never a scored one. So is an attempt whose agent did not send the requested effort: its own session transcript shows a main-loop assistant message sent with another effort, with a value that is not an effort level, or with none (even when the agent then failed; a model with no effort selector may send none), or, on an otherwise valid run whose main loop answered, records no main-loop assistant message or cannot be found (read from the run's config root, or the session's pinned `plugins.config_dir` on hostloop and protocol, by the run's session id): `serving_substitution` with `meta.failure_rule: "effort_not_sent"`. Its envelope's `ok` ⇔ exit `0`; a refusal or a mid-run stop is the shared error envelope, its reason in `error.message` (category `usage` for a refusal, `runtime` for a stop), with the payload keys (`flow`, `variant`, `scheduled`, `scored`, `failed`, `exitCode`, and a dry run's `dryRun`) riding on it; `failed` counts failed (case, rep) slots, so a failed reference freeze or `summary.json` write exits `1` without adding to it. Its `--dry-run` `plan.cost` is the same object `eval --dry-run` emits at `plan.cost` (covered keys `jobs`, `meanUsd`, `p50Usd`, `p95Usd`, `worstObservedUsd`, `lowerBound`, `unpriced`, `pricedRuns`, `thinnest`; `p95Usd` and `worstObservedUsd` are pessimistic figures, never bounds, and `worstObservedUsd` is not the `--max-budget-usd` gate's figure), priced on `eval --dry-run`'s basis exactly — the scenario's runs on this machine at its effective tier and baseline, turn 1, every `hillclimb:` run excluded — so each covered key means the same on both commands. `jobs` is the slots the pass would still run after resuming. A scored row records how its judge ran in `meta.judge_transport` (`assertions[].judgeTransport`'s shape, `{isolation, cliVersion?, strictMcp?}`), or `meta.judge_transports` listing each distinct one when its asserts were judged differently; neither is present when no host judge recorded one, so a round judged under a different isolation can be told apart from a regression. `skill_invoked` is `1` or `0` for whether the run invoked the tracked skill; a blank cell means not measured (no tracked skill, or a record that could not tell), never "not invoked". The tracked skill is matched by the id the agent registers for it, `<plugin>:<name>`, the name being a skill directory's name, or a `SKILL.md`'s frontmatter `name` where the agent reads one (else its directory's name), with every character outside `[a-zA-Z0-9_-]` replaced by `-` (`skills/my.skill/` registers `<plugin>:my-skill`), among the skills the agent loads: `skills/*/` whenever `skills/` exists, each path in the plugin manifest's `skills` field (a string or an array of paths relative to the plugin root, each `.` or starting with `./`, as the agent's manifest schema requires, and each a skill directory or a directory of them), and a root `SKILL.md` only when the manifest has no `skills` field at all (`"skills": []` is a field, naming nothing) and there is no `skills/` dir; a `SKILL.md` that is not a regular file or is over 1 MiB is skipped, as the agent skips it. A plugin with one such skill is tracked; with several, `--skill <name>` (a skill directory's name or its registered name) picks one, and without it the column is omitted. Each scored row records the tracked id in `meta.skill_tracked`, absent when the column is omitted. A `--skill` selection enters the harness sha as `skill:<registered name>` and `--approve-harness` records it as `_state.json` `harness_skill`, so a changed, added or dropped `--skill` is a harness-gate refusal that names the change; without `--skill` the sha is unchanged. `--skill` is not remembered between passes: the runner command must pass the same one every time. A scenario's declared `metrics` are grade keys on every scored row of the flow, over the union of the cases' declarations: `<id>` only when measured — an unmeasured metric is omitted, never 0 — and `<id>_present` (1 measured; 0 on a case that does not declare it, on an agent-caused failure, or when unavailable, with the reason in `meta.metrics_unavailable`). `hillclimb state-template` declares each `<id>_present` with the other `_present` companions after `pass`, and each metric last as `{id, kind: "float", label, better}`, plus `scale` and `min` when the scenario declares them; every scored row records each flow metric's declaration signature in `meta.metric_sigs`. `hillclimb check`'s headroom takes a float headline's good end from `scale` (higher is better) or `min`, 0 when absent (lower is better), and without a `scale` says so instead of computing one; a row that lacks a declared metric (its `meta.metric_sigs` lacks the id) is a note, not an error — it predates the metric when a row after it carries it, and no scenario declares the metric any more when it comes after the last row that does (variants in order, then file order) — a float outside `[min, scale]` is a warning, and so are a case's rows carrying more than one `meta.assert_sig` (graded under different assertion sets) and rows whose `meta.assert_sig` is not their scenario's now — the scenario target `hillclimb check [<scenario.yaml | dir/>]` was given, else the scenario files the last `--approve-harness` hashed (the `.yaml` keys of `_state.json` `harness_files` whose stem is a case's id), else those `harness_paths` lists (a recorded file is a case's scenario when it has a `prompt:` and its stem is the case's id; a note when nothing is recorded, or names a recorded scenario not found from the current directory or that does not parse, or a case the target holds no scenario for, and skips that case; the remedy's `regrade` target is one that hashes exactly the files the flow was approved over — the scenarios' directory when it holds exactly them, else the case's own file or its directory, else `<scenarios>`). `hillclimb check` exits `0` clean, `1` on any error finding (no warning or note changes it), `2` usage (a target that does not load included); `hillclimb state-template` exits `0`, or `2` on usage or a refusal (a metric id declared differently in two scenarios included); `hillclimb regrade` re-evaluates every selected row from its kept run before any judge call — its deterministic asserts and `expect_denied` hosts with `verify-run`'s evaluation (a changed, added or removed one applied in a default re-grade), its metrics re-measured there (a row no judge re-grades included: a case with no judged assert, an agent-failed row — signatures and `<id>_present: 0`, never a value — and a fill row that needs no comparison; every rewritten row of a case that declares a metric, a re-judged one included, is marked `meta.regrade_remeasured: true`, and each variant's payload entry counts the rows re-measured with no judge call in `remeasured` and those re-evaluated in `reevaluated`); a row whose re-evaluation changes nothing stays byte for byte. It exits `0` when every selected row was rewritten or there was nothing to do, `1` when a row was listed instead (no scenario file for its case in the target, without `--case`; its kept run gone or refused, its kept work dir gone while its case declares a metric or a filesystem assert needs it, an assert the recorded `workspace_fixture` satisfies on its own; an assert whose literal the run scrubbed, on a row graded under another assertion list, since an edit inside a scrubbed literal cannot be applied; a re-judge of an assert scrubbed in the run that this process's scrub does not reproduce, without `--allow-scrub-change`; a part of a row's judge input not proven scrubbed with a set covering its run's, without `--allow-scrub-change`; judged evidence that cannot be recomposed; judged evidence that changed since its grade, without `--rejudge`; current evidence less redacted than the graded document, in both modes unless `--allow-scrub-change`; a scenario with no judged assert left to re-grade; in a fill an assert list, a rubric or a deterministic outcome that does not line up with the scenario as graded, a kept outcome judged against a reference that has changed since, an unreadable regrade file, or a fill that would move `pass`; its re-grade judge-invalid, misaligned or stopped; an open `judge_invalid` slot, which `hillclimb run <target> --flow <dir> --variant <v> --case <id> --reps <rep+1>` re-runs — it is never moved into `results.jsonl`; an assert unchanged since the run that re-evaluates differently from the kept run is never listed: the row keeps the run's own outcome and records `meta.regrade_kept_live`), `1` also for a failure after the first judge call (naming the variants already rewritten), and `2` on usage or a refusal before any judge call (a host `claude` that cannot run the judge isolated, the gate, a held lock, a scenario metric declared differently in two scenarios or from the flow's rows' `meta.metric_sigs`, as on `run`, an evidence refusal anywhere — nothing is written then); `hillclimb freeze-ref` exits `0` when no case was refused (an entry already complete is reported, not refused), `1` when a case was (no good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a different prompt, a missing compose key whose recorded run is gone, a missing compose key whose run's scrub set this process cannot prove it covers (re-run the variant), a store write that failed, or a run whose judged document cannot be composed, differs from the one its live judge read, or has no live fingerprint to check it against), and `2` on usage (a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, its variant's lock held by a live run). `lint` exits `127` when `python3` is missing (spawn error), and `1` — never `0` — when the scenario loader rejected a file but its findings could not be handed to the linter (an unwritable temp directory); `replay` exits
940
954
  `2` on a **whole-cassette operational failure** — anything `readCassette` rejects (unreadable, invalid
941
955
  shape, unsupported version, unrecognized assertion key) or any per-file throw, plus the batch loop's
942
956
  own source-resolution failures (`--assert-from`/`--reassert` drift, scenario-parse errors, `--write`
@@ -1157,6 +1171,7 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
1157
1171
  `assertions[]` entry, only `assertionIndex`, `docMatchesLive`, `pass`, `judgeInvalid`, `judgeModel`,
1158
1172
  `judgeModelRequested`, `judgeCostUsd` and `semanticClaims` (its other keys follow the `RunResult` assertion entry, which is not
1159
1173
  pinned field by field); and the error envelope's `error.code` values (`doc_drift`, `unchecked_content`,
1174
+ `task_unverifiable`, `rubric_unverifiable`, `evidence_unverifiable`, `reference_unverifiable`,
1160
1175
  `no_semantic_asserts`),
1161
1176
  `refusals[]` and post-write-failure `runs[]`. The enums are covered as sets: `docMatchesLive`, `change`,
1162
1177
  `authoredCapture.source`, a section's `kind`, `error.code`. **Adding a key or an enum value is MINOR** — a
@@ -458,6 +458,7 @@ export async function cmdHillclimb(args, deps) {
458
458
  approveHarness: p.flags["--approve-harness"] === true,
459
459
  allowDocDrift: p.flags["--allow-doc-drift"] === true,
460
460
  allowUnchecked: p.flags["--allow-unchecked"] === true,
461
+ allowScrubChange: p.flags["--allow-scrub-change"] === true,
461
462
  }, {
462
463
  cwd: process.cwd(),
463
464
  env: process.env,