cowork-harness 4.3.0 → 4.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +4 -4
- package/.claude/skills/cowork-harness/references/assertion-catalog.md +1 -1
- package/.claude/skills/cowork-harness/references/assertions-guide.md +1 -1
- package/.claude/skills/cowork-harness/references/authoring.md +1 -1
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +1 -1
- package/.claude/skills/cowork-harness/references/debugging.md +5 -2
- package/.claude/skills/cowork-harness/references/eval.md +1 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/gotchas.md +1 -1
- package/.claude/skills/cowork-harness/references/hillclimb-recipe.md +1 -1
- package/.claude/skills/cowork-harness/references/hillclimb.md +14 -4
- package/.claude/skills/cowork-harness/references/measurement.md +1 -1
- package/.claude/skills/cowork-harness/references/run-record-replay.md +5 -2
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/semantic-judging.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +93 -0
- package/DESIGN.md +1 -1
- package/README.md +3 -3
- package/SPEC.md +23 -8
- package/dist/hillclimb/cli.js +1 -0
- package/dist/hillclimb/freeze-ref.js +21 -4
- package/dist/hillclimb/regrade.js +61 -10
- package/dist/hillclimb/usage.js +11 -5
- package/dist/refs/cli-usage.js +5 -3
- package/dist/refs/cli.js +33 -4
- package/dist/refs/compose.js +4 -1
- package/dist/refs/store.js +5 -3
- package/dist/run/cassette.js +4 -0
- package/dist/run/chat-result.js +2 -0
- package/dist/run/execute.js +13 -0
- package/dist/run/pairwise-prepass.js +37 -8
- package/dist/run/regrade-usage.js +4 -3
- package/dist/run/regrade.js +168 -4
- package/dist/scrub-set.js +196 -0
- package/docs/ci.md +1 -1
- package/docs/cli.md +73 -15
- package/docs/companion-skill.md +2 -2
- package/docs/hillclimb.md +13 -2
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/regrade.json +45 -6
- package/schema/run-result.json +15 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, asserting artifacts, egress, or sub-agent dispatch, measuring how long each tool call took (toolDurations / trace), or debugging a failed run or verdict from its result.json or transcript. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. Also for comparing two versions of a skill before merging an edit — did it make answers worse? (`eval`: paired A/B, pinned models) — or improving one round by round (`hillclimb`, the `/claude-api hillclimb` runner). NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold / critique / stats / eval / hillclimb commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 4.
|
|
7
|
-
tracks-harness: cowork-harness 4.
|
|
6
|
+
version: 4.4.0
|
|
7
|
+
tracks-harness: cowork-harness 4.4.0 (baseline desktop-2.19675.0)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -26,7 +26,7 @@ allowlist). This skill exists mostly to keep you out of those traps — the *Inv
|
|
|
26
26
|
full landmine catalog in [`references/gotchas.md`](references/gotchas.md) are the highest-value part.
|
|
27
27
|
Read them.
|
|
28
28
|
|
|
29
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.
|
|
29
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 4.4.0` (baseline
|
|
30
30
|
> `desktop-2.19675.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
31
31
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
32
32
|
|
|
@@ -43,7 +43,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
43
43
|
|
|
44
44
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
45
45
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
46
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.
|
|
46
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 4.4.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^4.4.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^4.4.0"`. **Pin `@^4.4.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
47
47
|
|
|
48
48
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
49
49
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertion catalog
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Every `assert:` key with its semantics, and the
|
|
4
4
|
verdict-signal table. Which keys survive `replay` is in [`scenario-schema.md`](./scenario-schema.md#replay-class);
|
|
5
5
|
the scenario and session YAML fields are there too.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Assertions guide
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when choosing assertion keys: the two orthogonal axes and the goal → key map. The full catalog is `assertion-catalog.md`.
|
|
4
4
|
|
|
5
5
|
### Assertions: two orthogonal axes
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Authoring a scenario
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when composing a `scenarios/*.yaml`: session vs scenario, discovery, the fidelity tier, the answer path, `web_fetch`, and scaffold + lint.
|
|
4
4
|
|
|
5
5
|
## Part I — AUTHOR a scenario
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^4` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "4.
|
|
20
|
+
(e.g. `version: "4.4.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -82,7 +82,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
82
82
|
GitHub-hosted runners, no token/Docker/agent:
|
|
83
83
|
|
|
84
84
|
```yaml
|
|
85
|
-
- run: npm i -g "cowork-harness@^4.
|
|
85
|
+
- run: npm i -g "cowork-harness@^4.4.0"
|
|
86
86
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
87
87
|
# no silent false-greens. WITHOUT --strict this
|
|
88
88
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -398,7 +398,7 @@ jobs:
|
|
|
398
398
|
with: { node-version: '24' }
|
|
399
399
|
- uses: actions/setup-python@v5
|
|
400
400
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
401
|
-
- run: npm i -g "cowork-harness@^4.
|
|
401
|
+
- run: npm i -g "cowork-harness@^4.4.0"
|
|
402
402
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
403
403
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
404
404
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -427,7 +427,7 @@ jobs:
|
|
|
427
427
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
428
428
|
fi
|
|
429
429
|
- if: steps.guard.outputs.live == 'true'
|
|
430
|
-
run: npm i -g "cowork-harness@^4.
|
|
430
|
+
run: npm i -g "cowork-harness@^4.4.0"
|
|
431
431
|
- if: steps.guard.outputs.live == 'true'
|
|
432
432
|
run: cowork-harness run scenarios/ --output-format json
|
|
433
433
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Debugging a run
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when a run misbehaved or a green looks wrong: triage, the observability output, and `chat`.
|
|
4
4
|
|
|
5
5
|
## Part III — Debug
|
|
6
6
|
|
|
@@ -40,7 +40,10 @@ rubric changed (or you want another judge model) on a run you already paid for,
|
|
|
40
40
|
kept run without re-running the agent — unlike the tools above it is not token-free (the judge call is its
|
|
41
41
|
spend) — writes the grade beside the run, and says whether the judge read the same document the live judge did.
|
|
42
42
|
Content the live judge never read (a widened `evidence_files` / `include_subagent_text` / `include_fork_results` scope, or a larger
|
|
43
|
-
`--authored-total-bytes`) is refused unless you pass `--allow-unchecked`.
|
|
43
|
+
`--authored-total-bytes`) is refused unless you pass `--allow-unchecked`. A `task_unverifiable` /
|
|
44
|
+
`rubric_unverifiable` / `evidence_unverifiable` / `reference_unverifiable` refusal means a part of the judge's input
|
|
45
|
+
cannot be proven scrubbed with the run's scrub set (typically a pre-4.4 run with an edited rubric): re-run the
|
|
46
|
+
case, or pass `--allow-scrub-change` after checking `COWORK_HARNESS_SCRUB_VALUES` / `_KEYS`.
|
|
44
47
|
|
|
45
48
|
**microvm: "control-protocol write failed" with `env: 'claude': No such file or directory` in the agent
|
|
46
49
|
stderr** usually means the VM never finished provisioning (the agent never reached PATH). Check
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `eval` — paired before/after comparison of a skill edit (EXPERIMENTAL)
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full guide is
|
|
4
4
|
[docs/eval.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/eval.md); this is the part you
|
|
5
5
|
need while running it.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`).
|
|
4
4
|
|
|
5
5
|
> **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
|
|
6
6
|
> plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Gotchas
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The full "✓ passed ≠ correct" landmine catalog.
|
|
4
4
|
|
|
5
5
|
## Gotchas — the "✓ passed ≠ correct" landmines
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Recipe 7 — Climb a skill with `/claude-api hillclimb` and the harness as its runner
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). This page is the loop's procedure, step by step, in the order of the
|
|
5
5
|
`/claude-api hillclimb` guide. Every mechanic (flags, refusals, the gate, `regrade`, `freeze-ref`, exit codes,
|
|
6
6
|
row keys) is in [`hillclimb.md`](hillclimb.md); the setup and the full list of differences from the guide's own
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `hillclimb` — the runner for a `/claude-api hillclimb` loop
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). It needs a `cowork-harness` whose
|
|
4
4
|
`hillclimb --help` lists `--skill` (help goes to stderr). The command reference is
|
|
5
5
|
[docs/cli.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/cli.md); this is the part a loop needs
|
|
6
6
|
while it runs. It covers `run`, `check`, `state-template`, `freeze-ref` and `regrade`.
|
|
@@ -274,7 +274,17 @@ every climb there is finished.
|
|
|
274
274
|
markers than the graded one (a secret the run scrubbed that this process does not), or a changed authored file
|
|
275
275
|
whose graded fingerprint has no marker count, is listed even under `--rejudge`, by path only, with no judge call.
|
|
276
276
|
Set the run's `COWORK_HARNESS_SCRUB_VALUES` / `COWORK_HARNESS_SCRUB_KEYS` and regrade again, or, after checking
|
|
277
|
-
them, `--rejudge --allow-
|
|
277
|
+
them, `--rejudge --allow-scrub-change` grades it anyway (`--allow-doc-drift` does not).
|
|
278
|
+
- **Never a part the run scrubbed, unproven.** A run records a keyed fingerprint of its scrub set (`result.json`
|
|
279
|
+
`scrubSet`; the key is `scrubset.key` beside the runs root). When this process's set provably covers it, every part of
|
|
280
|
+
the judge's input is covered. Otherwise each part must equal the run's own scrubbed record: the pairwise `## Task`
|
|
281
|
+
line (the prompt), each rubric line, each evidence note, and each reference, by its bytes whatever it is called
|
|
282
|
+
(its send must hash to what a live comparison of the run under the same compose key was sent, `refSentSha256`; a
|
|
283
|
+
pre-4.4 grade's stored text must be one that comparison recorded). A reference re-frozen
|
|
284
|
+
since, or a `--fill-refs` reference the run never judged, proves nothing. A row with a part proven neither way is listed, with no judge call. A run from before 4.4 has no fingerprint, so
|
|
285
|
+
its new or edited rubric text is listed (one stderr line says so); so is a run from another machine, or one whose
|
|
286
|
+
token has rotated since. Re-run the case, or pass `--allow-scrub-change` after checking the scrub settings: the
|
|
287
|
+
regrade file then records `scrubAcceptedBy`. Neither `--rejudge` nor `--allow-doc-drift` implies it.
|
|
278
288
|
- **`--rejudge`:** every judged assert of every selected row is re-judged, with the flow's references as they are
|
|
279
289
|
now. Use it after a judge change the triggers above do not see. Not with `--fill-refs`.
|
|
280
290
|
- **`--fill-refs`:** only the `semantic_pairwise` comparisons a row lacks are judged (a reference frozen after the
|
|
@@ -293,7 +303,7 @@ every climb there is finished.
|
|
|
293
303
|
|
|
294
304
|
Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant with rows), `--case ID`
|
|
295
305
|
(repeatable), `--judge-model ID`, `--fill-refs`, `--rejudge`, `--approve-harness`, `--allow-doc-drift`, `--allow-unchecked`,
|
|
296
|
-
`--output-format text|json`, `--dotenv FILE`, `--run-dir DIR`.
|
|
306
|
+
`--allow-scrub-change`, `--output-format text|json`, `--dotenv FILE`, `--run-dir DIR`.
|
|
297
307
|
|
|
298
308
|
- **Everything that can refuse does so before the first judge call**, and then writes nothing: a host `claude`
|
|
299
309
|
that cannot run the judge isolated (asked only when a judge would be called — again under the locks when a
|
|
@@ -323,7 +333,7 @@ Flags: `--flow DIR`, `--variant all|baseline|v<N>` (default `all`: every variant
|
|
|
323
333
|
its literal (one scrubbed value for another) cannot be told from no edit: on a row whose `meta.assert_sig` is not
|
|
324
334
|
the scenario's now, it lists the row, untouched, on every regrade until the case is re-run. One this process cannot
|
|
325
335
|
reproduce is kept unchanged too — never re-evaluated or re-judged over scrubbed evidence; a re-judge it would need
|
|
326
|
-
lists the row (same scrub settings, or `--allow-
|
|
336
|
+
lists the row (same scrub settings, or `--allow-scrub-change` after checking — then the judge sees the RAW rubric
|
|
327
337
|
against the scrubbed evidence, so its grade may not match the live run's) — and named on stderr (an edited one
|
|
328
338
|
takes a re-run). A row no judge re-grades is re-evaluated too: a case with no judged assert, a row whose judged
|
|
329
339
|
asserts all keep their entries, an agent-failed row (it gains the metric signature and `<id>_present: 0`, never a
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Measurement
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it before comparing runs: `--repeat`, `--ablate-skill`, and the hygiene that keeps a batch valid.
|
|
4
4
|
|
|
5
5
|
### Measure — before/after, with/without (`--repeat`, `--ablate-skill`)
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Run, record and lock
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). Read it when running a scenario, recording or placing a cassette, reading verdict signals, checking a background run, or choosing CI lanes.
|
|
4
4
|
|
|
5
5
|
## Part II — RUN, RECORD & LOCK
|
|
6
6
|
|
|
@@ -20,7 +20,10 @@ reworded gate or a `choose:` the run never offered fails here in ~1s instead of
|
|
|
20
20
|
never calls the semantic judge, so a `semantic_matches` or `semantic_pairwise` assert is not re-graded by it; after a rubric change,
|
|
21
21
|
`cowork-harness regrade <run-dir> --scenario <scenario.yaml>` re-grades those against the kept run (the judge call
|
|
22
22
|
is the only spend) and reports whether the judge read the same document the live judge did (widening the evidence
|
|
23
|
-
scope needs `--allow-unchecked`: content the live judge never read is refused otherwise). A
|
|
23
|
+
scope needs `--allow-unchecked`: content the live judge never read is refused otherwise). A part of the judge's
|
|
24
|
+
input that cannot be proven scrubbed with the run's scrub set is refused too: on a run from before 4.4 (no
|
|
25
|
+
`scrubSet` in `result.json`), from another machine, or after a token rotated, unchanged rubric lines re-grade but a
|
|
26
|
+
new or edited one is refused until the case is re-run or `--allow-scrub-change` is passed. A run dir moved or
|
|
24
27
|
downloaded from where it ran is read from where it is; a COPY beside its still-present original is refused (its
|
|
25
28
|
`result.json` names the original's files), so grade the original, or re-run the scenario. Or skip
|
|
26
29
|
the discovery/encode/record dance entirely and answer gates **live during the recording** with
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, replay class, web_fetch, authoring gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 4.4.0`
|
|
4
4
|
(baseline `desktop-2.19675.0`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# `semantic_matches` — the full semantics
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 4.
|
|
3
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`). The one-line summary is in
|
|
4
4
|
[assertion-catalog.md](assertion-catalog.md); this is the whole contract, split out so the catalog stays within the
|
|
5
5
|
agent's single-read size.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 4.
|
|
5
|
+
Tracks `cowork-harness 4.4.0` (baseline `desktop-2.19675.0`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,99 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [4.4.0] — 2026-10-05
|
|
10
|
+
|
|
11
|
+
A security fix: `regrade` and `hillclimb regrade` no longer send a judge a value the run scrubbed. A run now records
|
|
12
|
+
a keyed fingerprint of its scrub set, and a re-grade sends only the parts of the judge's input it can prove are
|
|
13
|
+
scrubbed with a set covering the run's. **A run recorded before 4.4 has no fingerprint**, so re-grading it with new
|
|
14
|
+
or edited judge input (rubric text, an evidence note, a reference re-frozen since) is refused by `regrade` and
|
|
15
|
+
listed by `hillclimb regrade` until the case is re-run or `--allow-scrub-change` is passed.
|
|
16
|
+
|
|
17
|
+
### Security
|
|
18
|
+
|
|
19
|
+
- **A re-grade no longer sends a judge a value the run scrubbed.** Before, `regrade` and `hillclimb regrade`
|
|
20
|
+
(`--rejudge` and `--fill-refs` included) rebuilt the judge's input with the scrub set of the process re-grading.
|
|
21
|
+
When that set was smaller than the run's, a value the run had scrubbed reached the host `claude` judge raw, and
|
|
22
|
+
nothing was listed or refused: the `semantic_pairwise` `## Task` line (the scenario prompt), a new or edited
|
|
23
|
+
rubric line (core `regrade` had no rubric check; `hillclimb regrade` checked only asserts matched to the run's),
|
|
24
|
+
and the frozen pairwise reference, which was sent as stored on every lane, live included. The judge's rationale,
|
|
25
|
+
which can quote its input, is then stored in the regrade file and in `results.jsonl`. Now:
|
|
26
|
+
- A run records a keyed fingerprint of its scrub set, `result.json` `scrubSet` (`{v, keyId, values}`): one
|
|
27
|
+
HMAC-SHA256 per scrubbed string under a random per-installation key, `scrubset.key` beside the runs root (mode
|
|
28
|
+
0600; `~/.cowork-harness/scrubset.key` for the default root), so a runs root uploaded or shared carries no key.
|
|
29
|
+
It holds no value, and without the key there is nothing to brute-force.
|
|
30
|
+
- Before any judge call, a re-grade proves each part it sends: covered when the run's set is a subset of this
|
|
31
|
+
process's under the same key (a grown set included), else by equality with the run's own scrubbed record —
|
|
32
|
+
the task line (checked only when a pairwise comparison will be judged), each rubric line, each evidence-health
|
|
33
|
+
and scratch note, and each reference, by its bytes whatever it is called: its send now must hash to what a live
|
|
34
|
+
comparison of the run under the same compose key was sent (`pairwise[].refSentSha256`, new), or, for a grade
|
|
35
|
+
recorded before 4.4 (whose judge got the stored text raw), the stored text must be one such a comparison
|
|
36
|
+
recorded (`refDocSha256`). A reference re-frozen since, or one the run never judged (a `--fill-refs` column),
|
|
37
|
+
proves nothing. A part proven neither way is refused by `regrade` (exit 2, `error.code` `task_unverifiable`,
|
|
38
|
+
`rubric_unverifiable`, `evidence_unverifiable` or `reference_unverifiable`, in `refusals[]` beside `doc_drift`)
|
|
39
|
+
and listed by `hillclimb regrade` (exit 1),
|
|
40
|
+
naming the parts, never their text: each `refusals[]` entry carries `scrubSet` (why the set is not proven),
|
|
41
|
+
`scrubSetDetail` (the run's own reason, when it recorded one), and `rubricLines`, `evidenceSections` or
|
|
42
|
+
`references`.
|
|
43
|
+
- The new `--allow-scrub-change` flag (both commands) sends such a part anyway, recorded as `scrubAcceptedBy` on
|
|
44
|
+
the regrade file and `runs[]`. `--allow-doc-drift`, `--allow-unchecked` and `--rejudge` never imply it; an
|
|
45
|
+
accepted drift whose authored file lost scrub markers is part of what is proven.
|
|
46
|
+
- Every pairwise comparison scrubs the stored reference with the grading process's set before the judge reads it,
|
|
47
|
+
live runs included, and records how many markers that added (`pairwise[].refRedactions`). A reference document's
|
|
48
|
+
sidecar records `scrubCount`, a count only, and a reference frozen with fewer strings is noted
|
|
49
|
+
(`ref_scrub_weaker`).
|
|
50
|
+
- `hillclimb freeze-ref` and a baseline pass add a missing compose key to a reference only when this process's set
|
|
51
|
+
provably covers the run's; otherwise the case is refused with "re-run the variant". `ref freeze` freezes an
|
|
52
|
+
unchecked document (`--allow-unchecked`) under the same proof, or with `--allow-scrub-change` (which
|
|
53
|
+
`ref verify` refuses, like the other freeze-only flags).
|
|
54
|
+
- An unusable installation key (a symlink, a file readable by others, another owner, an empty file, not a key) is
|
|
55
|
+
warned about once, naming its path and the defect, and recorded on the result as `scrubSetUnavailable`, so a
|
|
56
|
+
re-grade names that cause instead of calling the run older than 4.4. An existing key file is never removed or
|
|
57
|
+
replaced. A new key is written whole before it is linked into place; on a filesystem with no hard links it is
|
|
58
|
+
created exclusively in place instead.
|
|
59
|
+
- An invalid pairwise judge reply that quotes its input is scrubbed before it is warned on stderr or stored, and
|
|
60
|
+
every other warning a pairwise comparison prints (a reference's name, a model id) is scrubbed too.
|
|
61
|
+
|
|
62
|
+
### Upgrade notes
|
|
63
|
+
|
|
64
|
+
- **Cassettes: no re-record needed.** Nothing under `src/runtime`, `src/hostloop`, `src/staging`, `src/session.ts`,
|
|
65
|
+
`baselines/` or `docker/` changed, `latest` still resolves to `desktop-2.19675.0`, and `CASSETTE_VERSION` is still
|
|
66
|
+
14. `src/run/cassette.ts` changed only to leave the new `scrubSet` / `scrubSetUnavailable` result fields unset on
|
|
67
|
+
a replay; a cassette's recorded content and fingerprints are unchanged.
|
|
68
|
+
- **Re-grading a run from before 4.4.** Such a run records no scrub-set fingerprint, so nothing can prove its set
|
|
69
|
+
is covered. Its unchanged task and rubric lines still re-grade, and so does a reference still byte-identical to the
|
|
70
|
+
one its live grade recorded. New or edited rubric text, a changed evidence note, or a reference re-frozen since (or
|
|
71
|
+
never judged by the run, as a `--fill-refs` column is) is refused by `regrade` and listed by
|
|
72
|
+
`hillclimb regrade`, with one summary line naming the remedy: re-run the case (a new run records its set), or
|
|
73
|
+
pass `--allow-scrub-change` after checking the scrub settings. The same applies to a run recorded on another
|
|
74
|
+
machine, and to a run whose auth token has rotated since, since a named key's value is part of the set.
|
|
75
|
+
- **Keep `scrubset.key` out of version control.** The first live run under a runs root creates the installation key
|
|
76
|
+
beside it: `~/.cowork-harness/scrubset.key` for the default root, but the current directory for `--run-dir runs`.
|
|
77
|
+
When the runs root sits in a repository, add `scrubset.key` to its `.gitignore`.
|
|
78
|
+
|
|
79
|
+
### Changed
|
|
80
|
+
|
|
81
|
+
- **`--allow-doc-drift` no longer unlocks `hillclimb regrade`'s scrubbed-literal or less-redacted listings;
|
|
82
|
+
`--allow-scrub-change` replaces it for those.** A row listed because an assert's scrubbed literal cannot be
|
|
83
|
+
reproduced, or because its evidence would be less redacted than the graded document, was released by
|
|
84
|
+
`--allow-doc-drift` in 4.3.0 (with `--rejudge` for the latter). Passing `--allow-doc-drift` now leaves such a row
|
|
85
|
+
listed; pass `--allow-scrub-change` (with `--rejudge` for a less-redacted row) after checking the scrub settings.
|
|
86
|
+
4.3.0's re-grade surfaces are new and experimental, so this change to the flag's meaning ships in a minor release.
|
|
87
|
+
- **What the scrub-set proof does not cover.** Content accepted with `--allow-unchecked`, and the documents of
|
|
88
|
+
`unknown` / `live_refused` asserts (no live fingerprint to compare with), are still protected only by the
|
|
89
|
+
re-grading process's scrub set, as before; `regrade` warns about both before any judge call. So is an authored
|
|
90
|
+
file a drift accepted with `--allow-doc-drift` changed on a run whose set is not proven covered, when its scrub
|
|
91
|
+
markers did not drop: it is sent scrubbed with the current set only.
|
|
92
|
+
|
|
93
|
+
### Fixed
|
|
94
|
+
|
|
95
|
+
- **Corrections to the 4.3.0 notes.** 4.3.0 said a `hillclimb regrade` re-judge over a scrubbed literal "lists the
|
|
96
|
+
row, with `--allow-doc-drift` the only override", and the docs described the drift and less-redacted checks as
|
|
97
|
+
what keeps a value the run scrubbed from the judge. Those checks covered the judged documents only: the pairwise
|
|
98
|
+
task line, an edited or added rubric line and the frozen reference were not covered, and core `regrade` had no
|
|
99
|
+
rubric check. The `ref freeze` docs said the run under test "gets the same redaction" as the stored reference;
|
|
100
|
+
that held for host paths, not for the secret scrub. All of these are fixed above.
|
|
101
|
+
|
|
9
102
|
## [4.3.0] — 2026-10-04
|
|
10
103
|
|
|
11
104
|
Full support for Claude Code's `/claude-api hillclimb` loop, with `cowork-harness hillclimb` as its runner,
|
package/DESIGN.md
CHANGED
|
@@ -208,7 +208,7 @@ The policy that produces those `allow`/`deny` responses is the **Decider** seam
|
|
|
208
208
|
> Cowork system-prompt fingerprint all unchanged from `1.20186.0`, with the staged VM ELF re-synced
|
|
209
209
|
> 2.1.202 → 2.1.205 — and the live pass of that era was deliberately **not** restamped onto it.)
|
|
210
210
|
|
|
211
|
-
> **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **
|
|
211
|
+
> **Scope of that claim.** `2026-10-03 / desktop-2.19675.0` is the baseline carrying the latest live pass, run against agent **2.1.286** (the staged Linux ELF for `container` and `microvm`, and the staged native macOS build `f2326db61802` for `hostloop`); the `protocol` tier runs the HOST `claude`, **2.1.289** for the final pass. No baselines have shipped since. **The 4.4.0 pass, 2026-10-05, on the release-prep commit `e9c516df`** (same baseline and staged agent; host `claude` 2.1.289 for `protocol`): (i) `vitest list --staticParse=false` listed all 23 live tests with no `SKIPPED` warning when the token was exported, and 6 with four warnings when none was resolvable; (ii) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip; (iii) the auto-memory check on all three tiers, `Tests 2 passed (2)` each; (iv) the companion skill's router check, three prompts on `container`: the authoring prompt read `authoring.md` and wrote a scenario that passes `lint`, the tool-timing prompt read `measurement.md` and answered with `trace … --view tool-durations`, and the debugging prompt read no reference and answered with correct clarifying questions, so `debugging.md` routing was not exercised in this pass; (v) the re-grade scrub-set check: a live run on the default runs root created `scrubset.key` beside the runs directory (mode 0600) and recorded `scrubSet` (`v: 1`) with the configured scrub literal absent from `result.json`; a re-grade of that run with an edited rubric line and a smaller scrub set was refused as `rubric_unverifiable` (exit 2) naming only the edited line, with no judge call and the run directory byte-identical; and a re-grade of a run recorded before 4.4 with an edited rubric line was refused with the `legacy` reason and the one-line remedy, again with no judge call and the run directory byte-identical. Live spend about **$2.83** recorded. **The 4.3.0 pass, on the release commit `8a5e479f`:** (1) the live suite on `protocol`, `container` and `hostloop`: 23 passed, one pre-registered skip (the auto-memory prerequisites check, which runs only under `COWORK_LIVE_REQUIRE=1`); before it ran, `vitest list --staticParse=false` listed all 23 with no `SKIPPED` warning, and 6 with four warnings when no token was resolvable, so a silent skip would have shown; (2) the auto-memory check on all three tiers, `Tests 2 passed (2)` each, every init frame without `memory_paths`; (3) the graders' pinned effort: every judged assert that called the judge in the live hillclimb acceptance runs (45, Opus 4.8: 25 `semantic_matches`, 20 `semantic_pairwise`; host `claude` 2.1.288 and 2.1.289) records `judgeTransport.effort: "high"`; the five pairwise asserts without a `judgeTransport` are the reference variant's own rows, which call no judge. **Earlier in the same release, on the release candidate's tree:** (4) a `microvm` smoke on a freshly booted VM (agent 2.1.286); (5) `boundary-check`, 6/6 constraints enforced; (6) an `eval` A/A at `--reps 4` on `container`, every row `no detectable change`, its report rebuilt byte-identical; (7) the companion skill's router check, three prompts on `container`: routing was correct on all three; one answer was wrong because a reference showed `artifact_text`'s `contains` without its list shape (fixed in the references, with a test that keeps every list-typed field shown as a list), and the same prompt's verdict went red on a false `host_path_leak`, because a reference spelled out the host roots that check looks for and the agent echoed it (fixed in the skill's files and in the check, which no longer counts a literal that comes from the plugin's or a local skill's own staged files); re-run after the fixes, the prompt routed correctly, wrote a scenario that passes `lint` and loads, and carried no `host_path_leak`; (8) the `hostloop` detached-agent kill check: the first pass found the workspace sidecar container, its `docker run` client and its network surviving SIGINT, which was fixed, and the re-run passed on a normal run, SIGINT to the harness process and SIGINT to the whole process group (exit 130, nothing left behind); (9) `vm delete`'s usage refusal; (10) `chat` under a real terminal, on `protocol`: Ctrl-C mid-turn exits 130 in about 2 s with nothing left running and no `result.json`; `/exit` followed by Ctrl-C exits 130 after writing the turn's `result.json`; Ctrl-C at the idle prompt ends the session normally (exit 0, `result.json` written); SIGHUP mid-turn exits 129 with nothing left running; and closing the terminal window mid-turn (run by hand) leaves no harness or agent process and no `result.json`. Total live spend about **$11.20**, including the re-runs. **Scope-out, so this is not read as more than it is.** (a) A live pass verifies observed behaviour, not the whole spawn contract by construction, and these checks are model-dependent, so a single red is evidence of model variance until a re-run says otherwise. (b) **CI does not live-validate anything.** Its "scenario suite (… live inference)" job is skipped as a whole without an `ANTHROPIC_API_KEY` repository secret, and none is set, so it shows as *skipped*. Before 2026-09 it instead ran with every real step skipped and reported *success*, so an older green check there means nothing ran. This note, not CI, is the live evidence. Separately and not a live matter: all four committed cassettes were **re-recorded** on 2026-10-03 against `2.19675.0` with auto-memory off, each on its original model: `example-pdf-skill` and `test/fixtures/tool-call-dispatch/dispatch-shell.cassette.json` (`container`, agent 2.1.286), `hostloop-computer-links` (`hostloop`, agent 2.1.286) and `example-multiselect-gate` (`protocol`, which runs the host CLI, here 2.1.287), about $0.75 in all. Their init frames carry no `memory_paths`. The recorder's own scrub (now applied to every recording) removed, across the four, the account's model menu, a subscription account's `rate_limit_info` (from `example-pdf-skill`, `dispatch-shell` and `hostloop-computer-links`) and the agent's sub-agent hand-back frame (from `dispatch-shell`). `verify-cassettes` exits 0 on all four, with one accepted `unscanned` entry (`example-pdf-skill`'s uploaded artifact body, too large to commit). Re-stamp this paragraph, naming the baseline, whenever a live pass is actually re-run.
|
|
212
212
|
|
|
213
213
|
> The staged agent ELF is unchanged (2.1.181) across the 1.14271.0→1.15200.0 asar bump, and 2.1.187 across the 1.15200.0→1.15962.0 bump. The live scenario suite (`protocol` + `container` tiers) was re-run against the 1.15200.0 baseline; the 1.15962.0 bump was verified via asar analysis (content byte-identical: host-loop generator, system prompt, identity string, gates, and egress domains all unchanged) plus a full local test suite pass. The 1.15962.1→1.17377.1 bump moved the staged agent to **2.1.197** and added `api.claude.ai` to the egress allowlist; re-verified via `sync` (no unknown deltas) plus a manual asar spot-check of the reconstructed prompt content (substantively unchanged — see the Parity entry in CHANGELOG.md) and a full live scenario-suite pass (`protocol` + `container` tiers).
|
|
214
214
|
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ npm ci && npm run build
|
|
|
36
36
|
node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
(Installing globally — `npm install -g "cowork-harness@^4.
|
|
39
|
+
(Installing globally — `npm install -g "cowork-harness@^4.4.0"` — gives you the `cowork-harness` CLI for your own
|
|
40
40
|
scenarios and cassettes; the bundled example above also replays from a global install — see the `$(npm root -g)` path below.)
|
|
41
41
|
|
|
42
42
|
Full setup → [Quick start](./docs/cli.md#quick-start).
|
|
@@ -49,8 +49,8 @@ Three ways to use this project. Each row is the whole hook — follow the link f
|
|
|
49
49
|
|
|
50
50
|
| I want to… | Start here | Needs |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.
|
|
53
|
-
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.
|
|
52
|
+
| **Run scenarios myself** from a terminal | **[docs/cli.md](./docs/cli.md)**<br><br>`npm i -g "cowork-harness@^4.4.0"`<br>`cowork-harness replay examples/replays/example-pdf-skill.cassette.json` | Node ≥ 22. The replay demo above is token-free and needs nothing else; live tiers above `protocol` need Docker + a staged agent binary |
|
|
53
|
+
| **Have Claude Code drive it** for me | **[docs/companion-skill.md](./docs/companion-skill.md)**<br><br>`/plugin marketplace add yaniv-golan/cowork-harness`<br>`/plugin install cowork-harness@cowork-harness` | Claude Code. The skill self-bootstraps the CLI via `npx "cowork-harness@^4.4.0"` |
|
|
54
54
|
| **Gate my skill in CI** | **[docs/ci.md](./docs/ci.md)**<br><br>`- uses: yaniv-golan/cowork-harness@v4`<br>` with: { command: replay, path: cassettes/ }` | Nothing for the token-free gate; the live lane needs a self-hosted runner with Docker + an agent binary |
|
|
55
55
|
|
|
56
56
|
Improving a skill round by round with Claude Code's `/claude-api hillclimb` loop? The harness is its runner, with
|
package/SPEC.md
CHANGED
|
@@ -606,7 +606,7 @@ there are three families:
|
|
|
606
606
|
`dryRun: true` with the estimate at `plan.cost`, where `eval --dry-run` puts it), `hillclimb check` (`reading`, `disclaimer`, `profile`, `findings`, `errors`, `notes`, `warnings`),
|
|
607
607
|
`hillclimb state-template` (`state`, `metrics_md`, and with `--flow` `metrics_md_file`; `notes` when a column was left
|
|
608
608
|
undeclared, experimental), `hillclimb freeze-ref` (payload experimental, §12: `frozen`, `added`, `exists`, `refused`,
|
|
609
|
-
`exitCode`), `hillclimb regrade` (payload experimental, §12: `flow`, `variants` (each `{variant, rewritten, judged, rebuilt, ownRefOnly, listedAfterJudge, judgeUsd?, judgeUnpriced, judgeStopped, listed, reevaluated, remeasured, agentFailed, evidenceChanged, regradeFiles, backup?}`), `exitCode`; so are `<variant>/regrade.md`, the `regrade-<sha16>.bak.jsonl` backups and the rows' `meta.regrade_*` and `meta.assert_sig` keys; a row a judged entry of which the current harness would show its judge different evidence than it was graded on is in `evidenceChanged` and is never silently kept or re-graded — listed and kept without `--rejudge`, re-judged on the current evidence with it, recording `evidence_changed` and both document hashes in `meta.regrade_evidence` — except a row whose current document would be less redacted than the graded one, listed in both modes unless `--allow-
|
|
609
|
+
`exitCode`), `hillclimb regrade` (payload experimental, §12: `flow`, `variants` (each `{variant, rewritten, judged, rebuilt, ownRefOnly, listedAfterJudge, judgeUsd?, judgeUnpriced, judgeStopped, listed, reevaluated, remeasured, agentFailed, evidenceChanged, regradeFiles, backup?}`), `exitCode`; so are `<variant>/regrade.md`, the `regrade-<sha16>.bak.jsonl` backups and the rows' `meta.regrade_*` and `meta.assert_sig` keys; a row a judged entry of which the current harness would show its judge different evidence than it was graded on is in `evidenceChanged` and is never silently kept or re-graded — listed and kept without `--rejudge`, re-judged on the current evidence with it, recording `evidence_changed` and both document hashes in `meta.regrade_evidence` — except a row whose current document would be less redacted than the graded one, listed in both modes unless `--allow-scrub-change` is also given; and a row with a part of its judge input not proven scrubbed with a set covering its run's, listed unless `--allow-scrub-change`), `verify-cassettes` (§11.1), `doctor` (§11.2), `rehash`,
|
|
610
610
|
`answer` (`gate`, `answers`), `fixture export` (exit `0` written, `2` usage or refusal; payload experimental,
|
|
611
611
|
§12: `message`, `written`, `skipped`, `refused` (`[]`), `notes`, `bytes`, `outputsDir`, `partial`, `result`; a
|
|
612
612
|
refusal is the error envelope, its message in `error.message`, carrying `refused[]` and whichever of the other
|
|
@@ -711,6 +711,19 @@ Two evidence refusals are decided before any judge call, for every run dir:
|
|
|
711
711
|
secrets) are never unchecked content, so a smaller budget — which only drops or truncates file content, and may
|
|
712
712
|
add a health note saying so — is never refused as unchecked; content it truncates makes that assert refuse its
|
|
713
713
|
own evidence. An assert whose evidence will be refused sends nothing and is not measured.
|
|
714
|
+
- **Scrub coverage.** A run records `scrubSet` (`{v, keyId, values}`: one HMAC-SHA256 per string its scrub used,
|
|
715
|
+
under a per-installation key, `scrubset.key`, beside the runs root). When this process's set provably covers the
|
|
716
|
+
run's (every recorded HMAC recomputed under the same key), every part of the judge's input is covered. Otherwise
|
|
717
|
+
each part sent must equal the run's own scrubbed record: the `semantic_pairwise` task line (only when a pairwise
|
|
718
|
+
comparison is judged), each rubric line, each evidence-health or scratch note, an accepted drift's authored file
|
|
719
|
+
that carries fewer scrub markers, and each reference: its send must hash to the `refSentSha256` any live comparison of the
|
|
720
|
+
run recorded under the same compose key, whatever the reference is called (a pre-4.4 grade, which sent the stored
|
|
721
|
+
text raw: the stored text's sha256 must equal such a comparison's `refDocSha256`). An
|
|
722
|
+
unproven part refuses unless `--allow-scrub-change` is passed, with `task_unverifiable`, `rubric_unverifiable`,
|
|
723
|
+
`evidence_unverifiable` or `reference_unverifiable`; accepted, the run's entry and file record `scrubAcceptedBy`.
|
|
724
|
+
Neither `--allow-doc-drift` nor `--allow-unchecked` implies it. A run with no `scrubSet` (harness < 4.4), another
|
|
725
|
+
key, or a set lacking a string the run scrubbed (a rotated token included) proves nothing: only parts equal to its
|
|
726
|
+
record are sent.
|
|
714
727
|
|
|
715
728
|
Not checked: a run in which no live assert recorded a `judgedDoc` has no rebuilt document to measure against — its
|
|
716
729
|
asserts are `unknown` or `live_refused`, neither drift- nor secret-checked, and are warned about before the judge
|
|
@@ -720,11 +733,12 @@ so its extra content can refuse. The envelope is scrubbed with the same secret s
|
|
|
720
733
|
**Exit codes:** `0` every re-graded assert passes · `1` any fails or is judge-invalid · `2`, with three meanings:
|
|
721
734
|
a usage error; a refusal before any judge call (a multi-turn, partial, replay or chat run dir, a pruned work dir,
|
|
722
735
|
a missing transcript sidecar, a run that did not record `authoredCapture` without `--authored-total-bytes`, an
|
|
723
|
-
alias judge model, a scenario with neither a `semantic_matches` nor a `semantic_pairwise` assert (`error.code: "no_semantic_asserts"`), and the
|
|
736
|
+
alias judge model, a scenario with neither a `semantic_matches` nor a `semantic_pairwise` assert (`error.code: "no_semantic_asserts"`), and the evidence and scrub refusals above); or a failure
|
|
724
737
|
writing a regrade file after earlier run dirs were graded. Each is the shared error envelope. The evidence
|
|
725
738
|
refusals are collected over every run dir and carry `error.code` — `doc_drift` when any dir drifted, else
|
|
726
|
-
`unchecked_content
|
|
727
|
-
|
|
739
|
+
`unchecked_content`, else the first of `task_unverifiable`, `rubric_unverifiable`, `evidence_unverifiable`,
|
|
740
|
+
`reference_unverifiable` — and a top-level `refusals[]`, one entry per run dir and code: `{runDir, code,
|
|
741
|
+
uncheckedCount?, uncheckedSections?, liveDocDrift?, scrubSet?, scrubSetDetail?, rubricLines?, evidenceSections?, references?}`; every other refusal stops at the first run dir that
|
|
728
742
|
fires it and carries no code. So `refusals[]` is complete only when no other refusal fires: a batch with a
|
|
729
743
|
drifted dir and a later dir refused for another reason (a pruned work dir, say) reports only the latter, with no
|
|
730
744
|
code and no `refusals[]`.
|
|
@@ -878,11 +892,11 @@ assertions (never user-authored themselves):
|
|
|
878
892
|
"results":[], // [] except record's post-run refusal: the refused run, beside the non-null error
|
|
879
893
|
"budget?": { /* §11 --max-budget-usd marker — present when a pre-flight ran */ },
|
|
880
894
|
"error": { "category": "usage|unanswered|boundary|runtime|internal", "message": "string", "hint?": "string",
|
|
881
|
-
"code?": "budget_exceeded|doc_drift|unchecked_content|no_semantic_asserts", "budget?": { /* §11 --max-budget-usd */ } } }
|
|
895
|
+
"code?": "budget_exceeded|doc_drift|unchecked_content|task_unverifiable|rubric_unverifiable|evidence_unverifiable|reference_unverifiable|no_semantic_asserts", "budget?": { /* §11 --max-budget-usd */ } } }
|
|
882
896
|
```
|
|
883
897
|
`error.code` narrows a category, never replaces it: `budget_exceeded` is the `--max-budget-usd` refusal (§11);
|
|
884
|
-
`doc_drift` and `unchecked_content` are `regrade`'s evidence refusals
|
|
885
|
-
top-level `refusals[]` (see `regrade` above); `no_semantic_asserts` is `regrade`'s `usage` refusal of a scenario
|
|
898
|
+
`doc_drift` and `unchecked_content` are `regrade`'s evidence refusals and the four `*_unverifiable` codes its scrub
|
|
899
|
+
refusals, whose error envelope also carries a top-level `refusals[]` (see `regrade` above); `no_semantic_asserts` is `regrade`'s `usage` refusal of a scenario
|
|
886
900
|
with neither a `semantic_matches` nor a `semantic_pairwise` assert (nothing to re-grade).
|
|
887
901
|
Categories come from TYPED errors (`UnansweredError`→`unanswered`, `BoundaryError`→`boundary`).
|
|
888
902
|
`results` is `[]` with one exception: when `record` refuses to write a cassette after the agent finished,
|
|
@@ -936,7 +950,7 @@ prints the same exclusion warning the run prints. `verify-run` follows the same
|
|
|
936
950
|
not exist or is a file, or a scenario file that does not load, is `usage`; a directory holding no completed
|
|
937
951
|
run stays `runtime`. `answer` splits the same way: a directory or gate that is not there is `usage`; a gate
|
|
938
952
|
request that exists but cannot be read or parsed, or an answer that cannot be written, is `runtime`.
|
|
939
|
-
**Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `eval` (and `eval report`) exits `0` when the comparison completed — whatever drops the rows show, unless `--fail-on` was given; `1` for a drop at the `--fail-on` level (`possible` or `confirmed`; no gating without the flag — under `possible` a drop includes an `insufficient_refusals` row, one the candidate's excess `semantic_matches` refusals for unavailable evidence took below the rep threshold, which would otherwise hide a drop as `insufficient`), every row `insufficient`, or a judge model that differed across reps; `2` for a usage error or any refusal before the first run (an alias model, an eval dir inside a git work tree, identical arms, the answer-key guard, a scenario input a run would refuse, an unreachable `--fail-on confirmed`, a `--max-budget-usd` refusal — `error.code: "budget_exceeded"`); `3` when an arm snapshot could not be copied or failed its staging preflight. Its `--output-format json` envelope is `{tool, version, command:"eval", ok, evalDir, arms, pins, sections, summary, cost, stoppedEarly, error}`, with `ok` ⇔ exit `0`. `eval --dry-run` runs no agent and creates no eval dir: it exits `0` with the plan (`{tool, version, command:"eval", ok:true, dryRun:true, plan, budget?, error:null}`), `2` for any refusal the real eval would make before its first run (the budget refusal included), `3` as above, or when its temp dir is inside a git work tree or git cannot tell whether it is (set TMPDIR) — its arm snapshots go to a temp dir that is removed afterwards. `hillclimb run` exits `0` when every attempted (case, rep) was scored, `1` when any attempt failed (an `errors.jsonl` row, or a scored row whose trace or copies could not be written), the pass stopped mid-run (rows already written are kept), a baseline pass could not freeze a pairwise case's reference, or `summary.json` could not be written after the pass, and `2` for any refusal before spending: a usage error, the harness gate, no usable agent credential at a selected case's tier (`doctor`'s check, as on `eval`; at `protocol` a login only in the real config dir too, since hillclimb runs protocol with a managed config dir), an alias model, a session that does not declare exactly one `plugins.local_plugins` entry (or an inline session, or cases tuning different plugins), `on_unanswered: prompt`, the `protocol` tier without a managed config dir, a flow-dir or `_state.json` problem, split or duplicate case ids, an unknown `--case`, the answer-key guard, a `harness_paths` entry inside the tuned plugin, an unknown, untracked or ambiguous `--skill` (the message lists the plugin's skills), a tracked skill (or none) that differs from a `meta.skill_tracked` recorded by the variant's existing rows, a requested model or effort for a case that differs from the `meta.model_requested` / `meta.effort` the variant's rows for that case recorded (scored or error rows) (a row recording neither is held to its served `model`), an `--effort` level the pinned model does not offer (or any effort on a model with no effort selector, dated ids included), `xhigh` or `max` with `extended_thinking: false`, `extended_thinking: false` on a model whose baseline entry sets `disallowThinkingDisabled`, a missing or incomplete variant snapshot, a scenario input a run would refuse, a `semantic_pairwise` reference that is missing, damaged or exposed through a mount (the gate a run applies), a host `claude` that cannot run the judge or LLM decider isolated (checked when a scenario uses `semantic_matches` or `semantic_pairwise`, or `on_unanswered: llm` with no decider channel, as on `eval`), or a decider with `--concurrency` above 1, or one scenario metric id declared differently in two scenarios, or declared differently from the signature the flow's existing rows carry for it in `meta.metric_sigs` (adding or removing a metric is allowed; removing one warns); `--dry-run` exits `0` unless such a refusal fires; the harness gate does not refuse a dry run, which reports its status instead. Under `--case`, the per-case refusals (the session, its model pins, `on_unanswered: prompt`, the `protocol` tier, a scenario input, a `semantic_pairwise` reference, the isolation check, and the mounts the answer-key guard checks) cover the selected cases only; the flow-level ones cover every case: every scenario file must parse, every case's baseline must load, the harness gate hashes every case's scenario and session file (an unselected case's inputs it cannot read — missing, a directory, unreadable — or that its unparseable session cannot name, are left out of it; the gate hashes each path's name, so the sha still moves), the answer-key guard keeps every case's scenario and session file unreadable, and the one-plugin rule covers every case whose session parses (an unselected one whose session is inline or does not parse is skipped, with a note). `hillclimb regrade --case` follows the same rule, with the same harness sha. `--timeout-s` bounds the whole attempt: the agent phase through the scenario's `timeout_ms` (lowered to it), and the judge phase through a deadline no judge may start after; an attempt whose agent phase reaches its bound, or whose judge phase reaches the deadline, is an `errors.jsonl` row (`timeout`), never a scored one. So is an attempt whose agent did not send the requested effort: its own session transcript shows a main-loop assistant message sent with another effort, with a value that is not an effort level, or with none (even when the agent then failed; a model with no effort selector may send none), or, on an otherwise valid run whose main loop answered, records no main-loop assistant message or cannot be found (read from the run's config root, or the session's pinned `plugins.config_dir` on hostloop and protocol, by the run's session id): `serving_substitution` with `meta.failure_rule: "effort_not_sent"`. Its envelope's `ok` ⇔ exit `0`; a refusal or a mid-run stop is the shared error envelope, its reason in `error.message` (category `usage` for a refusal, `runtime` for a stop), with the payload keys (`flow`, `variant`, `scheduled`, `scored`, `failed`, `exitCode`, and a dry run's `dryRun`) riding on it; `failed` counts failed (case, rep) slots, so a failed reference freeze or `summary.json` write exits `1` without adding to it. Its `--dry-run` `plan.cost` is the same object `eval --dry-run` emits at `plan.cost` (covered keys `jobs`, `meanUsd`, `p50Usd`, `p95Usd`, `worstObservedUsd`, `lowerBound`, `unpriced`, `pricedRuns`, `thinnest`; `p95Usd` and `worstObservedUsd` are pessimistic figures, never bounds, and `worstObservedUsd` is not the `--max-budget-usd` gate's figure), priced on `eval --dry-run`'s basis exactly — the scenario's runs on this machine at its effective tier and baseline, turn 1, every `hillclimb:` run excluded — so each covered key means the same on both commands. `jobs` is the slots the pass would still run after resuming. A scored row records how its judge ran in `meta.judge_transport` (`assertions[].judgeTransport`'s shape, `{isolation, cliVersion?, strictMcp?}`), or `meta.judge_transports` listing each distinct one when its asserts were judged differently; neither is present when no host judge recorded one, so a round judged under a different isolation can be told apart from a regression. `skill_invoked` is `1` or `0` for whether the run invoked the tracked skill; a blank cell means not measured (no tracked skill, or a record that could not tell), never "not invoked". The tracked skill is matched by the id the agent registers for it, `<plugin>:<name>`, the name being a skill directory's name, or a `SKILL.md`'s frontmatter `name` where the agent reads one (else its directory's name), with every character outside `[a-zA-Z0-9_-]` replaced by `-` (`skills/my.skill/` registers `<plugin>:my-skill`), among the skills the agent loads: `skills/*/` whenever `skills/` exists, each path in the plugin manifest's `skills` field (a string or an array of paths relative to the plugin root, each `.` or starting with `./`, as the agent's manifest schema requires, and each a skill directory or a directory of them), and a root `SKILL.md` only when the manifest has no `skills` field at all (`"skills": []` is a field, naming nothing) and there is no `skills/` dir; a `SKILL.md` that is not a regular file or is over 1 MiB is skipped, as the agent skips it. A plugin with one such skill is tracked; with several, `--skill <name>` (a skill directory's name or its registered name) picks one, and without it the column is omitted. Each scored row records the tracked id in `meta.skill_tracked`, absent when the column is omitted. A `--skill` selection enters the harness sha as `skill:<registered name>` and `--approve-harness` records it as `_state.json` `harness_skill`, so a changed, added or dropped `--skill` is a harness-gate refusal that names the change; without `--skill` the sha is unchanged. `--skill` is not remembered between passes: the runner command must pass the same one every time. A scenario's declared `metrics` are grade keys on every scored row of the flow, over the union of the cases' declarations: `<id>` only when measured — an unmeasured metric is omitted, never 0 — and `<id>_present` (1 measured; 0 on a case that does not declare it, on an agent-caused failure, or when unavailable, with the reason in `meta.metrics_unavailable`). `hillclimb state-template` declares each `<id>_present` with the other `_present` companions after `pass`, and each metric last as `{id, kind: "float", label, better}`, plus `scale` and `min` when the scenario declares them; every scored row records each flow metric's declaration signature in `meta.metric_sigs`. `hillclimb check`'s headroom takes a float headline's good end from `scale` (higher is better) or `min`, 0 when absent (lower is better), and without a `scale` says so instead of computing one; a row that lacks a declared metric (its `meta.metric_sigs` lacks the id) is a note, not an error — it predates the metric when a row after it carries it, and no scenario declares the metric any more when it comes after the last row that does (variants in order, then file order) — a float outside `[min, scale]` is a warning, and so are a case's rows carrying more than one `meta.assert_sig` (graded under different assertion sets) and rows whose `meta.assert_sig` is not their scenario's now — the scenario target `hillclimb check [<scenario.yaml | dir/>]` was given, else the scenario files the last `--approve-harness` hashed (the `.yaml` keys of `_state.json` `harness_files` whose stem is a case's id), else those `harness_paths` lists (a recorded file is a case's scenario when it has a `prompt:` and its stem is the case's id; a note when nothing is recorded, or names a recorded scenario not found from the current directory or that does not parse, or a case the target holds no scenario for, and skips that case; the remedy's `regrade` target is one that hashes exactly the files the flow was approved over — the scenarios' directory when it holds exactly them, else the case's own file or its directory, else `<scenarios>`). `hillclimb check` exits `0` clean, `1` on any error finding (no warning or note changes it), `2` usage (a target that does not load included); `hillclimb state-template` exits `0`, or `2` on usage or a refusal (a metric id declared differently in two scenarios included); `hillclimb regrade` re-evaluates every selected row from its kept run before any judge call — its deterministic asserts and `expect_denied` hosts with `verify-run`'s evaluation (a changed, added or removed one applied in a default re-grade), its metrics re-measured there (a row no judge re-grades included: a case with no judged assert, an agent-failed row — signatures and `<id>_present: 0`, never a value — and a fill row that needs no comparison; every rewritten row of a case that declares a metric, a re-judged one included, is marked `meta.regrade_remeasured: true`, and each variant's payload entry counts the rows re-measured with no judge call in `remeasured` and those re-evaluated in `reevaluated`); a row whose re-evaluation changes nothing stays byte for byte. It exits `0` when every selected row was rewritten or there was nothing to do, `1` when a row was listed instead (no scenario file for its case in the target, without `--case`; its kept run gone or refused, its kept work dir gone while its case declares a metric or a filesystem assert needs it, an assert the recorded `workspace_fixture` satisfies on its own; an assert whose literal the run scrubbed, on a row graded under another assertion list, since an edit inside a scrubbed literal cannot be applied; a re-judge of an assert scrubbed in the run that this process's scrub does not reproduce, without `--allow-doc-drift`; judged evidence that cannot be recomposed; judged evidence that changed since its grade, without `--rejudge`; current evidence less redacted than the graded document, in both modes unless `--allow-doc-drift`; a scenario with no judged assert left to re-grade; in a fill an assert list, a rubric or a deterministic outcome that does not line up with the scenario as graded, a kept outcome judged against a reference that has changed since, an unreadable regrade file, or a fill that would move `pass`; its re-grade judge-invalid, misaligned or stopped; an open `judge_invalid` slot, which `hillclimb run <target> --flow <dir> --variant <v> --case <id> --reps <rep+1>` re-runs — it is never moved into `results.jsonl`; an assert unchanged since the run that re-evaluates differently from the kept run is never listed: the row keeps the run's own outcome and records `meta.regrade_kept_live`), `1` also for a failure after the first judge call (naming the variants already rewritten), and `2` on usage or a refusal before any judge call (a host `claude` that cannot run the judge isolated, the gate, a held lock, a scenario metric declared differently in two scenarios or from the flow's rows' `meta.metric_sigs`, as on `run`, an evidence refusal anywhere — nothing is written then); `hillclimb freeze-ref` exits `0` when no case was refused (an entry already complete is reported, not refused), `1` when a case was (no good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a different prompt, a missing compose key whose recorded run is gone, a store write that failed, or a run whose judged document cannot be composed, differs from the one its live judge read, or has no live fingerprint to check it against), and `2` on usage (a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, its variant's lock held by a live run). `lint` exits `127` when `python3` is missing (spawn error), and `1` — never `0` — when the scenario loader rejected a file but its findings could not be handed to the linter (an unwritable temp directory); `replay` exits
|
|
953
|
+
**Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `eval` (and `eval report`) exits `0` when the comparison completed — whatever drops the rows show, unless `--fail-on` was given; `1` for a drop at the `--fail-on` level (`possible` or `confirmed`; no gating without the flag — under `possible` a drop includes an `insufficient_refusals` row, one the candidate's excess `semantic_matches` refusals for unavailable evidence took below the rep threshold, which would otherwise hide a drop as `insufficient`), every row `insufficient`, or a judge model that differed across reps; `2` for a usage error or any refusal before the first run (an alias model, an eval dir inside a git work tree, identical arms, the answer-key guard, a scenario input a run would refuse, an unreachable `--fail-on confirmed`, a `--max-budget-usd` refusal — `error.code: "budget_exceeded"`); `3` when an arm snapshot could not be copied or failed its staging preflight. Its `--output-format json` envelope is `{tool, version, command:"eval", ok, evalDir, arms, pins, sections, summary, cost, stoppedEarly, error}`, with `ok` ⇔ exit `0`. `eval --dry-run` runs no agent and creates no eval dir: it exits `0` with the plan (`{tool, version, command:"eval", ok:true, dryRun:true, plan, budget?, error:null}`), `2` for any refusal the real eval would make before its first run (the budget refusal included), `3` as above, or when its temp dir is inside a git work tree or git cannot tell whether it is (set TMPDIR) — its arm snapshots go to a temp dir that is removed afterwards. `hillclimb run` exits `0` when every attempted (case, rep) was scored, `1` when any attempt failed (an `errors.jsonl` row, or a scored row whose trace or copies could not be written), the pass stopped mid-run (rows already written are kept), a baseline pass could not freeze a pairwise case's reference, or `summary.json` could not be written after the pass, and `2` for any refusal before spending: a usage error, the harness gate, no usable agent credential at a selected case's tier (`doctor`'s check, as on `eval`; at `protocol` a login only in the real config dir too, since hillclimb runs protocol with a managed config dir), an alias model, a session that does not declare exactly one `plugins.local_plugins` entry (or an inline session, or cases tuning different plugins), `on_unanswered: prompt`, the `protocol` tier without a managed config dir, a flow-dir or `_state.json` problem, split or duplicate case ids, an unknown `--case`, the answer-key guard, a `harness_paths` entry inside the tuned plugin, an unknown, untracked or ambiguous `--skill` (the message lists the plugin's skills), a tracked skill (or none) that differs from a `meta.skill_tracked` recorded by the variant's existing rows, a requested model or effort for a case that differs from the `meta.model_requested` / `meta.effort` the variant's rows for that case recorded (scored or error rows) (a row recording neither is held to its served `model`), an `--effort` level the pinned model does not offer (or any effort on a model with no effort selector, dated ids included), `xhigh` or `max` with `extended_thinking: false`, `extended_thinking: false` on a model whose baseline entry sets `disallowThinkingDisabled`, a missing or incomplete variant snapshot, a scenario input a run would refuse, a `semantic_pairwise` reference that is missing, damaged or exposed through a mount (the gate a run applies), a host `claude` that cannot run the judge or LLM decider isolated (checked when a scenario uses `semantic_matches` or `semantic_pairwise`, or `on_unanswered: llm` with no decider channel, as on `eval`), or a decider with `--concurrency` above 1, or one scenario metric id declared differently in two scenarios, or declared differently from the signature the flow's existing rows carry for it in `meta.metric_sigs` (adding or removing a metric is allowed; removing one warns); `--dry-run` exits `0` unless such a refusal fires; the harness gate does not refuse a dry run, which reports its status instead. Under `--case`, the per-case refusals (the session, its model pins, `on_unanswered: prompt`, the `protocol` tier, a scenario input, a `semantic_pairwise` reference, the isolation check, and the mounts the answer-key guard checks) cover the selected cases only; the flow-level ones cover every case: every scenario file must parse, every case's baseline must load, the harness gate hashes every case's scenario and session file (an unselected case's inputs it cannot read — missing, a directory, unreadable — or that its unparseable session cannot name, are left out of it; the gate hashes each path's name, so the sha still moves), the answer-key guard keeps every case's scenario and session file unreadable, and the one-plugin rule covers every case whose session parses (an unselected one whose session is inline or does not parse is skipped, with a note). `hillclimb regrade --case` follows the same rule, with the same harness sha. `--timeout-s` bounds the whole attempt: the agent phase through the scenario's `timeout_ms` (lowered to it), and the judge phase through a deadline no judge may start after; an attempt whose agent phase reaches its bound, or whose judge phase reaches the deadline, is an `errors.jsonl` row (`timeout`), never a scored one. So is an attempt whose agent did not send the requested effort: its own session transcript shows a main-loop assistant message sent with another effort, with a value that is not an effort level, or with none (even when the agent then failed; a model with no effort selector may send none), or, on an otherwise valid run whose main loop answered, records no main-loop assistant message or cannot be found (read from the run's config root, or the session's pinned `plugins.config_dir` on hostloop and protocol, by the run's session id): `serving_substitution` with `meta.failure_rule: "effort_not_sent"`. Its envelope's `ok` ⇔ exit `0`; a refusal or a mid-run stop is the shared error envelope, its reason in `error.message` (category `usage` for a refusal, `runtime` for a stop), with the payload keys (`flow`, `variant`, `scheduled`, `scored`, `failed`, `exitCode`, and a dry run's `dryRun`) riding on it; `failed` counts failed (case, rep) slots, so a failed reference freeze or `summary.json` write exits `1` without adding to it. Its `--dry-run` `plan.cost` is the same object `eval --dry-run` emits at `plan.cost` (covered keys `jobs`, `meanUsd`, `p50Usd`, `p95Usd`, `worstObservedUsd`, `lowerBound`, `unpriced`, `pricedRuns`, `thinnest`; `p95Usd` and `worstObservedUsd` are pessimistic figures, never bounds, and `worstObservedUsd` is not the `--max-budget-usd` gate's figure), priced on `eval --dry-run`'s basis exactly — the scenario's runs on this machine at its effective tier and baseline, turn 1, every `hillclimb:` run excluded — so each covered key means the same on both commands. `jobs` is the slots the pass would still run after resuming. A scored row records how its judge ran in `meta.judge_transport` (`assertions[].judgeTransport`'s shape, `{isolation, cliVersion?, strictMcp?}`), or `meta.judge_transports` listing each distinct one when its asserts were judged differently; neither is present when no host judge recorded one, so a round judged under a different isolation can be told apart from a regression. `skill_invoked` is `1` or `0` for whether the run invoked the tracked skill; a blank cell means not measured (no tracked skill, or a record that could not tell), never "not invoked". The tracked skill is matched by the id the agent registers for it, `<plugin>:<name>`, the name being a skill directory's name, or a `SKILL.md`'s frontmatter `name` where the agent reads one (else its directory's name), with every character outside `[a-zA-Z0-9_-]` replaced by `-` (`skills/my.skill/` registers `<plugin>:my-skill`), among the skills the agent loads: `skills/*/` whenever `skills/` exists, each path in the plugin manifest's `skills` field (a string or an array of paths relative to the plugin root, each `.` or starting with `./`, as the agent's manifest schema requires, and each a skill directory or a directory of them), and a root `SKILL.md` only when the manifest has no `skills` field at all (`"skills": []` is a field, naming nothing) and there is no `skills/` dir; a `SKILL.md` that is not a regular file or is over 1 MiB is skipped, as the agent skips it. A plugin with one such skill is tracked; with several, `--skill <name>` (a skill directory's name or its registered name) picks one, and without it the column is omitted. Each scored row records the tracked id in `meta.skill_tracked`, absent when the column is omitted. A `--skill` selection enters the harness sha as `skill:<registered name>` and `--approve-harness` records it as `_state.json` `harness_skill`, so a changed, added or dropped `--skill` is a harness-gate refusal that names the change; without `--skill` the sha is unchanged. `--skill` is not remembered between passes: the runner command must pass the same one every time. A scenario's declared `metrics` are grade keys on every scored row of the flow, over the union of the cases' declarations: `<id>` only when measured — an unmeasured metric is omitted, never 0 — and `<id>_present` (1 measured; 0 on a case that does not declare it, on an agent-caused failure, or when unavailable, with the reason in `meta.metrics_unavailable`). `hillclimb state-template` declares each `<id>_present` with the other `_present` companions after `pass`, and each metric last as `{id, kind: "float", label, better}`, plus `scale` and `min` when the scenario declares them; every scored row records each flow metric's declaration signature in `meta.metric_sigs`. `hillclimb check`'s headroom takes a float headline's good end from `scale` (higher is better) or `min`, 0 when absent (lower is better), and without a `scale` says so instead of computing one; a row that lacks a declared metric (its `meta.metric_sigs` lacks the id) is a note, not an error — it predates the metric when a row after it carries it, and no scenario declares the metric any more when it comes after the last row that does (variants in order, then file order) — a float outside `[min, scale]` is a warning, and so are a case's rows carrying more than one `meta.assert_sig` (graded under different assertion sets) and rows whose `meta.assert_sig` is not their scenario's now — the scenario target `hillclimb check [<scenario.yaml | dir/>]` was given, else the scenario files the last `--approve-harness` hashed (the `.yaml` keys of `_state.json` `harness_files` whose stem is a case's id), else those `harness_paths` lists (a recorded file is a case's scenario when it has a `prompt:` and its stem is the case's id; a note when nothing is recorded, or names a recorded scenario not found from the current directory or that does not parse, or a case the target holds no scenario for, and skips that case; the remedy's `regrade` target is one that hashes exactly the files the flow was approved over — the scenarios' directory when it holds exactly them, else the case's own file or its directory, else `<scenarios>`). `hillclimb check` exits `0` clean, `1` on any error finding (no warning or note changes it), `2` usage (a target that does not load included); `hillclimb state-template` exits `0`, or `2` on usage or a refusal (a metric id declared differently in two scenarios included); `hillclimb regrade` re-evaluates every selected row from its kept run before any judge call — its deterministic asserts and `expect_denied` hosts with `verify-run`'s evaluation (a changed, added or removed one applied in a default re-grade), its metrics re-measured there (a row no judge re-grades included: a case with no judged assert, an agent-failed row — signatures and `<id>_present: 0`, never a value — and a fill row that needs no comparison; every rewritten row of a case that declares a metric, a re-judged one included, is marked `meta.regrade_remeasured: true`, and each variant's payload entry counts the rows re-measured with no judge call in `remeasured` and those re-evaluated in `reevaluated`); a row whose re-evaluation changes nothing stays byte for byte. It exits `0` when every selected row was rewritten or there was nothing to do, `1` when a row was listed instead (no scenario file for its case in the target, without `--case`; its kept run gone or refused, its kept work dir gone while its case declares a metric or a filesystem assert needs it, an assert the recorded `workspace_fixture` satisfies on its own; an assert whose literal the run scrubbed, on a row graded under another assertion list, since an edit inside a scrubbed literal cannot be applied; a re-judge of an assert scrubbed in the run that this process's scrub does not reproduce, without `--allow-scrub-change`; a part of a row's judge input not proven scrubbed with a set covering its run's, without `--allow-scrub-change`; judged evidence that cannot be recomposed; judged evidence that changed since its grade, without `--rejudge`; current evidence less redacted than the graded document, in both modes unless `--allow-scrub-change`; a scenario with no judged assert left to re-grade; in a fill an assert list, a rubric or a deterministic outcome that does not line up with the scenario as graded, a kept outcome judged against a reference that has changed since, an unreadable regrade file, or a fill that would move `pass`; its re-grade judge-invalid, misaligned or stopped; an open `judge_invalid` slot, which `hillclimb run <target> --flow <dir> --variant <v> --case <id> --reps <rep+1>` re-runs — it is never moved into `results.jsonl`; an assert unchanged since the run that re-evaluates differently from the kept run is never listed: the row keeps the run's own outcome and records `meta.regrade_kept_live`), `1` also for a failure after the first judge call (naming the variants already rewritten), and `2` on usage or a refusal before any judge call (a host `claude` that cannot run the judge isolated, the gate, a held lock, a scenario metric declared differently in two scenarios or from the flow's rows' `meta.metric_sigs`, as on `run`, an evidence refusal anywhere — nothing is written then); `hillclimb freeze-ref` exits `0` when no case was refused (an entry already complete is reported, not refused), `1` when a case was (no good row whose run is under the runs root and delivered its output, a damaged entry, a reference frozen for a different prompt, a missing compose key whose recorded run is gone, a missing compose key whose run's scrub set this process cannot prove it covers (re-run the variant), a store write that failed, or a run whose judged document cannot be composed, differs from the one its live judge read, or has no live fingerprint to check it against), and `2` on usage (a bad `--variant`, no flow dir, a variant with no `results.jsonl`, no selected case with `semantic_pairwise`, its variant's lock held by a live run). `lint` exits `127` when `python3` is missing (spawn error), and `1` — never `0` — when the scenario loader rejected a file but its findings could not be handed to the linter (an unwritable temp directory); `replay` exits
|
|
940
954
|
`2` on a **whole-cassette operational failure** — anything `readCassette` rejects (unreadable, invalid
|
|
941
955
|
shape, unsupported version, unrecognized assertion key) or any per-file throw, plus the batch loop's
|
|
942
956
|
own source-resolution failures (`--assert-from`/`--reassert` drift, scenario-parse errors, `--write`
|
|
@@ -1157,6 +1171,7 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
|
|
|
1157
1171
|
`assertions[]` entry, only `assertionIndex`, `docMatchesLive`, `pass`, `judgeInvalid`, `judgeModel`,
|
|
1158
1172
|
`judgeModelRequested`, `judgeCostUsd` and `semanticClaims` (its other keys follow the `RunResult` assertion entry, which is not
|
|
1159
1173
|
pinned field by field); and the error envelope's `error.code` values (`doc_drift`, `unchecked_content`,
|
|
1174
|
+
`task_unverifiable`, `rubric_unverifiable`, `evidence_unverifiable`, `reference_unverifiable`,
|
|
1160
1175
|
`no_semantic_asserts`),
|
|
1161
1176
|
`refusals[]` and post-write-failure `runs[]`. The enums are covered as sets: `docMatchesLive`, `change`,
|
|
1162
1177
|
`authoredCapture.source`, a section's `kind`, `error.code`. **Adding a key or an enum value is MINOR** — a
|
package/dist/hillclimb/cli.js
CHANGED
|
@@ -458,6 +458,7 @@ export async function cmdHillclimb(args, deps) {
|
|
|
458
458
|
approveHarness: p.flags["--approve-harness"] === true,
|
|
459
459
|
allowDocDrift: p.flags["--allow-doc-drift"] === true,
|
|
460
460
|
allowUnchecked: p.flags["--allow-unchecked"] === true,
|
|
461
|
+
allowScrubChange: p.flags["--allow-scrub-change"] === true,
|
|
461
462
|
}, {
|
|
462
463
|
cwd: process.cwd(),
|
|
463
464
|
env: process.env,
|