cowork-harness 1.8.0 → 1.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 1.8.0
7
- tracks-harness: cowork-harness 1.8.0 (baseline desktop-1.24012.1)
6
+ version: 1.9.0
7
+ tracks-harness: cowork-harness 1.9.0 (baseline desktop-1.24012.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.8.0` (baseline
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.9.0` (baseline
26
26
  > `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
@@ -39,16 +39,28 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
41
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
42
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.8.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.8.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.8.0"`. **Pin `@>=1.8.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
42
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.9.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.9.0"`. **Pin `@>=1.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
43
43
 
44
- What the ≥ 1.8.0 floor gates, by release:
44
+ What the ≥ 1.9.0 floor gates, by release:
45
+
46
+ - **1.9.0 (critique report schema + shareable, screenshot-safe output):**
47
+ `schema/critique-report.json` (descriptive, test-pinned — parse the report against it, not prose;
48
+ NOT §12-frozen while critique is EXPERIMENTAL), `gradedSkill` in the report (text header + JSON;
49
+ multi-skill pairing key — pair by `(gradedSkillHash, gradedSkill)`, since the hash alone keys the whole
50
+ plugin and cross-pairs its skills). Critique's text report and stderr diagnostics now **collapse host
51
+ paths to `~`** (a shared report or screenshot no longer leaks the username + filesystem layout; the JSON
52
+ report and persisted-artifact paths stay raw for machine consumers), a **missing or non-directory skill
53
+ folder is refused pre-spawn** (fast exit 2, no stray run dir left behind), and the report header + stderr
54
+ prefixes now read **`critique:`** (were `skill-critique:`).
45
55
 
46
56
  - **1.8.0 (critique hardening + hostloop unpin):** `critique --skill <name>` (multi-skill-plugin
47
57
  grading — a plugin root with no `--skill` is refused pre-spend; the evidence package gains the
48
58
  invoked skill's `agents/<name>.md` + bounded `references/*.md` CONTENT), run-dir artifacts
49
59
  (`critique-report.json` always, `critique-evidence-package.txt` when the evaluator ran,
50
60
  `critique-salvage.json` on exit 2) + `--out <path>`, per-critique `costUsd` across all four
51
- workloads, `findingFingerprint` per item (cross-INPUT clustering; skillHash stays the cross-FIX
61
+ workloads, `findingFingerprint` per item (cross-INPUT clustering; HIGH-precision
62
+ LOW-recall — a match proves reproduction, a mismatch does NOT prove non-reproduction, the same
63
+ finding reworded fingerprints differently; skillHash stays the cross-FIX
52
64
  key), gate-answer `--answer` echo in the report, per-item-tolerant evaluator parse
53
65
  (`droppedEvaluatorItems`), `skillMdTruncated` (a readable-but-oversized SKILL.md is flagged
54
66
  "graded a cut copy" — distinct from missing/unreadable), `subagents[].webSearches` +
@@ -436,7 +448,9 @@ cassette — has its own recipe:
436
448
  input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
437
449
  pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
438
450
  produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
439
- kept run predates the current skill). **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
451
+ kept run predates the current skill). **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
452
+ skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
453
+ critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
440
454
  pairing step there silently groups on an absent key instead of erroring — check the field is present, or
441
455
  require ≥ 1.5.0. See `docs/debugging.md`
442
456
  (repo-only) for the full loop.
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.9.0` (baseline `desktop-1.24012.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
13
13
  ```
14
14
 
15
15
  The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
16
- release; pin an exact version (e.g. `version: "1.8.0"`) for reproducible CI.
16
+ release; pin an exact version (e.g. `version: "1.9.0"`) for reproducible CI.
17
17
 
18
18
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
19
19
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
57
57
  GitHub-hosted runners, no token/Docker/agent:
58
58
 
59
59
  ```yaml
60
- - run: npm i -g "cowork-harness@>=1.8.0"
60
+ - run: npm i -g "cowork-harness@>=1.9.0"
61
61
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
62
62
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
63
63
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -197,7 +197,7 @@ jobs:
197
197
  with: { node-version: '20' }
198
198
  - uses: actions/setup-python@v5
199
199
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
200
- - run: npm i -g "cowork-harness@>=1.8.0"
200
+ - run: npm i -g "cowork-harness@>=1.9.0"
201
201
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
202
202
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
203
203
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -226,7 +226,7 @@ jobs:
226
226
  echo "live=true" >> "$GITHUB_OUTPUT"
227
227
  fi
228
228
  - if: steps.guard.outputs.live == 'true'
229
- run: npm i -g "cowork-harness@>=1.8.0"
229
+ run: npm i -g "cowork-harness@>=1.9.0"
230
230
  - if: steps.guard.outputs.live == 'true'
231
231
  run: cowork-harness run scenarios/ --output-format json
232
232
  env:
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`).
3
+ Self-contained reference. Tracks `cowork-harness 1.9.0` (baseline `desktop-1.24012.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.8.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.9.0`
4
4
  (baseline `desktop-1.24012.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 1.8.0` (baseline `desktop-1.24012.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 1.9.0` (baseline `desktop-1.24012.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
package/CHANGELOG.md CHANGED
@@ -6,6 +6,85 @@ All notable changes to this project are documented here. The format is based on
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [1.9.0] — 2026-07-24
10
+
11
+ ### Added
12
+
13
+ - **`schema/critique-report.json`** — a descriptive, test-pinned schema for `critique`'s JSON report /
14
+ `critique-report.json` artifact, so automation consumers (budget pacers gating on `costUsd.complete`,
15
+ harvesters) parse field names/shapes from a schema instead of prose. Deliberately **not** a SPEC
16
+ §12-frozen surface (unlike `doctor.json`) — critique is EXPERIMENTAL and additive field changes are
17
+ expected; the schema says so in its own description, and a two-way sync test pins it against the
18
+ actual report builder on every branch (findings / infraFailure / evaluatorError).
19
+ - **`gradedSkill` in the critique report** (text header + JSON): the resolved `skills/<name>` the
20
+ packager graded under `--skill`/auto-selection. Load-bearing for multi-skill plugins:
21
+ `gradedSkillHash` keys the whole mounted plugin, so pairing by hash alone cross-pairs critiques of
22
+ *different* skills — pair by `(gradedSkillHash, gradedSkill)`. Docs updated accordingly.
23
+
24
+ ### Docs
25
+
26
+ - Truncation→DROPPED mechanics: a finding whose `evidence` quotes SKILL.md text past the packaging cap
27
+ fails citation-resolution and lands in DROPPED (the check runs against the *cut* copy) — documented
28
+ next to `skillMdTruncated` so a back-half DROPPED skew on an oversized skill has its cause named.
29
+ - `findingFingerprint` direction-of-inference: high-precision, LOW-RECALL — a match proves
30
+ reproduction; a mismatch does NOT prove non-reproduction (the same finding reworded fingerprints
31
+ differently). The Reproduction section says so before anyone concludes "didn't reproduce".
32
+ - Evaluator cost share: the two evaluator passes were ~3/4 of a measured e2e total — the
33
+ calibrate-then-`--evaluator-model` strategy is now in the cost section and `critique --help`, with
34
+ the armor-verification-is-default-evaluator-only caveat.
35
+ - SPEC §12 now names the critique report explicitly under **NOT covered**: `schema/critique-report.json`
36
+ is descriptive (parse against it, not prose) but not the compatibility contract while critique is
37
+ EXPERIMENTAL — its surface-baseline presence is for change visibility, and it is the promotion
38
+ candidate on the `doctor.json` template once critique stabilizes. README, llms.txt, and the shipped
39
+ skill point at the schema and the `(gradedSkillHash, gradedSkill)` pairing rule.
40
+
41
+ ### Fixed
42
+
43
+ - **critique's "Attached inputs" evidence no longer reports `(none)` when `mounts.json` is corrupt.**
44
+ `listAttachedInputs` derived connected-folder names from `loadVmPathContext`, which returns `null` for
45
+ BOTH an absent `mounts.json` (legitimately no mounts) and a present-but-unparseable one (the folder map
46
+ is UNKNOWN). Uploads already distinguished these (the ENOENT-vs-read-fault split), but folders — which
47
+ have no fixed-layout fallback — silently collapsed to `[]`, so a corrupt `mounts.json` rendered `(none)`,
48
+ telling the evaluator "the agent correctly saw no connected folder" when the truth was unknown. It now
49
+ surfaces a corrupt `mounts.json` as an explicit UNKNOWN note, completing the same
50
+ confabulation-vs-correct guard the uploads path already applies.
51
+ - **`diff` no longer treats two tool inputs that differ only past the ~2000-char cap as the same call.**
52
+ `canonicalizeInput` truncated a tool input's canonical JSON to a 2000-char cap and used that truncated
53
+ string as the tool-sequence equality key, so two `Write`/`Bash` calls sharing a long identical prefix
54
+ but differing only in the dropped tail compared as `op: "same"` in `diffToolSequence` — a false "no
55
+ change" that could flip the advisory `diff` exit code to 0. A truncated key now folds in a
56
+ `#<len>·<sha16>` hash of the full canonical string, so the key depends on the entire content while the
57
+ visible prefix stays readable in the hunk; both diff sides canonicalize identically, so the comparison
58
+ stays consistent.
59
+ - **`COWORK_VM_GATEWAY` is now validated as a canonical IPv4 literal.** The L2-microVM gateway override was
60
+ interpolated verbatim into a root-run guest `iptables -A OUTPUT -d <gateway>` command (via `sh -c`), so a
61
+ malformed or hostile value could inject shell syntax into privileged provisioning — or, more mundanely,
62
+ leave the firewall in an unknown state. `vmGatewayIp()` now rejects anything that isn't a canonical IPv4
63
+ literal (digits-and-dots only, octets 0–255, no leading zeros), failing loud instead of reaching the
64
+ shell. Operator-set env var, so this is defense-in-depth; no valid gateway value is affected.
65
+ - **The `agent.stderr.log` sink is now flushed before the teardown secret-scrub reads it.** The stderr sink
66
+ was piped fire-and-forget and never awaited, so bytes still buffered when `scrubRawRunLogs` read the file
67
+ could land raw *afterwards* — a persisted-secret leak in a narrow teardown window. `LiveAgentSession` now
68
+ pipes it with `{ end: false }` and ends+awaits it in the same session-teardown drain that already flushes
69
+ `events.jsonl` / `control-out.jsonl`, so the session generator resolves only after the sink is fully
70
+ flushed — the scrub always sees the complete log.
71
+ - **`critique` no longer prints raw host paths in its report or diagnostics.** The text report's `run dir:`
72
+ line, the `inspect <dir>` hints, the write-failure diagnostics, and the echoed skill-folder path all
73
+ printed absolute `$HOME`-rooted paths, so a shared report or screenshot leaked the username + filesystem
74
+ layout (it landed in a video frame). critique was the one rendering path in the CLI that never called
75
+ `tildeify`, while `skill`/`run` scrub unconditionally. Every human-facing path is now collapsed to `~`;
76
+ the JSON report and persisted-artifact paths stay raw (machine data a consumer feeds back to a tool). The
77
+ `--demo` rejection is unchanged — it was never the fix — but its reason now notes the report already
78
+ collapses paths.
79
+ - **`critique` fails fast on a missing or non-directory skill folder.** A typo'd/absent positional folder
80
+ previously minted a session and spawned the task turn before infra-failing (exit 2), leaving a stray run
81
+ dir behind. `resolveCritiquedSkillDir` now `existsSync`/`isDirectory`-checks the folder up front — before
82
+ any session is minted or spawned — so a bad path exits 2 immediately with nothing left on disk. A
83
+ present-but-`SKILL.md`-less folder still defers to the packager's degraded flow, unchanged.
84
+ - **`critique`'s report header and stderr diagnostics now say `critique:` (were `skill-critique:`).** A
85
+ leftover label from the `scripts/skill-critique.ts` instrument; the invoked command is `critique`.
86
+ Cosmetic, no schema change.
87
+
9
88
  ## [1.8.0] — 2026-07-23
10
89
 
11
90
  ### Added
package/README.md CHANGED
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
91
91
 
92
92
  > **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
93
93
  > <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
94
- > From a global install (`npm i -g "cowork-harness@>=1.8.0"`), point at the package root instead:
94
+ > From a global install (`npm i -g "cowork-harness@>=1.9.0"`), point at the package root instead:
95
95
  > `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
96
96
  > (or copy the cassette into your own project and pass that path).
97
97
 
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
101
101
  > - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
102
102
  > - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
103
103
  > - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
104
- > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.8.0"`.
104
+ > - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.9.0"`.
105
105
 
106
106
  Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
107
107
  test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
126
126
  claude plugin install cowork-harness@cowork-harness
127
127
  ```
128
128
 
129
- The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.8.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
129
+ The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.9.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
130
130
 
131
131
  It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
132
132
 
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
147
147
  To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
148
148
  skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
149
149
  assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
150
- above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.8.0"` — see
150
+ above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.9.0"` — see
151
151
  [above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
152
152
 
153
153
  ### Prerequisites for anything above `protocol` fidelity
@@ -690,7 +690,7 @@ jobs:
690
690
  anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
691
691
  ```
692
692
 
693
- Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.8.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
693
+ Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.9.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
694
694
 
695
695
  The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **seven-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
696
696
 
@@ -752,7 +752,7 @@ Most runs need **none** of these — the defaults are correct. They're grouped b
752
752
  - **Staleness boundary** (the git-tracked-files-only default and its `git add`/hard-fail rationale are explained in [Test a local skill](#test-a-local-skill-in-one-command)): a **non-repo** source dir falls back to a raw walk instead; OS-junk like `.DS_Store` is always excluded; a partial exclusion emits a `::notice:: [stage]`. `COWORK_HARNESS_GITSET=0` opts out to a raw walk for every dir (and copies untracked files too); `COWORK_HARNESS_DEBUG_SKILLHASH=1` dumps the exact file set feeding the staleness hash on a mismatch (and flags OS-junk) so a drift source is one line. Declare per-plugin non-runtime paths in a `.cowork-hashignore` file (or the session `staleness.hash_ignore`). `COWORK_HARNESS_AGENT_SCOPE=skill` (opt-in) refines a scenario's `skills:` scope so a **skill-named** `agents/<name>.md` re-stales only that skill's cassettes instead of the whole fleet (generic agents stay shared; stamped into the fingerprint, so flipping it is a one-time re-record like `GITSET`).
753
753
  - **`skill` / `chat` defaults:** `COWORK_HARNESS_FIDELITY` sets the default fidelity tier for ad-hoc `skill`/`chat` runs (a `--fidelity` flag or a scenario's `fidelity:` still wins); `COWORK_HARNESS_MODEL` sets the default model; `COWORK_HARNESS_OUTPUT_FORMAT` (`text`|`json`) sets the default output format. Each is overridden by the matching explicit flag. **Caveat:** `chat` only accepts `protocol`/`container`/`hostloop` — `microvm` and `cowork` are rejected (no Lima/auto-pick plumbing in the interactive REPL), so a `COWORK_HARNESS_FIDELITY` set to either is rejected loudly for `chat` even though `skill` accepts the full tier set.
754
754
  - **Secret scrubbing:** `COWORK_HARNESS_SCRUB_KEYS=<KEY1,KEY2>` adds extra env-var names whose values are redacted from logs (beyond the known auth tokens + `ANTHROPIC_CUSTOM_HEADERS`); `COWORK_HARNESS_SCRUB_VALUES=<v1,v2>` redacts literal values regardless of env. **Committed-cassette redaction:** `COWORK_HARNESS_REDACT_PATTERNS=<rx1,rx2>` / `COWORK_HARNESS_REDACT_KEYS=<k1,k2>` extend the privacy layer that scrubs recorded `controlOut` before a cassette is written for commit.
755
- - L2 microVM: `COWORK_VM_GATEWAY` overrides the Lima host-proxy gateway IP (default `192.168.5.2`); `COWORK_VM_PROXY_PORT` pins the egress-proxy port (unset, the host binds an OS-assigned free port and threads that same value into the guest firewall + `HTTP(S)_PROXY`). The Lima instance is named `cowork-vm-<config-hash>` (a config change → a fresh VM); `COWORK_LIMA_INSTANCE` pins a fixed name, and `vm prune` removes orphaned ones.
755
+ - L2 microVM: `COWORK_VM_GATEWAY` overrides the Lima host-proxy gateway IP (default `192.168.5.2`; must be a canonical IPv4 literal — an invalid value is rejected, since it is interpolated into the guest firewall rule); `COWORK_VM_PROXY_PORT` pins the egress-proxy port (unset, the host binds an OS-assigned free port and threads that same value into the guest firewall + `HTTP(S)_PROXY`). The Lima instance is named `cowork-vm-<config-hash>` (a config change → a fresh VM); `COWORK_LIMA_INSTANCE` pins a fixed name, and `vm prune` removes orphaned ones.
756
756
  - **Advanced / internal escape hatches** (rarely needed): `PYTHON` overrides the interpreter for `lint` / scenario tooling (default `python3`); `COWORK_HARNESS_DEBUG=1` surfaces which `.env` files were loaded; `COWORK_HARNESS_CLAUDE_BIN=<path>` points the `--decider-llm` transport at a specific `claude` binary; `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` lets the harness use the newest sibling agent binary when the baseline-pinned version is missing (a fidelity compromise — off by default) — a same-major.minor **patch** bump of the staged NATIVE binary is auto-accepted without this flag (the native binary carries no sha256 pin, so a patch drift is safe by default; it prints a loud stderr note naming the pinned and substituted versions). At `hostloop`, and at `cowork` **only when it resolves to host-loop** on the synced baseline, the staged **VM ELF** is auto-accepted on a patch bump too, because on that path it's a non-executed parity mount into the bash sidecar; a `cowork` baseline that resolves to VM-loop instead executes the ELF directly, so it keeps the strict sha-pinned exact-version match, same as `container`/`microvm`, which always keep it (the ELF is the executed agent there), so the flag remains required for any ELF drift on those tiers/paths, and for a major/minor gap everywhere; `COWORK_MANAGED_CONFIG=1` forces the managed-config path on `protocol`; `COWORK_HARNESS_ALLOW_MISSING_PROMPT=1` downgrades a missing prompt asset to a warning; `COWORK_HARNESS_SAFE_STAGING_PREFIX=<a,b>` whitelists prefixes under which a delete in `outputs/` is allowed (otherwise delete-in-outputs fails loud); `COWORK_HARNESS_NO_HEARTBEAT=1` disables the idle-run heartbeat, `COWORK_HARNESS_HEARTBEAT_MS` tunes its interval (default 30000ms / 30s); `COWORK_HARNESS_CLI=/path/to/cli.js` overrides which built CLI the Python `cowork` pytest lane drives (see python/README.md).
757
757
  - Pin `baseline: desktop-<ver>` and `model:` in a session for byte-stable runs; use `latest` to track.
758
758
 
@@ -801,7 +801,7 @@ This repo is built to be driven by agents, not just read by humans:
801
801
  - **[AGENTS.md](./AGENTS.md)** — the canonical agent-instructions file (architecture seams, the build gate, invariants, ethos). Read it before changing code. Also indexed in **[llms.txt](./llms.txt)**.
802
802
  - **Companion skill** — [`.claude/skills/cowork-harness/`](./.claude/skills/cowork-harness/SKILL.md) teaches an agent to drive the harness; install it via the marketplace (see [above](#drive-it-from-claude-code-companion-skill)).
803
803
  - **Machine-readable interfaces** — stable `--output-format json` envelope on stdout, deterministic exit codes (`0`/`1`/`2`/`3`, with a couple of documented per-command exceptions — see [SPEC.md](./SPEC.md) for the full table), and `--help` on every command.
804
- - **JSON Schemas** — [`schema/scenario.schema.json`](./schema/scenario.schema.json) and [`schema/session.schema.json`](./schema/session.schema.json) describe every field of the YAML you author (generated from the source schemas; `npm run schema`). [`schema/protocol.v1.json`](./schema/protocol.v1.json) (hand-authored) schemas the harness's own control-channel wire protocol, with a golden vector pack at [`fixtures/protocol/v1/`](./fixtures/protocol/v1/) — see [docs/protocol.md](./docs/protocol.md).
804
+ - **JSON Schemas** — [`schema/scenario.schema.json`](./schema/scenario.schema.json) and [`schema/session.schema.json`](./schema/session.schema.json) describe every field of the YAML you author (generated from the source schemas; `npm run schema`). [`schema/protocol.v1.json`](./schema/protocol.v1.json) (hand-authored) schemas the harness's own control-channel wire protocol, with a golden vector pack at [`fixtures/protocol/v1/`](./fixtures/protocol/v1/) — see [docs/protocol.md](./docs/protocol.md). [`schema/critique-report.json`](./schema/critique-report.json) describes `critique`'s JSON report / `critique-report.json` artifact for automation consumers (budget pacers, harvesters) — **descriptive, not §12-frozen** while critique is EXPERIMENTAL (see [SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)).
805
805
 
806
806
  `AGENTS.md`, `SPEC.md`, and `DESIGN.md` **are** shipped in the npm package (see `package.json` `files`) —
807
807
  a global install has them locally too, not just on GitHub.
package/SPEC.md CHANGED
@@ -793,6 +793,13 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
793
793
  the `doctor`/`verify-cassettes`/RunResult envelopes above, these have no `schema/*.json` and may change
794
794
  (fields, rule ids, the artifact-write-back finding shape) while the analyzers stabilize. Parse at your
795
795
  own risk until they are promoted to a covered surface.
796
+ - **The `critique` report** (`--output-format json` / the `critique-report.json` run-dir artifact) —
797
+ NOT yet frozen, and — uniquely — it DOES have a schema file: `schema/critique-report.json` is
798
+ **descriptive** (authoritative field names/shapes, test-pinned against the builder) so automation can
799
+ parse against a schema rather than prose, but it is NOT this section's compatibility contract —
800
+ critique is EXPERIMENTAL and additive field changes may land in any minor release (the schema's own
801
+ `description` says so). Its presence in the surface-drift baseline is for change *visibility*, not
802
+ coverage. It is the promotion CANDIDATE once critique stabilizes, on the `doctor.json` template.
796
803
  - **`docs/internal/**`** — untracked working notes.
797
804
  - **The reconstructed system-prompt append text** — a paraphrase by design (see
798
805
  [docs/fidelity-gaps.md](./docs/fidelity-gaps.md)); behaviorally equivalent, not byte-stable.
@@ -242,6 +242,7 @@ export class LiveAgentSession {
242
242
  proc;
243
243
  events;
244
244
  controlOut;
245
+ errLog;
245
246
  timeline;
246
247
  lineIndex = 0;
247
248
  reqById = new Map();
@@ -275,8 +276,11 @@ export class LiveAgentSession {
275
276
  this.events = createWriteStream(join(outDir, "events.jsonl"), { flags: "a" });
276
277
  this.controlOut = createWriteStream(join(outDir, "control-out.jsonl"), { flags: "a" });
277
278
  this.timeline = new TimelineWriter(outDir);
278
- const errLog = createWriteStream(join(outDir, "agent.stderr.log"), { flags: "a" });
279
- this.proc.stderr.pipe(errLog);
279
+ // #60: keep a reference and pipe with `{ end: false }` so WE own the close — the drain in `start()`'s
280
+ // finally ends this stream and AWAITS its flush before the generator resolves, so a subsequent
281
+ // `scrubRawRunLogs` can't read the file while bytes are still buffered (a raw-secret-after-scrub leak).
282
+ this.errLog = createWriteStream(join(outDir, "agent.stderr.log"), { flags: "a" });
283
+ this.proc.stderr.pipe(this.errLog, { end: false });
280
284
  // keep a bounded stderr tail and capture the exit code/signal so a child that dies nonzero
281
285
  // (with no structured {type:"result"} error) is surfaced as a typed error event, not a silent stop.
282
286
  this.proc.stderr.on("data", (d) => {
@@ -404,11 +408,14 @@ export class LiveAgentSession {
404
408
  this.closing = true; // discard any late write() from translate()'s async hook/mcp paths
405
409
  await this.drainAll(); // queue empty + all callbacks confirmed before ending either stream
406
410
  // AWAIT the stream flush before the generator resolves. executeScenario reads/scans/scrubs
407
- // events.jsonl + control-out.jsonl immediately after `drive()` returns; a fire-and-forget end()
408
- // races the final buffered writes. end(cb) fires the callback on 'finish' (fully flushed).
411
+ // events.jsonl + control-out.jsonl + agent.stderr.log immediately after `drive()` returns; a
412
+ // fire-and-forget end() races the final buffered writes. end(cb) fires the callback on 'finish'
413
+ // (fully flushed). #60: agent.stderr.log is ended here too — piped with `{ end: false }` (see the
414
+ // ctor) so its last buffered bytes flush BEFORE the scrub reads it, never landing raw afterwards.
409
415
  await Promise.all([
410
416
  new Promise((res) => this.events.end(() => res())),
411
417
  new Promise((res) => this.controlOut.end(() => res())),
418
+ new Promise((res) => this.errLog.end(() => res())),
412
419
  new Promise((res) => this.timeline.end(() => res())),
413
420
  ]);
414
421
  }
@@ -17,7 +17,8 @@ import { fileURLToPath } from "node:url";
17
17
  import { lookupSkillFlag } from "../run/skill-flag-surface.js";
18
18
  import { gradedAliasPath, turnArtifactPath } from "../run/turn-layout.js";
19
19
  import { renderKnownLimitations } from "./limitations.js";
20
- import { existsSync, readFileSync, copyFileSync, writeFileSync, readdirSync } from "node:fs";
20
+ import { tildeify } from "../io.js";
21
+ import { existsSync, readFileSync, copyFileSync, writeFileSync, readdirSync, statSync } from "node:fs";
21
22
  import { randomUUID } from "node:crypto";
22
23
  import { writeSync } from "node:fs";
23
24
  import { basename, join } from "node:path";
@@ -98,6 +99,7 @@ Not accepted (each errors with its reason rather than being silently ignored):
98
99
  --repeat + companions fixed two-turn protocol — loop critique itself; pair by fingerprint.skillHash
99
100
  --ablate-skill grading a skill you removed is incoherent
100
101
  --quiet/--verbose/--compact/--demo/--dry-run inner-turn rendering or preview — no effect on the report
102
+ (which already collapses host paths to ~)
101
103
 
102
104
  Repeating a flag: --upload/--folder/--plugin/--marketplace/--enable/--answer accumulate (that is how you
103
105
  pass several). Every other value-taking flag is single-valued and repeating it is a USAGE ERROR rather
@@ -107,8 +109,10 @@ Repeating a flag: --upload/--folder/--plugin/--marketplace/--enable/--answer acc
107
109
  COST AND PREREQUISITES — read before running:
108
110
  * Each critique is FOUR model workloads: two graded runs (task + reflection) at the chosen tier and two
109
111
  evaluator passes over an evidence package of up to ${MAX_PACKAGE_BYTES / 1024}KB.
110
- * The evaluator defaults to ${DEFAULT_EVALUATOR_MODEL} — the most expensive tier. Override it if
111
- that is not what you want.
112
+ * The evaluator defaults to ${DEFAULT_EVALUATOR_MODEL} — the most expensive tier — and the two
113
+ evaluator passes DOMINATE spend (~3/4 of a measured e2e total). For a batch: calibrate with run 1's
114
+ costUsd (gate on costUsd.complete), then consider a cheaper --evaluator-model for the sweep — noting
115
+ the armor's injection-resistance is verified for the DEFAULT evaluator only.
112
116
  * container needs Docker/Lima; hostloop needs Docker (the bash/web_fetch sidecar) PLUS the staged native
113
117
  agent binary, and writes to the real host FS (a writable --folder requires --allow-host-writes). Both
114
118
  tiers need an authenticated \`claude\` CLI on PATH.
@@ -361,6 +365,22 @@ function parseArgs(argv) {
361
365
  * folder → same hash — so generation pairing keeps working; it is a per-plugin key, not per-skill.
362
366
  * Exported for unit tests. */
363
367
  export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
368
+ // Fail-fast on a typo'd / absent path BEFORE the caller mints a session and spawns the task turn — a
369
+ // missing folder otherwise only surfaces as a mid-run mount failure that leaves a stray run dir behind.
370
+ // This lives here (not in parseArgs) on purpose: parseArgs is unit-tested with fictitious paths, whereas
371
+ // this resolver is only ever called with a real folder. One statSync in a try/catch also covers a broken
372
+ // symlink; the existsSync guard just gives the common typo the clearer "not found" message.
373
+ if (!existsSync(skillFolder))
374
+ throw new Error(`skill folder not found: ${tildeify(skillFolder)}`);
375
+ let stat;
376
+ try {
377
+ stat = statSync(skillFolder);
378
+ }
379
+ catch {
380
+ throw new Error(`skill folder not found: ${tildeify(skillFolder)}`);
381
+ }
382
+ if (!stat.isDirectory())
383
+ throw new Error(`not a directory: ${tildeify(skillFolder)}`);
364
384
  const agentsMdFor = (name) => {
365
385
  const p = join(skillFolder, "agents", `${name}.md`);
366
386
  return existsSync(p) ? p : undefined;
@@ -380,7 +400,7 @@ export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
380
400
  const candidate = join(skillFolder, "skills", skillSelector);
381
401
  if (!existsSync(join(candidate, "SKILL.md"))) {
382
402
  const available = listPluginSkills();
383
- throw new Error(`--skill ${skillSelector}: no skills/${skillSelector}/SKILL.md under ${skillFolder}` +
403
+ throw new Error(`--skill ${skillSelector}: no skills/${skillSelector}/SKILL.md under ${tildeify(skillFolder)}` +
384
404
  (available.length ? ` — available skills: ${available.join(", ")}` : ` — no skills/<name>/SKILL.md found at all`));
385
405
  }
386
406
  return { skillDir: candidate, agentsMdPath: agentsMdFor(skillSelector) };
@@ -391,7 +411,7 @@ export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
391
411
  if (skills.length === 1)
392
412
  return { skillDir: join(skillFolder, "skills", skills[0]), agentsMdPath: agentsMdFor(skills[0]), autoSelectedSkill: skills[0] };
393
413
  if (skills.length > 1)
394
- throw new Error(`${skillFolder} is a multi-skill plugin root (no root SKILL.md; skills: ${skills.join(", ")}) — ` +
414
+ throw new Error(`${tildeify(skillFolder)} is a multi-skill plugin root (no root SKILL.md; skills: ${skills.join(", ")}) — ` +
395
415
  `pass --skill <name> so critique grades the INVOKED skill's SKILL.md instead of a missing root one`);
396
416
  return { skillDir: skillFolder }; // no SKILL.md anywhere — the packager's existing missing/degraded flow reports it
397
417
  }
@@ -702,10 +722,10 @@ export const VERDICT_PROVENANCE = {
702
722
  export function buildTextReport(state) {
703
723
  const { skillFolder, prompt, sessionId, outDir, taskResult, gradedOutcome, gradedSkillHash, selfReportStatus, items, evaluatorModel, requestedModel, evaluatorError, infraFailure, turn1ResultDegraded, turn1SliceDegraded, skillMdStatus, } = state;
704
724
  const out = [];
705
- out.push(`skill-critique: ${skillFolder}`);
725
+ out.push(`critique: ${tildeify(skillFolder)}`);
706
726
  out.push(` probe: ${prompt}`);
707
727
  out.push(` session: ${sessionId}`);
708
- out.push(` run dir: ${outDir}`);
728
+ out.push(` run dir: ${tildeify(outDir)}`);
709
729
  out.push(` fidelity: ${state.fidelity}` +
710
730
  (state.gradedEffectiveFidelity
711
731
  ? ` (graded turn recorded ${state.gradedEffectiveFidelity}${state.gradedBaseline ? `, baseline ${state.gradedBaseline}` : ""})`
@@ -722,6 +742,8 @@ export function buildTextReport(state) {
722
742
  // reflection turn's result.json for them).
723
743
  if (gradedOutcome)
724
744
  out.push(` graded outcome: ${gradedOutcome}`);
745
+ if (state.gradedSkill)
746
+ out.push(` graded skill: ${state.gradedSkill} (pair by skillHash + this name — skillHash keys the whole mounted plugin)`);
725
747
  if (gradedSkillHash)
726
748
  out.push(` graded skillHash: ${gradedSkillHash.slice(0, 12)}`);
727
749
  if (evaluatorModel)
@@ -750,7 +772,7 @@ export function buildTextReport(state) {
750
772
  // The dominant real-world cause of a "missing" SKILL.md is pointing critique at a MULTI-SKILL PLUGIN
751
773
  // root (skills/<name>/SKILL.md, no root SKILL.md) — name the cause and the fix, not just the symptom.
752
774
  if (skillMdStatus === "missing")
753
- out.push(` NOTE: if ${skillFolder} is a multi-skill plugin root, pass --skill <name> (or point critique at <plugin>/skills/<name> directly) so the invoked skill's SKILL.md is graded.`);
775
+ out.push(` NOTE: if ${tildeify(skillFolder)} is a multi-skill plugin root, pass --skill <name> (or point critique at <plugin>/skills/<name> directly) so the invoked skill's SKILL.md is graded.`);
754
776
  out.push(` verdict scope: advisory self-run — NOT an independent attestation (never gate a skill on it)`);
755
777
  out.push("");
756
778
  const dropped = state.droppedEvaluatorItems;
@@ -768,12 +790,12 @@ export function buildTextReport(state) {
768
790
  }
769
791
  if (infraFailure) {
770
792
  out.push(`INFRASTRUCTURE/PROTOCOL FAILURE (reflection turn): ${infraFailure}`);
771
- out.push(`The evaluator was NOT invoked — this is a broken discovery run, not a critique. Re-run, or inspect ${outDir} directly.`);
793
+ out.push(`The evaluator was NOT invoked — this is a broken discovery run, not a critique. Re-run, or inspect ${tildeify(outDir)} directly.`);
772
794
  return out.join("\n");
773
795
  }
774
796
  if (evaluatorError) {
775
797
  out.push(`EVALUATOR FAILED: ${evaluatorError}`);
776
- out.push(`No critique items were produced. Re-run, or inspect ${outDir} directly.`);
798
+ out.push(`No critique items were produced. Re-run, or inspect ${tildeify(outDir)} directly.`);
777
799
  return out.join("\n");
778
800
  }
779
801
  const byBucket = new Map();
@@ -828,6 +850,7 @@ export function buildJsonReport(state) {
828
850
  gradedEffectiveFidelity: state.gradedEffectiveFidelity,
829
851
  gradedBaseline: state.gradedBaseline,
830
852
  costUsd: state.costUsd,
853
+ gradedSkill: state.gradedSkill,
831
854
  skillInvocationObserved: state.skillInvocationObserved,
832
855
  gateAnswers: state.gateAnswers,
833
856
  taskResult,
@@ -894,7 +917,7 @@ function writeRunArtifact(outDir, name, content) {
894
917
  writeFileSync(join(outDir, name), content);
895
918
  }
896
919
  catch (e) {
897
- process.stderr.write(`skill-critique: could not write ${name} under ${outDir}: ${String(e)}\n`);
920
+ process.stderr.write(`critique: could not write ${name} under ${tildeify(outDir)}: ${String(e)}\n`);
898
921
  }
899
922
  }
900
923
  /** Persist the run-dir artifacts every critique leaves behind (all best-effort):
@@ -925,7 +948,7 @@ function writeOutFile(outPath, state, outputFormat) {
925
948
  writeFileSync(outPath, content);
926
949
  }
927
950
  catch (e) {
928
- process.stderr.write(`skill-critique: --out ${outPath} could not be written: ${String(e)}\n`);
951
+ process.stderr.write(`critique: --out ${tildeify(outPath)} could not be written: ${String(e)}\n`);
929
952
  }
930
953
  }
931
954
  export function buildTaskTurnArgs(opts, sessionId) {
@@ -1000,7 +1023,7 @@ async function main(argv = process.argv.slice(2)) {
1000
1023
  return;
1001
1024
  }
1002
1025
  if (resolvedSkill.autoSelectedSkill)
1003
- process.stderr.write(`::notice:: [critique] ${opts.skillFolder} is a single-skill plugin — grading skills/${resolvedSkill.autoSelectedSkill}/SKILL.md (pass --skill to be explicit)\n`);
1026
+ process.stderr.write(`::notice:: [critique] ${tildeify(opts.skillFolder)} is a single-skill plugin — grading skills/${resolvedSkill.autoSelectedSkill}/SKILL.md (pass --skill to be explicit)\n`);
1004
1027
  const sessionId = `crit-${randomUUID()}`;
1005
1028
  try {
1006
1029
  // 1. Task turn.
@@ -1017,7 +1040,7 @@ async function main(argv = process.argv.slice(2)) {
1017
1040
  ]
1018
1041
  .filter(Boolean)
1019
1042
  .join("; ");
1020
- process.stderr.write(`skill-critique: could not determine the task run's directory (no envelope outDir and no [status] line)${diag ? ` [${diag}]` : ""}.\n` +
1043
+ process.stderr.write(`critique: could not determine the task run's directory (no envelope outDir and no [status] line)${diag ? ` [${diag}]` : ""}.\n` +
1021
1044
  `--- task stdout ---\n${task.stdout}\n--- task stderr (tail) ---\n${task.stderr.slice(-4000)}\n`);
1022
1045
  process.exit(EXIT_INSTRUMENT_FAILURE);
1023
1046
  return;
@@ -1216,6 +1239,7 @@ async function main(argv = process.argv.slice(2)) {
1216
1239
  gradedEffectiveFidelity,
1217
1240
  gradedBaseline,
1218
1241
  costUsd,
1242
+ gradedSkill: gradedSkillName,
1219
1243
  skillInvocationObserved,
1220
1244
  gateAnswers: gateAnswers?.length ? gateAnswers : undefined,
1221
1245
  taskResult,
@@ -1254,7 +1278,7 @@ async function main(argv = process.argv.slice(2)) {
1254
1278
  process.exit(EXIT_INSTRUMENT_FAILURE);
1255
1279
  }
1256
1280
  catch (e) {
1257
- process.stderr.write(`skill-critique: unexpected failure: ${e.stack ?? String(e)}\n`);
1281
+ process.stderr.write(`critique: unexpected failure: ${e.stack ?? String(e)}\n`);
1258
1282
  process.exit(EXIT_INSTRUMENT_FAILURE); // an unexpected throw means no critique was produced
1259
1283
  }
1260
1284
  // FINDINGS never gate: any classification — including a task run that ERRORED, which is itself a
@@ -95,6 +95,13 @@ function listAttachedInputs(runDir) {
95
95
  const loaded = loadVmPathContext(runDir);
96
96
  const uploadsDir = loaded?.ctx.uploadsHostDir ?? join(runDir, "work", "session", "mnt", "uploads");
97
97
  const folderNames = loaded ? Array.from(loaded.ctx.folders.keys()).sort() : [];
98
+ // `loadVmPathContext` returns null for BOTH an absent mounts.json (legitimately no recorded mount
99
+ // context — folders default to []) AND a present-but-corrupt one (the folder map is UNKNOWN, not empty).
100
+ // Uploads has a fixed-layout fallback dir it can still probe; connected FOLDERS have none, so a corrupt
101
+ // mounts.json would silently render "(none)" — telling the evaluator "the agent correctly saw no
102
+ // connected folder" when the truth is UNKNOWN. That is the same confabulation-vs-correct false-clean the
103
+ // uploads path guards against below (the ENOENT-vs-read-fault split); mirror it for folders. #14
104
+ const mountsCorrupt = loaded === null && existsSync(join(runDir, "mounts.json"));
98
105
  const lines = [];
99
106
  try {
100
107
  const uploadNames = readdirSync(uploadsDir, { withFileTypes: true })
@@ -124,6 +131,9 @@ function listAttachedInputs(runDir) {
124
131
  }
125
132
  for (const name of folderNames)
126
133
  lines.push(`${name} (connected folder)`);
134
+ if (mountsCorrupt)
135
+ lines.push(`(connected-folder context could not be read: mounts.json present but unparseable — ` +
136
+ `folder attachment presence UNKNOWN, not confirmed absent)`);
127
137
  return lines.join("\n");
128
138
  }
129
139
  /** Byte-bound a text section, appending a loud (never silent) truncation marker so the evaluator knows the
package/dist/run/diff.js CHANGED
@@ -1,6 +1,7 @@
1
1
  // Run/cassette diff engine with normalization. Compares two runs, two cassettes, or a run and a
2
2
  // cassette (both reduce to an event stream + result metadata through parseMessage/buildTrace, the same
3
3
  // typed model `trace` uses — so this stays correct as the SDK schema evolves).
4
+ import { createHash } from "node:crypto";
4
5
  import { DEFAULT_SCAN_PATTERNS } from "../scan.js";
5
6
  import { diffFileSigsPaths } from "./cassette.js";
6
7
  const HOST_PATH_RE = DEFAULT_SCAN_PATTERNS.find((p) => p.cls === "path").re;
@@ -56,7 +57,13 @@ const CANON_CAP = 2000;
56
57
  /** Bounded, key-aware canonicalization of a tool's `input` (or any structured value) for tool-sequence
57
58
  * comparison — masks volatile string spans AND volatile key names, then caps the length. The 100-char
58
59
  * `summarize()` in trace-view.ts is display-lossy by design; this needs its own, larger, comparison-safe
59
- * cap (still bounded — never an unbounded dump into a diff hunk). */
60
+ * cap (still bounded — never an unbounded dump into a diff hunk).
61
+ *
62
+ * #41: when the value exceeds the cap, the truncated prefix ALONE is not a safe equality key — two inputs
63
+ * sharing the first `CANON_CAP` chars but differing only in the dropped tail would compare equal (a
64
+ * false "same" tool call in `diffToolSequence`). So a truncated key folds in a hash of the FULL canonical
65
+ * string: the prefix stays human-readable in a diff hunk, the `#<len>·<sha16>` suffix makes the key
66
+ * depend on the entire content. Both diff sides run this identically, so the comparison stays consistent. */
60
67
  export function canonicalizeInput(input, normalize = true) {
61
68
  let s;
62
69
  try {
@@ -65,7 +72,10 @@ export function canonicalizeInput(input, normalize = true) {
65
72
  catch {
66
73
  s = String(input);
67
74
  }
68
- return s.length > CANON_CAP ? s.slice(0, CANON_CAP) + "…" : s;
75
+ if (s.length <= CANON_CAP)
76
+ return s;
77
+ const fullHash = createHash("sha256").update(s).digest("hex").slice(0, 16);
78
+ return `${s.slice(0, CANON_CAP)}…#${s.length}·${fullHash}`;
69
79
  }
70
80
  /** LCS over full-row equality (name AND canon must match for "same"), then a post-pass over each gap
71
81
  * between LCS matches: a 1-vs-1 gap with the SAME tool name is a "changed" input, not a remove+add pair
@@ -1664,10 +1664,11 @@ function scrubFileInPlace(path, secrets) {
1664
1664
  * executeScenario's outermost `finally` (and the chat lane's teardown) so every exit path AFTER the
1665
1665
  * agent session exists scrubs — success, the unanswered-gate salvage rethrow, and any rethrown fault
1666
1666
  * (agent crash, infra error, hostloop snapshot failure). Deliberately NOT total coverage: a throw
1667
- * before that try has no raw logs yet; a SIGKILL of the harness process skips any finally; and the
1668
- * agent.stderr.log write stream is never awaited, so bytes still buffered at scrub time can land raw
1669
- * afterwards (closing that needs an awaitable stderr-sink close on the session — a follow-up, not
1670
- * this seam). Exported for tests. */
1667
+ * before that try has no raw logs yet, and a SIGKILL of the harness process skips any finally. #60: the
1668
+ * agent.stderr.log sink IS now flushed before this runs — `LiveAgentSession` pipes it with `{ end: false }`
1669
+ * and ends+awaits it in `start()`'s drain, so no buffered stderr byte lands raw after the scrub reads the
1670
+ * file (the session generator resolves only after that flush, and this runs after it resolves). Exported
1671
+ * for tests. */
1671
1672
  export function scrubRawRunLogs(outDir, secrets) {
1672
1673
  scrubFileInPlace(join(outDir, "events.jsonl"), secrets);
1673
1674
  scrubFileInPlace(join(outDir, "control-out.jsonl"), secrets);
@@ -91,7 +91,7 @@ export const SKILL_FLAG_SURFACE = [
91
91
  arity: 0,
92
92
  critique: {
93
93
  kind: "reject",
94
- reason: "critique produces its own report; inner-turn rendering flags have no effect on it",
94
+ reason: "critique produces its own report (host paths in it are already collapsed to ~); inner-turn rendering flags have no effect on it",
95
95
  },
96
96
  })),
97
97
  {
@@ -51,7 +51,19 @@ export function instanceName(baseline) {
51
51
  * result of this one helper into both.
52
52
  */
53
53
  export function vmGatewayIp() {
54
- return process.env.COWORK_VM_GATEWAY ?? "192.168.5.2";
54
+ const raw = process.env.COWORK_VM_GATEWAY ?? "192.168.5.2";
55
+ // #95: this value is interpolated into a root-run iptables command inside the guest
56
+ // (guestFirewallScript → `iptables -A OUTPUT -d ${gatewayIp}` executed via `sh -c`). Validate it as a
57
+ // canonical IPv4 literal and reject everything else, so a malformed or hostile override can never inject
58
+ // shell syntax into privileged provisioning. Defense-in-depth: the var is operator-set, but an
59
+ // unvalidated string reaching a root `sh -c` should be impossible by construction, not by trust. IPv4
60
+ // only — the Apple VZ user-network gateway is IPv4, and the digits-and-dots grammar excludes every shell
61
+ // metacharacter (`;`, `$`, backtick, whitespace, …).
62
+ const octets = raw.split(".");
63
+ const canonicalIPv4 = octets.length === 4 && octets.every((o) => /^\d{1,3}$/.test(o) && Number(o) <= 255 && String(Number(o)) === o);
64
+ if (!canonicalIPv4)
65
+ throw new Error(`COWORK_VM_GATEWAY must be a canonical IPv4 literal (e.g. 192.168.5.2); got ${JSON.stringify(raw)}`);
66
+ return raw;
55
67
  }
56
68
  export function vmStatus(instance) {
57
69
  const r = spawnSync(limaPath(), ["list", instance, "--format", "{{.Status}}"], { encoding: "utf8" });
package/docs/critique.md CHANGED
@@ -124,7 +124,7 @@ ignored.
124
124
  | `--session-id` / `--resume` | critique mints and manages its own session — the reflection turn *is* a resume of it |
125
125
  | `--repeat` + companions | fixed two-turn protocol; loop `critique` itself and pair by `fingerprint.skillHash` |
126
126
  | `--ablate-skill` | grading a skill you removed is incoherent |
127
- | `--quiet`/`-q` / `--verbose` / `--compact` / `--demo` / `--dry-run` | inner-turn rendering or preview — no effect on the report |
127
+ | `--quiet`/`-q` / `--verbose` / `--compact` / `--demo` / `--dry-run` | inner-turn rendering or preview — no effect on the report (which already collapses host paths to `~`) |
128
128
 
129
129
  **Repeating a flag.** `--upload`, `--folder`, `--plugin`, `--marketplace`, `--enable` and `--answer` accumulate,
130
130
  so repeating them is how you pass several. Every other value-taking flag is single-valued and repeating it is
@@ -144,8 +144,11 @@ adjudicable". So:
144
144
  auto-selects with a notice.
145
145
  - **Selection only:** the positional folder is still what both turns mount (session identity is
146
146
  unchanged), and **`fingerprint.skillHash` is unchanged by `--skill`** — it keys the *mounted folder*,
147
- so it pairs generations per-plugin, not per-skill. Pair per-skill runs by `--label` if you need finer
148
- grouping.
147
+ so it pairs generations per-plugin, not per-skill. **Workflow implication: pairing critiques of a
148
+ multi-skill plugin by skillHash alone CROSS-PAIRS different skills** — pair by
149
+ **(`gradedSkillHash`, `gradedSkill`)**; the report's `gradedSkill` field carries the resolved
150
+ `skills/<name>` (`--skill` or the auto-selection). `--label` remains available for coarser
151
+ generation tags.
149
152
  - The report carries an advisory **`skillInvocationObserved`**: `false` means the graded run's own
150
153
  `skillActivity` never mentions the selected skill — the critique may be grading a run that did not
151
154
  actually invoke it.
@@ -170,6 +173,12 @@ It does **not** record their contents — see Known limitations.
170
173
  evaluator passes.
171
174
  - The evaluator defaults to the most expensive tier. Override with `--evaluator-model <id>` or
172
175
  **`COWORK_HARNESS_EVALUATOR_MODEL`**.
176
+ - **The evaluator passes dominate spend** — on a measured end-to-end (trivial probe, default evaluator)
177
+ the two evaluator passes were ~3/4 of the total. For a wide batch: calibrate with run 1's `costUsd`
178
+ (gate on `costUsd.complete` — `false` means the total undercounts), then consider a cheaper
179
+ `--evaluator-model` for the sweep. Caveat: the armor's injection-resistance is verified for the
180
+ shipped **default** evaluator model only — changing it voids that specific verification (matters when
181
+ critiquing skills you did not write).
173
182
  - **container** needs Docker/Lima; **hostloop** needs Docker (the bash/web_fetch sidecar) **plus** the
174
183
  staged native agent binary, and writes to the real host filesystem — a writable `--folder` there requires
175
184
  `--allow-host-writes`. Both tiers need an authenticated `claude` CLI on PATH.
@@ -227,6 +236,14 @@ malformed evaluator items (the surviving findings are then not necessarily the c
227
236
  readable-but-oversized SKILL.md is flagged **`skillMdTruncated`** ("the evaluator graded a cut copy") —
228
237
  distinct from missing/unreadable, which alone force the mechanical `"already-covered"` downgrade.
229
238
 
239
+ **Truncation has a second, sharper consequence than the not-adjudicable steer: DROPPED findings.**
240
+ Citation validation checks each finding's `evidence` excerpt verbatim against the *packaged* (cut)
241
+ copy — so a finding that quotes text past the cut cannot resolve and lands in **DROPPED**, even when
242
+ the quote is a perfectly accurate excerpt of the real file. A skill well over the cap should expect a
243
+ not-adjudicable/DROPPED skew concentrated on its back half; if you see back-half findings in DROPPED
244
+ on a `skillMdTruncated` run, this is why — front-load the operative guidance, split the skill, or
245
+ treat those items as leads to re-check by hand.
246
+
230
247
  ### Run-dir artifacts
231
248
 
232
249
  Beyond stdout, every critique leaves durable artifacts at the run-dir root (best-effort writes —
@@ -240,6 +257,10 @@ Beyond stdout, every critique leaves durable artifacts at the run-dir root (best
240
257
 
241
258
  These artifacts (and the report's JSON shape) are part of critique's **EXPERIMENTAL** surface — useful
242
259
  and stable in practice, but not yet a frozen SPEC §12 covered surface; field additions are expected.
260
+ The report's field names and shapes are authoritatively described by
261
+ [`schema/critique-report.json`](../schema/critique-report.json) (descriptive + test-pinned against the
262
+ actual builder, unlike the §12-frozen `doctor.json`), so automation consumers — budget pacers gating on
263
+ `costUsd.complete`, harvesters pairing on `gradedSkill` — parse against a schema, not prose.
243
264
 
244
265
  ## Reproduction — the ≥2-run discipline
245
266
 
@@ -258,8 +279,15 @@ Then pair/cluster across the reports:
258
279
  - **Same finding across runs/inputs?** cluster by each item's **`findingFingerprint`** (sha over the
259
280
  normalized idea + classification + recommendedAction, deliberately excluding the input-specific
260
281
  `evidence` excerpt — so the same finding matches across different decks/transcripts).
282
+ - **The fingerprint is high-precision, LOW-RECALL — read the direction correctly.** `idea` is
283
+ model-authored free text, so the same underlying finding *reworded* across runs fingerprints
284
+ differently. A **match proves** reproduction; a **mismatch does NOT prove** non-reproduction — before
285
+ concluding "didn't reproduce", skim the unmatched items for rewordings of the same substance.
261
286
  - A finding that recurs across ≥2 runs with the same `findingFingerprint` meets the reproduction bar;
262
- a one-off is a lead, not a conclusion.
287
+ a fingerprint one-off is a lead — possibly a real one-off, possibly a reworded repeat.
288
+ - **Multi-skill plugins: never pair by `gradedSkillHash` alone.** The hash keys the whole mounted
289
+ plugin, so it cross-pairs critiques of *different* skills in the same plugin — pair by
290
+ **(`gradedSkillHash`, `gradedSkill`)**; `gradedSkill` is the report's resolved `skills/<name>`.
263
291
  - To make the graded runs deterministic across repeats, copy the report's echoed `--answer` lines
264
292
  (the graded run's resolved gate answers) into the next invocation.
265
293
 
package/docs/scenario.md CHANGED
@@ -952,7 +952,8 @@ with `COWORK_LIMA_INSTANCE`.
952
952
  - **A run errors with "not mounted — VM not provisioned for this harness config"** — the VM predates a
953
953
  config change (its mounts don't match). Recreate it: `cowork-harness vm delete && cowork-harness vm init`.
954
954
  - **Egress allowed/denied looks wrong** — the guest firewall and the proxy URL must point at the same
955
- gateway. The default Apple-VZ user-network gateway is `192.168.5.2`; override with `COWORK_VM_GATEWAY`,
955
+ gateway. The default Apple-VZ user-network gateway is `192.168.5.2`; override with `COWORK_VM_GATEWAY`
956
+ (a canonical IPv4 literal — an invalid value is rejected, as it feeds the guest iptables rule),
956
957
  and the proxy port with `COWORK_VM_PROXY_PORT` (unset, the host binds an OS-assigned free port;
957
958
  `8899` is only the guest-config fallback when a VM is spawned without an explicit port — not the
958
959
  effective default of a normal run). The harness threads one resolved
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
16
16
 
17
17
  Run it with:
18
18
 
19
- > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.8.0"`. (`replay` itself needs nothing else — no token, no Docker.)
19
+ > Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.9.0"`. (`replay` itself needs nothing else — no token, no Docker.)
20
20
 
21
21
  ```sh
22
22
  cowork-harness replay examples/replays/example-pdf-skill.cassette.json
package/llms.txt CHANGED
@@ -17,6 +17,7 @@ It is a fidelity fixture, not the Desktop runtime. The CLI binary is `cowork-har
17
17
  - [examples/README.md](examples/README.md): worked, copyable example scenarios + sessions + skills (from a source checkout — the npm package ships only examples/replays/; protocol/container tiers, token-free replay, prerequisites)
18
18
  - [schema/scenario.schema.json](schema/scenario.schema.json): JSON Schema for scenario files (machine-readable)
19
19
  - [schema/session.schema.json](schema/session.schema.json): JSON Schema for session files (machine-readable)
20
+ - [schema/critique-report.json](schema/critique-report.json): descriptive schema for `critique`'s JSON report / `critique-report.json` artifact (EXPERIMENTAL — not §12-frozen; field additions expected)
20
21
  - [.claude/skills/cowork-harness/references/ci-recipe.md](.claude/skills/cowork-harness/references/ci-recipe.md): copy-paste GitHub Actions — token-free replay PR gate + nightly live lane
21
22
 
22
23
  ## Concepts & internals
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "cowork-harness",
3
- "version": "1.8.0",
3
+ "version": "1.9.0",
4
4
  "description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -0,0 +1,221 @@
1
+ {
2
+ "$schema": "http://json-schema.org/draft-07/schema#",
3
+ "$id": "critique-report.json",
4
+ "title": "cowork-harness critique --output-format json report (and critique-report.json run-dir artifact)",
5
+ "description": "EXPERIMENTAL, descriptive schema — NOT a SPEC §12-frozen surface: additive field changes are expected between minor versions (unlike doctor.json). It exists so automation consumers (budget pacers, harvesters) have authoritative field names/shapes instead of prose. Exactly one of infraFailure / evaluatorError / evaluatorModel-with-findings describes the outcome; items is [] on both failure branches.",
6
+ "type": "object",
7
+ "required": ["skillFolder", "prompt", "sessionId", "outDir", "fidelity", "selfReportStatus", "verdictProvenance", "items"],
8
+ "additionalProperties": false,
9
+ "properties": {
10
+ "skillFolder": {
11
+ "type": "string",
12
+ "description": "the positional folder BOTH turns mounted (a plain skill, or a plugin root)"
13
+ },
14
+ "prompt": {
15
+ "type": "string"
16
+ },
17
+ "sessionId": {
18
+ "type": "string"
19
+ },
20
+ "outDir": {
21
+ "type": "string",
22
+ "description": "the kept run dir (turns/1 = graded, turns/2 = reflection; critique-* artifacts at its root)"
23
+ },
24
+ "fidelity": {
25
+ "type": "string",
26
+ "description": "the tier critique pinned for both turns (cowork is refused, so requested == resolved)",
27
+ "enum": ["container", "hostloop"]
28
+ },
29
+ "gradedEffectiveFidelity": {
30
+ "type": "string",
31
+ "description": "best-effort: the tier the graded turn's own result.json RECORDS (should equal fidelity; both surfaced so a mismatch is visible)"
32
+ },
33
+ "gradedBaseline": {
34
+ "type": "string",
35
+ "description": "best-effort: the graded turn's fingerprint.baseline (Desktop appVersion)"
36
+ },
37
+ "costUsd": {
38
+ "type": "object",
39
+ "description": "per-critique cost across the FOUR model workloads. GATE ON `complete` BEFORE TRUSTING totalUsd: false means one or more workloads were unpriced and the total UNDERCOUNTS true spend.",
40
+ "required": ["totalUsd", "complete"],
41
+ "additionalProperties": false,
42
+ "properties": {
43
+ "taskTurnUsd": {
44
+ "type": "number"
45
+ },
46
+ "reflectionTurnUsd": {
47
+ "type": "number"
48
+ },
49
+ "evaluatorPass1Usd": {
50
+ "type": "number"
51
+ },
52
+ "evaluatorPass2Usd": {
53
+ "type": "number"
54
+ },
55
+ "totalUsd": {
56
+ "type": "number"
57
+ },
58
+ "complete": {
59
+ "type": "boolean"
60
+ }
61
+ }
62
+ },
63
+ "gradedSkill": {
64
+ "type": "string",
65
+ "description": "the resolved skills/<name> the packager graded (--skill or single-skill auto-selection); ABSENT for a plain skill folder. Multi-skill-plugin pairing key: skillHash keys the whole mounted plugin, so pair by (gradedSkillHash, gradedSkill), never skillHash alone"
66
+ },
67
+ "skillInvocationObserved": {
68
+ "type": "boolean",
69
+ "description": "advisory: whether the graded run's own skillActivity mentions gradedSkill; false = the critique may be grading a run that never invoked the selected skill"
70
+ },
71
+ "gateAnswers": {
72
+ "type": "array",
73
+ "description": "the graded run's resolved gate answers, for a deterministic follow-up run",
74
+ "items": {
75
+ "type": "object",
76
+ "required": ["question", "answer", "answeredBy"],
77
+ "additionalProperties": false,
78
+ "properties": {
79
+ "question": {
80
+ "type": "string"
81
+ },
82
+ "answer": {
83
+ "type": "string"
84
+ },
85
+ "answeredBy": {
86
+ "type": "string"
87
+ }
88
+ }
89
+ }
90
+ },
91
+ "taskResult": {
92
+ "type": "string",
93
+ "description": "the graded turn's own result — error is a GRADEABLE outcome, not an instrument failure",
94
+ "enum": ["success", "error"]
95
+ },
96
+ "gradedOutcome": {
97
+ "type": "string",
98
+ "description": "the graded turn's result.json outcome (e.g. delivered_clean)"
99
+ },
100
+ "gradedSkillHash": {
101
+ "type": "string",
102
+ "description": "content-exact hash of the MOUNTED folder — the cross-FIX pairing key (per-plugin under --skill)"
103
+ },
104
+ "selfReportStatus": {
105
+ "type": "string",
106
+ "description": "unavailable = pass 2 was skipped; findings are pass-1-only",
107
+ "enum": ["captured", "unavailable"]
108
+ },
109
+ "evaluatorIntegrity": {
110
+ "type": "object",
111
+ "description": "mechanical canary check per pass; false = that pass ignored a trusted instruction — an empty critique may be adversarial silencing, not a clean skill",
112
+ "required": ["pass1Canary"],
113
+ "additionalProperties": false,
114
+ "properties": {
115
+ "pass1Canary": {
116
+ "type": "boolean"
117
+ },
118
+ "pass2Canary": {
119
+ "type": "boolean"
120
+ }
121
+ }
122
+ },
123
+ "droppedEvaluatorItems": {
124
+ "type": "object",
125
+ "description": "malformed items the per-item-tolerant parse dropped; non-zero means the surviving findings under-represent the full reply",
126
+ "required": ["pass1"],
127
+ "additionalProperties": false,
128
+ "properties": {
129
+ "pass1": {
130
+ "type": "number"
131
+ },
132
+ "pass2": {
133
+ "type": "number"
134
+ }
135
+ }
136
+ },
137
+ "turn1ResultDegraded": {
138
+ "type": "boolean",
139
+ "description": "the graded turn's canonical result was corrupted/never archived — result-derived sections are empty DEFAULTS, treat as unknown"
140
+ },
141
+ "turn1SliceDegraded": {
142
+ "type": "boolean",
143
+ "description": "the turn-1 transcript slice's boundary could not be trusted — treat gaps as unknown"
144
+ },
145
+ "skillMdStatus": {
146
+ "type": "string",
147
+ "description": "readability of the packaged SKILL.md; missing/unreadable force the mechanical already-covered -> not-adjudicable downgrade",
148
+ "enum": ["readable", "missing", "unreadable"]
149
+ },
150
+ "skillMdTruncated": {
151
+ "type": "boolean",
152
+ "description": "SKILL.md was READABLE but over its cap — the evaluator graded a CUT copy. Consequences: claims about past-cut content are steered to not-adjudicable, and a finding QUOTING past-cut text fails citation-resolution and lands in DROPPED"
153
+ },
154
+ "verdictProvenance": {
155
+ "type": "object",
156
+ "description": "advisory scoping: the verdict is a self-run discovery lead, never an independent attestation",
157
+ "required": ["kind", "advisory", "caveat"],
158
+ "additionalProperties": false,
159
+ "properties": {
160
+ "kind": {
161
+ "type": "string",
162
+ "enum": ["self-run"]
163
+ },
164
+ "advisory": {
165
+ "type": "boolean"
166
+ },
167
+ "caveat": {
168
+ "type": "string"
169
+ }
170
+ }
171
+ },
172
+ "infraFailure": {
173
+ "type": "string",
174
+ "description": "instrument failure (task turn killed / reflection protocol broke) — no critique was produced; items is []; process exits 2"
175
+ },
176
+ "evaluatorError": {
177
+ "type": "string",
178
+ "description": "the evaluator threw (embeds the raw reply) — no critique was produced; items is []; process exits 2; see critique-salvage.json"
179
+ },
180
+ "evaluatorModel": {
181
+ "type": "string",
182
+ "description": "the transport-RESOLVED evaluator model (never the requested alias); present only when the evaluator completed"
183
+ },
184
+ "items": {
185
+ "type": "array",
186
+ "description": "the citation-validated findings from both passes. citationResolved:false = the cited excerpt did not resolve verbatim against the evidence package (DROPPED section — transparency only, do not act on as-is)",
187
+ "items": {
188
+ "type": "object",
189
+ "required": ["source", "idea", "classification", "evidence", "recommendedAction"],
190
+ "additionalProperties": false,
191
+ "properties": {
192
+ "source": {
193
+ "type": "string",
194
+ "enum": ["evaluator", "self-report"]
195
+ },
196
+ "idea": {
197
+ "type": "string"
198
+ },
199
+ "classification": {
200
+ "type": "string",
201
+ "enum": ["grounded-and-actionable", "grounded-but-not-worth-it", "confabulated", "already-covered", "not-adjudicable"]
202
+ },
203
+ "evidence": {
204
+ "type": "string",
205
+ "description": "the model's cited excerpt — verbatim-checked against the evidence package"
206
+ },
207
+ "recommendedAction": {
208
+ "type": "string"
209
+ },
210
+ "citationResolved": {
211
+ "type": "boolean"
212
+ },
213
+ "findingFingerprint": {
214
+ "type": "string",
215
+ "description": "sha256/16 over normalized idea+classification+recommendedAction (evidence excluded). Cross-INPUT clustering key — HIGH-PRECISION, LOW-RECALL: idea is model-authored free text, so a MATCH proves the same finding recurred; a MISMATCH does NOT prove non-reproduction (the same finding reworded fingerprints differently)"
216
+ }
217
+ }
218
+ }
219
+ }
220
+ }
221
+ }