cowork-harness 1.8.0 → 1.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +21 -7
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/CHANGELOG.md +79 -0
- package/README.md +7 -7
- package/SPEC.md +7 -0
- package/dist/agent/session.js +11 -4
- package/dist/critique/command.js +39 -15
- package/dist/critique/package-evidence.js +10 -0
- package/dist/run/diff.js +12 -2
- package/dist/run/execute.js +5 -4
- package/dist/run/skill-flag-surface.js +1 -1
- package/dist/runtime/lima.js +13 -1
- package/docs/critique.md +32 -4
- package/docs/scenario.md +2 -1
- package/examples/replays/README.md +1 -1
- package/llms.txt +1 -0
- package/package.json +1 -1
- package/schema/critique-report.json +221 -0
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.9.0
|
|
7
|
+
tracks-harness: cowork-harness 1.9.0 (baseline desktop-1.24012.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.9.0` (baseline
|
|
26
26
|
> `desktop-1.24012.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,16 +39,28 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.9.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.9.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.9.0"`. **Pin `@>=1.9.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
43
|
|
|
44
|
-
What the ≥ 1.
|
|
44
|
+
What the ≥ 1.9.0 floor gates, by release:
|
|
45
|
+
|
|
46
|
+
- **1.9.0 (critique report schema + shareable, screenshot-safe output):**
|
|
47
|
+
`schema/critique-report.json` (descriptive, test-pinned — parse the report against it, not prose;
|
|
48
|
+
NOT §12-frozen while critique is EXPERIMENTAL), `gradedSkill` in the report (text header + JSON;
|
|
49
|
+
multi-skill pairing key — pair by `(gradedSkillHash, gradedSkill)`, since the hash alone keys the whole
|
|
50
|
+
plugin and cross-pairs its skills). Critique's text report and stderr diagnostics now **collapse host
|
|
51
|
+
paths to `~`** (a shared report or screenshot no longer leaks the username + filesystem layout; the JSON
|
|
52
|
+
report and persisted-artifact paths stay raw for machine consumers), a **missing or non-directory skill
|
|
53
|
+
folder is refused pre-spawn** (fast exit 2, no stray run dir left behind), and the report header + stderr
|
|
54
|
+
prefixes now read **`critique:`** (were `skill-critique:`).
|
|
45
55
|
|
|
46
56
|
- **1.8.0 (critique hardening + hostloop unpin):** `critique --skill <name>` (multi-skill-plugin
|
|
47
57
|
grading — a plugin root with no `--skill` is refused pre-spend; the evidence package gains the
|
|
48
58
|
invoked skill's `agents/<name>.md` + bounded `references/*.md` CONTENT), run-dir artifacts
|
|
49
59
|
(`critique-report.json` always, `critique-evidence-package.txt` when the evaluator ran,
|
|
50
60
|
`critique-salvage.json` on exit 2) + `--out <path>`, per-critique `costUsd` across all four
|
|
51
|
-
workloads, `findingFingerprint` per item (cross-INPUT clustering;
|
|
61
|
+
workloads, `findingFingerprint` per item (cross-INPUT clustering; HIGH-precision
|
|
62
|
+
LOW-recall — a match proves reproduction, a mismatch does NOT prove non-reproduction, the same
|
|
63
|
+
finding reworded fingerprints differently; skillHash stays the cross-FIX
|
|
52
64
|
key), gate-answer `--answer` echo in the report, per-item-tolerant evaluator parse
|
|
53
65
|
(`droppedEvaluatorItems`), `skillMdTruncated` (a readable-but-oversized SKILL.md is flagged
|
|
54
66
|
"graded a cut copy" — distinct from missing/unreadable), `subagents[].webSearches` +
|
|
@@ -436,7 +448,9 @@ cassette — has its own recipe:
|
|
|
436
448
|
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
437
449
|
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
438
450
|
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
439
|
-
kept run predates the current skill). **
|
|
451
|
+
kept run predates the current skill). **Multi-skill plugin caveat (post-1.7.0 CLIs with `--skill`):**
|
|
452
|
+
skillHash keys the whole MOUNTED plugin, so on a multi-skill plugin the hash alone cross-pairs
|
|
453
|
+
critiques of DIFFERENT skills — pair by the report's `(gradedSkillHash, gradedSkill)` pair. **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
|
|
440
454
|
pairing step there silently groups on an absent key instead of erroring — check the field is present, or
|
|
441
455
|
require ≥ 1.5.0. See `docs/debugging.md`
|
|
442
456
|
(repo-only) for the full loop.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.9.0` (baseline `desktop-1.24012.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.9.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.9.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.9.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.9.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.9.0`
|
|
4
4
|
(baseline `desktop-1.24012.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 1.
|
|
5
|
+
Tracks `cowork-harness 1.9.0` (baseline `desktop-1.24012.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,85 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.9.0] — 2026-07-24
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`schema/critique-report.json`** — a descriptive, test-pinned schema for `critique`'s JSON report /
|
|
14
|
+
`critique-report.json` artifact, so automation consumers (budget pacers gating on `costUsd.complete`,
|
|
15
|
+
harvesters) parse field names/shapes from a schema instead of prose. Deliberately **not** a SPEC
|
|
16
|
+
§12-frozen surface (unlike `doctor.json`) — critique is EXPERIMENTAL and additive field changes are
|
|
17
|
+
expected; the schema says so in its own description, and a two-way sync test pins it against the
|
|
18
|
+
actual report builder on every branch (findings / infraFailure / evaluatorError).
|
|
19
|
+
- **`gradedSkill` in the critique report** (text header + JSON): the resolved `skills/<name>` the
|
|
20
|
+
packager graded under `--skill`/auto-selection. Load-bearing for multi-skill plugins:
|
|
21
|
+
`gradedSkillHash` keys the whole mounted plugin, so pairing by hash alone cross-pairs critiques of
|
|
22
|
+
*different* skills — pair by `(gradedSkillHash, gradedSkill)`. Docs updated accordingly.
|
|
23
|
+
|
|
24
|
+
### Docs
|
|
25
|
+
|
|
26
|
+
- Truncation→DROPPED mechanics: a finding whose `evidence` quotes SKILL.md text past the packaging cap
|
|
27
|
+
fails citation-resolution and lands in DROPPED (the check runs against the *cut* copy) — documented
|
|
28
|
+
next to `skillMdTruncated` so a back-half DROPPED skew on an oversized skill has its cause named.
|
|
29
|
+
- `findingFingerprint` direction-of-inference: high-precision, LOW-RECALL — a match proves
|
|
30
|
+
reproduction; a mismatch does NOT prove non-reproduction (the same finding reworded fingerprints
|
|
31
|
+
differently). The Reproduction section says so before anyone concludes "didn't reproduce".
|
|
32
|
+
- Evaluator cost share: the two evaluator passes were ~3/4 of a measured e2e total — the
|
|
33
|
+
calibrate-then-`--evaluator-model` strategy is now in the cost section and `critique --help`, with
|
|
34
|
+
the armor-verification-is-default-evaluator-only caveat.
|
|
35
|
+
- SPEC §12 now names the critique report explicitly under **NOT covered**: `schema/critique-report.json`
|
|
36
|
+
is descriptive (parse against it, not prose) but not the compatibility contract while critique is
|
|
37
|
+
EXPERIMENTAL — its surface-baseline presence is for change visibility, and it is the promotion
|
|
38
|
+
candidate on the `doctor.json` template once critique stabilizes. README, llms.txt, and the shipped
|
|
39
|
+
skill point at the schema and the `(gradedSkillHash, gradedSkill)` pairing rule.
|
|
40
|
+
|
|
41
|
+
### Fixed
|
|
42
|
+
|
|
43
|
+
- **critique's "Attached inputs" evidence no longer reports `(none)` when `mounts.json` is corrupt.**
|
|
44
|
+
`listAttachedInputs` derived connected-folder names from `loadVmPathContext`, which returns `null` for
|
|
45
|
+
BOTH an absent `mounts.json` (legitimately no mounts) and a present-but-unparseable one (the folder map
|
|
46
|
+
is UNKNOWN). Uploads already distinguished these (the ENOENT-vs-read-fault split), but folders — which
|
|
47
|
+
have no fixed-layout fallback — silently collapsed to `[]`, so a corrupt `mounts.json` rendered `(none)`,
|
|
48
|
+
telling the evaluator "the agent correctly saw no connected folder" when the truth was unknown. It now
|
|
49
|
+
surfaces a corrupt `mounts.json` as an explicit UNKNOWN note, completing the same
|
|
50
|
+
confabulation-vs-correct guard the uploads path already applies.
|
|
51
|
+
- **`diff` no longer treats two tool inputs that differ only past the ~2000-char cap as the same call.**
|
|
52
|
+
`canonicalizeInput` truncated a tool input's canonical JSON to a 2000-char cap and used that truncated
|
|
53
|
+
string as the tool-sequence equality key, so two `Write`/`Bash` calls sharing a long identical prefix
|
|
54
|
+
but differing only in the dropped tail compared as `op: "same"` in `diffToolSequence` — a false "no
|
|
55
|
+
change" that could flip the advisory `diff` exit code to 0. A truncated key now folds in a
|
|
56
|
+
`#<len>·<sha16>` hash of the full canonical string, so the key depends on the entire content while the
|
|
57
|
+
visible prefix stays readable in the hunk; both diff sides canonicalize identically, so the comparison
|
|
58
|
+
stays consistent.
|
|
59
|
+
- **`COWORK_VM_GATEWAY` is now validated as a canonical IPv4 literal.** The L2-microVM gateway override was
|
|
60
|
+
interpolated verbatim into a root-run guest `iptables -A OUTPUT -d <gateway>` command (via `sh -c`), so a
|
|
61
|
+
malformed or hostile value could inject shell syntax into privileged provisioning — or, more mundanely,
|
|
62
|
+
leave the firewall in an unknown state. `vmGatewayIp()` now rejects anything that isn't a canonical IPv4
|
|
63
|
+
literal (digits-and-dots only, octets 0–255, no leading zeros), failing loud instead of reaching the
|
|
64
|
+
shell. Operator-set env var, so this is defense-in-depth; no valid gateway value is affected.
|
|
65
|
+
- **The `agent.stderr.log` sink is now flushed before the teardown secret-scrub reads it.** The stderr sink
|
|
66
|
+
was piped fire-and-forget and never awaited, so bytes still buffered when `scrubRawRunLogs` read the file
|
|
67
|
+
could land raw *afterwards* — a persisted-secret leak in a narrow teardown window. `LiveAgentSession` now
|
|
68
|
+
pipes it with `{ end: false }` and ends+awaits it in the same session-teardown drain that already flushes
|
|
69
|
+
`events.jsonl` / `control-out.jsonl`, so the session generator resolves only after the sink is fully
|
|
70
|
+
flushed — the scrub always sees the complete log.
|
|
71
|
+
- **`critique` no longer prints raw host paths in its report or diagnostics.** The text report's `run dir:`
|
|
72
|
+
line, the `inspect <dir>` hints, the write-failure diagnostics, and the echoed skill-folder path all
|
|
73
|
+
printed absolute `$HOME`-rooted paths, so a shared report or screenshot leaked the username + filesystem
|
|
74
|
+
layout (it landed in a video frame). critique was the one rendering path in the CLI that never called
|
|
75
|
+
`tildeify`, while `skill`/`run` scrub unconditionally. Every human-facing path is now collapsed to `~`;
|
|
76
|
+
the JSON report and persisted-artifact paths stay raw (machine data a consumer feeds back to a tool). The
|
|
77
|
+
`--demo` rejection is unchanged — it was never the fix — but its reason now notes the report already
|
|
78
|
+
collapses paths.
|
|
79
|
+
- **`critique` fails fast on a missing or non-directory skill folder.** A typo'd/absent positional folder
|
|
80
|
+
previously minted a session and spawned the task turn before infra-failing (exit 2), leaving a stray run
|
|
81
|
+
dir behind. `resolveCritiquedSkillDir` now `existsSync`/`isDirectory`-checks the folder up front — before
|
|
82
|
+
any session is minted or spawned — so a bad path exits 2 immediately with nothing left on disk. A
|
|
83
|
+
present-but-`SKILL.md`-less folder still defers to the packager's degraded flow, unchanged.
|
|
84
|
+
- **`critique`'s report header and stderr diagnostics now say `critique:` (were `skill-critique:`).** A
|
|
85
|
+
leftover label from the `scripts/skill-critique.ts` instrument; the invoked command is `critique`.
|
|
86
|
+
Cosmetic, no schema change.
|
|
87
|
+
|
|
9
88
|
## [1.8.0] — 2026-07-23
|
|
10
89
|
|
|
11
90
|
### Added
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.9.0"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.9.0"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.9.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.9.0"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -690,7 +690,7 @@ jobs:
|
|
|
690
690
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
691
691
|
```
|
|
692
692
|
|
|
693
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.
|
|
693
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.9.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
694
694
|
|
|
695
695
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **seven-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
696
696
|
|
|
@@ -752,7 +752,7 @@ Most runs need **none** of these — the defaults are correct. They're grouped b
|
|
|
752
752
|
- **Staleness boundary** (the git-tracked-files-only default and its `git add`/hard-fail rationale are explained in [Test a local skill](#test-a-local-skill-in-one-command)): a **non-repo** source dir falls back to a raw walk instead; OS-junk like `.DS_Store` is always excluded; a partial exclusion emits a `::notice:: [stage]`. `COWORK_HARNESS_GITSET=0` opts out to a raw walk for every dir (and copies untracked files too); `COWORK_HARNESS_DEBUG_SKILLHASH=1` dumps the exact file set feeding the staleness hash on a mismatch (and flags OS-junk) so a drift source is one line. Declare per-plugin non-runtime paths in a `.cowork-hashignore` file (or the session `staleness.hash_ignore`). `COWORK_HARNESS_AGENT_SCOPE=skill` (opt-in) refines a scenario's `skills:` scope so a **skill-named** `agents/<name>.md` re-stales only that skill's cassettes instead of the whole fleet (generic agents stay shared; stamped into the fingerprint, so flipping it is a one-time re-record like `GITSET`).
|
|
753
753
|
- **`skill` / `chat` defaults:** `COWORK_HARNESS_FIDELITY` sets the default fidelity tier for ad-hoc `skill`/`chat` runs (a `--fidelity` flag or a scenario's `fidelity:` still wins); `COWORK_HARNESS_MODEL` sets the default model; `COWORK_HARNESS_OUTPUT_FORMAT` (`text`|`json`) sets the default output format. Each is overridden by the matching explicit flag. **Caveat:** `chat` only accepts `protocol`/`container`/`hostloop` — `microvm` and `cowork` are rejected (no Lima/auto-pick plumbing in the interactive REPL), so a `COWORK_HARNESS_FIDELITY` set to either is rejected loudly for `chat` even though `skill` accepts the full tier set.
|
|
754
754
|
- **Secret scrubbing:** `COWORK_HARNESS_SCRUB_KEYS=<KEY1,KEY2>` adds extra env-var names whose values are redacted from logs (beyond the known auth tokens + `ANTHROPIC_CUSTOM_HEADERS`); `COWORK_HARNESS_SCRUB_VALUES=<v1,v2>` redacts literal values regardless of env. **Committed-cassette redaction:** `COWORK_HARNESS_REDACT_PATTERNS=<rx1,rx2>` / `COWORK_HARNESS_REDACT_KEYS=<k1,k2>` extend the privacy layer that scrubs recorded `controlOut` before a cassette is written for commit.
|
|
755
|
-
- L2 microVM: `COWORK_VM_GATEWAY` overrides the Lima host-proxy gateway IP (default `192.168.5.2
|
|
755
|
+
- L2 microVM: `COWORK_VM_GATEWAY` overrides the Lima host-proxy gateway IP (default `192.168.5.2`; must be a canonical IPv4 literal — an invalid value is rejected, since it is interpolated into the guest firewall rule); `COWORK_VM_PROXY_PORT` pins the egress-proxy port (unset, the host binds an OS-assigned free port and threads that same value into the guest firewall + `HTTP(S)_PROXY`). The Lima instance is named `cowork-vm-<config-hash>` (a config change → a fresh VM); `COWORK_LIMA_INSTANCE` pins a fixed name, and `vm prune` removes orphaned ones.
|
|
756
756
|
- **Advanced / internal escape hatches** (rarely needed): `PYTHON` overrides the interpreter for `lint` / scenario tooling (default `python3`); `COWORK_HARNESS_DEBUG=1` surfaces which `.env` files were loaded; `COWORK_HARNESS_CLAUDE_BIN=<path>` points the `--decider-llm` transport at a specific `claude` binary; `COWORK_HARNESS_ALLOW_AGENT_FALLBACK=1` lets the harness use the newest sibling agent binary when the baseline-pinned version is missing (a fidelity compromise — off by default) — a same-major.minor **patch** bump of the staged NATIVE binary is auto-accepted without this flag (the native binary carries no sha256 pin, so a patch drift is safe by default; it prints a loud stderr note naming the pinned and substituted versions). At `hostloop`, and at `cowork` **only when it resolves to host-loop** on the synced baseline, the staged **VM ELF** is auto-accepted on a patch bump too, because on that path it's a non-executed parity mount into the bash sidecar; a `cowork` baseline that resolves to VM-loop instead executes the ELF directly, so it keeps the strict sha-pinned exact-version match, same as `container`/`microvm`, which always keep it (the ELF is the executed agent there), so the flag remains required for any ELF drift on those tiers/paths, and for a major/minor gap everywhere; `COWORK_MANAGED_CONFIG=1` forces the managed-config path on `protocol`; `COWORK_HARNESS_ALLOW_MISSING_PROMPT=1` downgrades a missing prompt asset to a warning; `COWORK_HARNESS_SAFE_STAGING_PREFIX=<a,b>` whitelists prefixes under which a delete in `outputs/` is allowed (otherwise delete-in-outputs fails loud); `COWORK_HARNESS_NO_HEARTBEAT=1` disables the idle-run heartbeat, `COWORK_HARNESS_HEARTBEAT_MS` tunes its interval (default 30000ms / 30s); `COWORK_HARNESS_CLI=/path/to/cli.js` overrides which built CLI the Python `cowork` pytest lane drives (see python/README.md).
|
|
757
757
|
- Pin `baseline: desktop-<ver>` and `model:` in a session for byte-stable runs; use `latest` to track.
|
|
758
758
|
|
|
@@ -801,7 +801,7 @@ This repo is built to be driven by agents, not just read by humans:
|
|
|
801
801
|
- **[AGENTS.md](./AGENTS.md)** — the canonical agent-instructions file (architecture seams, the build gate, invariants, ethos). Read it before changing code. Also indexed in **[llms.txt](./llms.txt)**.
|
|
802
802
|
- **Companion skill** — [`.claude/skills/cowork-harness/`](./.claude/skills/cowork-harness/SKILL.md) teaches an agent to drive the harness; install it via the marketplace (see [above](#drive-it-from-claude-code-companion-skill)).
|
|
803
803
|
- **Machine-readable interfaces** — stable `--output-format json` envelope on stdout, deterministic exit codes (`0`/`1`/`2`/`3`, with a couple of documented per-command exceptions — see [SPEC.md](./SPEC.md) for the full table), and `--help` on every command.
|
|
804
|
-
- **JSON Schemas** — [`schema/scenario.schema.json`](./schema/scenario.schema.json) and [`schema/session.schema.json`](./schema/session.schema.json) describe every field of the YAML you author (generated from the source schemas; `npm run schema`). [`schema/protocol.v1.json`](./schema/protocol.v1.json) (hand-authored) schemas the harness's own control-channel wire protocol, with a golden vector pack at [`fixtures/protocol/v1/`](./fixtures/protocol/v1/) — see [docs/protocol.md](./docs/protocol.md).
|
|
804
|
+
- **JSON Schemas** — [`schema/scenario.schema.json`](./schema/scenario.schema.json) and [`schema/session.schema.json`](./schema/session.schema.json) describe every field of the YAML you author (generated from the source schemas; `npm run schema`). [`schema/protocol.v1.json`](./schema/protocol.v1.json) (hand-authored) schemas the harness's own control-channel wire protocol, with a golden vector pack at [`fixtures/protocol/v1/`](./fixtures/protocol/v1/) — see [docs/protocol.md](./docs/protocol.md). [`schema/critique-report.json`](./schema/critique-report.json) describes `critique`'s JSON report / `critique-report.json` artifact for automation consumers (budget pacers, harvesters) — **descriptive, not §12-frozen** while critique is EXPERIMENTAL (see [SPEC.md §12](./SPEC.md#12-versioning--the-10-compatibility-contract)).
|
|
805
805
|
|
|
806
806
|
`AGENTS.md`, `SPEC.md`, and `DESIGN.md` **are** shipped in the npm package (see `package.json` `files`) —
|
|
807
807
|
a global install has them locally too, not just on GitHub.
|
package/SPEC.md
CHANGED
|
@@ -793,6 +793,13 @@ Covered-surface changes follow semver as of `1.0.0` — see [RELEASING.md](./REL
|
|
|
793
793
|
the `doctor`/`verify-cassettes`/RunResult envelopes above, these have no `schema/*.json` and may change
|
|
794
794
|
(fields, rule ids, the artifact-write-back finding shape) while the analyzers stabilize. Parse at your
|
|
795
795
|
own risk until they are promoted to a covered surface.
|
|
796
|
+
- **The `critique` report** (`--output-format json` / the `critique-report.json` run-dir artifact) —
|
|
797
|
+
NOT yet frozen, and — uniquely — it DOES have a schema file: `schema/critique-report.json` is
|
|
798
|
+
**descriptive** (authoritative field names/shapes, test-pinned against the builder) so automation can
|
|
799
|
+
parse against a schema rather than prose, but it is NOT this section's compatibility contract —
|
|
800
|
+
critique is EXPERIMENTAL and additive field changes may land in any minor release (the schema's own
|
|
801
|
+
`description` says so). Its presence in the surface-drift baseline is for change *visibility*, not
|
|
802
|
+
coverage. It is the promotion CANDIDATE once critique stabilizes, on the `doctor.json` template.
|
|
796
803
|
- **`docs/internal/**`** — untracked working notes.
|
|
797
804
|
- **The reconstructed system-prompt append text** — a paraphrase by design (see
|
|
798
805
|
[docs/fidelity-gaps.md](./docs/fidelity-gaps.md)); behaviorally equivalent, not byte-stable.
|
package/dist/agent/session.js
CHANGED
|
@@ -242,6 +242,7 @@ export class LiveAgentSession {
|
|
|
242
242
|
proc;
|
|
243
243
|
events;
|
|
244
244
|
controlOut;
|
|
245
|
+
errLog;
|
|
245
246
|
timeline;
|
|
246
247
|
lineIndex = 0;
|
|
247
248
|
reqById = new Map();
|
|
@@ -275,8 +276,11 @@ export class LiveAgentSession {
|
|
|
275
276
|
this.events = createWriteStream(join(outDir, "events.jsonl"), { flags: "a" });
|
|
276
277
|
this.controlOut = createWriteStream(join(outDir, "control-out.jsonl"), { flags: "a" });
|
|
277
278
|
this.timeline = new TimelineWriter(outDir);
|
|
278
|
-
|
|
279
|
-
this
|
|
279
|
+
// #60: keep a reference and pipe with `{ end: false }` so WE own the close — the drain in `start()`'s
|
|
280
|
+
// finally ends this stream and AWAITS its flush before the generator resolves, so a subsequent
|
|
281
|
+
// `scrubRawRunLogs` can't read the file while bytes are still buffered (a raw-secret-after-scrub leak).
|
|
282
|
+
this.errLog = createWriteStream(join(outDir, "agent.stderr.log"), { flags: "a" });
|
|
283
|
+
this.proc.stderr.pipe(this.errLog, { end: false });
|
|
280
284
|
// keep a bounded stderr tail and capture the exit code/signal so a child that dies nonzero
|
|
281
285
|
// (with no structured {type:"result"} error) is surfaced as a typed error event, not a silent stop.
|
|
282
286
|
this.proc.stderr.on("data", (d) => {
|
|
@@ -404,11 +408,14 @@ export class LiveAgentSession {
|
|
|
404
408
|
this.closing = true; // discard any late write() from translate()'s async hook/mcp paths
|
|
405
409
|
await this.drainAll(); // queue empty + all callbacks confirmed before ending either stream
|
|
406
410
|
// AWAIT the stream flush before the generator resolves. executeScenario reads/scans/scrubs
|
|
407
|
-
// events.jsonl + control-out.jsonl immediately after `drive()` returns; a
|
|
408
|
-
// races the final buffered writes. end(cb) fires the callback on 'finish'
|
|
411
|
+
// events.jsonl + control-out.jsonl + agent.stderr.log immediately after `drive()` returns; a
|
|
412
|
+
// fire-and-forget end() races the final buffered writes. end(cb) fires the callback on 'finish'
|
|
413
|
+
// (fully flushed). #60: agent.stderr.log is ended here too — piped with `{ end: false }` (see the
|
|
414
|
+
// ctor) so its last buffered bytes flush BEFORE the scrub reads it, never landing raw afterwards.
|
|
409
415
|
await Promise.all([
|
|
410
416
|
new Promise((res) => this.events.end(() => res())),
|
|
411
417
|
new Promise((res) => this.controlOut.end(() => res())),
|
|
418
|
+
new Promise((res) => this.errLog.end(() => res())),
|
|
412
419
|
new Promise((res) => this.timeline.end(() => res())),
|
|
413
420
|
]);
|
|
414
421
|
}
|
package/dist/critique/command.js
CHANGED
|
@@ -17,7 +17,8 @@ import { fileURLToPath } from "node:url";
|
|
|
17
17
|
import { lookupSkillFlag } from "../run/skill-flag-surface.js";
|
|
18
18
|
import { gradedAliasPath, turnArtifactPath } from "../run/turn-layout.js";
|
|
19
19
|
import { renderKnownLimitations } from "./limitations.js";
|
|
20
|
-
import {
|
|
20
|
+
import { tildeify } from "../io.js";
|
|
21
|
+
import { existsSync, readFileSync, copyFileSync, writeFileSync, readdirSync, statSync } from "node:fs";
|
|
21
22
|
import { randomUUID } from "node:crypto";
|
|
22
23
|
import { writeSync } from "node:fs";
|
|
23
24
|
import { basename, join } from "node:path";
|
|
@@ -98,6 +99,7 @@ Not accepted (each errors with its reason rather than being silently ignored):
|
|
|
98
99
|
--repeat + companions fixed two-turn protocol — loop critique itself; pair by fingerprint.skillHash
|
|
99
100
|
--ablate-skill grading a skill you removed is incoherent
|
|
100
101
|
--quiet/--verbose/--compact/--demo/--dry-run inner-turn rendering or preview — no effect on the report
|
|
102
|
+
(which already collapses host paths to ~)
|
|
101
103
|
|
|
102
104
|
Repeating a flag: --upload/--folder/--plugin/--marketplace/--enable/--answer accumulate (that is how you
|
|
103
105
|
pass several). Every other value-taking flag is single-valued and repeating it is a USAGE ERROR rather
|
|
@@ -107,8 +109,10 @@ Repeating a flag: --upload/--folder/--plugin/--marketplace/--enable/--answer acc
|
|
|
107
109
|
COST AND PREREQUISITES — read before running:
|
|
108
110
|
* Each critique is FOUR model workloads: two graded runs (task + reflection) at the chosen tier and two
|
|
109
111
|
evaluator passes over an evidence package of up to ${MAX_PACKAGE_BYTES / 1024}KB.
|
|
110
|
-
* The evaluator defaults to ${DEFAULT_EVALUATOR_MODEL} — the most expensive tier
|
|
111
|
-
|
|
112
|
+
* The evaluator defaults to ${DEFAULT_EVALUATOR_MODEL} — the most expensive tier — and the two
|
|
113
|
+
evaluator passes DOMINATE spend (~3/4 of a measured e2e total). For a batch: calibrate with run 1's
|
|
114
|
+
costUsd (gate on costUsd.complete), then consider a cheaper --evaluator-model for the sweep — noting
|
|
115
|
+
the armor's injection-resistance is verified for the DEFAULT evaluator only.
|
|
112
116
|
* container needs Docker/Lima; hostloop needs Docker (the bash/web_fetch sidecar) PLUS the staged native
|
|
113
117
|
agent binary, and writes to the real host FS (a writable --folder requires --allow-host-writes). Both
|
|
114
118
|
tiers need an authenticated \`claude\` CLI on PATH.
|
|
@@ -361,6 +365,22 @@ function parseArgs(argv) {
|
|
|
361
365
|
* folder → same hash — so generation pairing keeps working; it is a per-plugin key, not per-skill.
|
|
362
366
|
* Exported for unit tests. */
|
|
363
367
|
export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
|
|
368
|
+
// Fail-fast on a typo'd / absent path BEFORE the caller mints a session and spawns the task turn — a
|
|
369
|
+
// missing folder otherwise only surfaces as a mid-run mount failure that leaves a stray run dir behind.
|
|
370
|
+
// This lives here (not in parseArgs) on purpose: parseArgs is unit-tested with fictitious paths, whereas
|
|
371
|
+
// this resolver is only ever called with a real folder. One statSync in a try/catch also covers a broken
|
|
372
|
+
// symlink; the existsSync guard just gives the common typo the clearer "not found" message.
|
|
373
|
+
if (!existsSync(skillFolder))
|
|
374
|
+
throw new Error(`skill folder not found: ${tildeify(skillFolder)}`);
|
|
375
|
+
let stat;
|
|
376
|
+
try {
|
|
377
|
+
stat = statSync(skillFolder);
|
|
378
|
+
}
|
|
379
|
+
catch {
|
|
380
|
+
throw new Error(`skill folder not found: ${tildeify(skillFolder)}`);
|
|
381
|
+
}
|
|
382
|
+
if (!stat.isDirectory())
|
|
383
|
+
throw new Error(`not a directory: ${tildeify(skillFolder)}`);
|
|
364
384
|
const agentsMdFor = (name) => {
|
|
365
385
|
const p = join(skillFolder, "agents", `${name}.md`);
|
|
366
386
|
return existsSync(p) ? p : undefined;
|
|
@@ -380,7 +400,7 @@ export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
|
|
|
380
400
|
const candidate = join(skillFolder, "skills", skillSelector);
|
|
381
401
|
if (!existsSync(join(candidate, "SKILL.md"))) {
|
|
382
402
|
const available = listPluginSkills();
|
|
383
|
-
throw new Error(`--skill ${skillSelector}: no skills/${skillSelector}/SKILL.md under ${skillFolder}` +
|
|
403
|
+
throw new Error(`--skill ${skillSelector}: no skills/${skillSelector}/SKILL.md under ${tildeify(skillFolder)}` +
|
|
384
404
|
(available.length ? ` — available skills: ${available.join(", ")}` : ` — no skills/<name>/SKILL.md found at all`));
|
|
385
405
|
}
|
|
386
406
|
return { skillDir: candidate, agentsMdPath: agentsMdFor(skillSelector) };
|
|
@@ -391,7 +411,7 @@ export function resolveCritiquedSkillDir(skillFolder, skillSelector) {
|
|
|
391
411
|
if (skills.length === 1)
|
|
392
412
|
return { skillDir: join(skillFolder, "skills", skills[0]), agentsMdPath: agentsMdFor(skills[0]), autoSelectedSkill: skills[0] };
|
|
393
413
|
if (skills.length > 1)
|
|
394
|
-
throw new Error(`${skillFolder} is a multi-skill plugin root (no root SKILL.md; skills: ${skills.join(", ")}) — ` +
|
|
414
|
+
throw new Error(`${tildeify(skillFolder)} is a multi-skill plugin root (no root SKILL.md; skills: ${skills.join(", ")}) — ` +
|
|
395
415
|
`pass --skill <name> so critique grades the INVOKED skill's SKILL.md instead of a missing root one`);
|
|
396
416
|
return { skillDir: skillFolder }; // no SKILL.md anywhere — the packager's existing missing/degraded flow reports it
|
|
397
417
|
}
|
|
@@ -702,10 +722,10 @@ export const VERDICT_PROVENANCE = {
|
|
|
702
722
|
export function buildTextReport(state) {
|
|
703
723
|
const { skillFolder, prompt, sessionId, outDir, taskResult, gradedOutcome, gradedSkillHash, selfReportStatus, items, evaluatorModel, requestedModel, evaluatorError, infraFailure, turn1ResultDegraded, turn1SliceDegraded, skillMdStatus, } = state;
|
|
704
724
|
const out = [];
|
|
705
|
-
out.push(`
|
|
725
|
+
out.push(`critique: ${tildeify(skillFolder)}`);
|
|
706
726
|
out.push(` probe: ${prompt}`);
|
|
707
727
|
out.push(` session: ${sessionId}`);
|
|
708
|
-
out.push(` run dir: ${outDir}`);
|
|
728
|
+
out.push(` run dir: ${tildeify(outDir)}`);
|
|
709
729
|
out.push(` fidelity: ${state.fidelity}` +
|
|
710
730
|
(state.gradedEffectiveFidelity
|
|
711
731
|
? ` (graded turn recorded ${state.gradedEffectiveFidelity}${state.gradedBaseline ? `, baseline ${state.gradedBaseline}` : ""})`
|
|
@@ -722,6 +742,8 @@ export function buildTextReport(state) {
|
|
|
722
742
|
// reflection turn's result.json for them).
|
|
723
743
|
if (gradedOutcome)
|
|
724
744
|
out.push(` graded outcome: ${gradedOutcome}`);
|
|
745
|
+
if (state.gradedSkill)
|
|
746
|
+
out.push(` graded skill: ${state.gradedSkill} (pair by skillHash + this name — skillHash keys the whole mounted plugin)`);
|
|
725
747
|
if (gradedSkillHash)
|
|
726
748
|
out.push(` graded skillHash: ${gradedSkillHash.slice(0, 12)}`);
|
|
727
749
|
if (evaluatorModel)
|
|
@@ -750,7 +772,7 @@ export function buildTextReport(state) {
|
|
|
750
772
|
// The dominant real-world cause of a "missing" SKILL.md is pointing critique at a MULTI-SKILL PLUGIN
|
|
751
773
|
// root (skills/<name>/SKILL.md, no root SKILL.md) — name the cause and the fix, not just the symptom.
|
|
752
774
|
if (skillMdStatus === "missing")
|
|
753
|
-
out.push(` NOTE: if ${skillFolder} is a multi-skill plugin root, pass --skill <name> (or point critique at <plugin>/skills/<name> directly) so the invoked skill's SKILL.md is graded.`);
|
|
775
|
+
out.push(` NOTE: if ${tildeify(skillFolder)} is a multi-skill plugin root, pass --skill <name> (or point critique at <plugin>/skills/<name> directly) so the invoked skill's SKILL.md is graded.`);
|
|
754
776
|
out.push(` verdict scope: advisory self-run — NOT an independent attestation (never gate a skill on it)`);
|
|
755
777
|
out.push("");
|
|
756
778
|
const dropped = state.droppedEvaluatorItems;
|
|
@@ -768,12 +790,12 @@ export function buildTextReport(state) {
|
|
|
768
790
|
}
|
|
769
791
|
if (infraFailure) {
|
|
770
792
|
out.push(`INFRASTRUCTURE/PROTOCOL FAILURE (reflection turn): ${infraFailure}`);
|
|
771
|
-
out.push(`The evaluator was NOT invoked — this is a broken discovery run, not a critique. Re-run, or inspect ${outDir} directly.`);
|
|
793
|
+
out.push(`The evaluator was NOT invoked — this is a broken discovery run, not a critique. Re-run, or inspect ${tildeify(outDir)} directly.`);
|
|
772
794
|
return out.join("\n");
|
|
773
795
|
}
|
|
774
796
|
if (evaluatorError) {
|
|
775
797
|
out.push(`EVALUATOR FAILED: ${evaluatorError}`);
|
|
776
|
-
out.push(`No critique items were produced. Re-run, or inspect ${outDir} directly.`);
|
|
798
|
+
out.push(`No critique items were produced. Re-run, or inspect ${tildeify(outDir)} directly.`);
|
|
777
799
|
return out.join("\n");
|
|
778
800
|
}
|
|
779
801
|
const byBucket = new Map();
|
|
@@ -828,6 +850,7 @@ export function buildJsonReport(state) {
|
|
|
828
850
|
gradedEffectiveFidelity: state.gradedEffectiveFidelity,
|
|
829
851
|
gradedBaseline: state.gradedBaseline,
|
|
830
852
|
costUsd: state.costUsd,
|
|
853
|
+
gradedSkill: state.gradedSkill,
|
|
831
854
|
skillInvocationObserved: state.skillInvocationObserved,
|
|
832
855
|
gateAnswers: state.gateAnswers,
|
|
833
856
|
taskResult,
|
|
@@ -894,7 +917,7 @@ function writeRunArtifact(outDir, name, content) {
|
|
|
894
917
|
writeFileSync(join(outDir, name), content);
|
|
895
918
|
}
|
|
896
919
|
catch (e) {
|
|
897
|
-
process.stderr.write(`
|
|
920
|
+
process.stderr.write(`critique: could not write ${name} under ${tildeify(outDir)}: ${String(e)}\n`);
|
|
898
921
|
}
|
|
899
922
|
}
|
|
900
923
|
/** Persist the run-dir artifacts every critique leaves behind (all best-effort):
|
|
@@ -925,7 +948,7 @@ function writeOutFile(outPath, state, outputFormat) {
|
|
|
925
948
|
writeFileSync(outPath, content);
|
|
926
949
|
}
|
|
927
950
|
catch (e) {
|
|
928
|
-
process.stderr.write(`
|
|
951
|
+
process.stderr.write(`critique: --out ${tildeify(outPath)} could not be written: ${String(e)}\n`);
|
|
929
952
|
}
|
|
930
953
|
}
|
|
931
954
|
export function buildTaskTurnArgs(opts, sessionId) {
|
|
@@ -1000,7 +1023,7 @@ async function main(argv = process.argv.slice(2)) {
|
|
|
1000
1023
|
return;
|
|
1001
1024
|
}
|
|
1002
1025
|
if (resolvedSkill.autoSelectedSkill)
|
|
1003
|
-
process.stderr.write(`::notice:: [critique] ${opts.skillFolder} is a single-skill plugin — grading skills/${resolvedSkill.autoSelectedSkill}/SKILL.md (pass --skill to be explicit)\n`);
|
|
1026
|
+
process.stderr.write(`::notice:: [critique] ${tildeify(opts.skillFolder)} is a single-skill plugin — grading skills/${resolvedSkill.autoSelectedSkill}/SKILL.md (pass --skill to be explicit)\n`);
|
|
1004
1027
|
const sessionId = `crit-${randomUUID()}`;
|
|
1005
1028
|
try {
|
|
1006
1029
|
// 1. Task turn.
|
|
@@ -1017,7 +1040,7 @@ async function main(argv = process.argv.slice(2)) {
|
|
|
1017
1040
|
]
|
|
1018
1041
|
.filter(Boolean)
|
|
1019
1042
|
.join("; ");
|
|
1020
|
-
process.stderr.write(`
|
|
1043
|
+
process.stderr.write(`critique: could not determine the task run's directory (no envelope outDir and no [status] line)${diag ? ` [${diag}]` : ""}.\n` +
|
|
1021
1044
|
`--- task stdout ---\n${task.stdout}\n--- task stderr (tail) ---\n${task.stderr.slice(-4000)}\n`);
|
|
1022
1045
|
process.exit(EXIT_INSTRUMENT_FAILURE);
|
|
1023
1046
|
return;
|
|
@@ -1216,6 +1239,7 @@ async function main(argv = process.argv.slice(2)) {
|
|
|
1216
1239
|
gradedEffectiveFidelity,
|
|
1217
1240
|
gradedBaseline,
|
|
1218
1241
|
costUsd,
|
|
1242
|
+
gradedSkill: gradedSkillName,
|
|
1219
1243
|
skillInvocationObserved,
|
|
1220
1244
|
gateAnswers: gateAnswers?.length ? gateAnswers : undefined,
|
|
1221
1245
|
taskResult,
|
|
@@ -1254,7 +1278,7 @@ async function main(argv = process.argv.slice(2)) {
|
|
|
1254
1278
|
process.exit(EXIT_INSTRUMENT_FAILURE);
|
|
1255
1279
|
}
|
|
1256
1280
|
catch (e) {
|
|
1257
|
-
process.stderr.write(`
|
|
1281
|
+
process.stderr.write(`critique: unexpected failure: ${e.stack ?? String(e)}\n`);
|
|
1258
1282
|
process.exit(EXIT_INSTRUMENT_FAILURE); // an unexpected throw means no critique was produced
|
|
1259
1283
|
}
|
|
1260
1284
|
// FINDINGS never gate: any classification — including a task run that ERRORED, which is itself a
|
|
@@ -95,6 +95,13 @@ function listAttachedInputs(runDir) {
|
|
|
95
95
|
const loaded = loadVmPathContext(runDir);
|
|
96
96
|
const uploadsDir = loaded?.ctx.uploadsHostDir ?? join(runDir, "work", "session", "mnt", "uploads");
|
|
97
97
|
const folderNames = loaded ? Array.from(loaded.ctx.folders.keys()).sort() : [];
|
|
98
|
+
// `loadVmPathContext` returns null for BOTH an absent mounts.json (legitimately no recorded mount
|
|
99
|
+
// context — folders default to []) AND a present-but-corrupt one (the folder map is UNKNOWN, not empty).
|
|
100
|
+
// Uploads has a fixed-layout fallback dir it can still probe; connected FOLDERS have none, so a corrupt
|
|
101
|
+
// mounts.json would silently render "(none)" — telling the evaluator "the agent correctly saw no
|
|
102
|
+
// connected folder" when the truth is UNKNOWN. That is the same confabulation-vs-correct false-clean the
|
|
103
|
+
// uploads path guards against below (the ENOENT-vs-read-fault split); mirror it for folders. #14
|
|
104
|
+
const mountsCorrupt = loaded === null && existsSync(join(runDir, "mounts.json"));
|
|
98
105
|
const lines = [];
|
|
99
106
|
try {
|
|
100
107
|
const uploadNames = readdirSync(uploadsDir, { withFileTypes: true })
|
|
@@ -124,6 +131,9 @@ function listAttachedInputs(runDir) {
|
|
|
124
131
|
}
|
|
125
132
|
for (const name of folderNames)
|
|
126
133
|
lines.push(`${name} (connected folder)`);
|
|
134
|
+
if (mountsCorrupt)
|
|
135
|
+
lines.push(`(connected-folder context could not be read: mounts.json present but unparseable — ` +
|
|
136
|
+
`folder attachment presence UNKNOWN, not confirmed absent)`);
|
|
127
137
|
return lines.join("\n");
|
|
128
138
|
}
|
|
129
139
|
/** Byte-bound a text section, appending a loud (never silent) truncation marker so the evaluator knows the
|
package/dist/run/diff.js
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
// Run/cassette diff engine with normalization. Compares two runs, two cassettes, or a run and a
|
|
2
2
|
// cassette (both reduce to an event stream + result metadata through parseMessage/buildTrace, the same
|
|
3
3
|
// typed model `trace` uses — so this stays correct as the SDK schema evolves).
|
|
4
|
+
import { createHash } from "node:crypto";
|
|
4
5
|
import { DEFAULT_SCAN_PATTERNS } from "../scan.js";
|
|
5
6
|
import { diffFileSigsPaths } from "./cassette.js";
|
|
6
7
|
const HOST_PATH_RE = DEFAULT_SCAN_PATTERNS.find((p) => p.cls === "path").re;
|
|
@@ -56,7 +57,13 @@ const CANON_CAP = 2000;
|
|
|
56
57
|
/** Bounded, key-aware canonicalization of a tool's `input` (or any structured value) for tool-sequence
|
|
57
58
|
* comparison — masks volatile string spans AND volatile key names, then caps the length. The 100-char
|
|
58
59
|
* `summarize()` in trace-view.ts is display-lossy by design; this needs its own, larger, comparison-safe
|
|
59
|
-
* cap (still bounded — never an unbounded dump into a diff hunk).
|
|
60
|
+
* cap (still bounded — never an unbounded dump into a diff hunk).
|
|
61
|
+
*
|
|
62
|
+
* #41: when the value exceeds the cap, the truncated prefix ALONE is not a safe equality key — two inputs
|
|
63
|
+
* sharing the first `CANON_CAP` chars but differing only in the dropped tail would compare equal (a
|
|
64
|
+
* false "same" tool call in `diffToolSequence`). So a truncated key folds in a hash of the FULL canonical
|
|
65
|
+
* string: the prefix stays human-readable in a diff hunk, the `#<len>·<sha16>` suffix makes the key
|
|
66
|
+
* depend on the entire content. Both diff sides run this identically, so the comparison stays consistent. */
|
|
60
67
|
export function canonicalizeInput(input, normalize = true) {
|
|
61
68
|
let s;
|
|
62
69
|
try {
|
|
@@ -65,7 +72,10 @@ export function canonicalizeInput(input, normalize = true) {
|
|
|
65
72
|
catch {
|
|
66
73
|
s = String(input);
|
|
67
74
|
}
|
|
68
|
-
|
|
75
|
+
if (s.length <= CANON_CAP)
|
|
76
|
+
return s;
|
|
77
|
+
const fullHash = createHash("sha256").update(s).digest("hex").slice(0, 16);
|
|
78
|
+
return `${s.slice(0, CANON_CAP)}…#${s.length}·${fullHash}`;
|
|
69
79
|
}
|
|
70
80
|
/** LCS over full-row equality (name AND canon must match for "same"), then a post-pass over each gap
|
|
71
81
|
* between LCS matches: a 1-vs-1 gap with the SAME tool name is a "changed" input, not a remove+add pair
|
package/dist/run/execute.js
CHANGED
|
@@ -1664,10 +1664,11 @@ function scrubFileInPlace(path, secrets) {
|
|
|
1664
1664
|
* executeScenario's outermost `finally` (and the chat lane's teardown) so every exit path AFTER the
|
|
1665
1665
|
* agent session exists scrubs — success, the unanswered-gate salvage rethrow, and any rethrown fault
|
|
1666
1666
|
* (agent crash, infra error, hostloop snapshot failure). Deliberately NOT total coverage: a throw
|
|
1667
|
-
* before that try has no raw logs yet
|
|
1668
|
-
* agent.stderr.log
|
|
1669
|
-
*
|
|
1670
|
-
* this
|
|
1667
|
+
* before that try has no raw logs yet, and a SIGKILL of the harness process skips any finally. #60: the
|
|
1668
|
+
* agent.stderr.log sink IS now flushed before this runs — `LiveAgentSession` pipes it with `{ end: false }`
|
|
1669
|
+
* and ends+awaits it in `start()`'s drain, so no buffered stderr byte lands raw after the scrub reads the
|
|
1670
|
+
* file (the session generator resolves only after that flush, and this runs after it resolves). Exported
|
|
1671
|
+
* for tests. */
|
|
1671
1672
|
export function scrubRawRunLogs(outDir, secrets) {
|
|
1672
1673
|
scrubFileInPlace(join(outDir, "events.jsonl"), secrets);
|
|
1673
1674
|
scrubFileInPlace(join(outDir, "control-out.jsonl"), secrets);
|
|
@@ -91,7 +91,7 @@ export const SKILL_FLAG_SURFACE = [
|
|
|
91
91
|
arity: 0,
|
|
92
92
|
critique: {
|
|
93
93
|
kind: "reject",
|
|
94
|
-
reason: "critique produces its own report; inner-turn rendering flags have no effect on it",
|
|
94
|
+
reason: "critique produces its own report (host paths in it are already collapsed to ~); inner-turn rendering flags have no effect on it",
|
|
95
95
|
},
|
|
96
96
|
})),
|
|
97
97
|
{
|
package/dist/runtime/lima.js
CHANGED
|
@@ -51,7 +51,19 @@ export function instanceName(baseline) {
|
|
|
51
51
|
* result of this one helper into both.
|
|
52
52
|
*/
|
|
53
53
|
export function vmGatewayIp() {
|
|
54
|
-
|
|
54
|
+
const raw = process.env.COWORK_VM_GATEWAY ?? "192.168.5.2";
|
|
55
|
+
// #95: this value is interpolated into a root-run iptables command inside the guest
|
|
56
|
+
// (guestFirewallScript → `iptables -A OUTPUT -d ${gatewayIp}` executed via `sh -c`). Validate it as a
|
|
57
|
+
// canonical IPv4 literal and reject everything else, so a malformed or hostile override can never inject
|
|
58
|
+
// shell syntax into privileged provisioning. Defense-in-depth: the var is operator-set, but an
|
|
59
|
+
// unvalidated string reaching a root `sh -c` should be impossible by construction, not by trust. IPv4
|
|
60
|
+
// only — the Apple VZ user-network gateway is IPv4, and the digits-and-dots grammar excludes every shell
|
|
61
|
+
// metacharacter (`;`, `$`, backtick, whitespace, …).
|
|
62
|
+
const octets = raw.split(".");
|
|
63
|
+
const canonicalIPv4 = octets.length === 4 && octets.every((o) => /^\d{1,3}$/.test(o) && Number(o) <= 255 && String(Number(o)) === o);
|
|
64
|
+
if (!canonicalIPv4)
|
|
65
|
+
throw new Error(`COWORK_VM_GATEWAY must be a canonical IPv4 literal (e.g. 192.168.5.2); got ${JSON.stringify(raw)}`);
|
|
66
|
+
return raw;
|
|
55
67
|
}
|
|
56
68
|
export function vmStatus(instance) {
|
|
57
69
|
const r = spawnSync(limaPath(), ["list", instance, "--format", "{{.Status}}"], { encoding: "utf8" });
|
package/docs/critique.md
CHANGED
|
@@ -124,7 +124,7 @@ ignored.
|
|
|
124
124
|
| `--session-id` / `--resume` | critique mints and manages its own session — the reflection turn *is* a resume of it |
|
|
125
125
|
| `--repeat` + companions | fixed two-turn protocol; loop `critique` itself and pair by `fingerprint.skillHash` |
|
|
126
126
|
| `--ablate-skill` | grading a skill you removed is incoherent |
|
|
127
|
-
| `--quiet`/`-q` / `--verbose` / `--compact` / `--demo` / `--dry-run` | inner-turn rendering or preview — no effect on the report |
|
|
127
|
+
| `--quiet`/`-q` / `--verbose` / `--compact` / `--demo` / `--dry-run` | inner-turn rendering or preview — no effect on the report (which already collapses host paths to `~`) |
|
|
128
128
|
|
|
129
129
|
**Repeating a flag.** `--upload`, `--folder`, `--plugin`, `--marketplace`, `--enable` and `--answer` accumulate,
|
|
130
130
|
so repeating them is how you pass several. Every other value-taking flag is single-valued and repeating it is
|
|
@@ -144,8 +144,11 @@ adjudicable". So:
|
|
|
144
144
|
auto-selects with a notice.
|
|
145
145
|
- **Selection only:** the positional folder is still what both turns mount (session identity is
|
|
146
146
|
unchanged), and **`fingerprint.skillHash` is unchanged by `--skill`** — it keys the *mounted folder*,
|
|
147
|
-
so it pairs generations per-plugin, not per-skill.
|
|
148
|
-
|
|
147
|
+
so it pairs generations per-plugin, not per-skill. **Workflow implication: pairing critiques of a
|
|
148
|
+
multi-skill plugin by skillHash alone CROSS-PAIRS different skills** — pair by
|
|
149
|
+
**(`gradedSkillHash`, `gradedSkill`)**; the report's `gradedSkill` field carries the resolved
|
|
150
|
+
`skills/<name>` (`--skill` or the auto-selection). `--label` remains available for coarser
|
|
151
|
+
generation tags.
|
|
149
152
|
- The report carries an advisory **`skillInvocationObserved`**: `false` means the graded run's own
|
|
150
153
|
`skillActivity` never mentions the selected skill — the critique may be grading a run that did not
|
|
151
154
|
actually invoke it.
|
|
@@ -170,6 +173,12 @@ It does **not** record their contents — see Known limitations.
|
|
|
170
173
|
evaluator passes.
|
|
171
174
|
- The evaluator defaults to the most expensive tier. Override with `--evaluator-model <id>` or
|
|
172
175
|
**`COWORK_HARNESS_EVALUATOR_MODEL`**.
|
|
176
|
+
- **The evaluator passes dominate spend** — on a measured end-to-end (trivial probe, default evaluator)
|
|
177
|
+
the two evaluator passes were ~3/4 of the total. For a wide batch: calibrate with run 1's `costUsd`
|
|
178
|
+
(gate on `costUsd.complete` — `false` means the total undercounts), then consider a cheaper
|
|
179
|
+
`--evaluator-model` for the sweep. Caveat: the armor's injection-resistance is verified for the
|
|
180
|
+
shipped **default** evaluator model only — changing it voids that specific verification (matters when
|
|
181
|
+
critiquing skills you did not write).
|
|
173
182
|
- **container** needs Docker/Lima; **hostloop** needs Docker (the bash/web_fetch sidecar) **plus** the
|
|
174
183
|
staged native agent binary, and writes to the real host filesystem — a writable `--folder` there requires
|
|
175
184
|
`--allow-host-writes`. Both tiers need an authenticated `claude` CLI on PATH.
|
|
@@ -227,6 +236,14 @@ malformed evaluator items (the surviving findings are then not necessarily the c
|
|
|
227
236
|
readable-but-oversized SKILL.md is flagged **`skillMdTruncated`** ("the evaluator graded a cut copy") —
|
|
228
237
|
distinct from missing/unreadable, which alone force the mechanical `"already-covered"` downgrade.
|
|
229
238
|
|
|
239
|
+
**Truncation has a second, sharper consequence than the not-adjudicable steer: DROPPED findings.**
|
|
240
|
+
Citation validation checks each finding's `evidence` excerpt verbatim against the *packaged* (cut)
|
|
241
|
+
copy — so a finding that quotes text past the cut cannot resolve and lands in **DROPPED**, even when
|
|
242
|
+
the quote is a perfectly accurate excerpt of the real file. A skill well over the cap should expect a
|
|
243
|
+
not-adjudicable/DROPPED skew concentrated on its back half; if you see back-half findings in DROPPED
|
|
244
|
+
on a `skillMdTruncated` run, this is why — front-load the operative guidance, split the skill, or
|
|
245
|
+
treat those items as leads to re-check by hand.
|
|
246
|
+
|
|
230
247
|
### Run-dir artifacts
|
|
231
248
|
|
|
232
249
|
Beyond stdout, every critique leaves durable artifacts at the run-dir root (best-effort writes —
|
|
@@ -240,6 +257,10 @@ Beyond stdout, every critique leaves durable artifacts at the run-dir root (best
|
|
|
240
257
|
|
|
241
258
|
These artifacts (and the report's JSON shape) are part of critique's **EXPERIMENTAL** surface — useful
|
|
242
259
|
and stable in practice, but not yet a frozen SPEC §12 covered surface; field additions are expected.
|
|
260
|
+
The report's field names and shapes are authoritatively described by
|
|
261
|
+
[`schema/critique-report.json`](../schema/critique-report.json) (descriptive + test-pinned against the
|
|
262
|
+
actual builder, unlike the §12-frozen `doctor.json`), so automation consumers — budget pacers gating on
|
|
263
|
+
`costUsd.complete`, harvesters pairing on `gradedSkill` — parse against a schema, not prose.
|
|
243
264
|
|
|
244
265
|
## Reproduction — the ≥2-run discipline
|
|
245
266
|
|
|
@@ -258,8 +279,15 @@ Then pair/cluster across the reports:
|
|
|
258
279
|
- **Same finding across runs/inputs?** cluster by each item's **`findingFingerprint`** (sha over the
|
|
259
280
|
normalized idea + classification + recommendedAction, deliberately excluding the input-specific
|
|
260
281
|
`evidence` excerpt — so the same finding matches across different decks/transcripts).
|
|
282
|
+
- **The fingerprint is high-precision, LOW-RECALL — read the direction correctly.** `idea` is
|
|
283
|
+
model-authored free text, so the same underlying finding *reworded* across runs fingerprints
|
|
284
|
+
differently. A **match proves** reproduction; a **mismatch does NOT prove** non-reproduction — before
|
|
285
|
+
concluding "didn't reproduce", skim the unmatched items for rewordings of the same substance.
|
|
261
286
|
- A finding that recurs across ≥2 runs with the same `findingFingerprint` meets the reproduction bar;
|
|
262
|
-
a one-off is a lead,
|
|
287
|
+
a fingerprint one-off is a lead — possibly a real one-off, possibly a reworded repeat.
|
|
288
|
+
- **Multi-skill plugins: never pair by `gradedSkillHash` alone.** The hash keys the whole mounted
|
|
289
|
+
plugin, so it cross-pairs critiques of *different* skills in the same plugin — pair by
|
|
290
|
+
**(`gradedSkillHash`, `gradedSkill`)**; `gradedSkill` is the report's resolved `skills/<name>`.
|
|
263
291
|
- To make the graded runs deterministic across repeats, copy the report's echoed `--answer` lines
|
|
264
292
|
(the graded run's resolved gate answers) into the next invocation.
|
|
265
293
|
|
package/docs/scenario.md
CHANGED
|
@@ -952,7 +952,8 @@ with `COWORK_LIMA_INSTANCE`.
|
|
|
952
952
|
- **A run errors with "not mounted — VM not provisioned for this harness config"** — the VM predates a
|
|
953
953
|
config change (its mounts don't match). Recreate it: `cowork-harness vm delete && cowork-harness vm init`.
|
|
954
954
|
- **Egress allowed/denied looks wrong** — the guest firewall and the proxy URL must point at the same
|
|
955
|
-
gateway. The default Apple-VZ user-network gateway is `192.168.5.2`; override with `COWORK_VM_GATEWAY
|
|
955
|
+
gateway. The default Apple-VZ user-network gateway is `192.168.5.2`; override with `COWORK_VM_GATEWAY`
|
|
956
|
+
(a canonical IPv4 literal — an invalid value is rejected, as it feeds the guest iptables rule),
|
|
956
957
|
and the proxy port with `COWORK_VM_PROXY_PORT` (unset, the host binds an OS-assigned free port;
|
|
957
958
|
`8899` is only the guest-config fallback when a VM is spawned without an explicit port — not the
|
|
958
959
|
effective default of a normal run). The harness threads one resolved
|
|
@@ -16,7 +16,7 @@ DOES exercise a real gate exchange, see `example-multiselect-gate.cassette.json`
|
|
|
16
16
|
|
|
17
17
|
Run it with:
|
|
18
18
|
|
|
19
|
-
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.
|
|
19
|
+
> Assumes the `cowork-harness` CLI is available — from a source checkout run `npm ci && npm run build && npm link` first, or `npm i -g "cowork-harness@>=1.9.0"`. (`replay` itself needs nothing else — no token, no Docker.)
|
|
20
20
|
|
|
21
21
|
```sh
|
|
22
22
|
cowork-harness replay examples/replays/example-pdf-skill.cassette.json
|
package/llms.txt
CHANGED
|
@@ -17,6 +17,7 @@ It is a fidelity fixture, not the Desktop runtime. The CLI binary is `cowork-har
|
|
|
17
17
|
- [examples/README.md](examples/README.md): worked, copyable example scenarios + sessions + skills (from a source checkout — the npm package ships only examples/replays/; protocol/container tiers, token-free replay, prerequisites)
|
|
18
18
|
- [schema/scenario.schema.json](schema/scenario.schema.json): JSON Schema for scenario files (machine-readable)
|
|
19
19
|
- [schema/session.schema.json](schema/session.schema.json): JSON Schema for session files (machine-readable)
|
|
20
|
+
- [schema/critique-report.json](schema/critique-report.json): descriptive schema for `critique`'s JSON report / `critique-report.json` artifact (EXPERIMENTAL — not §12-frozen; field additions expected)
|
|
20
21
|
- [.claude/skills/cowork-harness/references/ci-recipe.md](.claude/skills/cowork-harness/references/ci-recipe.md): copy-paste GitHub Actions — token-free replay PR gate + nightly live lane
|
|
21
22
|
|
|
22
23
|
## Concepts & internals
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "cowork-harness",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.9.0",
|
|
4
4
|
"description": "Scriptable, CI-friendly harness for Claude Cowork's runtime contract for testing skills across scenarios \u2014 same agent, mounts, egress allowlist, permission protocol, and sandbox limitations.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -0,0 +1,221 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$schema": "http://json-schema.org/draft-07/schema#",
|
|
3
|
+
"$id": "critique-report.json",
|
|
4
|
+
"title": "cowork-harness critique --output-format json report (and critique-report.json run-dir artifact)",
|
|
5
|
+
"description": "EXPERIMENTAL, descriptive schema — NOT a SPEC §12-frozen surface: additive field changes are expected between minor versions (unlike doctor.json). It exists so automation consumers (budget pacers, harvesters) have authoritative field names/shapes instead of prose. Exactly one of infraFailure / evaluatorError / evaluatorModel-with-findings describes the outcome; items is [] on both failure branches.",
|
|
6
|
+
"type": "object",
|
|
7
|
+
"required": ["skillFolder", "prompt", "sessionId", "outDir", "fidelity", "selfReportStatus", "verdictProvenance", "items"],
|
|
8
|
+
"additionalProperties": false,
|
|
9
|
+
"properties": {
|
|
10
|
+
"skillFolder": {
|
|
11
|
+
"type": "string",
|
|
12
|
+
"description": "the positional folder BOTH turns mounted (a plain skill, or a plugin root)"
|
|
13
|
+
},
|
|
14
|
+
"prompt": {
|
|
15
|
+
"type": "string"
|
|
16
|
+
},
|
|
17
|
+
"sessionId": {
|
|
18
|
+
"type": "string"
|
|
19
|
+
},
|
|
20
|
+
"outDir": {
|
|
21
|
+
"type": "string",
|
|
22
|
+
"description": "the kept run dir (turns/1 = graded, turns/2 = reflection; critique-* artifacts at its root)"
|
|
23
|
+
},
|
|
24
|
+
"fidelity": {
|
|
25
|
+
"type": "string",
|
|
26
|
+
"description": "the tier critique pinned for both turns (cowork is refused, so requested == resolved)",
|
|
27
|
+
"enum": ["container", "hostloop"]
|
|
28
|
+
},
|
|
29
|
+
"gradedEffectiveFidelity": {
|
|
30
|
+
"type": "string",
|
|
31
|
+
"description": "best-effort: the tier the graded turn's own result.json RECORDS (should equal fidelity; both surfaced so a mismatch is visible)"
|
|
32
|
+
},
|
|
33
|
+
"gradedBaseline": {
|
|
34
|
+
"type": "string",
|
|
35
|
+
"description": "best-effort: the graded turn's fingerprint.baseline (Desktop appVersion)"
|
|
36
|
+
},
|
|
37
|
+
"costUsd": {
|
|
38
|
+
"type": "object",
|
|
39
|
+
"description": "per-critique cost across the FOUR model workloads. GATE ON `complete` BEFORE TRUSTING totalUsd: false means one or more workloads were unpriced and the total UNDERCOUNTS true spend.",
|
|
40
|
+
"required": ["totalUsd", "complete"],
|
|
41
|
+
"additionalProperties": false,
|
|
42
|
+
"properties": {
|
|
43
|
+
"taskTurnUsd": {
|
|
44
|
+
"type": "number"
|
|
45
|
+
},
|
|
46
|
+
"reflectionTurnUsd": {
|
|
47
|
+
"type": "number"
|
|
48
|
+
},
|
|
49
|
+
"evaluatorPass1Usd": {
|
|
50
|
+
"type": "number"
|
|
51
|
+
},
|
|
52
|
+
"evaluatorPass2Usd": {
|
|
53
|
+
"type": "number"
|
|
54
|
+
},
|
|
55
|
+
"totalUsd": {
|
|
56
|
+
"type": "number"
|
|
57
|
+
},
|
|
58
|
+
"complete": {
|
|
59
|
+
"type": "boolean"
|
|
60
|
+
}
|
|
61
|
+
}
|
|
62
|
+
},
|
|
63
|
+
"gradedSkill": {
|
|
64
|
+
"type": "string",
|
|
65
|
+
"description": "the resolved skills/<name> the packager graded (--skill or single-skill auto-selection); ABSENT for a plain skill folder. Multi-skill-plugin pairing key: skillHash keys the whole mounted plugin, so pair by (gradedSkillHash, gradedSkill), never skillHash alone"
|
|
66
|
+
},
|
|
67
|
+
"skillInvocationObserved": {
|
|
68
|
+
"type": "boolean",
|
|
69
|
+
"description": "advisory: whether the graded run's own skillActivity mentions gradedSkill; false = the critique may be grading a run that never invoked the selected skill"
|
|
70
|
+
},
|
|
71
|
+
"gateAnswers": {
|
|
72
|
+
"type": "array",
|
|
73
|
+
"description": "the graded run's resolved gate answers, for a deterministic follow-up run",
|
|
74
|
+
"items": {
|
|
75
|
+
"type": "object",
|
|
76
|
+
"required": ["question", "answer", "answeredBy"],
|
|
77
|
+
"additionalProperties": false,
|
|
78
|
+
"properties": {
|
|
79
|
+
"question": {
|
|
80
|
+
"type": "string"
|
|
81
|
+
},
|
|
82
|
+
"answer": {
|
|
83
|
+
"type": "string"
|
|
84
|
+
},
|
|
85
|
+
"answeredBy": {
|
|
86
|
+
"type": "string"
|
|
87
|
+
}
|
|
88
|
+
}
|
|
89
|
+
}
|
|
90
|
+
},
|
|
91
|
+
"taskResult": {
|
|
92
|
+
"type": "string",
|
|
93
|
+
"description": "the graded turn's own result — error is a GRADEABLE outcome, not an instrument failure",
|
|
94
|
+
"enum": ["success", "error"]
|
|
95
|
+
},
|
|
96
|
+
"gradedOutcome": {
|
|
97
|
+
"type": "string",
|
|
98
|
+
"description": "the graded turn's result.json outcome (e.g. delivered_clean)"
|
|
99
|
+
},
|
|
100
|
+
"gradedSkillHash": {
|
|
101
|
+
"type": "string",
|
|
102
|
+
"description": "content-exact hash of the MOUNTED folder — the cross-FIX pairing key (per-plugin under --skill)"
|
|
103
|
+
},
|
|
104
|
+
"selfReportStatus": {
|
|
105
|
+
"type": "string",
|
|
106
|
+
"description": "unavailable = pass 2 was skipped; findings are pass-1-only",
|
|
107
|
+
"enum": ["captured", "unavailable"]
|
|
108
|
+
},
|
|
109
|
+
"evaluatorIntegrity": {
|
|
110
|
+
"type": "object",
|
|
111
|
+
"description": "mechanical canary check per pass; false = that pass ignored a trusted instruction — an empty critique may be adversarial silencing, not a clean skill",
|
|
112
|
+
"required": ["pass1Canary"],
|
|
113
|
+
"additionalProperties": false,
|
|
114
|
+
"properties": {
|
|
115
|
+
"pass1Canary": {
|
|
116
|
+
"type": "boolean"
|
|
117
|
+
},
|
|
118
|
+
"pass2Canary": {
|
|
119
|
+
"type": "boolean"
|
|
120
|
+
}
|
|
121
|
+
}
|
|
122
|
+
},
|
|
123
|
+
"droppedEvaluatorItems": {
|
|
124
|
+
"type": "object",
|
|
125
|
+
"description": "malformed items the per-item-tolerant parse dropped; non-zero means the surviving findings under-represent the full reply",
|
|
126
|
+
"required": ["pass1"],
|
|
127
|
+
"additionalProperties": false,
|
|
128
|
+
"properties": {
|
|
129
|
+
"pass1": {
|
|
130
|
+
"type": "number"
|
|
131
|
+
},
|
|
132
|
+
"pass2": {
|
|
133
|
+
"type": "number"
|
|
134
|
+
}
|
|
135
|
+
}
|
|
136
|
+
},
|
|
137
|
+
"turn1ResultDegraded": {
|
|
138
|
+
"type": "boolean",
|
|
139
|
+
"description": "the graded turn's canonical result was corrupted/never archived — result-derived sections are empty DEFAULTS, treat as unknown"
|
|
140
|
+
},
|
|
141
|
+
"turn1SliceDegraded": {
|
|
142
|
+
"type": "boolean",
|
|
143
|
+
"description": "the turn-1 transcript slice's boundary could not be trusted — treat gaps as unknown"
|
|
144
|
+
},
|
|
145
|
+
"skillMdStatus": {
|
|
146
|
+
"type": "string",
|
|
147
|
+
"description": "readability of the packaged SKILL.md; missing/unreadable force the mechanical already-covered -> not-adjudicable downgrade",
|
|
148
|
+
"enum": ["readable", "missing", "unreadable"]
|
|
149
|
+
},
|
|
150
|
+
"skillMdTruncated": {
|
|
151
|
+
"type": "boolean",
|
|
152
|
+
"description": "SKILL.md was READABLE but over its cap — the evaluator graded a CUT copy. Consequences: claims about past-cut content are steered to not-adjudicable, and a finding QUOTING past-cut text fails citation-resolution and lands in DROPPED"
|
|
153
|
+
},
|
|
154
|
+
"verdictProvenance": {
|
|
155
|
+
"type": "object",
|
|
156
|
+
"description": "advisory scoping: the verdict is a self-run discovery lead, never an independent attestation",
|
|
157
|
+
"required": ["kind", "advisory", "caveat"],
|
|
158
|
+
"additionalProperties": false,
|
|
159
|
+
"properties": {
|
|
160
|
+
"kind": {
|
|
161
|
+
"type": "string",
|
|
162
|
+
"enum": ["self-run"]
|
|
163
|
+
},
|
|
164
|
+
"advisory": {
|
|
165
|
+
"type": "boolean"
|
|
166
|
+
},
|
|
167
|
+
"caveat": {
|
|
168
|
+
"type": "string"
|
|
169
|
+
}
|
|
170
|
+
}
|
|
171
|
+
},
|
|
172
|
+
"infraFailure": {
|
|
173
|
+
"type": "string",
|
|
174
|
+
"description": "instrument failure (task turn killed / reflection protocol broke) — no critique was produced; items is []; process exits 2"
|
|
175
|
+
},
|
|
176
|
+
"evaluatorError": {
|
|
177
|
+
"type": "string",
|
|
178
|
+
"description": "the evaluator threw (embeds the raw reply) — no critique was produced; items is []; process exits 2; see critique-salvage.json"
|
|
179
|
+
},
|
|
180
|
+
"evaluatorModel": {
|
|
181
|
+
"type": "string",
|
|
182
|
+
"description": "the transport-RESOLVED evaluator model (never the requested alias); present only when the evaluator completed"
|
|
183
|
+
},
|
|
184
|
+
"items": {
|
|
185
|
+
"type": "array",
|
|
186
|
+
"description": "the citation-validated findings from both passes. citationResolved:false = the cited excerpt did not resolve verbatim against the evidence package (DROPPED section — transparency only, do not act on as-is)",
|
|
187
|
+
"items": {
|
|
188
|
+
"type": "object",
|
|
189
|
+
"required": ["source", "idea", "classification", "evidence", "recommendedAction"],
|
|
190
|
+
"additionalProperties": false,
|
|
191
|
+
"properties": {
|
|
192
|
+
"source": {
|
|
193
|
+
"type": "string",
|
|
194
|
+
"enum": ["evaluator", "self-report"]
|
|
195
|
+
},
|
|
196
|
+
"idea": {
|
|
197
|
+
"type": "string"
|
|
198
|
+
},
|
|
199
|
+
"classification": {
|
|
200
|
+
"type": "string",
|
|
201
|
+
"enum": ["grounded-and-actionable", "grounded-but-not-worth-it", "confabulated", "already-covered", "not-adjudicable"]
|
|
202
|
+
},
|
|
203
|
+
"evidence": {
|
|
204
|
+
"type": "string",
|
|
205
|
+
"description": "the model's cited excerpt — verbatim-checked against the evidence package"
|
|
206
|
+
},
|
|
207
|
+
"recommendedAction": {
|
|
208
|
+
"type": "string"
|
|
209
|
+
},
|
|
210
|
+
"citationResolved": {
|
|
211
|
+
"type": "boolean"
|
|
212
|
+
},
|
|
213
|
+
"findingFingerprint": {
|
|
214
|
+
"type": "string",
|
|
215
|
+
"description": "sha256/16 over normalized idea+classification+recommendedAction (evidence excluded). Cross-INPUT clustering key — HIGH-PRECISION, LOW-RECALL: idea is model-authored free text, so a MATCH proves the same finding recurred; a MISMATCH does NOT prove non-reproduction (the same finding reworded fingerprints differently)"
|
|
216
|
+
}
|
|
217
|
+
}
|
|
218
|
+
}
|
|
219
|
+
}
|
|
220
|
+
}
|
|
221
|
+
}
|