cowork-harness 1.4.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +28 -10
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +16 -3
- package/CHANGELOG.md +73 -0
- package/README.md +7 -6
- package/SPEC.md +1 -1
- package/dist/cli.js +122 -75
- package/dist/critique/armor.js +44 -0
- package/dist/critique/command.js +676 -0
- package/dist/critique/evaluator.js +388 -0
- package/dist/critique/evidence.js +175 -0
- package/dist/critique/package-evidence.js +220 -0
- package/dist/run/cassette.js +30 -14
- package/dist/run/chat-result.js +1 -0
- package/dist/run/diff.js +11 -1
- package/dist/run/envelope.js +5 -1
- package/dist/run/execute.js +9 -4
- package/dist/run/outcome.js +15 -0
- package/dist/run/repeat-flags.js +70 -0
- package/docs/README.md +1 -0
- package/docs/critique.md +117 -0
- package/docs/debugging.md +62 -4
- package/docs/gotchas.md +3 -2
- package/docs/scenario.md +3 -2
- package/docs/stats.md +43 -1
- package/examples/replays/README.md +1 -1
- package/llms.txt +2 -1
- package/package.json +3 -4
- package/schema/run-result.json +160 -54
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 1.
|
|
7
|
-
tracks-harness: cowork-harness 1.
|
|
6
|
+
version: 1.5.0
|
|
7
|
+
tracks-harness: cowork-harness 1.5.0 (baseline desktop-1.22209.3)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,7 +22,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 1.5.0` (baseline
|
|
26
26
|
> `desktop-1.22209.3`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
@@ -39,10 +39,20 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
41
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
42
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.
|
|
43
|
-
|
|
44
|
-
What the ≥ 1.
|
|
45
|
-
|
|
42
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 1.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=1.5.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=1.5.0"`. **Pin `@>=1.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
43
|
+
|
|
44
|
+
What the ≥ 1.5.0 floor gates, by release:
|
|
45
|
+
|
|
46
|
+
- **1.5.0:** `critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL) — runs a skill, asks the agent
|
|
47
|
+
what confused it, then grades that self-report against a frozen record of the run: a blinded evaluator
|
|
48
|
+
plus mechanical citation checking. Findings NEVER gate (exit 0); exit 2 means no critique was produced.
|
|
49
|
+
`skill --repeat N` (2-100) with `--min-pass-rate`/`--stop-on-diverge`/`--max-budget-usd`/
|
|
50
|
+
`--allow-budget-stop` brings the variance rollup to the exploratory lane. `result.json` gains
|
|
51
|
+
`outcome` (`errored`/`no_deliverable`/`delivered_with_verdict_fail`/`delivered_clean`). The
|
|
52
|
+
`skill`/`probe-dispatch` lanes now emit `fingerprint.skillHash`/`skillCommit` at all — before 1.5.0
|
|
53
|
+
they emitted NEITHER, so a generation-pairing step over those lanes silently grouped on an absent key.
|
|
54
|
+
`diff`'s exit code now honours its own documented gateable signal (tools/artifacts/meta; a
|
|
55
|
+
transcript-only difference no longer fails it).
|
|
46
56
|
- **core set (pre-0.21.0 vintage, or mixed):** `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes` incl. the `--allow-domain`/`--allow-email`/`--allow-patterns-file` allows (`--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, and `/help` in the REPL.
|
|
47
57
|
- **0.21.0:** `verify-cassettes --allow-path` (`path` — local absolute filesystem paths — is the scanner's 4th class), and `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field.
|
|
48
58
|
- **0.22.0:** `computer_links_resolve`.
|
|
@@ -86,9 +96,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
86
96
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
87
97
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
88
98
|
|
|
99
|
+
> **"repo-only" in this skill means "not bundled with the installed SKILL"** — not "unavailable". An
|
|
100
|
+
> **npm** install ships `docs/`, `README.md` and `SPEC.md` in the tarball, so try
|
|
101
|
+
> `node_modules/cowork-harness/docs/<name>.md` before assuming a pointer dangles. A **plugin**
|
|
102
|
+
> install loads a trimmed source-only cache where those pointers genuinely do dangle.
|
|
103
|
+
|
|
89
104
|
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
90
105
|
lint-skill · analyze-skill · probe-dispatch ·
|
|
91
|
-
verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
106
|
+
verify-run · trace · inspect · diff · critique · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
92
107
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
93
108
|
|
|
94
109
|
**Two different `scaffold` tools — don't confuse them.** The native `cowork-harness scaffold <run-id>`
|
|
@@ -366,7 +381,10 @@ cassette — has its own recipe:
|
|
|
366
381
|
input + result are captured, not just errored ones. When iterating, tag generations with `--label` and
|
|
367
382
|
pair a critique only with a `result.json` whose `fingerprint.skillHash` **matches** the skill that
|
|
368
383
|
produced it (`inspect`/the run-index row surface a short `skillHash` prefix; `verify-run` warns when a
|
|
369
|
-
kept run predates the current skill).
|
|
384
|
+
kept run predates the current skill). **On a pre-1.5.0 CLI the `skill` lane emits no `fingerprint.skillHash` at all**, so a
|
|
385
|
+
pairing step there silently groups on an absent key instead of erroring — check the field is present, or
|
|
386
|
+
require ≥ 1.5.0. See `docs/debugging.md`
|
|
387
|
+
(repo-only) for the full loop.
|
|
370
388
|
|
|
371
389
|
#### Interpreting verdict signals
|
|
372
390
|
|
|
@@ -462,7 +480,7 @@ Docker, no re-record.
|
|
|
462
480
|
| Situation | Symptom | Reach for (in order) |
|
|
463
481
|
|---|---|---|
|
|
464
482
|
| **The skill misbehaved** | wrong output, an unexpected gate, a denied tool, an opaque crash | `inspect` — what did it produce? · `trace <run-dir> --view <view>` — what did it actually do (tools, gates, sub-agent tree)? · `verify-run` — re-assert cheaply when only an assertion is wrong · `diff <old-run> <new-run>` — what changed since it worked · `chat` — reproduce it by hand |
|
|
465
|
-
| **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
|
|
483
|
+
| **A green you don't trust** | an assert that may have tested nothing, a stale cassette, an auto-answered or decided gate | `replay --explain` — the evidence trail behind each *passing* assert · `lint` — assertions on the wrong CI lane / mixed-class keys · `verify-cassettes` — privacy + staleness over committed cassettes · the Gotchas landmine catalog — how a check passes vacuously · `run --repeat N` / `skill --repeat N` — did it pass, or pass once? · `stats` — flaky or expensive over time |
|
|
466
484
|
|
|
467
485
|
A failed run also records `errorSource` (where the failure originated) and `stderrLogPath` (the captured
|
|
468
486
|
agent stderr) — read those before re-running; a re-record rarely tells you more than the captured stderr
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 1.5.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -13,7 +13,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
13
13
|
```
|
|
14
14
|
|
|
15
15
|
The Action's `version` input defaults to `latest` — intentional so a copy-pasted recipe tracks the current
|
|
16
|
-
release; pin an exact version (e.g. `version: "1.
|
|
16
|
+
release; pin an exact version (e.g. `version: "1.5.0"`) for reproducible CI.
|
|
17
17
|
|
|
18
18
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
19
19
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -57,7 +57,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
57
57
|
GitHub-hosted runners, no token/Docker/agent:
|
|
58
58
|
|
|
59
59
|
```yaml
|
|
60
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
60
|
+
- run: npm i -g "cowork-harness@>=1.5.0"
|
|
61
61
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
62
62
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
63
63
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -197,7 +197,7 @@ jobs:
|
|
|
197
197
|
with: { node-version: '20' }
|
|
198
198
|
- uses: actions/setup-python@v5
|
|
199
199
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
200
|
-
- run: npm i -g "cowork-harness@>=1.
|
|
200
|
+
- run: npm i -g "cowork-harness@>=1.5.0"
|
|
201
201
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
202
202
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
203
203
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -226,7 +226,7 @@ jobs:
|
|
|
226
226
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
227
227
|
fi
|
|
228
228
|
- if: steps.guard.outputs.live == 'true'
|
|
229
|
-
run: npm i -g "cowork-harness@>=1.
|
|
229
|
+
run: npm i -g "cowork-harness@>=1.5.0"
|
|
230
230
|
- if: steps.guard.outputs.live == 'true'
|
|
231
231
|
run: cowork-harness run scenarios/ --output-format json
|
|
232
232
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 1.5.0`
|
|
4
4
|
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
|
|
5
|
-
front-matter (currently
|
|
5
|
+
front-matter (currently 1.5.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -170,7 +170,17 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
170
170
|
|
|
171
171
|
1. **Verify before you trust.** A green run is not a correct run, and a skill's self-reported finding (a
|
|
172
172
|
self-critique appendix, "I extracted X") is not real until its cited evidence is found in the run's own
|
|
173
|
-
output.
|
|
173
|
+
output. **Reproduce before acting on a finding:** `cowork-harness skill <folder> "<prompt>" --repeat 5 --label gen-1`
|
|
174
|
+
runs the same skill+prompt N times (2-100) and prints a variance rollup instead of a single pass/fail —
|
|
175
|
+
`--repeat` works on the `skill` lane, not just `run`. A single green run proves it passed *once*.
|
|
176
|
+
Companions: `--min-pass-rate`, `--stop-on-diverge`, `--max-budget-usd`, `--allow-budget-stop`. It rejects `--session-id`/
|
|
177
|
+
`--resume` (both pin one run dir) and `--decider-cmd`/`--decider-dir` (a driving agent x N is not a
|
|
178
|
+
measurement).
|
|
179
|
+
The full loop (harvest -> reproduce -> fix -> prove freshness -> compare) is written out end-to-end in
|
|
180
|
+
docs/debugging.md under "The whole loop, end to end". The harness now SHIPS a grader — `cowork-harness critique <skill-folder> --prompt "<probe>"` runs the
|
|
181
|
+
skill, asks the agent what confused it, and grades that self-report against a frozen record of the run
|
|
182
|
+
(blinded evaluator + mechanical citation checking). See docs/critique.md for cost and limits. If you
|
|
183
|
+
prefer to build your own grader, the substrate is still here:
|
|
174
184
|
- `result.json` → `finalMessage` (the skill's own answer/critique) + `toolResults[]` (tool outputs).
|
|
175
185
|
- `cowork-harness trace <run-dir> --output-format json` → the tool-call stream. Add `--full-results` so
|
|
176
186
|
a **successful** call's full input + result are captured (the default view slices them to ~100/120
|
|
@@ -181,7 +191,10 @@ Hardening a skill is a loop: run → read what it did → fix → run again. Two
|
|
|
181
191
|
verdict.
|
|
182
192
|
2. **Don't cross-pair generations.** When you run the same skill across fixes, never pair a *pre-fix*
|
|
183
193
|
`result.json` with a *post-fix* critique. The authoritative version key is `fingerprint.skillHash` —
|
|
184
|
-
content-exact, on every live run
|
|
194
|
+
content-exact, on every live `run`/`skill` run that mounts a skill or plugin — **but only on ≥ 1.5.0; earlier CLIs emit
|
|
195
|
+
no `skillHash` on the `skill` lane at all, so verify the field is present before pairing on it** (a run that mounts nothing
|
|
196
|
+
records none; the `chat` lane records no fingerprint), changes on any tracked edit. **Group/pair on it**
|
|
197
|
+
(`inspect` and the
|
|
185
198
|
run-index row surface a short prefix). Add `--label <tag>` for a human-readable generation name
|
|
186
199
|
(skillHash is the correctness key; the label is ergonomics). `cowork-harness verify-run <run-dir>
|
|
187
200
|
<scenario.yaml>` is the native staleness guard: it **warns** when a kept run predates the current
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,79 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [1.5.0] — 2026-07-20
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`cowork-harness critique <skill-folder> --prompt "<probe>"` (EXPERIMENTAL).** Runs a skill, asks the
|
|
14
|
+
agent what confused it, then grades that self-report against a frozen record of what actually happened —
|
|
15
|
+
a byte-boundary evidence snapshot taken *before* the reflection turn, a first evaluator pass that is
|
|
16
|
+
structurally blind to the self-report (the text is never put in its prompt, not merely ignored), and
|
|
17
|
+
mechanical citation checking that drops any claim not quoting the evidence verbatim. **A discovery
|
|
18
|
+
instrument, never a gate:** findings of any classification exit 0 (including a task run that itself
|
|
19
|
+
errored — that is a finding about the skill); exit 2 is reserved for a usage error or an instrument
|
|
20
|
+
failure, where no critique was produced. Previously a maintainer-only script that could not run from an installed package at all.
|
|
21
|
+
Costs four model workloads per critique — see [docs/critique.md](./docs/critique.md).
|
|
22
|
+
- **Evidence-package armoring.** The self-report was already fenced, but the evidence package — which
|
|
23
|
+
carries a third-party SKILL.md verbatim into both evaluator prompts — was not, so hostile skill content
|
|
24
|
+
could steer the grader directly. Untrusted content now sits inside per-run nonce markers and only
|
|
25
|
+
nonce-tagged headings outside them count as instructions; a skill cannot pre-author the nonce. Verified
|
|
26
|
+
by a red-team probe across three models: the structural-forgery payload that steered all three now
|
|
27
|
+
matches control. Content that merely *argues* is a documented residual — fencing separates planes, it
|
|
28
|
+
cannot stop persuasion.
|
|
29
|
+
- **`skill --repeat <N>`** (2–100), with `--min-pass-rate` / `--stop-on-diverge` / `--max-budget-usd` /
|
|
30
|
+
`--allow-budget-stop`. The variance rollup already existed but was `run`-only, because the flags were
|
|
31
|
+
parsed inline in the `run` command — the exploratory lane, where an iterate-across-fixes loop actually
|
|
32
|
+
lives, rejected them as unknown. Both lanes now share one parse (`run/repeat-flags.ts`) and one batch
|
|
33
|
+
engine, so rollup shape, JSON envelope, and batch verdict match. `skill --repeat` additionally rejects
|
|
34
|
+
`--session-id`/`--resume` (both pin a single run dir, so iterations would overwrite each other rather
|
|
35
|
+
than produce N independent samples) and `--decider-dir`/`--decider-cmd` (the reproducibility invariant
|
|
36
|
+
`run` already enforced).
|
|
37
|
+
- **`result.json` `outcome`** — a one-field rollup of the `result` × `verdict.pass` × exit-code matrix:
|
|
38
|
+
`errored` / `no_deliverable` / `delivered_with_verdict_fail` / `delivered_clean`. Those three signals
|
|
39
|
+
legitimately disagree (a fail-severity signal flips the verdict while `result` stays `"success"`), and a
|
|
40
|
+
consumer driving a loop had to reconstruct "did this iteration deliver something usable?" from all
|
|
41
|
+
three. A pure function of fields the run already carries, so it cannot disagree with them; the granular
|
|
42
|
+
fields stay authoritative. Absent whenever `verdict` is absent. Note `delivered_*` means "no
|
|
43
|
+
stall/question signal fired", not positive evidence a deliverable exists — check `artifacts`.
|
|
44
|
+
- **Generation-pairing `jq` recipes** in [docs/stats.md](./docs/stats.md) — pass-rate/cost per generation,
|
|
45
|
+
a per-generation verdict-signal histogram, and the before/after of a single fix, grouped on
|
|
46
|
+
`skillHash`/`runLabel` over `index.jsonl`.
|
|
47
|
+
|
|
48
|
+
### Changed
|
|
49
|
+
|
|
50
|
+
- **`diff`'s exit code now honors its own documented contract.** `--help` has always said "transcript is
|
|
51
|
+
advisory … tools/artifacts/meta are the gateable signal", but `identical` conjoined all four views, so
|
|
52
|
+
two live runs of the SAME skill exited 1 and the signal could not separate "behaviour changed" from
|
|
53
|
+
model stochasticity. `identical` now means the gateable views agree; transcript drift is reported
|
|
54
|
+
separately (`transcriptDiffers` in JSON, rendered in text when you ask for that view) so it stays
|
|
55
|
+
visible. `diff` also carries `skillHash` in the meta view now — a diff across a fix could not previously
|
|
56
|
+
name which two generations it compared.
|
|
57
|
+
|
|
58
|
+
### Fixed
|
|
59
|
+
|
|
60
|
+
- **`fingerprint.skillHash` and `skillCommit` are now recorded on the `skill` and `probe-dispatch` lanes.**
|
|
61
|
+
Both resolved their skill dirs by re-reading the session *file*, so the lanes that mount via an in-memory
|
|
62
|
+
session — and pass the `"(inline)"` sentinel as the session path — emitted no `skillHash` and a null
|
|
63
|
+
`skillCommit`, even though the mounts were already in scope at the call site. The resolved session object
|
|
64
|
+
is now threaded through instead. Consequences, all previously broken on those lanes: `result.json` carries
|
|
65
|
+
the content-exact generation key, the run-index `skillHash` column populates (so a harvest step can group
|
|
66
|
+
runs by generation), and the run's own "pair critiques by `fingerprint.skillHash`" tip — which is printed
|
|
67
|
+
*only* on the `skill` lane — is no longer advertising a field that lane could not emit. Scoped to the
|
|
68
|
+
sentinel branch: the file-based path is untouched, so recorded cassettes and the
|
|
69
|
+
staleness / `verify-run` recomputes are unaffected. A session that mounts nothing still yields no hash —
|
|
70
|
+
there is nothing to hash.
|
|
71
|
+
|
|
72
|
+
### Changed
|
|
73
|
+
|
|
74
|
+
- **Reflective skill-critique prompt v2** (`REFLECTION_PROMPT_VERSION` 1 → 2; maintainer instrument, not a
|
|
75
|
+
shipped surface). Adds a sub-agent question (were any dispatched, and was the skill clear about when to
|
|
76
|
+
dispatch, what context to hand them, and what to expect back); replaces the "change ONE thing" cap with
|
|
77
|
+
exhaustive solicitation, since a separate evaluator already triages and drops ungrounded findings, so
|
|
78
|
+
capping at the source loses signal for no quality gain; drops a "fidelity tier" example that is
|
|
79
|
+
cowork-harness vocabulary a third-party skill's agent never encountered; and bounds the pass-2 self-report
|
|
80
|
+
now that the prompt invites longer replies.
|
|
81
|
+
|
|
9
82
|
## [1.4.0] — 2026-07-19
|
|
10
83
|
|
|
11
84
|
### Added
|
package/README.md
CHANGED
|
@@ -91,7 +91,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
91
91
|
|
|
92
92
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
93
93
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
94
|
-
> From a global install (`npm i -g "cowork-harness@>=1.
|
|
94
|
+
> From a global install (`npm i -g "cowork-harness@>=1.5.0"`), point at the package root instead:
|
|
95
95
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
96
96
|
> (or copy the cassette into your own project and pass that path).
|
|
97
97
|
|
|
@@ -101,7 +101,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
101
101
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
102
102
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
103
103
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
104
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.
|
|
104
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@>=1.5.0"`.
|
|
105
105
|
|
|
106
106
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
107
107
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -126,7 +126,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
126
126
|
claude plugin install cowork-harness@cowork-harness
|
|
127
127
|
```
|
|
128
128
|
|
|
129
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.
|
|
129
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@>=1.5.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 20). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
130
130
|
|
|
131
131
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
132
132
|
|
|
@@ -147,7 +147,7 @@ A global install is enough for CI `lint`, reading the teaching skill, and replay
|
|
|
147
147
|
To `run` the worked examples live or copy them as a starting point, use a source checkout. (The marketplace
|
|
148
148
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
149
149
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
150
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.
|
|
150
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@>=1.5.0"` — see
|
|
151
151
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
152
152
|
|
|
153
153
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -392,7 +392,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
392
392
|
|
|
393
393
|
| Command | What it does | Reach for it when… |
|
|
394
394
|
|---|---|---|
|
|
395
|
-
| `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
|
|
395
|
+
| `skill <folder> "<prompt>"` | Run a local skill/plugin folder once against the staged agent; `--compact`/`--demo` trim output for shareable screenshots/GIFs; `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift); `--label <tag>` stamps a generation name for the iterate-across-fixes loop (surfaced in `result.json`/index/`inspect`); `--repeat N` (2-100) runs the same skill+prompt N times and aggregates a variance rollup — "did this finding reproduce, or did it pass once?"; `--allow-missing-capability` stops a capability false-negative on the lean `core` image from failing the verdict (open-ended equivalent of asserting `allow_missing_capability: true`) | ad-hoc "is the skill alive / does it do X?" — the fast inner loop |
|
|
396
396
|
| `run <scenario.yaml \| dir/>` | Run authored scenarios with `assert:` + a CI-ready exit code; a decider can answer unscripted gates; `--repeat`/`--matrix` add variance runs / a compatibility matrix (detail below); `--ablate-skill` runs the same prompt with the skill removed (a negative control for skill-lift) | you want a repeatable, **asserted regression test** — to **measure flakiness** instead of trusting one green — or to **test a compatibility matrix** (multiple baselines/models/skill variants) in one run |
|
|
397
397
|
| `chat <folder> [prompt]` | Interactive multi-turn REPL against a skill (TTY); optional seed prompt is sent as the first turn. `--upload <file>` / `--folder <dir>` (repeatable) attach files/project folders; `--verbose` shows thinking blocks + tool inputs; `--fidelity protocol\|container\|hostloop` (no `microvm`/`cowork` in the REPL); `--allow-host-writes` consents to a writable `hostloop`-fidelity connected folder (same consent as clicking "connect folder" in Desktop); `--raw` skips the control protocol for native `docker run -it` (rejects `--upload`/`--folder`/`--plugin`/`--fidelity`) | debugging a multi-turn flow by hand |
|
|
398
398
|
| `record` / `replay` | **Record a live run once → replay it token-free, Docker-free thereafter** (key flags below; `replay --explain` prints the evidence behind every passing assert) | **token-free, Docker-free CI** from a once-recorded run |
|
|
@@ -415,6 +415,7 @@ Skill testing is the headline use, but the tool is a general harness over the Co
|
|
|
415
415
|
| `boundary-check [baseline] [--session <file>]` | Prove the **L1 Docker** sandbox enforces Cowork's limitations (sealed FS + default-deny egress; `container`/`hostloop` share this sandbox — `microvm`'s guest firewall is not probed here); `--session` folds a session's `egress.extra_allow` into the probe allowlist | verifying the harness's own fidelity |
|
|
416
416
|
| `sync` / `list` | Derive/refresh (`sync [--diff] [--allow-empty\|--force]`) & list platform baselines from the Desktop install | after Claude Desktop updates (baselines ship, so it's optional otherwise) |
|
|
417
417
|
| `diff <a> <b>` | Compare two baselines, two runs, two cassettes, or a run+cassette — kind auto-detected by content (view/normalization flags below). Token-free, no live Desktop/Docker needed | "what changed between two runs/cassettes/baselines?" |
|
|
418
|
+
| `critique <skill-folder>` | **EXPERIMENTAL.** Run a skill, ask the agent what confused it, then grade that self-report against a frozen record of the run — blinded evaluator, mechanical citation checking. Discovery instrument, never a gate (findings always exit 0). Costs four model workloads; see [docs/critique.md](./docs/critique.md) | "what confused the agent about my skill?" |
|
|
418
419
|
| `doctor [--tier <t>]` | Read-only prerequisite check, per tier (Docker + agent image for `container`/`hostloop`/`cowork`; **Lima** for `microvm`; plus staged agent, token, baseline); prints the exact `docker build` line if the agent image is missing | "can I run the live tiers — what's missing?" before a first live run |
|
|
419
420
|
| `prune [<runs-dir>] [--keep-last <n>] [--pinned-older-than <N>d\|h\|m]` | Prune accumulated run dirs (keeps the N most recent per scenario; pinned `--session-id` runs are never pruned unless `--pinned-older-than` opts in to reclaiming stale ones by last-activity age); the optional positional overrides the runs root; `--dry-run` | the machine-global runs root has grown and you want space back |
|
|
420
421
|
| `rehash <dir/>` | Migrate cassette fingerprints to the current format version when the content is provably unchanged (`--dry-run`); no re-record needed | a cassette-format bump flagged committed fixtures as stale |
|
|
@@ -670,7 +671,7 @@ jobs:
|
|
|
670
671
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
671
672
|
```
|
|
672
673
|
|
|
673
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.
|
|
674
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — intentional so recipes track the current release; pin an exact version for reproducible CI. The companion skill's `cowork-harness@>=1.5.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any advisory finding); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). See [`action.yml`](./action.yml) for the full input reference.
|
|
674
675
|
|
|
675
676
|
The provided [GitHub Actions workflow](.github/workflows/ci.yml) runs a **six-stage pipeline**. The **unit** stage is the token-free gate you can copy into your skill repo; the `action-self-test`, `python`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
676
677
|
|
package/SPEC.md
CHANGED
|
@@ -603,7 +603,7 @@ unused — reserving it now keeps a later addition additive rather than a renumb
|
|
|
603
603
|
`0`/`1`/`2`/`3` space. Exit-code space is **per-command**, not global (`status` uses `0`/`1`/`2`/`3`
|
|
604
604
|
with its own meanings); this reservation applies only to the `run`/`skill` family.
|
|
605
605
|
|
|
606
|
-
**Per-command exceptions:** `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
|
|
606
|
+
**Per-command exceptions:** `critique` **never gates on findings** — it exits `0` for any finding of any classification, and even when the task run it graded ERRORED (that is a finding about the skill, not a broken instrument). It exits `2` only for a usage error or an **instrument failure**: the turn was killed, the reflection protocol broke, or the evaluator was never invoked *or threw* — i.e. no critique was produced. Do not gate CI on `critique`; that inverts its design. `lint` exits `127` when `python3` is missing (spawn error); `replay` exits
|
|
607
607
|
`2` on a **whole-cassette operational failure** — anything `readCassette` rejects (unreadable, invalid
|
|
608
608
|
shape, unsupported version, unrecognized assertion key) or any per-file throw, plus the batch loop's
|
|
609
609
|
own source-resolution failures (`--assert-from`/`--reassert` drift, scenario-parse errors, `--write`
|