cowork-harness 2.4.0 → 2.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +5 -5
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +72 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +5 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +2 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +64 -0
- package/CHANGELOG.md +203 -0
- package/README.md +6 -6
- package/SPEC.md +12 -0
- package/dist/assert.js +42 -1
- package/dist/cli.js +8 -0
- package/dist/critique/command.js +202 -22
- package/dist/critique/evaluator.js +9 -3
- package/dist/critique/package-evidence.js +39 -5
- package/dist/run/cassette.js +23 -1
- package/dist/run/chat-result.js +1 -0
- package/dist/run/execute.js +31 -1
- package/dist/run/probe-dispatch.js +8 -2
- package/dist/run/provenance.js +6 -9
- package/dist/run/run.js +197 -8
- package/dist/run/tier-vacuous-tools.js +69 -0
- package/dist/run/tool-name-canonicalization.js +70 -0
- package/dist/types.js +36 -3
- package/docs/cassette.md +2 -0
- package/docs/critique.md +52 -0
- package/docs/debugging.md +3 -1
- package/docs/scenario.md +8 -5
- package/docs/subagents.md +5 -3
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/critique-report.json +41 -18
- package/schema/run-result.json +47 -1
- package/schema/scenario.schema.json +13 -3
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.
|
|
7
|
-
tracks-harness: cowork-harness 2.
|
|
6
|
+
version: 2.5.0
|
|
7
|
+
tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
|
|
29
29
|
> `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -147,7 +147,7 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
147
147
|
| Tier | What it gives you | Use when |
|
|
148
148
|
|---|---|---|
|
|
149
149
|
| `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
|
|
150
|
-
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm
|
|
150
|
+
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
|
|
151
151
|
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
|
|
152
152
|
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
|
|
153
153
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "2.
|
|
20
|
+
(e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.5.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^2.
|
|
345
|
+
- run: npm i -g "cowork-harness@^2.5.0"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^2.
|
|
374
|
+
run: npm i -g "cowork-harness@^2.5.0"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.
|
|
3
|
+
Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
|
|
|
34
34
|
**total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
|
|
35
35
|
UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
|
|
36
36
|
|
|
37
|
+
## An exit-2 report — which turn failed, and whether it was really infrastructure
|
|
38
|
+
|
|
39
|
+
Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
|
|
40
|
+
carry that; read all three before touching anything.
|
|
41
|
+
|
|
42
|
+
| Field | Meaning |
|
|
43
|
+
|---|---|
|
|
44
|
+
| `infraFailure` | the reason |
|
|
45
|
+
| `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
|
|
46
|
+
| `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
|
|
47
|
+
| `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
|
|
48
|
+
|
|
49
|
+
**Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
|
|
50
|
+
unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
|
|
51
|
+
agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
|
|
52
|
+
and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
|
|
53
|
+
**`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
|
|
54
|
+
turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
|
|
55
|
+
unrecognized kind renders as infrastructure):
|
|
56
|
+
|
|
57
|
+
- `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
|
|
58
|
+
- `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
|
|
59
|
+
|
|
60
|
+
**A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
|
|
61
|
+
full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
|
|
62
|
+
a broken instrument.
|
|
63
|
+
|
|
64
|
+
**Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
|
|
65
|
+
verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
|
|
66
|
+
"the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
|
|
67
|
+
`--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
|
|
68
|
+
reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
|
|
69
|
+
is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
|
|
70
|
+
*conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
|
|
71
|
+
|
|
72
|
+
For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
|
|
73
|
+
gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
|
|
74
|
+
(or the scenario's `on_unanswered: llm`).
|
|
75
|
+
|
|
76
|
+
## Which model was graded — `gradedModels`
|
|
77
|
+
|
|
78
|
+
**The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
|
|
79
|
+
session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
|
|
80
|
+
which may not be the model you are otherwise working under, and nothing about the run announces it.
|
|
81
|
+
|
|
82
|
+
`gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
|
|
83
|
+
is the only record of which model produced the behaviour being graded — distinct from the evaluator's
|
|
84
|
+
resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
|
|
85
|
+
separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
|
|
86
|
+
`--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
|
|
87
|
+
read `gradedModels` back to confirm it took.
|
|
88
|
+
|
|
89
|
+
It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
|
|
90
|
+
messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
|
|
91
|
+
(crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
|
|
92
|
+
can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
|
|
93
|
+
|
|
37
94
|
## The report's item shape — no `title`, no `summary`
|
|
38
95
|
|
|
39
96
|
Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
|
|
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
|
|
|
80
137
|
specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
|
|
81
138
|
`not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
|
|
82
139
|
|
|
140
|
+
## Read `referencesAccessed`, not `referencesRead`
|
|
141
|
+
|
|
142
|
+
`referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
|
|
143
|
+
`Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
|
|
144
|
+
`referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
|
|
145
|
+
(`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
|
|
146
|
+
|
|
147
|
+
Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
|
|
148
|
+
entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
|
|
149
|
+
then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
|
|
150
|
+
weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
|
|
151
|
+
nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
|
|
152
|
+
"none".
|
|
153
|
+
|
|
83
154
|
## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
|
|
84
155
|
|
|
85
156
|
`result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
|
|
4
4
|
(baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -305,6 +305,10 @@ same set live from the schema.
|
|
|
305
305
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
306
306
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
307
307
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
308
|
+
| **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
|
|
309
|
+
| **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
|
|
310
|
+
| `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
|
|
311
|
+
| `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
|
|
308
312
|
| `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
|
|
309
313
|
| `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
|
|
310
314
|
| `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.
|
|
5
|
+
Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -35,6 +35,7 @@
|
|
|
35
35
|
"no_hook_blocked",
|
|
36
36
|
"no_lost_write_back",
|
|
37
37
|
"no_mcp_error",
|
|
38
|
+
"no_observed_reference_access",
|
|
38
39
|
"no_path_denied",
|
|
39
40
|
"no_scratchpad_leak",
|
|
40
41
|
"no_skill_triggered",
|
|
@@ -46,6 +47,7 @@
|
|
|
46
47
|
"question_context",
|
|
47
48
|
"question_options",
|
|
48
49
|
"questions_count_max",
|
|
50
|
+
"reference_read",
|
|
49
51
|
"replay_protocol_fidelity",
|
|
50
52
|
"result",
|
|
51
53
|
"self_heal_ran",
|
|
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
|
|
|
80
80
|
"tool_result_not_matches",
|
|
81
81
|
"tool_called",
|
|
82
82
|
"tool_not_called",
|
|
83
|
+
# Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
|
|
84
|
+
# tool inputs), so they are content keys, not live-only.
|
|
85
|
+
"reference_read",
|
|
86
|
+
"no_observed_reference_access",
|
|
83
87
|
"subagent_tool_used",
|
|
84
88
|
"subagent_tool_absent",
|
|
85
89
|
"subagent_dispatched",
|
|
@@ -294,6 +298,8 @@ REGEX_KEYS = {
|
|
|
294
298
|
"hook_blocked",
|
|
295
299
|
"tool_result_matches",
|
|
296
300
|
"tool_result_not_matches",
|
|
301
|
+
"reference_read",
|
|
302
|
+
"no_observed_reference_access",
|
|
297
303
|
}
|
|
298
304
|
VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
|
|
299
305
|
VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
|
|
@@ -583,6 +589,44 @@ def lint_doc(doc, path, raw_lines):
|
|
|
583
589
|
# warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
|
|
584
590
|
# declare the combination incompatible, so authoring it is a bug even if a tool-free run could
|
|
585
591
|
# accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
|
|
592
|
+
# `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
|
|
593
|
+
# vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
|
|
594
|
+
# (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
|
|
595
|
+
# and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
|
|
596
|
+
# passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
|
|
597
|
+
# must never be flagged here. The harness refuses these at load; this catches them before any run.
|
|
598
|
+
_TIER_VACUOUS = {
|
|
599
|
+
"hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
|
|
600
|
+
"container": {"mcp__workspace__bash": "Bash"},
|
|
601
|
+
"microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
|
|
602
|
+
# `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
|
|
603
|
+
# host CLI registry — machine-dependent and about a different product.
|
|
604
|
+
}
|
|
605
|
+
# BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
|
|
606
|
+
# USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
|
|
607
|
+
for _key in ("tool_not_called", "subagent_tool_absent"):
|
|
608
|
+
for _v in _assert_values(items, _key):
|
|
609
|
+
if not isinstance(_v, str) or "*" in _v or "?" in _v:
|
|
610
|
+
continue # a glob is not a literal claim about one tool
|
|
611
|
+
_repl = _TIER_VACUOUS.get(fidelity, {})
|
|
612
|
+
if _v not in _repl:
|
|
613
|
+
continue
|
|
614
|
+
_instead = _repl[_v]
|
|
615
|
+
findings.append(
|
|
616
|
+
Finding(
|
|
617
|
+
"WARN",
|
|
618
|
+
"tool-not-called-tier-vacuous",
|
|
619
|
+
f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
|
|
620
|
+
f"`{_v}` at all, so this can never be violated and verifies nothing.",
|
|
621
|
+
(
|
|
622
|
+
f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
|
|
623
|
+
if _instead
|
|
624
|
+
else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
|
|
625
|
+
),
|
|
626
|
+
path,
|
|
627
|
+
)
|
|
628
|
+
)
|
|
629
|
+
|
|
586
630
|
# linter stays offline — the message carries the gate fact instead of reading a baseline).
|
|
587
631
|
if "transcript_no_host_path" in assert_keys:
|
|
588
632
|
if fidelity in ("hostloop", "protocol"):
|
|
@@ -1016,6 +1060,26 @@ def lint_doc(doc, path, raw_lines):
|
|
|
1016
1060
|
)
|
|
1017
1061
|
)
|
|
1018
1062
|
|
|
1063
|
+
# Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
|
|
1064
|
+
# with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
|
|
1065
|
+
# that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
|
|
1066
|
+
# this only fires on the case that is unambiguously self-defeating.
|
|
1067
|
+
read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
|
|
1068
|
+
unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
|
|
1069
|
+
both_refs = sorted(read_pats & unread_pats)
|
|
1070
|
+
if both_refs:
|
|
1071
|
+
findings.append(
|
|
1072
|
+
Finding(
|
|
1073
|
+
"ERROR",
|
|
1074
|
+
"reference-access-contradiction",
|
|
1075
|
+
f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
|
|
1076
|
+
"(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
|
|
1077
|
+
"Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
|
|
1078
|
+
"another is not, give the two keys different patterns.",
|
|
1079
|
+
path,
|
|
1080
|
+
)
|
|
1081
|
+
)
|
|
1082
|
+
|
|
1019
1083
|
# W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
|
|
1020
1084
|
findings.extend(_lint_regex_quoting(path, raw_lines))
|
|
1021
1085
|
|
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,209 @@ All notable changes to this project are documented here. The format is based on
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [2.5.0] — 2026-08-28
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- **`cowork-harness assertions --list` gains a "Skill references (progressive disclosure)" family** for the
|
|
14
|
+
two new keys, so they are not appended to a flat dump nobody reads.
|
|
15
|
+
|
|
16
|
+
- **A guard that a committed cassette's `tool_not_called` is actually violable by its own recording.**
|
|
17
|
+
The tool surface at a tier is **gate-conditional**: at `container`, `mcp__workspace__web_fetch` is
|
|
18
|
+
offered only when `coworkWebFetchViaApi` is on (17 of 138 measured container runs), and with the gate
|
|
19
|
+
off `WebFetch` is offered instead. `examples/replays/example-pdf-skill.cassette.json` asserts
|
|
20
|
+
`tool_not_called: "mcp__workspace__web_fetch"` and passes today only because it was recorded gate-ON —
|
|
21
|
+
a re-record with the gate off would leave it naming a tool the run could never have called, so it would
|
|
22
|
+
keep passing while verifying nothing, and nothing in the suite would notice. `verify-cassettes`'s
|
|
23
|
+
`replaced-builtin` note keys on the recorded *inventory*, never on the assertions, and only covers
|
|
24
|
+
built-ins being replaced, never the inverse.
|
|
25
|
+
|
|
26
|
+
The guard checks every committed cassette's `tool_called`/`tool_not_called` against that cassette's own
|
|
27
|
+
frozen init inventory, through the same glob engine the evaluator uses. It deliberately does **not**
|
|
28
|
+
cover `subagent_tool_absent` (judged against the per-dispatch `declaredTools`, a different inventory)
|
|
29
|
+
and cannot see alias-class vacuity (`tool_not_called: "Task"` names a tool that is in every inventory
|
|
30
|
+
yet never emitted — the agent binary canonicalizes `Task` to `Agent`); both are recorded as separate
|
|
31
|
+
work rather than left implied.
|
|
32
|
+
|
|
33
|
+
- **`referencesAccessed` — reference access through EVERY tool channel, not just the `Read` tool.**
|
|
34
|
+
`referencesRead` counts one channel, and `critique`'s headline invited a reader to conclude the agent
|
|
35
|
+
had opened no reference — a claim about *reading* that a one-channel count cannot support. An agent
|
|
36
|
+
that `cat`s, `grep`s or globs a reference has reached it just as much. The new `RunResult` field (and
|
|
37
|
+
its `subagents[]` twin) records each file with the channel(s) it was reached through: `read`
|
|
38
|
+
(`Read.file_path`), `grep` (its `path` input), and `bash` (a `Bash`/`mcp__workspace__bash`
|
|
39
|
+
command naming the path).
|
|
40
|
+
|
|
41
|
+
All four channels apply the **same** `skillReferenceReadPath()` predicate, so a token only counts when
|
|
42
|
+
it is rooted in the mounted plugin — the agent's own `node scripts/build.js` is not a skill-script
|
|
43
|
+
access, and a non-skill filename never reaches `result.json` under a field claiming it is skill
|
|
44
|
+
content. Redirection targets and every argument of a write verb (`rm`/`mv`/`mkdir`/`touch`/`chmod`/`tee`) is excluded, as is a verb that only inspects metadata (`ls`/`test`/`stat`/`echo`): those are files the command
|
|
45
|
+
wrote or destroyed.
|
|
46
|
+
|
|
47
|
+
It **deliberately under-approximates**. A `cd` into the skill dir followed by a bare
|
|
48
|
+
`cat references/x.md`, a heredoc body and a `$VAR`-built path are all invisible, and there are tests
|
|
49
|
+
pinning them as misses so a later widening is a visible change rather than a silent one.
|
|
50
|
+
|
|
51
|
+
**Presence is the cannot-verify channel**, which is the one way it differs from `referencesRead`:
|
|
52
|
+
`[]` means the drive ran and observed nothing (a real negative); **absent** means there was no
|
|
53
|
+
observable drive, and must never be read as "none". Present on live *and* replay — cassettes freeze
|
|
54
|
+
whole tool inputs, so the replay re-drive reconstructs every channel identically.
|
|
55
|
+
|
|
56
|
+
`referencesRead` is unchanged in meaning and is now documented as this field's `read`-channel
|
|
57
|
+
projection, produced by the same capture so the two cannot disagree.
|
|
58
|
+
|
|
59
|
+
Scope is **main agent ∪ sub-agents**, via a single `unionReferenceAccesses()` derivation shared by the
|
|
60
|
+
assertion keys, the critique report and `--probe dispatch` — a dispatcher-shaped skill does all its
|
|
61
|
+
reading a level down, so judging the top-level list alone would report "never reached" on a run where a
|
|
62
|
+
sub-agent read the file cover to cover. A **truncated cassette** (one that could never be driven)
|
|
63
|
+
reports cannot-verify rather than an empty list, for the same reason.
|
|
64
|
+
|
|
65
|
+
- **`reference_read` / `no_observed_reference_access` assertion keys.** Gate on whether a skill's
|
|
66
|
+
progressive disclosure actually works: a well-partitioned skill and one whose second half is dead look
|
|
67
|
+
identical from outside. Regex (unanchored, case-insensitive — the same shared helper every regex key
|
|
68
|
+
uses), main agent ∪ sub-agents, evaluated on replay as well as live.
|
|
69
|
+
|
|
70
|
+
The negative key is named `no_observed_reference_access`, not `no_reference_read`, because the
|
|
71
|
+
detector under-approximates by design: it proves nothing was *seen*, not that the file went unread.
|
|
72
|
+
Both keys **fail evidence-unavailable** when the run recorded no observable access list — including
|
|
73
|
+
the negative one, which is the direction that would otherwise pass vacuously off a missing field. The
|
|
74
|
+
scenario linter rejects asserting both with the same pattern.
|
|
75
|
+
|
|
76
|
+
- **`critique` names why the GRADED turn errored** — `gradedErrorReason` in the JSON report and inline in
|
|
77
|
+
the text NOTE. `taskResult: "error"` is a legitimate **gradeable** outcome and the critique still runs,
|
|
78
|
+
but the report said only "the task run ended in error", so an exhausted quota or a dropped connection
|
|
79
|
+
read as a defect in the skill under review.
|
|
80
|
+
|
|
81
|
+
- **`critique` reports which model produced the GRADED run** — `gradedModels` in the JSON report and
|
|
82
|
+
`graded model(s):` in the text header, read back from the graded turn's own `result.json` and filtered
|
|
83
|
+
of the agent's locally-fabricated `<synthetic>` entries. The report named the *evaluator's* resolved
|
|
84
|
+
model and no other, so the model that produced the behaviour being graded appeared nowhere. It cannot
|
|
85
|
+
be inferred from context either: the turns are a subprocess that inherits no model from whatever
|
|
86
|
+
invoked `critique`, so an omitted `--model` silently grades whatever the spawned agent defaults to. A
|
|
87
|
+
run with nothing recorded now says `graded model(s): unknown` rather than staying silent. The ids are
|
|
88
|
+
**observed, not requested** — read from the model stamped on the graded turn's assistant messages, not
|
|
89
|
+
from the flag.
|
|
90
|
+
- `isLiveModelId` (`src/types.ts`) — the agent-marker filter every consumer of `RunResult.models` must
|
|
91
|
+
apply, now declared once beside the field it governs. It replaces **three** divergent copies, two of
|
|
92
|
+
which disagreed: `src/run/provenance.ts` (which renders `provenance.model` on every JSON envelope)
|
|
93
|
+
matched the angle-bracket **shape**, `scripts/eval-gate.ts` the `<` **prefix** — so a malformed or
|
|
94
|
+
truncated marker such as `"<synthetic"` was dropped by one and rendered as if it were a model id by the
|
|
95
|
+
other. The shared rule is the shape.
|
|
96
|
+
|
|
97
|
+
### Changed
|
|
98
|
+
|
|
99
|
+
- **Skill and docs updated for the tier refusal.** `SKILL.md` previously told authors a tier-mismatched
|
|
100
|
+
`tool_not_called` "can silently void a web-fetch assertion" when moving a scenario between tiers — the
|
|
101
|
+
harness now refuses it at load instead, so that guidance described behaviour that no longer exists. The
|
|
102
|
+
skill's critique reference gains the `referencesAccessed` field, its channels, and the
|
|
103
|
+
`[]`-vs-absent cannot-verify rule.
|
|
104
|
+
|
|
105
|
+
- **`subagent_declared_but_unused` documents that it fires only on a dispatch that declares a tool list.**
|
|
106
|
+
It reads `subagents[].declaredTools`, populated from a `tools`/`allowedTools` key in the dispatch input.
|
|
107
|
+
The `Agent` tool carries neither, so that list is empty and the key passes on every such dispatch —
|
|
108
|
+
**0 of 1091 real dispatches** carry a non-empty list. The key is not wrong, but a green means "not
|
|
109
|
+
applicable here", never "no fabrication", and neither its description nor its docs said so.
|
|
110
|
+
|
|
111
|
+
- **BEHAVIOUR CHANGE: `tool_not_called` and `subagent_tool_absent` naming a tool the tier does not serve
|
|
112
|
+
are now REFUSED at scenario load.** `tool_not_called: "Bash"` at `hostloop` passed vacuously, always — that tier disallows the
|
|
113
|
+
built-in shell and aliases it to `mcp__workspace__bash`, so the run could never have called `Bash`
|
|
114
|
+
whatever the agent did. The assertion read as a guarantee and verified nothing. The inverse was equally
|
|
115
|
+
broken: `mcp__workspace__bash` at `container`, where the built-in is served instead.
|
|
116
|
+
|
|
117
|
+
The refusal is a `UsageError` at the point the tier resolves — after `fidelity: cowork` becomes a real
|
|
118
|
+
tier, and before any staging, image pull or spawn, so no model spend is wasted. The message names the
|
|
119
|
+
tool to assert instead, and names the sibling keys that behave differently (`tool_called` still fails
|
|
120
|
+
normally; `subagent_tool_absent` is judged against a per-dispatch inventory this check cannot
|
|
121
|
+
determine). `scenario.py lint` WARNs on the same set, before any run at all.
|
|
122
|
+
|
|
123
|
+
**The table is closed to tools the harness itself removes or registers**, and never derived from the
|
|
124
|
+
launch plan. `--tools` gates the built-in set alone — the agent binary's own help says so — while every
|
|
125
|
+
sandbox tier separately passes `--mcp-config`. A launch-set-derived check would therefore have rejected
|
|
126
|
+
`tool_not_called: "mcp__example-fs__write_file"` for a session using the `mcp.config` this repo ships an
|
|
127
|
+
example of, while that tool was registered and callable. Globs are never refused, and the table
|
|
128
|
+
under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the correct
|
|
129
|
+
side to err on when the verdict is a hard refusal.
|
|
130
|
+
|
|
131
|
+
There is **no opt-out**, deliberately. The repo's `allow_*` modifiers all cover cases where the harness
|
|
132
|
+
might be wrong about a real signal; a fired reject here cannot be a false positive, so there is no
|
|
133
|
+
legitimate scenario to rescue. If one is ever found, the table is wrong and the table should change.
|
|
134
|
+
|
|
135
|
+
**`subagent_tool_absent` is covered for the same reason `tool_not_called` is.** It reads the tools
|
|
136
|
+
sub-agents actually USED, not a per-dispatch declared list, so a tool the tier never serves makes it
|
|
137
|
+
vacuous in exactly the same way — corroborated by the run population, where sub-agent `Bash` calls
|
|
138
|
+
appear 20 times, all at `container` and never at `hostloop`. Covering one key and not the other would
|
|
139
|
+
refuse an assertion at hostloop while silently greening the sub-agent form of the identical claim.
|
|
140
|
+
|
|
141
|
+
`e2e/scenarios/canary-hostloop.yaml` carried exactly this defect and is fixed in the same change: its
|
|
142
|
+
`tool_not_called: Bash` never proved anything, and its `tool_called: mcp__workspace__*` already carries
|
|
143
|
+
the canary's stated purpose.
|
|
144
|
+
|
|
145
|
+
- **`critique`'s "no references were Read" headline now says what it observed.** It reads the wide
|
|
146
|
+
signal, names the channels it looked through, and states its own under-approximation in one short
|
|
147
|
+
clause instead of the load-bearing caveat that did all the work. Where the run recorded no observable
|
|
148
|
+
tool stream it makes **no claim** rather than rendering a clean negative. The evaluator's evidence
|
|
149
|
+
section moves with it (previously main-agent `Read`s only — a different population from the headline's,
|
|
150
|
+
so widening one without the other would have handed the evaluator a prompt contradicting the report),
|
|
151
|
+
and the grading prompt now tells the evaluator not to issue a finding whose only support is a path
|
|
152
|
+
missing from that list. `--probe dispatch` prints the wide list too.
|
|
153
|
+
|
|
154
|
+
### Fixed
|
|
155
|
+
|
|
156
|
+
- **`tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did.** The agent binary
|
|
157
|
+
canonicalizes a set of legacy tool names — `Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, and nine
|
|
158
|
+
more — while the spawn tool list still declares the **legacy** spelling. So the init inventory echoes
|
|
159
|
+
back `Task`, every actual dispatch is emitted as `Agent`, and a literal matcher could never connect the
|
|
160
|
+
two. Measured across 506 kept runs: `Task` offered **506** times and called **0**; `Agent` called **188**
|
|
161
|
+
times and offered **0**. The negative form was a permanent vacuous pass in the most common dispatch
|
|
162
|
+
assertion there is, at every tier.
|
|
163
|
+
|
|
164
|
+
All four tool keys (`tool_called`, `tool_not_called`, `subagent_tool_used`, `subagent_tool_absent`) now
|
|
165
|
+
match either spelling, through the one shared matcher. Globs see both too — the **recorded name** is
|
|
166
|
+
expanded rather than the author's pattern, so `tool_not_called: "Ta*"` and `"*"` are violated by a
|
|
167
|
+
recorded `Agent`, which rewriting the pattern would not have fixed.
|
|
168
|
+
|
|
169
|
+
Recorded data is unchanged: `toolCounts`, `context.tools` and every cassette keep exactly what the agent
|
|
170
|
+
reported, the same verbatim posture `RunResult.models` documents. Only matching is alias-aware, so
|
|
171
|
+
`tool_available: "Task"` still matches the inventory's literal `Task` and no committed cassette changes
|
|
172
|
+
meaning.
|
|
173
|
+
|
|
174
|
+
**Still vacuous, and now documented as such:** `tool_not_called` on a tool the *tier* never offers —
|
|
175
|
+
`Bash` at `hostloop`, or `mcp__workspace__bash` at `container`. That is a separate class with a separate
|
|
176
|
+
fix; the key's docs and the skill reference now warn about it rather than leaving it implied.
|
|
177
|
+
|
|
178
|
+
- **`critique` diagnosed an unanswered gate as an infrastructure failure, and named the wrong turn while
|
|
179
|
+
doing it.** A graded run that stopped at an `AskUserQuestion` gate with no scripted answer exits 2 with
|
|
180
|
+
a fully-formed `{ok:false, error:{category:"unanswered", …}}` envelope, but the report read only
|
|
181
|
+
`results[0].outDir` and answered *"task turn exited nonzero without a parseable result envelope — it
|
|
182
|
+
crashed before completing a gradeable task"* — under a header hardcoded to
|
|
183
|
+
`INFRASTRUCTURE/PROTOCOL FAILURE (reflection turn)`, though it was the **task** turn that stopped and
|
|
184
|
+
no reflection turn had been attempted. Both halves pointed at the wrong subsystem: a reader following
|
|
185
|
+
the report would audit Docker and the staged agent over what is one scenario flag.
|
|
186
|
+
|
|
187
|
+
The report now prefers the failed turn's own diagnosis, on both turns, carrying its message **and hint
|
|
188
|
+
verbatim** (and only once — `UnansweredError`'s message already contains its hint, so appending it
|
|
189
|
+
unconditionally printed the whole question, its options and its remedy tip twice). No remedy text is
|
|
190
|
+
synthesized on top: there are 36 `UnansweredError` sites and only one is "the skill asked an unscripted
|
|
191
|
+
question", so advice keyed on the category is wrong for nearly all of them.
|
|
192
|
+
|
|
193
|
+
Two new fields carry the classification — `infraFailurePhase` (`task turn` | `reflection turn`) and
|
|
194
|
+
`infraFailureKind` (the harness `ErrCategory`). The header keys on **which** category, not on merely
|
|
195
|
+
having one: `unanswered` / `usage` / `boundary` render as `RUN FAILED (<turn>, <kind>)`; everything
|
|
196
|
+
else renders as `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>)` and it **fails closed** — `internal` is the
|
|
197
|
+
CLI's catch-all for any unexpected throw (Docker down, container start failure, missing staged agent, a
|
|
198
|
+
harness bug), so an unrecognized category is treated as an instrument failure rather than assumed
|
|
199
|
+
ordinary.
|
|
200
|
+
|
|
201
|
+
- **A turn that exited 1 was reported as a bare exit code, so an exhausted quota looked like a crash.**
|
|
202
|
+
A turn that RAN and errored exits `1` with a full result envelope whose top-level `error` is `null` —
|
|
203
|
+
its cause lives in `results[0]`, not in an error object. Reading only the error object left the report
|
|
204
|
+
saying `reflection turn exited with code 1 (expected 0)` and nothing more, for a run whose actual cause
|
|
205
|
+
was an exhausted seven-day quota; it was findable only by opening `events.jsonl` by hand. The failed
|
|
206
|
+
turn's `resultErrorKind` / `errorSource` / `resultSubtype` are now read and rendered — `usage_limit`
|
|
207
|
+
as "the account's quota is exhausted; retry after the reset. This is NOT a harness or skill defect",
|
|
208
|
+
`transport` as a retryable tail-end drop — matching how the run renderer has always shown them.
|
|
209
|
+
`usage_limit` and `transport` join the ordinary/actionable set; `agent` does not, because for
|
|
210
|
+
critique's own protocol turn an agent-level failure *is* the instrument breaking.
|
|
211
|
+
|
|
9
212
|
## [2.4.0] — 2026-08-27
|
|
10
213
|
|
|
11
214
|
### Fidelity
|
package/README.md
CHANGED
|
@@ -120,7 +120,7 @@ node dist/cli.js replay examples/replays/example-pdf-skill.cassette.json
|
|
|
120
120
|
|
|
121
121
|
> **Installed globally instead?** Once linked/installed, the same command is `cowork-harness replay
|
|
122
122
|
> <cassette>` — but the relative path above only resolves from a source checkout's `examples/replays/`.
|
|
123
|
-
> From a global install (`npm i -g "cowork-harness@^2.
|
|
123
|
+
> From a global install (`npm i -g "cowork-harness@^2.5.0"`), point at the package root instead:
|
|
124
124
|
> `cowork-harness replay "$(npm root -g)/cowork-harness/examples/replays/example-pdf-skill.cassette.json"`
|
|
125
125
|
> (or copy the cassette into your own project and pass that path).
|
|
126
126
|
|
|
@@ -130,7 +130,7 @@ Live `run`/`skill` need the prerequisites in the next section — note the `prot
|
|
|
130
130
|
> - **Replay only (zero setup):** `cowork-harness replay <cassette>` — no token, no Docker, no agent. The command above.
|
|
131
131
|
> - **`protocol` (real model, no Docker):** needs only the auth token (item 3 below).
|
|
132
132
|
> - **Live `container` / `microvm` / `hostloop` / `cowork`:** needs Docker (or Lima for `microvm`), a staged agent, and the token — run `cowork-harness doctor` first.
|
|
133
|
-
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.
|
|
133
|
+
> - **Invocation:** from a source checkout, `node dist/cli.js <cmd>` (or `npm link` to get the `cowork-harness` command); from a global install, `cowork-harness <cmd>`; the companion skill falls back to `npx "cowork-harness@^2.5.0"`.
|
|
134
134
|
|
|
135
135
|
Two more worked examples worth knowing about: `examples/scenarios/protocol-smoke.yaml` (zero-Docker smoke
|
|
136
136
|
test) and `examples/scenarios/skill-loads.yaml` (container-tier acceptance check) — see
|
|
@@ -155,7 +155,7 @@ claude plugin marketplace add yaniv-golan/cowork-harness
|
|
|
155
155
|
claude plugin install cowork-harness@cowork-harness
|
|
156
156
|
```
|
|
157
157
|
|
|
158
|
-
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.
|
|
158
|
+
The skill **self-bootstraps the CLI**: if `cowork-harness` isn't on your PATH it falls back to `npx "cowork-harness@^2.5.0"` (a version floor that fails loud rather than silently fetching a too-old CLI; Node ≥ 22). Tiers above `protocol` still need Docker/Lima and a Claude Desktop agent binary — see the prerequisites below.
|
|
159
159
|
|
|
160
160
|
It also follows the open [Agent Skills](https://agentskills.io) spec, so it installs cross-editor (Cursor, Codex, OpenCode, …) via [`npx skills`](https://github.com/vercel-labs/skills) (Vercel Labs' CLI implementation of that spec):
|
|
161
161
|
|
|
@@ -180,7 +180,7 @@ global install puts nothing in your working directory. The matrix, answer-policy
|
|
|
180
180
|
ones that still need a source checkout. (The marketplace
|
|
181
181
|
skill install itself only pulls `.claude/skills/cowork-harness/` — SKILL.md + `references/` + `scenario.py`/
|
|
182
182
|
assertion keys, per `.claude-plugin/marketplace.json`'s `source` — not the rest of this table; the full set
|
|
183
|
-
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.
|
|
183
|
+
above becomes available once the skill's first command self-bootstraps `npx "cowork-harness@^2.5.0"` — see
|
|
184
184
|
[above](#drive-it-from-claude-code-companion-skill) — which pulls the same npm package as the global-install row.)
|
|
185
185
|
|
|
186
186
|
### Prerequisites for anything above `protocol` fidelity
|
|
@@ -297,7 +297,7 @@ cowork-harness lint examples/scenarios/*.yaml --strict --min-severity WARN
|
|
|
297
297
|
>
|
|
298
298
|
> | When | Assertions |
|
|
299
299
|
> |---|---|
|
|
300
|
-
> | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
|
|
300
|
+
> | Always | `transcript_*`, `tool_*`, `subagent_*`, `no_vm_path_file_op`, `dispatch_count_max`, `skill_triggered`/`no_skill_triggered`, `reference_read`/`no_observed_reference_access`, `max_cost_usd`/`max_tokens`/`tool_calls_max`/`max_turns` (against the *frozen recording's* spend, not fresh spend — a live `run` catches a real budget regression), `max_tool_errors`, `max_redundant_tool_calls`, `skill_available`, `connector_available`, `skill_tool_used`, `compaction_occurred`, `all_tasks_completed`, `task_count_min`, `task_status`, `no_scratchpad_leak`, `present_files_called`, `result`, the verdict modifiers |
|
|
301
301
|
> | Only if the cassette carries `controlOut` | `question_asked`, `question_options`, `question_context`, `questions_count_max`, `gate_answers_delivered`, `gate_answer_count_min`, `hook_blocked`, `no_hook_blocked`, `vm_path_denied`, `path_denied`, `no_path_denied` |
|
|
302
302
|
> | Only if the cassette carries an `artifacts` manifest | `file_exists`, `artifact_text`, `user_visible_artifact`, `artifact_json`, `computer_links_resolve`, `computer_links_resolve_if_present`, `no_unexpected_files`, `input_unmodified` |
|
|
303
303
|
> | Always skipped (live-only) | `file_absent`, `egress_*`, `expect_denied`, `no_delete_in_outputs`, `no_delete_in_mounts`, `self_heal_ran`, `transcript_no_host_path`, `no_mcp_error`, `max_peak_rss_bytes`, `semantic_matches`, `no_lost_write_back` — keep these in a periodic live `run` |
|
|
@@ -802,7 +802,7 @@ jobs:
|
|
|
802
802
|
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
803
803
|
```
|
|
804
804
|
|
|
805
|
-
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.
|
|
805
|
+
Every run writes a Markdown verdict table (scenario, pass/fail, signals, cost/turns when available, staleness findings, and the replay-skipped-assertions honesty line) to the job summary. Inputs: `command`, `path` (required), `version` (npm dist-tag/version, default `latest` — the recipes above pin `^2` instead, because leaving it at `latest` takes a CLI major the moment it is promoted even though your `uses:` ref never changed; pin an exact version for byte-reproducible CI. The companion skill's `cowork-harness@^2.5.0` floor guidance applies to ad-hoc CLI installs, not this input), `strict` (applies to `replay` (staleness findings), `lint`/`lint-skill` (WARN/INFO), and `analyze-skill` (any **`error`**-severity finding — advisory findings are precisely the class that does NOT gate); IGNORED — not forwarded — for `verify-cassettes`/`run`, which don't accept the flag), `fail-on-skill-drift` (**`replay`-only** — never forwarded to the analyzers), `extra-args`, `summary` (default `true`), `anthropic-api-key` (live lane only). Outputs: `ok` (`"true"`/`"false"`, mirrors the exit code), `envelope-path` (path to the raw JSON envelope, for post-processing), `summary-md` (the rendered verdict table, exposed as an output — not just written to `$GITHUB_STEP_SUMMARY` — because that file is scoped to this action's own invocation and a caller's later step gets a fresh, empty one). See [`action.yml`](https://github.com/yaniv-golan/cowork-harness/blob/main/action.yml) for the full input/output reference.
|
|
806
806
|
|
|
807
807
|
The provided [GitHub Actions workflow](https://github.com/yaniv-golan/cowork-harness/blob/main/.github/workflows/ci.yml) runs a **nine-stage pipeline**. The **build** + **test** stages are the token-free gate you can copy into your skill repo; the `floor`, `action-self-test`, `python`, `image-recipe`, `boundary`, `scenarios`, and `parity-drift` stages are this repo's own fidelity self-tests and are not directly portable (they build the harness's Docker image and run harness-specific e2e scenarios — see [`ci-recipe.md`](./.claude/skills/cowork-harness/references/ci-recipe.md) for the skill-repo template):
|
|
808
808
|
|