cowork-harness 2.3.0 → 2.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +16 -9
- package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
- package/.claude/skills/cowork-harness/references/critique.md +72 -1
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +5 -1
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +2 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +88 -0
- package/CHANGELOG.md +333 -0
- package/README.md +6 -6
- package/SPEC.md +12 -0
- package/dist/assert.js +70 -4
- package/dist/cli.js +8 -0
- package/dist/critique/command.js +202 -22
- package/dist/critique/evaluator.js +9 -3
- package/dist/critique/package-evidence.js +39 -5
- package/dist/hostloop/workspace-handler.js +12 -2
- package/dist/run/cassette.js +72 -2
- package/dist/run/chat-result.js +1 -0
- package/dist/run/execute.js +107 -9
- package/dist/run/probe-dispatch.js +8 -2
- package/dist/run/provenance.js +6 -9
- package/dist/run/run.js +197 -8
- package/dist/run/tier-vacuous-tools.js +69 -0
- package/dist/run/tool-name-canonicalization.js +70 -0
- package/dist/runtime/container.js +53 -2
- package/dist/runtime/hostloop.js +52 -8
- package/dist/sync/cowork-sync.js +26 -8
- package/dist/types.js +37 -4
- package/docs/cassette.md +2 -0
- package/docs/critique.md +52 -0
- package/docs/debugging.md +3 -1
- package/docs/fidelity-gaps.md +128 -5
- package/docs/scenario.md +61 -12
- package/docs/subagents.md +5 -3
- package/examples/replays/README.md +1 -1
- package/examples/replays/example-pdf-skill.cassette.json +161 -112
- package/examples/scenarios/example-pdf-skill.yaml +8 -1
- package/package.json +1 -1
- package/python/test_scenario_lint.py +39 -0
- package/schema/critique-report.json +41 -18
- package/schema/run-result.json +47 -1
- package/schema/scenario.schema.json +14 -4
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 2.
|
|
7
|
-
tracks-harness: cowork-harness 2.
|
|
6
|
+
version: 2.5.0
|
|
7
|
+
tracks-harness: cowork-harness 2.5.0 (baseline desktop-1.37937.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
25
25
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
26
26
|
the highest-value part. Read it.
|
|
27
27
|
|
|
28
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.
|
|
28
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 2.5.0` (baseline
|
|
29
29
|
> `desktop-1.37937.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
30
30
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
31
31
|
|
|
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
|
|
|
42
42
|
|
|
43
43
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
44
44
|
- **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
|
|
45
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.
|
|
45
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 2.5.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^2.5.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^2.5.0"`. **Pin `@^2.5.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
|
|
46
46
|
|
|
47
47
|
This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
|
|
48
48
|
OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
|
|
@@ -147,9 +147,9 @@ the folder basename (collision-resolved); there is no `to:` override. See `refer
|
|
|
147
147
|
| Tier | What it gives you | Use when |
|
|
148
148
|
|---|---|---|
|
|
149
149
|
| `protocol` | Fastest; no sandbox, no egress | Pure protocol/answer-shape tests. **Rejected** if the scenario asserts egress. |
|
|
150
|
-
| `container` | Real sandbox + real default-deny egress (**default**) | Most functional + boundary tests. |
|
|
151
|
-
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity | Testing untrusted code escape, not network behavior. |
|
|
152
|
-
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with native Bash
|
|
150
|
+
| `container` | Real sandbox + real default-deny egress (**default**). Models the **VM loop**: keeps the built-in `Bash`, but `WebFetch` is replaced by `mcp__workspace__web_fetch` — assert on that name, not `WebFetch`. (`run`/`record` only; `chat --fidelity container` still offers the built-in.) **The name is tier-specific**: `microvm` never offers it and `protocol` serves the operator's own host registry. A `tool_not_called`/`subagent_tool_absent` naming a tool its tier does not serve is **REFUSED at scenario load** (the message names what to write instead), so moving a scenario between tiers now errors rather than silently voiding the assertion — except at `protocol`, which is never judged | Most functional + boundary tests. |
|
|
151
|
+
| `microvm` | VM-grade escape **isolation** (macOS arm64). Egress transport is the *same allowlist proxy as `container`* — not better network fidelity. Unlike `container`, still offers the built-in `WebFetch` | Testing untrusted code escape, not network behavior. |
|
|
152
|
+
| `hostloop` / `cowork` | Production split-exec: the agent loop is a **native process on the host** (no container around the file tools — matching production), with **both** native `Bash` and `WebFetch` disabled and routed host-side via the workspace SDK-MCP server into a Docker VM sidecar (the VM loop replaces web_fetch only) | Highest-fidelity / parity runs. A writable connected folder needs `allow_host_writes: true` (see scenario-schema.md). |
|
|
153
153
|
|
|
154
154
|
Set the tier in the **scenario's `fidelity:` field**, not a flag — `run` rejects `--fidelity`
|
|
155
155
|
(it's a `skill`/`chat` flag; `run` takes fidelity only from the scenario). See
|
|
@@ -550,8 +550,15 @@ Recognize these before "fixing" a non-bug:
|
|
|
550
550
|
only when you thought to ask for it, and the runs that most need this are the ones where nobody did.
|
|
551
551
|
**Silent when the evidence cannot answer the question** (no workspace walk, or a tier that runs no
|
|
552
552
|
scratchpad walk, absent delivery telemetry, or a resumed turn) — "cannot tell" never reads as "clean".
|
|
553
|
-
**The fix is lane-dependent.** On `lane: local`, write deliverables
|
|
554
|
-
|
|
553
|
+
**The fix is lane-dependent — and so is the PATH.** On `lane: local`, write deliverables where the user
|
|
554
|
+
can see them, but do **not** hardcode the literal prefix `outputs/`: on the desktop-local host-loop lane
|
|
555
|
+
(what production runs) the file tools are ALREADY rooted at `outputs/`, so `outputs/x.md` doubles to
|
|
556
|
+
`outputs/outputs/x.md` and the user never sees it — a **bare filename** is correct there. At
|
|
557
|
+
`fidelity: container`/`microvm` (VM-loop, the harness default) the base is the session root instead, so a
|
|
558
|
+
bare name lands in the scratchpad and you want `{{workspaceFolder}}` or an explicit delivery. Addressing
|
|
559
|
+
a connected folder by name (`<folder>/x.md`) never reaches it on either lane — it builds a same-named
|
|
560
|
+
decoy inside `outputs`, reports success, and gives no signal. Measured 2026-08-27; see
|
|
561
|
+
[docs/scenario.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md), "Where a relative path actually lands". **On `lane: remote`, moving a file under `outputs/` does NOT help** —
|
|
555
562
|
nothing is delivered by location there, so only an explicit delivery counts. Assert
|
|
556
563
|
**`allow_undelivered_deliverables: true`** when the leftovers are intentional (intermediates, caches,
|
|
557
564
|
downloaded inputs) rather than a delivery gap.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
|
|
|
17
17
|
CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
|
|
18
18
|
copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
|
|
19
19
|
patch number to remember, and only wants a human decision at the next major. Pin an exact version
|
|
20
|
-
(e.g. `version: "2.
|
|
20
|
+
(e.g. `version: "2.5.0"`) instead when you want byte-reproducible CI.
|
|
21
21
|
|
|
22
22
|
Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
|
|
23
23
|
cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
|
|
@@ -67,7 +67,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
67
67
|
GitHub-hosted runners, no token/Docker/agent:
|
|
68
68
|
|
|
69
69
|
```yaml
|
|
70
|
-
- run: npm i -g "cowork-harness@^2.
|
|
70
|
+
- run: npm i -g "cowork-harness@^2.5.0"
|
|
71
71
|
- run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
|
|
72
72
|
# no silent false-greens. WITHOUT --strict this
|
|
73
73
|
# step cannot fail on a WARN-class rule (e.g.
|
|
@@ -342,7 +342,7 @@ jobs:
|
|
|
342
342
|
with: { node-version: '24' }
|
|
343
343
|
- uses: actions/setup-python@v5
|
|
344
344
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
345
|
-
- run: npm i -g "cowork-harness@^2.
|
|
345
|
+
- run: npm i -g "cowork-harness@^2.5.0"
|
|
346
346
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
347
347
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
348
348
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -371,7 +371,7 @@ jobs:
|
|
|
371
371
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
372
372
|
fi
|
|
373
373
|
- if: steps.guard.outputs.live == 'true'
|
|
374
|
-
run: npm i -g "cowork-harness@^2.
|
|
374
|
+
run: npm i -g "cowork-harness@^2.5.0"
|
|
375
375
|
- if: steps.guard.outputs.live == 'true'
|
|
376
376
|
run: cowork-harness run scenarios/ --output-format json
|
|
377
377
|
env:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Critique — the facts a plugin install can't otherwise reach
|
|
2
2
|
|
|
3
|
-
Tracks `cowork-harness 2.
|
|
3
|
+
Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`). This is **not** a trim of the full
|
|
4
4
|
[`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
|
|
5
5
|
flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
|
|
6
6
|
plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
|
|
@@ -34,6 +34,63 @@ computed over raw rows** — that exclusion also governs `stats --group-by skill
|
|
|
34
34
|
**total spend** still needs the raw rows (only the roll-up carries the evaluator passes); and a roll-up with `result:"error"` had an unpriced workload, so its totals
|
|
35
35
|
UNDERCOUNT. The index is the only cost record that survives run-dir pruning.
|
|
36
36
|
|
|
37
|
+
## An exit-2 report — which turn failed, and whether it was really infrastructure
|
|
38
|
+
|
|
39
|
+
Exit 2 means no findings were produced, so the report's only job is to say what went wrong. Three fields
|
|
40
|
+
carry that; read all three before touching anything.
|
|
41
|
+
|
|
42
|
+
| Field | Meaning |
|
|
43
|
+
|---|---|
|
|
44
|
+
| `infraFailure` | the reason |
|
|
45
|
+
| `infraFailurePhase` | `task turn` (the graded run) or `reflection turn` (critique's own protocol turn) |
|
|
46
|
+
| `infraFailureKind` | why it failed — a harness `ErrCategory` (error envelope, exit 2/3) **or** a `resultErrorKind` (`usage_limit`/`transport`/`agent`) from a turn that RAN and errored (exit 1, top-level `error: null`). **Absent** = killed, or no envelope |
|
|
47
|
+
| `gradedErrorReason` | on a `taskResult: "error"` run (still gradeable, exit 0): why the GRADED turn errored, so a quota exhaustion is not read as a skill defect |
|
|
48
|
+
|
|
49
|
+
**Do NOT read "has a kind" as "the instrument is fine".** The CLI's top-level catch turns every
|
|
50
|
+
unexpected throw into category **`internal`** — Docker down, container start failure, missing staged
|
|
51
|
+
agent, harness bug — and `runtime` carries a refused run dir. Only **`unanswered`, `usage`, `boundary`**
|
|
52
|
+
and, from the result-row taxonomy, **`usage_limit`** (quota exhausted — retry after reset) and
|
|
53
|
+
**`transport`** (a tail-end drop) are the caller's problem. `agent` is not: for critique's own protocol
|
|
54
|
+
turn that IS the instrument breaking. The header encodes exactly that split and fails closed (an
|
|
55
|
+
unrecognized kind renders as infrastructure):
|
|
56
|
+
|
|
57
|
+
- `RUN FAILED (<turn>, <kind>): …` → ordinary, actionable, instrument healthy.
|
|
58
|
+
- `INFRASTRUCTURE/PROTOCOL FAILURE (<turn>): …` → `internal`/`runtime`/`agent`/unknown/killed/no-envelope.
|
|
59
|
+
|
|
60
|
+
**A turn that exits 1 is not a crash.** It RAN and reported an errored result, with `error: null` and a
|
|
61
|
+
full `results[0]` — an exit-code-only reading of that path is what leaves an exhausted quota looking like
|
|
62
|
+
a broken instrument.
|
|
63
|
+
|
|
64
|
+
**Read the reason, not the category.** The reason carries the failed turn's own message *and* hint
|
|
65
|
+
verbatim. That matters most for `unanswered`, which is 36 distinct throw sites and only ONE of them is
|
|
66
|
+
"the skill asked an unscripted question" — the others are a mis-typed `--answer` label, malformed
|
|
67
|
+
`--answer-policy` YAML, a crashed or bad-JSON `--decider-cmd` helper, an out-of-set `--decider-llm`
|
|
68
|
+
reply, an unanswered dialog/elicit, even a self-declared harness bug. A remedy picked from the category
|
|
69
|
+
is wrong for nearly all of them; each site's own hint is written for its case. (Also note `--on-unanswered`
|
|
70
|
+
*conflicts* with `--decider-dir`/`--decider-cmd`, so it is not a blanket fallback.)
|
|
71
|
+
|
|
72
|
+
For the genuine unscripted-gate case: script it (`--answer`, `--answer-policy`), or, when the skill's
|
|
73
|
+
gates are LLM-authored and reworded every run so a literal regex will not match twice, use `--decider-llm`
|
|
74
|
+
(or the scenario's `on_unanswered: llm`).
|
|
75
|
+
|
|
76
|
+
## Which model was graded — `gradedModels`
|
|
77
|
+
|
|
78
|
+
**The two turns are a SUBPROCESS.** They inherit no model from whatever invoked `critique` — not your
|
|
79
|
+
session, not a project setting. With no `--model`, the graded run uses the spawned agent's own default,
|
|
80
|
+
which may not be the model you are otherwise working under, and nothing about the run announces it.
|
|
81
|
+
|
|
82
|
+
`gradedModels` (text header: `graded model(s):`) is read back from the graded turn's own `result.json` and
|
|
83
|
+
is the only record of which model produced the behaviour being graded — distinct from the evaluator's
|
|
84
|
+
resolved model, which is a **different workload with its own default** (`claude-opus-4-8`), reported
|
|
85
|
+
separately. An evaluator line naming a model you did not pass is therefore expected, not evidence your
|
|
86
|
+
`--model` was ignored. Pin with `--model <id>` whenever a critique will be compared against another, and
|
|
87
|
+
read `gradedModels` back to confirm it took.
|
|
88
|
+
|
|
89
|
+
It is **observed, not requested** — the ids come from the model stamped on the graded turn's assistant
|
|
90
|
+
messages, never from the flag. So `graded model(s): unknown` means no assistant message reached the run
|
|
91
|
+
(crash, kill, or a gate before the first reply); passing `--model` does not change that line. Past runs
|
|
92
|
+
can be checked without re-running: the same ids are in each kept run dir's `turns/1/result.json`.
|
|
93
|
+
|
|
37
94
|
## The report's item shape — no `title`, no `summary`
|
|
38
95
|
|
|
39
96
|
Each `items[]` entry's prose fields are **`idea`** and **`recommendedAction`** — there is no `title` field
|
|
@@ -80,6 +137,20 @@ false `already-covered` verdict. It is named in `corpusExcluded` instead, and an
|
|
|
80
137
|
specifically reports `skillMdStatus: "untracked"`, forcing the mechanical `already-covered` →
|
|
81
138
|
`not-adjudicable` downgrade. `git add` it (or commit before critiquing) if it should count as evidence.
|
|
82
139
|
|
|
140
|
+
## Read `referencesAccessed`, not `referencesRead`
|
|
141
|
+
|
|
142
|
+
`referencesRead` counts the **`Read` tool only**. An agent that reaches a reference with a `Bash cat`, a
|
|
143
|
+
`Grep` or a `Glob` leaves nothing in it, so its emptiness is **not** evidence the content went unread.
|
|
144
|
+
`referencesAccessed` is the wide signal — every file reached, with the channel each was reached through
|
|
145
|
+
(`read` / `grep` / `bash`) — and it is what the critique headline is computed from.
|
|
146
|
+
|
|
147
|
+
Two properties to carry: only the `read` channel is strong evidence the agent opened the file (a `bash`
|
|
148
|
+
entry means a command named the path); and detection **under-approximates** — a `cd` into the skill dir
|
|
149
|
+
then a bare relative `cat`, a heredoc body and a `$VAR`-built path are all invisible. So an absent path is
|
|
150
|
+
weak evidence, never proof. **Presence is the cannot-verify channel:** `[]` means the drive ran and saw
|
|
151
|
+
nothing (a real negative); an ABSENT field means there was no observable drive, and must never be read as
|
|
152
|
+
"none".
|
|
153
|
+
|
|
83
154
|
## `referencesRead` is main-agent-only — `noSkillFilesRead` is not
|
|
84
155
|
|
|
85
156
|
`result.json`'s top-level `referencesRead` lists **main-agent Reads only**. A dispatcher-style skill does
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 2.5.0`
|
|
4
4
|
(baseline `desktop-1.37937.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
|
|
5
5
|
[`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
|
|
6
6
|
|
|
@@ -305,6 +305,10 @@ same set live from the schema.
|
|
|
305
305
|
| `no_lost_write_back: true` | fails if the run authored an interactive HTML artifact (or a `.py`/`.js` generator of one) whose **relative** Submit/POST write-back is lost under Cowork (served from Cowork's own origin → resolves non-ok, a "Saved!" is silently false). Runs the shipped **static Tier A** analyzer over the files the run authored (diffed vs the pre-run manifest). A lost write-back on an **added** agent-authored source (`outputs/`, scratchpad) **fails**; a **pre-existing** file the skill only modified on a read-write mount is **advisory**; `-suspect` findings surface but pass. **Only `true` is valid**. **Live/verify-run only** — skipped-loud on replay; runs on every live sandbox tier including **microvm** (its outputs are snapshotted from the VM into the run dir); could-not-verify (fail-closed) on a `--resume` scratchpad or an unanalyzable candidate |
|
|
306
306
|
| `tool_called: <glob>` | a tool the agent ran matched this **glob** — `*` = any run, `?` = one char, exact when literal, anchored + case-sensitive. Exact name (`Write`) matches only that tool; `mcp__workspace__*` matches any workspace tool. GLOB, not regex (`.` is literal) — an empty glob, or one containing a regex/brace-expansion metacharacter (`.*`, `.+`, `\|`, `()`, `[]`, `+`, `^`, `$`, `{}`, `\d`/`\w`/`\s`/`\b`), is **rejected at load** (a hard schema error, not a runtime warning) — it would match no real tool name and pass a `_not_`/`_absent` assert vacuously. Applies whether the glob comes from an authored scenario or a recorded cassette's frozen assert. The bundled `scenario.py lint` does NOT perform this check — only the harness enforces it, at actual load (`run`/`skill`/`record`) |
|
|
307
307
|
| `tool_not_called: <glob>` | NO tool the agent ran matched this glob (`mcp__*` = "no MCP tool ran"). Same glob semantics as `tool_called`, including the empty/regex-ish rejection |
|
|
308
|
+
| **legacy tool names** | The agent binary canonicalizes legacy spellings (`Task`→`Agent`, `KillShell`/`KillBash`→`TaskStop`, …) while the spawn tool list still declares the LEGACY one — so the init inventory shows `Task` and every call is emitted as `Agent`. `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` match EITHER spelling, globs included. Before this, `tool_called: "Task"` could never pass and `tool_not_called: "Task"` always did. |
|
|
309
|
+
| **tier vacuity — now REFUSED** | `tool_not_called` **and `subagent_tool_absent`** naming a tool the tier does not serve is rejected at scenario load (`Bash`/`WebFetch`/`NotebookEdit` at `hostloop`; `mcp__workspace__bash` at `container`/`microvm`), because it could never be violated. `scenario.py lint` WARNs on the same set before any run. The table is CLOSED to tools the harness itself removes or registers — `--tools` gates only the built-in set while every tier separately passes `--mcp-config`, so a session-MCP tool name is offered without appearing in any tool list and is never refused. Globs are never refused. `subagent_tool_absent` is covered because it reads the tools sub-agents actually USED, not a per-dispatch declared list. Under-approximates on purpose: `REPL` at hostloop is vacuous too and is not caught, which is the right side to err on for a hard refusal. |
|
|
310
|
+
| `reference_read: <regex>` | a skill `references/`/`scripts/` file whose path matches this **regex** was ACCESSED — main agent or sub-agents, via `Read`, `Grep`/`Glob` (`path` input), or a `Bash`/`mcp__workspace__bash` command naming the path. Regex is **unanchored + case-insensitive** (the shared helper every regex key uses). **Under-approximates by design:** the path must be rooted in the mounted plugin, so a `cd` into the skill dir then a bare `cat references/x.md`, a heredoc body, and a `$VAR`-built path are invisible. Fails **evidence unavailable** when the run recorded no observable tool stream. Replay-capable (cassettes freeze whole tool inputs) |
|
|
311
|
+
| `no_observed_reference_access: <regex>` | no OBSERVED access matched the regex — the progressive-disclosure check: a reference the skill's routing never reaches. Named `observed` because detection under-approximates (see above), so it is **not proof the file went unread** — an agent that `cd`s and `cat`s it passes. Fails **evidence unavailable** rather than passing vacuously when no observable tool stream was recorded |
|
|
308
312
|
| `tool_result_contains: <str>` | a tool result includes the literal string (content / replay-checkable — substring match) |
|
|
309
313
|
| `tool_result_not_contains: <str>` | no tool result includes the literal string (content / replay-checkable; fails loud when tool results are absent) |
|
|
310
314
|
| `tool_result_matches: <regex>` | the regex sibling of `tool_result_contains` — case-insensitive regex matches at least one tool result; use for an error-signature family, not just one literal string |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way.
|
|
5
|
-
Tracks `cowork-harness 2.
|
|
5
|
+
Tracks `cowork-harness 2.5.0` (baseline `desktop-1.37937.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -35,6 +35,7 @@
|
|
|
35
35
|
"no_hook_blocked",
|
|
36
36
|
"no_lost_write_back",
|
|
37
37
|
"no_mcp_error",
|
|
38
|
+
"no_observed_reference_access",
|
|
38
39
|
"no_path_denied",
|
|
39
40
|
"no_scratchpad_leak",
|
|
40
41
|
"no_skill_triggered",
|
|
@@ -46,6 +47,7 @@
|
|
|
46
47
|
"question_context",
|
|
47
48
|
"question_options",
|
|
48
49
|
"questions_count_max",
|
|
50
|
+
"reference_read",
|
|
49
51
|
"replay_protocol_fidelity",
|
|
50
52
|
"result",
|
|
51
53
|
"self_heal_ran",
|
|
@@ -80,6 +80,10 @@ CONTENT_KEYS = {
|
|
|
80
80
|
"tool_result_not_matches",
|
|
81
81
|
"tool_called",
|
|
82
82
|
"tool_not_called",
|
|
83
|
+
# Replay re-derives these from the SAME frozen tool inputs the live run used (a cassette stores whole
|
|
84
|
+
# tool inputs), so they are content keys, not live-only.
|
|
85
|
+
"reference_read",
|
|
86
|
+
"no_observed_reference_access",
|
|
83
87
|
"subagent_tool_used",
|
|
84
88
|
"subagent_tool_absent",
|
|
85
89
|
"subagent_dispatched",
|
|
@@ -294,6 +298,8 @@ REGEX_KEYS = {
|
|
|
294
298
|
"hook_blocked",
|
|
295
299
|
"tool_result_matches",
|
|
296
300
|
"tool_result_not_matches",
|
|
301
|
+
"reference_read",
|
|
302
|
+
"no_observed_reference_access",
|
|
297
303
|
}
|
|
298
304
|
VALID_ON_UNANSWERED = {"fail", "prompt", "first", "llm"}
|
|
299
305
|
VALID_TIERS = ("protocol", "container", "microvm", "hostloop", "cowork")
|
|
@@ -488,6 +494,30 @@ def lint_doc(doc, path, raw_lines):
|
|
|
488
494
|
|
|
489
495
|
fidelity = (doc.get("fidelity") or "container")
|
|
490
496
|
lane = (doc.get("lane") or "local")
|
|
497
|
+
|
|
498
|
+
# W: no `fidelity:` — the default models the WRONG LANE.
|
|
499
|
+
# `container` (the schema default) models VM-loop; production runs HOST-LOOP, gate 1143815894 is
|
|
500
|
+
# force-ON in every shipped baseline. So an omitted key measures the scenario against a lane real
|
|
501
|
+
# users are not on: the file tools resolve a bare relative path differently, the shell starts
|
|
502
|
+
# somewhere else, and the offered tool set differs (measured 2026-08-27).
|
|
503
|
+
# Read the KEY, not the resolved value: `fidelity: container` is a deliberate choice and must not warn.
|
|
504
|
+
# DEPRECATION — `fidelity:` becomes REQUIRED in the next major; this is the warning window.
|
|
505
|
+
if "fidelity" not in doc:
|
|
506
|
+
findings.append(
|
|
507
|
+
Finding(
|
|
508
|
+
"WARN",
|
|
509
|
+
"fidelity-defaulted",
|
|
510
|
+
"no `fidelity:` — defaulting to `container`, which models the VM-LOOP lane. Production "
|
|
511
|
+
"runs HOST-LOOP by default (gate 1143815894), so this scenario is likely measured "
|
|
512
|
+
"against a lane your users are not on.",
|
|
513
|
+
"Name a tier: `fidelity: hostloop` to match production, `fidelity: cowork` to auto-pick "
|
|
514
|
+
"the way Cowork does, or `fidelity: container` to keep today's behaviour deliberately. "
|
|
515
|
+
"Switching tiers can COST you assertions: `no_scratchpad_leak` is container-only (an "
|
|
516
|
+
"error elsewhere) and `transcript_no_host_path` fails by design at hostloop/protocol. "
|
|
517
|
+
"The default is being removed — `fidelity:` becomes REQUIRED in the next major.",
|
|
518
|
+
path,
|
|
519
|
+
)
|
|
520
|
+
)
|
|
491
521
|
items = _assert_items(doc)
|
|
492
522
|
assert_keys = _all_assert_keys(items)
|
|
493
523
|
has_expect_denied = bool(doc.get("expect_denied"))
|
|
@@ -559,6 +589,44 @@ def lint_doc(doc, path, raw_lines):
|
|
|
559
589
|
# warns at run start, after authoring). Lint is deliberately STRICTER than the runtime: the docs
|
|
560
590
|
# declare the combination incompatible, so authoring it is a bug even if a tool-free run could
|
|
561
591
|
# accidentally pass. `cowork` gets a WARN naming the baseline-gate resolution dependency (the
|
|
592
|
+
# `tool_not_called` naming a tool the TIER does not serve can never be violated — it passes
|
|
593
|
+
# vacuously and verifies nothing. Expressible offline because the mapping is a harness constant
|
|
594
|
+
# (WORKSPACE_TOOL_ALIASES / VM_LOOP_TOOL_ALIASES), NOT a baseline read. Deliberately literals only,
|
|
595
|
+
# and deliberately a closed table: `--tools` gates the BUILT-IN set alone while every tier separately
|
|
596
|
+
# passes --mcp-config, so a session-MCP tool name is offered without appearing in any tool list and
|
|
597
|
+
# must never be flagged here. The harness refuses these at load; this catches them before any run.
|
|
598
|
+
_TIER_VACUOUS = {
|
|
599
|
+
"hostloop": {"Bash": "mcp__workspace__bash", "WebFetch": "mcp__workspace__web_fetch", "NotebookEdit": None},
|
|
600
|
+
"container": {"mcp__workspace__bash": "Bash"},
|
|
601
|
+
"microvm": {"mcp__workspace__bash": "Bash", "mcp__workspace__web_fetch": "WebFetch"},
|
|
602
|
+
# `protocol` absent on purpose: it passes no tool flags, so its surface is the operator's own
|
|
603
|
+
# host CLI registry — machine-dependent and about a different product.
|
|
604
|
+
}
|
|
605
|
+
# BOTH negative tool keys: `subagent_tool_absent` is judged against the tools sub-agents actually
|
|
606
|
+
# USED, not a per-dispatch declared list, so a tool the tier never serves makes it equally vacuous.
|
|
607
|
+
for _key in ("tool_not_called", "subagent_tool_absent"):
|
|
608
|
+
for _v in _assert_values(items, _key):
|
|
609
|
+
if not isinstance(_v, str) or "*" in _v or "?" in _v:
|
|
610
|
+
continue # a glob is not a literal claim about one tool
|
|
611
|
+
_repl = _TIER_VACUOUS.get(fidelity, {})
|
|
612
|
+
if _v not in _repl:
|
|
613
|
+
continue
|
|
614
|
+
_instead = _repl[_v]
|
|
615
|
+
findings.append(
|
|
616
|
+
Finding(
|
|
617
|
+
"WARN",
|
|
618
|
+
"tool-not-called-tier-vacuous",
|
|
619
|
+
f"`{_key}: {_v}` on `fidelity: {fidelity}` — that tier does not serve "
|
|
620
|
+
f"`{_v}` at all, so this can never be violated and verifies nothing.",
|
|
621
|
+
(
|
|
622
|
+
f"Assert `{_key}: {_instead}` instead — that is what the tier serves in its place."
|
|
623
|
+
if _instead
|
|
624
|
+
else f"The {fidelity} tier removes `{_v}` outright; drop this assertion."
|
|
625
|
+
),
|
|
626
|
+
path,
|
|
627
|
+
)
|
|
628
|
+
)
|
|
629
|
+
|
|
562
630
|
# linter stays offline — the message carries the gate fact instead of reading a baseline).
|
|
563
631
|
if "transcript_no_host_path" in assert_keys:
|
|
564
632
|
if fidelity in ("hostloop", "protocol"):
|
|
@@ -992,6 +1060,26 @@ def lint_doc(doc, path, raw_lines):
|
|
|
992
1060
|
)
|
|
993
1061
|
)
|
|
994
1062
|
|
|
1063
|
+
# Same shape for the reference-access pair: `reference_read: R` and `no_observed_reference_access: R`
|
|
1064
|
+
# with the IDENTICAL regex cannot both hold. Compared as raw pattern strings — two different regexes
|
|
1065
|
+
# that happen to match the same path are NOT a contradiction (the linter cannot know the paths), so
|
|
1066
|
+
# this only fires on the case that is unambiguously self-defeating.
|
|
1067
|
+
read_pats = {v for v in _assert_values(items, "reference_read") if isinstance(v, str)}
|
|
1068
|
+
unread_pats = {v for v in _assert_values(items, "no_observed_reference_access") if isinstance(v, str)}
|
|
1069
|
+
both_refs = sorted(read_pats & unread_pats)
|
|
1070
|
+
if both_refs:
|
|
1071
|
+
findings.append(
|
|
1072
|
+
Finding(
|
|
1073
|
+
"ERROR",
|
|
1074
|
+
"reference-access-contradiction",
|
|
1075
|
+
f"assert requires {both_refs} to be both accessed (`reference_read`) and never observed "
|
|
1076
|
+
"(`no_observed_reference_access`) — no run can satisfy that, so this would spend a run to fail.",
|
|
1077
|
+
"Drop whichever half the scenario does not mean. To check that ONE reference is reached while "
|
|
1078
|
+
"another is not, give the two keys different patterns.",
|
|
1079
|
+
path,
|
|
1080
|
+
)
|
|
1081
|
+
)
|
|
1082
|
+
|
|
995
1083
|
# W: double-quoted regex with a backslash (raw-text scan — the parser already ate it)
|
|
996
1084
|
findings.extend(_lint_regex_quoting(path, raw_lines))
|
|
997
1085
|
|