cowork-harness 0.30.0 → 0.32.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +18 -12
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -2
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  6. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
  7. package/.claude/skills/cowork-harness/scripts/scenario.py +342 -11
  8. package/CHANGELOG.md +323 -0
  9. package/README.md +18 -13
  10. package/SPEC.md +30 -10
  11. package/dist/assert.js +76 -1
  12. package/dist/baseline.js +55 -13
  13. package/dist/cli.js +260 -9
  14. package/dist/hostloop/canusetool-gate.js +2 -1
  15. package/dist/hostloop/pretooluse-path-hook.js +2 -1
  16. package/dist/run/analyze-skill.js +1093 -0
  17. package/dist/run/cassette.js +63 -12
  18. package/dist/run/chat-result.js +1 -0
  19. package/dist/run/doctor.js +12 -3
  20. package/dist/run/envelope.js +8 -4
  21. package/dist/run/execute.js +202 -25
  22. package/dist/run/latest-run.js +130 -0
  23. package/dist/run/probe-dispatch.js +81 -0
  24. package/dist/run/run.js +12 -2
  25. package/dist/run/scenario-tool.js +92 -10
  26. package/dist/run/subagent-reasoning.js +151 -0
  27. package/dist/run/verdict.js +23 -1
  28. package/dist/types.js +19 -0
  29. package/dist/vm-paths.js +13 -0
  30. package/docs/boundary.md +1 -1
  31. package/docs/cassette.md +21 -0
  32. package/docs/gotchas.md +5 -4
  33. package/docs/maintenance.md +1 -1
  34. package/docs/plugin-root.md +15 -1
  35. package/docs/scenario.md +22 -0
  36. package/docs/subagents.md +383 -0
  37. package/examples/replays/README.md +1 -1
  38. package/package.json +1 -1
  39. package/schema/run-result.json +122 -2
  40. package/schema/scenario.schema.json +29 -0
  41. package/schema/verify-cassettes.json +11 -6
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 0.30.0
7
- tracks-harness: cowork-harness 0.30.0 (baseline desktop-1.20186.1)
6
+ version: 0.32.0
7
+ tracks-harness: cowork-harness 0.32.0 (baseline desktop-1.20186.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.30.0` (baseline
26
- > `desktop-1.20186.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.32.0` (baseline
26
+ > `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -38,7 +38,7 @@ cowork-harness skill ./my-skill "do X" # run the skill once against the sta
38
38
  Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.30.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.30.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.30.0"`. **Pin `@>=0.30.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.30.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0).)
41
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.32.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.32.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.32.0"`. **Pin `@>=0.32.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.32.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0), the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction (all new in 0.31.0), `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict` (new in 0.32.0).)
42
42
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
43
43
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
44
44
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred) or `ANTHROPIC_API_KEY`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -63,13 +63,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
63
63
  the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
64
64
  `cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
65
65
  then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
66
- the loop 0.30.0's observability is built for; the *Triage* and *Inspecting a run's observability
66
+ the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
67
67
  output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
68
68
  `docs/debugging.md` — repo-only, not shipped with the installed skill).
69
69
  - **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
70
70
  TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
71
71
 
72
72
  Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
73
+ lint-skill · analyze-skill · probe-dispatch ·
73
74
  verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
74
75
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
75
76
 
@@ -421,9 +422,10 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
421
422
  the current list rather than relying on a fixed enumeration here.
422
423
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
423
424
  `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
424
- - **`result.json` carries the raw fields** the assertions read: `toolDurations`, `models`, `toolErrors`,
425
+ - **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
425
426
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
426
- `resolvedModel`/output/`attributedSkillId`, `outputTruncated`), `context` (tools/mcpServers/availableSkills), `tasks`,
427
+ `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
428
+ `context` (tools/mcpServers/availableSkills), `tasks`,
427
429
  `workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
428
430
  (`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
429
431
  evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
@@ -431,8 +433,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
431
433
  `bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
432
434
  `egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
433
435
  cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
434
- `outputTruncated` on a matched tool result). (Full per-field semantics: the README's "Observability
435
- fields" section — repo-only; `schema/run-result.json` is the machine source.)
436
+ `outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
437
+ `jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
438
+ `{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
439
+ semantics: the README's "Observability fields" section — repo-only; `schema/run-result.json` is the
440
+ machine source.)
436
441
  - **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
437
442
  **`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
438
443
  a re-record rarely tells you more than the captured stderr already does. Also check
@@ -674,8 +679,9 @@ assertion/replay-relevant ones).
674
679
  - `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
675
680
  - `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
676
681
  GitHub Actions pipeline.
677
- - `scripts/scenario.py` — `scaffold` a valid scenario skeleton and `lint` scenarios for the
678
- no-silent-false-green invariants (both usable as CI steps).
682
+ - `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
683
+ no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
684
+ (validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
679
685
  - Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
680
686
  alive* (Part II) above; the fuller recipe is in `docs/run-status.md` (repo-only, not shipped with the
681
687
  installed skill).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.30.0` (baseline `desktop-1.20186.0`).
3
+ Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -54,7 +54,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
54
54
  GitHub-hosted runners, no token/Docker/agent:
55
55
 
56
56
  ```yaml
57
- - run: npm i -g "cowork-harness@>=0.30.0"
57
+ - run: npm i -g "cowork-harness@>=0.32.0"
58
58
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
59
59
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
60
60
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -139,13 +139,19 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
139
139
  universal net (container-tier recordings can trip it too).
140
140
  - **Always-on scan gate** — `verify-cassettes` flags email / currency / bare-domain / local-path /
141
141
  machine-inventory matches it finds in the committed cassettes and **exits non-zero**, so "no leak" is
142
- a gate, not discipline.
142
+ a gate, not discipline. Non-zero is not one thing, though: exit `1` means verification RAN and found a
143
+ real finding (a PII match, a genuine staleness drift, or scenario-prompt drift); exit `3` means
144
+ verification could NOT complete (an `unverifiable-*`-class staleness finding, a cassette written by a
145
+ newer harness than this one understands, or a malformed/unreadable cassette). A plain `|| true` or `[
146
+ $? -ne 0 ]` tripwire treats both the same — if you need to tell "the gate caught something" apart from
147
+ "the gate couldn't run", branch on the exit code (or parse `--output-format json`'s per-file
148
+ `findings`/`staleness` vs `unverifiable`/`version`/`error` buckets).
143
149
  Suppress synthetic / public reference names (NVCA, Cooley GO, …) with `--allow <regex>`. (Multi-word
144
150
  proper names are NOT a default class — too noisy to gate on; add a pattern via config if your corpus
145
151
  needs it.)
146
152
 
147
153
  ```bash
148
- cowork-harness verify-cassettes cassettes/ # privacy scan + staleness, exit 1 on a finding
154
+ cowork-harness verify-cassettes cassettes/ # privacy scan + staleness — exit 1 = verified & failed, exit 3 = could not verify
149
155
  cowork-harness verify-cassettes cassettes/ --allow 'NVCA|Cooley GO|Acme'
150
156
  cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness only (skip the privacy scan); both run by default
151
157
  ```
@@ -188,7 +194,7 @@ jobs:
188
194
  with: { node-version: '20' }
189
195
  - uses: actions/setup-python@v5
190
196
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
191
- - run: npm i -g "cowork-harness@>=0.30.0"
197
+ - run: npm i -g "cowork-harness@>=0.32.0"
192
198
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
193
199
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
194
200
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -217,7 +223,7 @@ jobs:
217
223
  echo "live=true" >> "$GITHUB_OUTPUT"
218
224
  fi
219
225
  - if: steps.guard.outputs.live == 'true'
220
- run: npm i -g "cowork-harness@>=0.30.0"
226
+ run: npm i -g "cowork-harness@>=0.32.0"
221
227
  - if: steps.guard.outputs.live == 'true'
222
228
  run: cowork-harness run scenarios/ --output-format json
223
229
  env:
@@ -245,9 +251,12 @@ fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-
245
251
  JSON.
246
252
 
247
253
  `verify-cassettes` emits its **own** envelope (`{command, ok, coverage, results[]}` with per-file
248
- `findings`/`staleness`/`notes`/`version`/`error`), published as `schema/verify-cassettes.json` in the
249
- npm package. Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the JSON, not
250
- the human-readable text (which is explicitly NOT stable).
254
+ `findings`/`staleness`/`unverifiable`/`notes`/`version`/`error`), published as
255
+ `schema/verify-cassettes.json` in the npm package. `ok:false` doesn't say *why* — read the buckets, or
256
+ the exit code (`1` = `findings`/`staleness`/`scenarioDrift` populated, a real problem verified & found;
257
+ `3` = `unverifiable`/`version`/`error` populated, verification could not complete; a real finding wins
258
+ `1` if both are present). Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the
259
+ JSON, not the human-readable text (which is explicitly NOT stable).
251
260
 
252
261
  A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
253
262
  set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
@@ -280,8 +289,11 @@ does **not** imply the recording is still valid. Each replay result carries `sta
280
289
 
281
290
  (A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
282
291
  check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
283
- `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above is a hard fail —
284
- the gate is class-blind; notes never fail it.)
292
+ `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
293
+ gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
294
+ `format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
295
+ while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
296
+ fail it either way.)
285
297
 
286
298
  To gate in CI, pick the severity you want:
287
299
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.30.0` (baseline `desktop-1.20186.0`).
3
+ Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.30.0`
4
- (baseline `desktop-1.20186.0`). If your checkout is newer, prefer the live `docs/scenario.md`,
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.32.0`
4
+ (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -262,6 +262,7 @@ same set live from the schema.
262
262
  | `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
263
263
  | `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
264
264
  | `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
265
+ | `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
265
266
  | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
266
267
  | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
267
268
  | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
5
- front-matter (currently 0.29.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ front-matter (currently 0.32.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -45,6 +45,7 @@
45
45
  "skill_tool_used",
46
46
  "skill_triggered",
47
47
  "subagent_declared_but_unused",
48
+ "subagent_dispatch_healthy",
48
49
  "subagent_dispatched",
49
50
  "subagent_file_write",
50
51
  "subagent_output_contains",