cowork-harness 0.30.0 → 0.32.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +18 -12
- package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +342 -11
- package/CHANGELOG.md +323 -0
- package/README.md +18 -13
- package/SPEC.md +30 -10
- package/dist/assert.js +76 -1
- package/dist/baseline.js +55 -13
- package/dist/cli.js +260 -9
- package/dist/hostloop/canusetool-gate.js +2 -1
- package/dist/hostloop/pretooluse-path-hook.js +2 -1
- package/dist/run/analyze-skill.js +1093 -0
- package/dist/run/cassette.js +63 -12
- package/dist/run/chat-result.js +1 -0
- package/dist/run/doctor.js +12 -3
- package/dist/run/envelope.js +8 -4
- package/dist/run/execute.js +202 -25
- package/dist/run/latest-run.js +130 -0
- package/dist/run/probe-dispatch.js +81 -0
- package/dist/run/run.js +12 -2
- package/dist/run/scenario-tool.js +92 -10
- package/dist/run/subagent-reasoning.js +151 -0
- package/dist/run/verdict.js +23 -1
- package/dist/types.js +19 -0
- package/dist/vm-paths.js +13 -0
- package/docs/boundary.md +1 -1
- package/docs/cassette.md +21 -0
- package/docs/gotchas.md +5 -4
- package/docs/maintenance.md +1 -1
- package/docs/plugin-root.md +15 -1
- package/docs/scenario.md +22 -0
- package/docs/subagents.md +383 -0
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/run-result.json +122 -2
- package/schema/scenario.schema.json +29 -0
- package/schema/verify-cassettes.json +11 -6
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 0.
|
|
7
|
-
tracks-harness: cowork-harness 0.
|
|
6
|
+
version: 0.32.0
|
|
7
|
+
tracks-harness: cowork-harness 0.32.0 (baseline desktop-1.20186.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.
|
|
26
|
-
> `desktop-1.20186.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.32.0` (baseline
|
|
26
|
+
> `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -38,7 +38,7 @@ cowork-harness skill ./my-skill "do X" # run the skill once against the sta
|
|
|
38
38
|
Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.
|
|
41
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.32.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.32.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.32.0"`. **Pin `@>=0.32.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.32.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0), the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction (all new in 0.31.0), `analyze-skill`'s directory scan now covering a skill/plugin's full contract surface (recursive `agents/`/`references/`/`commands/`, plugin-root-aware, symlink-following) with line/block-scoped `analyze-skill: ignore-next-line`/`ignore-start`/`ignore-end` markers and multi-path/glob input, and `lint-skill`'s provable in-plugin `subagent_type` typo now a WARN that gates under `--strict` (new in 0.32.0).)
|
|
42
42
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
43
43
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
44
44
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred) or `ANTHROPIC_API_KEY`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -63,13 +63,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
63
63
|
the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
|
|
64
64
|
`cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
|
|
65
65
|
then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
|
|
66
|
-
the loop 0.
|
|
66
|
+
the loop 0.32.0's observability is built for; the *Triage* and *Inspecting a run's observability
|
|
67
67
|
output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
|
|
68
68
|
`docs/debugging.md` — repo-only, not shipped with the installed skill).
|
|
69
69
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
70
70
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
71
71
|
|
|
72
72
|
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
73
|
+
lint-skill · analyze-skill · probe-dispatch ·
|
|
73
74
|
verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
74
75
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
75
76
|
|
|
@@ -421,9 +422,10 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
421
422
|
the current list rather than relying on a fixed enumeration here.
|
|
422
423
|
- **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
|
|
423
424
|
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
|
|
424
|
-
- **`result.json` carries the raw fields** the assertions read: `toolDurations`, `models`, `toolErrors`,
|
|
425
|
+
- **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
|
|
425
426
|
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
426
|
-
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`
|
|
427
|
+
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
428
|
+
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
427
429
|
`workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
|
|
428
430
|
(`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
|
|
429
431
|
evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
|
|
@@ -431,8 +433,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
431
433
|
`bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
|
|
432
434
|
`egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
|
|
433
435
|
cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
|
|
434
|
-
`outputTruncated` on a matched tool result).
|
|
435
|
-
|
|
436
|
+
`outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
|
|
437
|
+
`jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
|
|
438
|
+
`{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
|
|
439
|
+
semantics: the README's "Observability fields" section — repo-only; `schema/run-result.json` is the
|
|
440
|
+
machine source.)
|
|
436
441
|
- **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
|
|
437
442
|
**`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
|
|
438
443
|
a re-record rarely tells you more than the captured stderr already does. Also check
|
|
@@ -674,8 +679,9 @@ assertion/replay-relevant ones).
|
|
|
674
679
|
- `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
|
|
675
680
|
- `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
|
|
676
681
|
GitHub Actions pipeline.
|
|
677
|
-
- `scripts/scenario.py` — `scaffold` a valid scenario skeleton
|
|
678
|
-
no-silent-false-green invariants (both usable as CI steps)
|
|
682
|
+
- `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
|
|
683
|
+
no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
|
|
684
|
+
(validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
|
|
679
685
|
- Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
|
|
680
686
|
alive* (Part II) above; the fuller recipe is in `docs/run-status.md` (repo-only, not shipped with the
|
|
681
687
|
installed skill).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -54,7 +54,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
54
54
|
GitHub-hosted runners, no token/Docker/agent:
|
|
55
55
|
|
|
56
56
|
```yaml
|
|
57
|
-
- run: npm i -g "cowork-harness@>=0.
|
|
57
|
+
- run: npm i -g "cowork-harness@>=0.32.0"
|
|
58
58
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
59
59
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
60
60
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -139,13 +139,19 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
139
139
|
universal net (container-tier recordings can trip it too).
|
|
140
140
|
- **Always-on scan gate** — `verify-cassettes` flags email / currency / bare-domain / local-path /
|
|
141
141
|
machine-inventory matches it finds in the committed cassettes and **exits non-zero**, so "no leak" is
|
|
142
|
-
a gate, not discipline.
|
|
142
|
+
a gate, not discipline. Non-zero is not one thing, though: exit `1` means verification RAN and found a
|
|
143
|
+
real finding (a PII match, a genuine staleness drift, or scenario-prompt drift); exit `3` means
|
|
144
|
+
verification could NOT complete (an `unverifiable-*`-class staleness finding, a cassette written by a
|
|
145
|
+
newer harness than this one understands, or a malformed/unreadable cassette). A plain `|| true` or `[
|
|
146
|
+
$? -ne 0 ]` tripwire treats both the same — if you need to tell "the gate caught something" apart from
|
|
147
|
+
"the gate couldn't run", branch on the exit code (or parse `--output-format json`'s per-file
|
|
148
|
+
`findings`/`staleness` vs `unverifiable`/`version`/`error` buckets).
|
|
143
149
|
Suppress synthetic / public reference names (NVCA, Cooley GO, …) with `--allow <regex>`. (Multi-word
|
|
144
150
|
proper names are NOT a default class — too noisy to gate on; add a pattern via config if your corpus
|
|
145
151
|
needs it.)
|
|
146
152
|
|
|
147
153
|
```bash
|
|
148
|
-
cowork-harness verify-cassettes cassettes/ # privacy scan + staleness
|
|
154
|
+
cowork-harness verify-cassettes cassettes/ # privacy scan + staleness — exit 1 = verified & failed, exit 3 = could not verify
|
|
149
155
|
cowork-harness verify-cassettes cassettes/ --allow 'NVCA|Cooley GO|Acme'
|
|
150
156
|
cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness only (skip the privacy scan); both run by default
|
|
151
157
|
```
|
|
@@ -188,7 +194,7 @@ jobs:
|
|
|
188
194
|
with: { node-version: '20' }
|
|
189
195
|
- uses: actions/setup-python@v5
|
|
190
196
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
191
|
-
- run: npm i -g "cowork-harness@>=0.
|
|
197
|
+
- run: npm i -g "cowork-harness@>=0.32.0"
|
|
192
198
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
193
199
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
194
200
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -217,7 +223,7 @@ jobs:
|
|
|
217
223
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
218
224
|
fi
|
|
219
225
|
- if: steps.guard.outputs.live == 'true'
|
|
220
|
-
run: npm i -g "cowork-harness@>=0.
|
|
226
|
+
run: npm i -g "cowork-harness@>=0.32.0"
|
|
221
227
|
- if: steps.guard.outputs.live == 'true'
|
|
222
228
|
run: cowork-harness run scenarios/ --output-format json
|
|
223
229
|
env:
|
|
@@ -245,9 +251,12 @@ fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-
|
|
|
245
251
|
JSON.
|
|
246
252
|
|
|
247
253
|
`verify-cassettes` emits its **own** envelope (`{command, ok, coverage, results[]}` with per-file
|
|
248
|
-
`findings`/`staleness`/`notes`/`version`/`error`), published as
|
|
249
|
-
|
|
250
|
-
the
|
|
254
|
+
`findings`/`staleness`/`unverifiable`/`notes`/`version`/`error`), published as
|
|
255
|
+
`schema/verify-cassettes.json` in the npm package. `ok:false` doesn't say *why* — read the buckets, or
|
|
256
|
+
the exit code (`1` = `findings`/`staleness`/`scenarioDrift` populated, a real problem verified & found;
|
|
257
|
+
`3` = `unverifiable`/`version`/`error` populated, verification could not complete; a real finding wins
|
|
258
|
+
`1` if both are present). Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the
|
|
259
|
+
JSON, not the human-readable text (which is explicitly NOT stable).
|
|
251
260
|
|
|
252
261
|
A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
|
|
253
262
|
set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
|
|
@@ -280,8 +289,11 @@ does **not** imply the recording is still valid. Each replay result carries `sta
|
|
|
280
289
|
|
|
281
290
|
(A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
|
|
282
291
|
check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
|
|
283
|
-
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above
|
|
284
|
-
|
|
292
|
+
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
|
|
293
|
+
gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
|
|
294
|
+
`format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
|
|
295
|
+
while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
|
|
296
|
+
fail it either way.)
|
|
285
297
|
|
|
286
298
|
To gate in CI, pick the severity you want:
|
|
287
299
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 0.32.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.
|
|
4
|
-
(baseline `desktop-1.20186.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.32.0`
|
|
4
|
+
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -262,6 +262,7 @@ same set live from the schema.
|
|
|
262
262
|
| `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
|
|
263
263
|
| `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
|
|
264
264
|
| `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
|
|
265
|
+
| `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
|
|
265
266
|
| `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
|
|
266
267
|
| `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
|
|
267
268
|
| `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
|
|
5
|
-
front-matter (currently 0.
|
|
5
|
+
front-matter (currently 0.32.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|