cowork-harness 0.30.0 → 0.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +18 -12
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
  3. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
  4. package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -2
  5. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  6. package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
  7. package/.claude/skills/cowork-harness/scripts/scenario.py +328 -10
  8. package/CHANGELOG.md +224 -0
  9. package/README.md +15 -10
  10. package/SPEC.md +30 -10
  11. package/dist/assert.js +76 -1
  12. package/dist/baseline.js +55 -13
  13. package/dist/cli.js +251 -9
  14. package/dist/hostloop/canusetool-gate.js +2 -1
  15. package/dist/hostloop/pretooluse-path-hook.js +2 -1
  16. package/dist/run/analyze-skill.js +619 -0
  17. package/dist/run/cassette.js +63 -12
  18. package/dist/run/chat-result.js +1 -0
  19. package/dist/run/doctor.js +12 -3
  20. package/dist/run/envelope.js +8 -4
  21. package/dist/run/execute.js +202 -25
  22. package/dist/run/latest-run.js +130 -0
  23. package/dist/run/probe-dispatch.js +81 -0
  24. package/dist/run/run.js +12 -2
  25. package/dist/run/scenario-tool.js +17 -0
  26. package/dist/run/subagent-reasoning.js +151 -0
  27. package/dist/run/verdict.js +23 -1
  28. package/dist/types.js +19 -0
  29. package/dist/vm-paths.js +13 -0
  30. package/docs/boundary.md +1 -1
  31. package/docs/cassette.md +21 -0
  32. package/docs/gotchas.md +5 -4
  33. package/docs/maintenance.md +1 -1
  34. package/docs/plugin-root.md +11 -1
  35. package/docs/scenario.md +22 -0
  36. package/docs/subagents.md +274 -0
  37. package/examples/replays/README.md +1 -1
  38. package/package.json +1 -1
  39. package/schema/run-result.json +122 -2
  40. package/schema/scenario.schema.json +29 -0
  41. package/schema/verify-cassettes.json +11 -6
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 0.30.0
7
- tracks-harness: cowork-harness 0.30.0 (baseline desktop-1.20186.1)
6
+ version: 0.31.0
7
+ tracks-harness: cowork-harness 0.31.0 (baseline desktop-1.20186.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
22
22
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
23
23
  the highest-value part. Read it.
24
24
 
25
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.30.0` (baseline
26
- > `desktop-1.20186.0`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
25
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.31.0` (baseline
26
+ > `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
27
27
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
28
28
 
29
29
  ## Preflight — make sure the harness can actually run
@@ -38,7 +38,7 @@ cowork-harness skill ./my-skill "do X" # run the skill once against the sta
38
38
  Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
39
39
 
40
40
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
41
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.30.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.30.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.30.0"`. **Pin `@>=0.30.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.30.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0).)
41
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.31.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.31.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.31.0"`. **Pin `@>=0.31.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.31.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0), the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction (all new in 0.31.0).)
42
42
  - **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
43
43
  - **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
44
44
  - **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred) or `ANTHROPIC_API_KEY`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
@@ -63,13 +63,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
63
63
  the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
64
64
  `cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
65
65
  then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
66
- the loop 0.30.0's observability is built for; the *Triage* and *Inspecting a run's observability
66
+ the loop 0.31.0's observability is built for; the *Triage* and *Inspecting a run's observability
67
67
  output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
68
68
  `docs/debugging.md` — repo-only, not shipped with the installed skill).
69
69
  - **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
70
70
  TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
71
71
 
72
72
  Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
73
+ lint-skill · analyze-skill · probe-dispatch ·
73
74
  verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
74
75
  list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
75
76
 
@@ -421,9 +422,10 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
421
422
  the current list rather than relying on a fixed enumeration here.
422
423
  - **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
423
424
  `tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
424
- - **`result.json` carries the raw fields** the assertions read: `toolDurations`, `models`, `toolErrors`,
425
+ - **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
425
426
  `redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
426
- `resolvedModel`/output/`attributedSkillId`, `outputTruncated`), `context` (tools/mcpServers/availableSkills), `tasks`,
427
+ `resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
428
+ `context` (tools/mcpServers/availableSkills), `tasks`,
427
429
  `workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
428
430
  (`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
429
431
  evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
@@ -431,8 +433,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
431
433
  `bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
432
434
  `egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
433
435
  cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
434
- `outputTruncated` on a matched tool result). (Full per-field semantics: the README's "Observability
435
- fields" section — repo-only; `schema/run-result.json` is the machine source.)
436
+ `outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
437
+ `jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
438
+ `{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
439
+ semantics: the README's "Observability fields" section — repo-only; `schema/run-result.json` is the
440
+ machine source.)
436
441
  - **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
437
442
  **`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
438
443
  a re-record rarely tells you more than the captured stderr already does. Also check
@@ -674,8 +679,9 @@ assertion/replay-relevant ones).
674
679
  - `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
675
680
  - `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
676
681
  GitHub Actions pipeline.
677
- - `scripts/scenario.py` — `scaffold` a valid scenario skeleton and `lint` scenarios for the
678
- no-silent-false-green invariants (both usable as CI steps).
682
+ - `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
683
+ no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
684
+ (validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
679
685
  - Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
680
686
  alive* (Part II) above; the fuller recipe is in `docs/run-status.md` (repo-only, not shipped with the
681
687
  installed skill).
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.30.0` (baseline `desktop-1.20186.0`).
3
+ Self-contained reference. Tracks `cowork-harness 0.31.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -54,7 +54,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
54
54
  GitHub-hosted runners, no token/Docker/agent:
55
55
 
56
56
  ```yaml
57
- - run: npm i -g "cowork-harness@>=0.30.0"
57
+ - run: npm i -g "cowork-harness@>=0.31.0"
58
58
  - run: cowork-harness lint scenarios/*.yaml # no silent false-greens
59
59
  - run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
60
60
  - run: cowork-harness replay cassettes/ # token-free content/structure
@@ -139,13 +139,19 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
139
139
  universal net (container-tier recordings can trip it too).
140
140
  - **Always-on scan gate** — `verify-cassettes` flags email / currency / bare-domain / local-path /
141
141
  machine-inventory matches it finds in the committed cassettes and **exits non-zero**, so "no leak" is
142
- a gate, not discipline.
142
+ a gate, not discipline. Non-zero is not one thing, though: exit `1` means verification RAN and found a
143
+ real finding (a PII match, a genuine staleness drift, or scenario-prompt drift); exit `3` means
144
+ verification could NOT complete (an `unverifiable-*`-class staleness finding, a cassette written by a
145
+ newer harness than this one understands, or a malformed/unreadable cassette). A plain `|| true` or `[
146
+ $? -ne 0 ]` tripwire treats both the same — if you need to tell "the gate caught something" apart from
147
+ "the gate couldn't run", branch on the exit code (or parse `--output-format json`'s per-file
148
+ `findings`/`staleness` vs `unverifiable`/`version`/`error` buckets).
143
149
  Suppress synthetic / public reference names (NVCA, Cooley GO, …) with `--allow <regex>`. (Multi-word
144
150
  proper names are NOT a default class — too noisy to gate on; add a pattern via config if your corpus
145
151
  needs it.)
146
152
 
147
153
  ```bash
148
- cowork-harness verify-cassettes cassettes/ # privacy scan + staleness, exit 1 on a finding
154
+ cowork-harness verify-cassettes cassettes/ # privacy scan + staleness — exit 1 = verified & failed, exit 3 = could not verify
149
155
  cowork-harness verify-cassettes cassettes/ --allow 'NVCA|Cooley GO|Acme'
150
156
  cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness only (skip the privacy scan); both run by default
151
157
  ```
@@ -188,7 +194,7 @@ jobs:
188
194
  with: { node-version: '20' }
189
195
  - uses: actions/setup-python@v5
190
196
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
191
- - run: npm i -g "cowork-harness@>=0.30.0"
197
+ - run: npm i -g "cowork-harness@>=0.31.0"
192
198
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
193
199
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
194
200
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -217,7 +223,7 @@ jobs:
217
223
  echo "live=true" >> "$GITHUB_OUTPUT"
218
224
  fi
219
225
  - if: steps.guard.outputs.live == 'true'
220
- run: npm i -g "cowork-harness@>=0.30.0"
226
+ run: npm i -g "cowork-harness@>=0.31.0"
221
227
  - if: steps.guard.outputs.live == 'true'
222
228
  run: cowork-harness run scenarios/ --output-format json
223
229
  env:
@@ -245,9 +251,12 @@ fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-
245
251
  JSON.
246
252
 
247
253
  `verify-cassettes` emits its **own** envelope (`{command, ok, coverage, results[]}` with per-file
248
- `findings`/`staleness`/`notes`/`version`/`error`), published as `schema/verify-cassettes.json` in the
249
- npm package. Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the JSON, not
250
- the human-readable text (which is explicitly NOT stable).
254
+ `findings`/`staleness`/`unverifiable`/`notes`/`version`/`error`), published as
255
+ `schema/verify-cassettes.json` in the npm package. `ok:false` doesn't say *why* — read the buckets, or
256
+ the exit code (`1` = `findings`/`staleness`/`scenarioDrift` populated, a real problem verified & found;
257
+ `3` = `unverifiable`/`version`/`error` populated, verification could not complete; a real finding wins
258
+ `1` if both are present). Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the
259
+ JSON, not the human-readable text (which is explicitly NOT stable).
251
260
 
252
261
  A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
253
262
  set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
@@ -280,8 +289,11 @@ does **not** imply the recording is still valid. Each replay result carries `sta
280
289
 
281
290
  (A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
282
291
  check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
283
- `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above is a hard fail —
284
- the gate is class-blind; notes never fail it.)
292
+ `·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
293
+ gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
294
+ `format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
295
+ while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
296
+ fail it either way.)
285
297
 
286
298
  To gate in CI, pick the severity you want:
287
299
 
@@ -1,6 +1,6 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 0.30.0` (baseline `desktop-1.20186.0`).
3
+ Self-contained reference. Tracks `cowork-harness 0.31.0` (baseline `desktop-1.20186.1`).
4
4
 
5
5
  ## Fidelity tiers (`fidelity:` in the scenario)
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, full gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.30.0`
4
- (baseline `desktop-1.20186.0`). If your checkout is newer, prefer the live `docs/scenario.md`,
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.31.0`
4
+ (baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
5
5
  `docs/session.md`, and `SPEC.md`.
6
6
 
7
7
  **Minimal scenario** — `prompt` is the only required field:
@@ -262,6 +262,7 @@ same set live from the schema.
262
262
  | `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
263
263
  | `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
264
264
  | `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
265
+ | `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
265
266
  | `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
266
267
  | `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
267
268
  | `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
5
- front-matter (currently 0.29.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ front-matter (currently 0.31.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -45,6 +45,7 @@
45
45
  "skill_tool_used",
46
46
  "skill_triggered",
47
47
  "subagent_declared_but_unused",
48
+ "subagent_dispatch_healthy",
48
49
  "subagent_dispatched",
49
50
  "subagent_file_write",
50
51
  "subagent_output_contains",
@@ -39,6 +39,7 @@ system PyYAML is preferred when present.
39
39
  from __future__ import annotations
40
40
 
41
41
  import argparse
42
+ import functools
42
43
  import json
43
44
  import re
44
45
  import sys
@@ -66,6 +67,7 @@ CONTENT_KEYS = {
66
67
  "subagent_output_contains",
67
68
  "no_vm_path_file_op",
68
69
  "subagent_file_write",
70
+ "subagent_dispatch_healthy",
69
71
  "dispatch_count_max",
70
72
  "skill_triggered",
71
73
  "no_skill_triggered",
@@ -741,6 +743,53 @@ def _finding_plugin_root_guarded(path, line, ctx_label):
741
743
  )
742
744
 
743
745
 
746
+ # A self-heal `find`'s `-path '<glob>'` / `-path "<glob>"` value (same-quote char, non-greedy so a
747
+ # quote inside the glob — unlikely in practice — doesn't get swallowed).
748
+ _FIND_PATH_VALUE = re.compile(r"""-path\s+(['"])(.*?)\1""")
749
+ # The skill segment out of a `*/skills/<name>/...` glob.
750
+ _FIND_PATH_SKILLS_SEG = re.compile(r"/skills/([A-Za-z0-9_.-]+)/")
751
+ # The plugin segment out of a `*/plugins/<name>/...` glob.
752
+ _FIND_PATH_PLUGINS_SEG = re.compile(r"/plugins/([A-Za-z0-9_.-]+)/")
753
+ # A generic `*/<name>/scripts` glob (no `skills/`/`plugins/` literal prefix) — the segment right
754
+ # before a `/scripts` path component. This is what a plugin-level self-heal targeting its own
755
+ # `<plugin>/scripts/...` layout looks like once the real mount path (e.g. `mnt/.local-plugins/...`)
756
+ # is glob-abbreviated to `*/<plugin>/scripts`.
757
+ _FIND_PATH_SCRIPTS_SEG = re.compile(r"/([A-Za-z0-9_.-]+)/scripts\b")
758
+
759
+
760
+ def _extract_find_path_token(find_cmd_line):
761
+ """Pull the skill/plugin-naming token out of a self-heal `find`'s `-path` glob (on the SAME
762
+ line as the `find`, matching how `_SELF_HEAL` itself is line-scoped), or None if there's no
763
+ `-path` clause or none of the recognized glob shapes match. A bare glob wildcard segment (e.g.
764
+ `*/scripts/*`, no name between the slashes) intentionally does not match the `[A-Za-z0-9_.-]+`
765
+ character class — that's a real absence of an extractable token, not a token, so it stays
766
+ conservative (INFO) rather than manufacturing a bogus `*` token to compare."""
767
+ m = _FIND_PATH_VALUE.search(find_cmd_line)
768
+ if not m:
769
+ return None
770
+ glob = m.group(2)
771
+ for pat in (_FIND_PATH_SKILLS_SEG, _FIND_PATH_PLUGINS_SEG, _FIND_PATH_SCRIPTS_SEG):
772
+ seg = pat.search(glob)
773
+ if seg:
774
+ return seg.group(1)
775
+ return None
776
+
777
+
778
+ def _finding_guard_pattern_mismatch(path, line, ctx_label, token, skill_name, plugin_name):
779
+ plugin_part = f" (plugin `{plugin_name}`)" if plugin_name else ""
780
+ return Finding(
781
+ "WARN",
782
+ "guard-pattern-mismatch",
783
+ f"`${{CLAUDE_PLUGIN_ROOT}}` used in an in-VM bash context ({ctx_label}); the block's self-heal "
784
+ f"`find` targets `{token}`, but this skill is `{skill_name}`{plugin_part} — the guard won't "
785
+ "discover THIS skill's mount (likely a copy-pasted self-heal).",
786
+ f"Fix the `find` pattern's `-path` to match this skill/plugin's own layout (`{skill_name}`"
787
+ f"{plugin_part}), not `{token}`.",
788
+ path,
789
+ line,
790
+ )
791
+
792
+
744
793
  def _finding_hook_host_write(path, line, what):
745
794
  return Finding(
746
795
  "WARN",
@@ -778,15 +827,46 @@ def _lint_skill_text(path, raw_lines, force_json=False):
778
827
  bash_token_lines = []
779
828
  bash_block_text = []
780
829
 
830
+ # Self-heal find-pattern guard: identity of the skill/plugin under lint, used to check a self-heal `find -path`
831
+ # actually names THIS skill or its enclosing plugin (not a copy-pasted mismatch). Only
832
+ # meaningful when linting an actual SKILL.md (force_json=False is exactly that case here — a
833
+ # hooks.json body never enters the "bash" ctx below, so this is otherwise unused).
834
+ skill_name = None
835
+ plugin_name = None
836
+ self_plugin_tokens = set()
837
+ if not force_json:
838
+ dir_name = Path(path).resolve().parent.name
839
+ fm_name = _agent_name_from_frontmatter(path, _require_yaml())
840
+ # Prefer the frontmatter `name:` (the skill's declared identity) for display when present;
841
+ # the parent-dir name is always in the match set as a cross-check (both count as "this
842
+ # skill" — a self-heal naming either is not a mismatch).
843
+ skill_name = fm_name or dir_name
844
+ self_plugin_tokens.add(dir_name)
845
+ if fm_name:
846
+ self_plugin_tokens.add(fm_name)
847
+ plugin_dir = _find_enclosing_plugin_dir(path)
848
+ if plugin_dir is not None:
849
+ plugin_name = _read_plugin_name(plugin_dir)
850
+ if plugin_name:
851
+ self_plugin_tokens.add(plugin_name)
852
+ self_plugin_tokens.add(Path(plugin_dir).name)
853
+
781
854
  def flush_bash():
782
855
  if bash_token_lines:
783
- healed = _SELF_HEAL.search("\n".join(bash_block_text)) is not None
856
+ self_heal_line = next((bl for bl in bash_block_text if _SELF_HEAL.search(bl)), None)
857
+ healed = self_heal_line is not None
858
+ token = _extract_find_path_token(self_heal_line) if healed else None
784
859
  for ln in bash_token_lines:
785
- findings.append(
786
- _finding_plugin_root_guarded(path, ln, "```bash block")
787
- if healed
788
- else _finding_plugin_root(path, ln, "```bash block")
789
- )
860
+ if not healed:
861
+ findings.append(_finding_plugin_root(path, ln, "```bash block"))
862
+ elif token is not None and token not in self_plugin_tokens:
863
+ findings.append(
864
+ _finding_guard_pattern_mismatch(
865
+ path, ln, "```bash block", token, skill_name, plugin_name
866
+ )
867
+ )
868
+ else:
869
+ findings.append(_finding_plugin_root_guarded(path, ln, "```bash block"))
790
870
  bash_token_lines.clear()
791
871
  bash_block_text.clear()
792
872
 
@@ -838,6 +918,210 @@ def _lint_skill_text(path, raw_lines, force_json=False):
838
918
  return findings
839
919
 
840
920
 
921
+ # --------------------------------------------------------------------------- #
922
+ # subagent_type static resolution
923
+ # --------------------------------------------------------------------------- #
924
+ #
925
+ # A pinned `subagent_type:` value that doesn't resolve to a real agent fails a definition lookup at
926
+ # dispatch time — but that's only discoverable via a live dispatch today. Resolve it statically from
927
+ # a plugin's own `.claude-plugin/plugin.json` (or `plugin.json`) + `agents/*.md` frontmatter instead.
928
+ #
929
+ # HONEST LIMIT: there is no harness registry of built-in agent types (the built-in set is
930
+ # agent-binary-version-dependent) — only `general-purpose` is harness-known. So an unresolved bare
931
+ # value is surfaced as INFO, never failed as WARN/ERROR; the linter can't disprove it's a real
932
+ # built-in. Do NOT add a committed built-in agent-type list here — that would silently go stale and
933
+ # either false-warn a real built-in or false-clear a typo.
934
+
935
+ _SUBAGENT_TYPE_RE = re.compile(r"subagent_type\s*[:=]\s*['\"]?([A-Za-z0-9_.:/-]+)['\"]?")
936
+
937
+
938
+ def _read_plugin_name(plugin_dir):
939
+ """Return the `name` field from `<plugin_dir>/.claude-plugin/plugin.json` (fallback
940
+ `<plugin_dir>/plugin.json`), or None if neither file exists or is parsable. Never raises."""
941
+ p = Path(plugin_dir)
942
+ for candidate in (p / ".claude-plugin" / "plugin.json", p / "plugin.json"):
943
+ if candidate.is_file():
944
+ try:
945
+ data = json.loads(candidate.read_text(encoding="utf-8"))
946
+ except Exception:
947
+ return None
948
+ name = data.get("name") if isinstance(data, dict) else None
949
+ return name if isinstance(name, str) and name.strip() else None
950
+ return None
951
+
952
+
953
+ _AGENT_FRONTMATTER = re.compile(r"^---\s*\n(.*?\n)---\s*(?:\n|$)", re.DOTALL)
954
+
955
+
956
+ def _agent_name_from_frontmatter(md_path, yaml_mod):
957
+ """Return a markdown file's `name:` frontmatter value, or None if there's no frontmatter, no
958
+ `name:` field, or it fails to parse. Originally for `agents/*.md` (caller falls back to the
959
+ filename stem there); also reused by the self-heal find-pattern guard for a SKILL.md, whose frontmatter has the same
960
+ `---\\nname: ...\\n---` shape — the parser itself is generic, only the name is agent-specific."""
961
+ try:
962
+ text = Path(md_path).read_text(encoding="utf-8")
963
+ except Exception:
964
+ return None
965
+ m = _AGENT_FRONTMATTER.match(text)
966
+ if not m:
967
+ return None
968
+ try:
969
+ data = yaml_mod.safe_load(m.group(1))
970
+ except Exception:
971
+ return None
972
+ if isinstance(data, dict):
973
+ name = data.get("name")
974
+ if isinstance(name, str) and name.strip():
975
+ return name.strip()
976
+ return None
977
+
978
+
979
+ def _resolve_plugin_agents(plugin_dir):
980
+ """Resolve in-plugin agent types: return the set of valid `<plugin>:<agent>` subagent types defined WITHIN plugin_dir.
981
+ Reads the plugin name from plugin.json and each agents/*.md's `name:` frontmatter (filename stem
982
+ fallback). Returns an empty set (never crashes) when no plugin.json is found — a bare SKILL.md
983
+ dir with no plugin manifest has nothing to resolve against."""
984
+ plugin_name = _read_plugin_name(plugin_dir)
985
+ if not plugin_name:
986
+ return set()
987
+ agents_dir = Path(plugin_dir) / "agents"
988
+ if not agents_dir.is_dir():
989
+ return set()
990
+ yaml = _require_yaml()
991
+ types = set()
992
+ for md in sorted(agents_dir.glob("*.md")):
993
+ agent_name = _agent_name_from_frontmatter(md, yaml) or md.stem
994
+ types.add(f"{plugin_name}:{agent_name}")
995
+ return types
996
+
997
+
998
+ def cmd_resolve_agent_types(args):
999
+ types = sorted(_resolve_plugin_agents(args.plugin_dir))
1000
+ if args.json:
1001
+ print(json.dumps(types))
1002
+ else:
1003
+ for t in types:
1004
+ print(t)
1005
+ return 0
1006
+
1007
+
1008
+ def _find_enclosing_plugin_dir(skill_md_path):
1009
+ """Resolve the enclosing plugin: walk up from a SKILL.md to the nearest ancestor dir containing
1010
+ `.claude-plugin/plugin.json` or `plugin.json` — that's the enclosing plugin. None if no ancestor
1011
+ has one (a bare SKILL.md dir with no plugin manifest anywhere above it)."""
1012
+ start = Path(skill_md_path).resolve().parent
1013
+ for anc in [start, *start.parents]:
1014
+ if (anc / ".claude-plugin" / "plugin.json").is_file() or (anc / "plugin.json").is_file():
1015
+ return anc
1016
+ return None
1017
+
1018
+
1019
+ def _finding_subagent_unresolvable(path, line, value):
1020
+ return Finding(
1021
+ "INFO",
1022
+ "subagent-type-unresolvable",
1023
+ f"pinned type `{value}` belongs to another plugin — can't confirm it resolves from here",
1024
+ "Verify it resolves in that plugin's own agents/ dir (e.g. `scenario.py resolve-agent-types "
1025
+ "<that-plugin-dir>`), or dispatch without pinning a cross-plugin type.",
1026
+ path,
1027
+ line,
1028
+ )
1029
+
1030
+
1031
+ def _finding_subagent_unknown(path, line, value):
1032
+ return Finding(
1033
+ "INFO",
1034
+ "subagent-type-unknown",
1035
+ f"pinned type `{value}` is not defined in this plugin and is not the `general-purpose` "
1036
+ "built-in — can't confirm statically (may be an agent-binary built-in)",
1037
+ "If it's meant to be an in-plugin agent, add `agents/<name>.md` with a `name:` frontmatter "
1038
+ "matching the pinned value (or rely on the filename-stem fallback). If it's a real built-in "
1039
+ "agent type, this INFO is expected — the linter has no built-in registry to check it against.",
1040
+ path,
1041
+ line,
1042
+ )
1043
+
1044
+
1045
+ def _finding_subagent_not_found_in_plugin(path, line, value, plugin_name, agent, sorted_agents):
1046
+ """A pinned `<this-plugin>:<agent>` value whose prefix names the RESOLVED plugin (this SKILL.md's
1047
+ own enclosing plugin) but whose agent isn't in its enumerated agents/*.md set. Unlike
1048
+ `subagent-type-unknown`, this can never be another binary's built-in — the namespace prefix
1049
+ already commits it to this plugin, and the plugin's agent set was fully enumerable — so it's a
1050
+ provable typo, not an unconfirmable unknown."""
1051
+ return Finding(
1052
+ "INFO",
1053
+ "subagent-type-not-found-in-plugin",
1054
+ f"pinned type `{value}` names this plugin (`{plugin_name}`) but `{agent}` is not among its "
1055
+ f"agents [{', '.join(sorted_agents)}] — likely a typo.",
1056
+ "Fix the agent name to match one of the listed agents, or add `agents/<agent>.md` with a "
1057
+ "`name:` frontmatter matching the pinned value (or rely on the filename-stem fallback) if it "
1058
+ "was meant to exist. Check `scenario.py resolve-agent-types <plugin-dir>` to confirm.",
1059
+ path,
1060
+ line,
1061
+ )
1062
+
1063
+
1064
+ def _classify_subagent_type(value, plugin_name, plugin_agent_types):
1065
+ """subagent_type severity ladder. Returns a Finding-builder (path, line, value) -> Finding, or None if
1066
+ clean. Severity is ALWAYS INFO — see the HONEST LIMIT note above this section; never WARN/ERROR.
1067
+
1068
+ Precedence:
1069
+ 1. Resolves in-plugin, or is `general-purpose` -> clean (None).
1070
+ 2. Has a `<prefix>:<agent>` shape whose prefix EQUALS the resolved plugin's own name AND the
1071
+ plugin's agent set was non-empty (i.e. enumerable) -> `subagent-type-not-found-in-plugin`.
1072
+ A value namespaced under this plugin's own name can never be another binary's built-in, so
1073
+ once the plugin is fully enumerated a miss here is a provable typo, not an unconfirmable
1074
+ unknown.
1075
+ 3. Has a `<prefix>:<agent>` shape whose prefix does NOT equal the resolved plugin's name (or
1076
+ there's no resolved plugin at all) -> `subagent-type-unresolvable` (belongs to another
1077
+ plugin, or an empty plugin_name means we truly can't tell whose namespace it is).
1078
+ 4. Otherwise (no colon, or same-plugin prefix but the plugin set was empty/unenumerable) ->
1079
+ `subagent-type-unknown` — genuinely can't confirm statically.
1080
+ """
1081
+ if value == "general-purpose" or value in plugin_agent_types:
1082
+ return None
1083
+ if ":" in value:
1084
+ prefix, agent = value.split(":", 1)
1085
+ if plugin_name is not None and prefix == plugin_name:
1086
+ if plugin_agent_types:
1087
+ sorted_agents = sorted(
1088
+ t.split(":", 1)[1] for t in plugin_agent_types if t.split(":", 1)[0] == plugin_name
1089
+ )
1090
+ return functools.partial(
1091
+ _finding_subagent_not_found_in_plugin,
1092
+ plugin_name=plugin_name,
1093
+ agent=agent,
1094
+ sorted_agents=sorted_agents,
1095
+ )
1096
+ # Plugin set is empty — couldn't enumerate (e.g. no agents/ dir) — genuinely can't confirm.
1097
+ return _finding_subagent_unknown
1098
+ return _finding_subagent_unresolvable
1099
+ return _finding_subagent_unknown
1100
+
1101
+
1102
+ def _lint_subagent_types(path, raw_lines):
1103
+ """Scan a SKILL.md's raw text (not limited to fenced blocks — a pinned `subagent_type`
1104
+ can appear in prose or YAML frontmatter) for pinned `subagent_type` values and classify each
1105
+ against the enclosing plugin's in-plugin agent set."""
1106
+ matches = []
1107
+ for i, line in enumerate(raw_lines, start=1):
1108
+ for m in _SUBAGENT_TYPE_RE.finditer(line):
1109
+ matches.append((i, m.group(1)))
1110
+ if not matches:
1111
+ return []
1112
+
1113
+ plugin_dir = _find_enclosing_plugin_dir(path)
1114
+ plugin_name = _read_plugin_name(plugin_dir) if plugin_dir else None
1115
+ plugin_agent_types = _resolve_plugin_agents(plugin_dir) if plugin_dir else set()
1116
+
1117
+ findings = []
1118
+ for line_no, value in matches:
1119
+ builder = _classify_subagent_type(value, plugin_name, plugin_agent_types)
1120
+ if builder is not None:
1121
+ findings.append(builder(path, line_no, value))
1122
+ return findings
1123
+
1124
+
841
1125
  def _resolve_skill_targets(arg):
842
1126
  """Return (skill_md_path_or_None, [hooks.json paths]) for a directory or file arg."""
843
1127
  p = Path(arg)
@@ -872,7 +1156,9 @@ def cmd_lint_skill(args):
872
1156
  continue
873
1157
  if md is not None:
874
1158
  n_files += 1
875
- all_findings.extend(_lint_skill_text(md, Path(md).read_text(encoding="utf-8").splitlines()))
1159
+ md_lines = Path(md).read_text(encoding="utf-8").splitlines()
1160
+ all_findings.extend(_lint_skill_text(md, md_lines))
1161
+ all_findings.extend(_lint_subagent_types(md, md_lines))
876
1162
  for hp in hooks:
877
1163
  n_files += 1
878
1164
  all_findings.extend(
@@ -883,7 +1169,13 @@ def cmd_lint_skill(args):
883
1169
  else:
884
1170
  _print_findings(all_findings, n_files, kind="skill file", clean_suffix=" — no Cowork host-loop footguns.")
885
1171
  has_error = any(x.severity == "ERROR" for x in all_findings)
886
- if has_error or (args.strict and all_findings):
1172
+ # --strict fails on WARN too, per its own --help text ("exit non-zero on WARN too, not just ERROR")
1173
+ # — but NEVER on INFO. The subagent_type ladder (subagent-type-unresolvable /
1174
+ # -not-found-in-plugin / -unknown) and plugin-root-guarded are always INFO by design (there is no
1175
+ # harness registry to disprove an unknown value against — see the subparser help above), so they
1176
+ # must always be surfaced, never failed, even under --strict.
1177
+ has_warn = any(x.severity == "WARN" for x in all_findings)
1178
+ if has_error or (args.strict and has_warn):
887
1179
  return 1
888
1180
  return 0
889
1181
 
@@ -1057,7 +1349,7 @@ def main(argv=None):
1057
1349
 
1058
1350
  lsp = sub.add_parser(
1059
1351
  "lint-skill",
1060
- help="lint SKILL.md bodies for two Cowork host-loop footguns (WARN-only, v1 narrow)",
1352
+ help="lint SKILL.md bodies for Cowork host-loop footguns + static subagent_type resolution",
1061
1353
  description=(
1062
1354
  "Inspect skill bodies (SKILL.md + any sibling hooks.json) for two antipatterns a paid "
1063
1355
  "Cowork host-loop run would expose:\n"
@@ -1068,7 +1360,15 @@ def main(argv=None):
1068
1360
  "ONLY a fenced ```bash/```sh/```shell block, a hooks-config JSON \"command\" value, or a "
1069
1361
  "Bash(...) directive. Host-side prose and Read/Grep directives (the correct way to read a "
1070
1362
  "reference via ${CLAUDE_PLUGIN_ROOT}/...) are left alone. False negatives are expected: a token "
1071
- "in an indented/unfenced shell snippet won't be caught."
1363
+ "in an indented/unfenced shell snippet won't be caught.\n\n"
1364
+ "Also statically resolves any pinned `subagent_type` value in the SKILL.md against the "
1365
+ "enclosing plugin's `agents/*.md` (see `resolve-agent-types`): a value that resolves in-plugin "
1366
+ "or is `general-purpose` is clean; a `<this-plugin>:<agent>` whose agent is missing from an "
1367
+ "enumerable plugin is a provable typo, reported as `subagent-type-not-found-in-plugin`; a "
1368
+ "`<other-plugin>:<agent>` is reported as `subagent-type-unresolvable`; any other unresolved "
1369
+ "value (including a same-plugin prefix the linter couldn't enumerate) is "
1370
+ "`subagent-type-unknown` — all three are always INFO, never WARN, since there is no harness "
1371
+ "registry of built-in agent types to disprove an unknown value against."
1072
1372
  ),
1073
1373
  formatter_class=argparse.RawDescriptionHelpFormatter,
1074
1374
  )
@@ -1077,6 +1377,24 @@ def main(argv=None):
1077
1377
  lsp.add_argument("--strict", action="store_true", help="exit non-zero on WARN too, not just ERROR")
1078
1378
  lsp.set_defaults(func=cmd_lint_skill)
1079
1379
 
1380
+ rap = sub.add_parser(
1381
+ "resolve-agent-types",
1382
+ help="print a plugin's valid <plugin>:<agent> subagent types (from plugin.json + agents/*.md)",
1383
+ description=(
1384
+ "Statically resolve the set of `<plugin>:<agent>` subagent types defined WITHIN a plugin "
1385
+ "dir: the plugin name comes from `.claude-plugin/plugin.json` (fallback `plugin.json`), "
1386
+ "each agent name comes from `agents/*.md`'s `name:` frontmatter (filename-stem fallback "
1387
+ "when a file has no `name:`). Prints an empty result (exit 0) for a dir with no "
1388
+ "plugin.json — nothing to resolve against. This is the token-free 'does "
1389
+ "`<plugin>:<agent>` resolve within this plugin?' answer that backs the `subagent_type` "
1390
+ "check folded into `lint-skill`."
1391
+ ),
1392
+ formatter_class=argparse.RawDescriptionHelpFormatter,
1393
+ )
1394
+ rap.add_argument("plugin_dir", help="plugin directory (containing .claude-plugin/plugin.json or plugin.json)")
1395
+ rap.add_argument("--json", action="store_true", help="emit the resolved types as a JSON array instead of one per line")
1396
+ rap.set_defaults(func=cmd_resolve_agent_types)
1397
+
1080
1398
  sp = sub.add_parser("scaffold", help="emit a valid scenario skeleton (self-linted)")
1081
1399
  sp.add_argument("--name", default="my-scenario", help="scenario name (default: my-scenario)")
1082
1400
  sp.add_argument("--prompt", help="the user turn (the prompt: block)")