cowork-harness 0.30.0 → 0.31.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude/skills/cowork-harness/SKILL.md +18 -12
- package/.claude/skills/cowork-harness/references/ci-recipe.md +23 -11
- package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +1 -1
- package/.claude/skills/cowork-harness/references/scenario-schema.md +3 -2
- package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
- package/.claude/skills/cowork-harness/scripts/assertion-keys.json +1 -0
- package/.claude/skills/cowork-harness/scripts/scenario.py +328 -10
- package/CHANGELOG.md +224 -0
- package/README.md +15 -10
- package/SPEC.md +30 -10
- package/dist/assert.js +76 -1
- package/dist/baseline.js +55 -13
- package/dist/cli.js +251 -9
- package/dist/hostloop/canusetool-gate.js +2 -1
- package/dist/hostloop/pretooluse-path-hook.js +2 -1
- package/dist/run/analyze-skill.js +619 -0
- package/dist/run/cassette.js +63 -12
- package/dist/run/chat-result.js +1 -0
- package/dist/run/doctor.js +12 -3
- package/dist/run/envelope.js +8 -4
- package/dist/run/execute.js +202 -25
- package/dist/run/latest-run.js +130 -0
- package/dist/run/probe-dispatch.js +81 -0
- package/dist/run/run.js +12 -2
- package/dist/run/scenario-tool.js +17 -0
- package/dist/run/subagent-reasoning.js +151 -0
- package/dist/run/verdict.js +23 -1
- package/dist/types.js +19 -0
- package/dist/vm-paths.js +13 -0
- package/docs/boundary.md +1 -1
- package/docs/cassette.md +21 -0
- package/docs/gotchas.md +5 -4
- package/docs/maintenance.md +1 -1
- package/docs/plugin-root.md +11 -1
- package/docs/scenario.md +22 -0
- package/docs/subagents.md +274 -0
- package/examples/replays/README.md +1 -1
- package/package.json +1 -1
- package/schema/run-result.json +122 -2
- package/schema/scenario.schema.json +29 -0
- package/schema/verify-cassettes.json +11 -6
|
@@ -3,8 +3,8 @@ name: cowork-harness
|
|
|
3
3
|
description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
|
|
4
4
|
metadata:
|
|
5
5
|
author: cowork-harness
|
|
6
|
-
version: 0.
|
|
7
|
-
tracks-harness: cowork-harness 0.
|
|
6
|
+
version: 0.31.0
|
|
7
|
+
tracks-harness: cowork-harness 0.31.0 (baseline desktop-1.20186.1)
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
# cowork-harness
|
|
@@ -22,8 +22,8 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
|
|
|
22
22
|
allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
|
|
23
23
|
the highest-value part. Read it.
|
|
24
24
|
|
|
25
|
-
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.
|
|
26
|
-
> `desktop-1.20186.
|
|
25
|
+
> **Version note:** the facts and `file:line` pointers here track `cowork-harness 0.31.0` (baseline
|
|
26
|
+
> `desktop-1.20186.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
|
|
27
27
|
> `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
|
|
28
28
|
|
|
29
29
|
## Preflight — make sure the harness can actually run
|
|
@@ -38,7 +38,7 @@ cowork-harness skill ./my-skill "do X" # run the skill once against the sta
|
|
|
38
38
|
Before the first command, confirm the CLI is reachable and **fail loud** (never fake a pass) when a tier's dependencies are missing:
|
|
39
39
|
|
|
40
40
|
- **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
|
|
41
|
-
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.
|
|
41
|
+
- **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 0.31.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@>=0.31.0" <cmd>` (Node ≥ 20), or install once with `npm i -g "cowork-harness@>=0.31.0"`. **Pin `@>=0.31.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published. (≥ 0.31.0 is what gates the commands/assertions this skill teaches: `assertions --list`, `scaffold <run-id>`, `trace --view dispatches`, `artifact_json` incl. the `in:` operator (passes when the resolved value deep-equals one of the listed members — value ∈ your list, not the reverse), `verify-cassettes`, batch `record <dir>`/`--rerecord-stale`, `record --concurrency <N>`, record-time redaction, multiSelect/`answer:`, `verify-run` answer-coverage, `record --max-artifact-bytes`, live record-time deciders, `verify-cassettes --allow-domain`/`--allow-email`/`--allow-path`/`--allow-patterns-file` (`path` — local absolute filesystem paths — is the scanner's 4th class, new in 0.21.0; `--allow-patterns-file <path>` is a FILE of patterns, one regex per line — not a path to allow, unlike `--allow <regex>`), scenario `skills:` staleness scoping with `COWORK_HARNESS_AGENT_SCOPE=skill`, `chat --plugin`, `/help` in the REPL, `hostloop`'s native host/VM process split with its `allow_host_writes:` consent field, `computer_links_resolve` (new in 0.22.0), `semantic_matches` (an LLM judge grades a fixed `rubric` of claims against the run's answer — **live-only**, so it is evidence-unavailable / skipped-loud on replay, never a vacuous pass; new in 0.27.0), **glob-matched** `tool_called`/`tool_not_called`/`subagent_tool_used`/`subagent_tool_absent` (a pattern like `mcp__workspace__*` matches any tool in the family; exact names match exactly; new in 0.28.0), the five path-gate assertion keys `no_vm_path_file_op`/`vm_path_denied`/`path_denied`/`no_path_denied`/`subagent_file_write`, the session-level `agent_env` knob, and resolved sub-agent identity on dispatch records (`resolvedAgentType`/`resolvedModel`; new in 0.30.0), the `lint-skill` (static host-loop footgun + `subagent_type` resolution linter) / `analyze-skill` (advisory `/sessions`-path static scan, `--strict`, `analyze-skill: ignore` marker) / `probe-dispatch` (single-dispatch mechanics probe) commands, `status --latest-for` (resolve a scenario's newest run dir by run time, not directory mtime), the `subagent_dispatch_healthy` composite assertion, and the persisted `result.json` fields `verdict`, `subagents[].referencesRead`/`subagents[].reasoning`, plus the `toolCounts`/`toolErrors`/`toolDurations` shape distinction (all new in 0.31.0).)
|
|
42
42
|
- **Agent binary (sandboxed live tiers — `container`/`microvm`/`hostloop`/`cowork`).** The staged Claude Code agent is **bind-mounted** from a local Claude Desktop install, or point `COWORK_AGENT_BINARY` at a `claude-code-vm/<ver>/claude` ELF. Nothing is bundled. `protocol` (L0) and `replay` need no staged agent; for the sandboxed tiers, no agent → no run; report that, don't skip silently.
|
|
43
43
|
- **Docker / Lima.** Only `--fidelity protocol` (L0) runs without them. `container` / `microvm` / `hostloop` / `cowork` need Docker (Lima for L2). If they're absent, drop to `--fidelity protocol` and **say so** — a green that never exercised the sandbox is not a sandbox pass.
|
|
44
44
|
- **Auth.** `CLAUDE_CODE_OAUTH_TOKEN` (preferred) or `ANTHROPIC_API_KEY`, via env or `.env`. Minting an OAuth token needs the **`claude` CLI** (`npm i -g @anthropic-ai/claude-code`, then `claude setup-token`).
|
|
@@ -63,13 +63,14 @@ reproducible regression (Part II), and **debug** a run that misbehaved or greene
|
|
|
63
63
|
the path, `trace <run-id>` finds it). **Localize the failure post-hoc** from that evidence:
|
|
64
64
|
`cowork-harness trace <run-dir>`'s views + the emitted `result.json` to see what the run actually did,
|
|
65
65
|
then `verify-run` to re-check a suspect assertion — all token-free, no Docker, no re-record. This is
|
|
66
|
-
the loop 0.
|
|
66
|
+
the loop 0.31.0's observability is built for; the *Triage* and *Inspecting a run's observability
|
|
67
67
|
output* sections in **Part III — Debug** are the detail (the fuller human-facing map lives in
|
|
68
68
|
`docs/debugging.md` — repo-only, not shipped with the installed skill).
|
|
69
69
|
- **Multi-turn / interactive reproduction** → `cowork-harness chat` (interactive; gates answered at the
|
|
70
70
|
TTY, **not** an asserted test — see *Debugging with `chat`* in **Part III — Debug**).
|
|
71
71
|
|
|
72
72
|
Full command set: `skill · run · chat · record · replay · verify-cassettes · rehash · prune · lint ·
|
|
73
|
+
lint-skill · analyze-skill · probe-dispatch ·
|
|
73
74
|
verify-run · trace · inspect · diff · stats · decide · gates · answer · scaffold · assertions --list · sync ·
|
|
74
75
|
list · boundary-check · status · vm <init|status|delete|prune> · doctor · init-redact`. Always check `cowork-harness <cmd> --help`.
|
|
75
76
|
|
|
@@ -421,9 +422,10 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
421
422
|
the current list rather than relying on a fixed enumeration here.
|
|
422
423
|
- **`cowork-harness stats [--metric <m>]`** — aggregate across the run index: `cost`, `duration`,
|
|
423
424
|
`tokens`, `cache-tokens`, `model-cost`, `turns`, `pass-rate`.
|
|
424
|
-
- **`result.json` carries the raw fields** the assertions read: `toolDurations`, `models`, `toolErrors`,
|
|
425
|
+
- **`result.json` carries the raw fields** the assertions read: `verdict`, `toolDurations`, `models`, `toolErrors`,
|
|
425
426
|
`redundantToolCalls`, `modelUsage`, `thinking`, `skillActivity`, `subagents[]` (prompt/`dispatchModel`/
|
|
426
|
-
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`
|
|
427
|
+
`resolvedModel`/output/`attributedSkillId`, `outputTruncated`, `referencesRead`, `reasoning`/`reasoningElided`),
|
|
428
|
+
`context` (tools/mcpServers/availableSkills), `tasks`,
|
|
427
429
|
`workspaceFiles`, `presentedFiles`, `hookEvents`, `mcpErrors`, `contextEvents`, `resources`
|
|
428
430
|
(`probeFailures` distinguishes a failed sample from a tier that was never sampleable). Provenance/
|
|
429
431
|
evidence-health fields: `command` (`run`/`skill`/`record`/`chat`/`replay` — finer than `mode`),
|
|
@@ -431,8 +433,11 @@ decide which assertions from *Assertions: two orthogonal axes* are worth adding)
|
|
|
431
433
|
`bySource` histogram), `evidenceErrors` (dropped/malformed telemetry lines per stream, incl.
|
|
432
434
|
`egressParse`), `fingerprint.frozen` (replay only — marks the shown staleness fingerprint as the
|
|
433
435
|
cassette's record-time value, not a fresh recompute), and `assertTextTruncated` (companion to
|
|
434
|
-
`outputTruncated` on a matched tool result).
|
|
435
|
-
|
|
436
|
+
`outputTruncated` on a matched tool result). Three separately-shaped rollups, easy to conflate in a
|
|
437
|
+
`jq` recipe: `toolCounts` is a flat `{tool: number}` call-count map, `toolErrors` is
|
|
438
|
+
`{tool: {calls, errors}}`, and `toolDurations` is `{tool: {calls, totalMs, maxMs}}`. (Full per-field
|
|
439
|
+
semantics: the README's "Observability fields" section — repo-only; `schema/run-result.json` is the
|
|
440
|
+
machine source.)
|
|
436
441
|
- **Opaque failure?** A failed run also records **`errorSource`** (where the failure originated) and
|
|
437
442
|
**`stderrLogPath`** (the captured agent stderr) — read those and `trace <run-dir>` *before* re-running;
|
|
438
443
|
a re-record rarely tells you more than the captured stderr already does. Also check
|
|
@@ -674,8 +679,9 @@ assertion/replay-relevant ones).
|
|
|
674
679
|
- `references/fidelity-and-answers.md` — fidelity tiers, answer paths, the determinism contract.
|
|
675
680
|
- `references/ci-recipe.md` — the packaged GitHub Action, replay-vs-live lane split, and the four-stage
|
|
676
681
|
GitHub Actions pipeline.
|
|
677
|
-
- `scripts/scenario.py` — `scaffold` a valid scenario skeleton
|
|
678
|
-
no-silent-false-green invariants (both usable as CI steps)
|
|
682
|
+
- `scripts/scenario.py` — `scaffold` a valid scenario skeleton, `lint` scenarios for the
|
|
683
|
+
no-silent-false-green invariants (both usable as CI steps), and `resolve-agent-types <plugin-dir>`
|
|
684
|
+
(validates a pinned `subagent_type` against the plugin's own `plugin.json` + `agents/*.md`).
|
|
679
685
|
- Checking a background run's status without `ps aux` — covered in *Checking whether a background run is
|
|
680
686
|
alive* (Part II) above; the fuller recipe is in `docs/run-status.md` (repo-only, not shipped with the
|
|
681
687
|
installed skill).
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# CI recipe — replay vs live lanes
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 0.31.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
**Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
|
|
6
6
|
job-summary reporter (verdict table, staleness findings, cost/turns when available):
|
|
@@ -54,7 +54,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
|
|
|
54
54
|
GitHub-hosted runners, no token/Docker/agent:
|
|
55
55
|
|
|
56
56
|
```yaml
|
|
57
|
-
- run: npm i -g "cowork-harness@>=0.
|
|
57
|
+
- run: npm i -g "cowork-harness@>=0.31.0"
|
|
58
58
|
- run: cowork-harness lint scenarios/*.yaml # no silent false-greens
|
|
59
59
|
- run: cowork-harness verify-cassettes cassettes/ # privacy + staleness
|
|
60
60
|
- run: cowork-harness replay cassettes/ # token-free content/structure
|
|
@@ -139,13 +139,19 @@ dollar figures). In a skill repo these cassettes get **committed**. So:
|
|
|
139
139
|
universal net (container-tier recordings can trip it too).
|
|
140
140
|
- **Always-on scan gate** — `verify-cassettes` flags email / currency / bare-domain / local-path /
|
|
141
141
|
machine-inventory matches it finds in the committed cassettes and **exits non-zero**, so "no leak" is
|
|
142
|
-
a gate, not discipline.
|
|
142
|
+
a gate, not discipline. Non-zero is not one thing, though: exit `1` means verification RAN and found a
|
|
143
|
+
real finding (a PII match, a genuine staleness drift, or scenario-prompt drift); exit `3` means
|
|
144
|
+
verification could NOT complete (an `unverifiable-*`-class staleness finding, a cassette written by a
|
|
145
|
+
newer harness than this one understands, or a malformed/unreadable cassette). A plain `|| true` or `[
|
|
146
|
+
$? -ne 0 ]` tripwire treats both the same — if you need to tell "the gate caught something" apart from
|
|
147
|
+
"the gate couldn't run", branch on the exit code (or parse `--output-format json`'s per-file
|
|
148
|
+
`findings`/`staleness` vs `unverifiable`/`version`/`error` buckets).
|
|
143
149
|
Suppress synthetic / public reference names (NVCA, Cooley GO, …) with `--allow <regex>`. (Multi-word
|
|
144
150
|
proper names are NOT a default class — too noisy to gate on; add a pattern via config if your corpus
|
|
145
151
|
needs it.)
|
|
146
152
|
|
|
147
153
|
```bash
|
|
148
|
-
cowork-harness verify-cassettes cassettes/ # privacy scan + staleness
|
|
154
|
+
cowork-harness verify-cassettes cassettes/ # privacy scan + staleness — exit 1 = verified & failed, exit 3 = could not verify
|
|
149
155
|
cowork-harness verify-cassettes cassettes/ --allow 'NVCA|Cooley GO|Acme'
|
|
150
156
|
cowork-harness verify-cassettes cassettes/ --skip-privacy # staleness only (skip the privacy scan); both run by default
|
|
151
157
|
```
|
|
@@ -188,7 +194,7 @@ jobs:
|
|
|
188
194
|
with: { node-version: '20' }
|
|
189
195
|
- uses: actions/setup-python@v5
|
|
190
196
|
with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
|
|
191
|
-
- run: npm i -g "cowork-harness@>=0.
|
|
197
|
+
- run: npm i -g "cowork-harness@>=0.31.0"
|
|
192
198
|
- run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
|
|
193
199
|
- run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
|
|
194
200
|
- run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
|
|
@@ -217,7 +223,7 @@ jobs:
|
|
|
217
223
|
echo "live=true" >> "$GITHUB_OUTPUT"
|
|
218
224
|
fi
|
|
219
225
|
- if: steps.guard.outputs.live == 'true'
|
|
220
|
-
run: npm i -g "cowork-harness@>=0.
|
|
226
|
+
run: npm i -g "cowork-harness@>=0.31.0"
|
|
221
227
|
- if: steps.guard.outputs.live == 'true'
|
|
222
228
|
run: cowork-harness run scenarios/ --output-format json
|
|
223
229
|
env:
|
|
@@ -245,9 +251,12 @@ fails or a run errors, so a plain `cowork-harness run scenarios/` is already CI-
|
|
|
245
251
|
JSON.
|
|
246
252
|
|
|
247
253
|
`verify-cassettes` emits its **own** envelope (`{command, ok, coverage, results[]}` with per-file
|
|
248
|
-
`findings`/`staleness`/`notes`/`version`/`error`), published as
|
|
249
|
-
|
|
250
|
-
the
|
|
254
|
+
`findings`/`staleness`/`unverifiable`/`notes`/`version`/`error`), published as
|
|
255
|
+
`schema/verify-cassettes.json` in the npm package. `ok:false` doesn't say *why* — read the buckets, or
|
|
256
|
+
the exit code (`1` = `findings`/`staleness`/`scenarioDrift` populated, a real problem verified & found;
|
|
257
|
+
`3` = `unverifiable`/`version`/`error` populated, verification could not complete; a real finding wins
|
|
258
|
+
`1` if both are present). Both envelope schemas are covered 1.0 contract surfaces (SPEC §12) — parse the
|
|
259
|
+
JSON, not the human-readable text (which is explicitly NOT stable).
|
|
251
260
|
|
|
252
261
|
A run writes to `~/.cowork-harness/runs/<name>/<sessionId>/` by default — outside any working tree. In CI,
|
|
253
262
|
set `COWORK_HARNESS_RUNS_DIR` (or pass `--run-dir`) to a workspace-relative path (e.g. `runs`) so an
|
|
@@ -280,8 +289,11 @@ does **not** imply the recording is still valid. Each replay result carries `sta
|
|
|
280
289
|
|
|
281
290
|
(A pre-`effectiveFidelity` cassette with an **explicit** tier is statically knowable — it passes the tier
|
|
282
291
|
check with a non-failing informational note in the `verify-cassettes` envelope's per-file `notes[]`, a
|
|
283
|
-
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above
|
|
284
|
-
|
|
292
|
+
`·`-prefixed row in text output. On `verify-cassettes` every staleness *finding* above still fails the
|
|
293
|
+
gate (`ok:false`) — but it's no longer class-blind on the EXIT CODE: a `baseline`/`skill`/`shared-root`/
|
|
294
|
+
`format`/`resolved-tier` class lands in the envelope's `staleness[]` (verified & failed — exit `1`),
|
|
295
|
+
while an `unverifiable-*` class lands in `unverifiable[]` (could not verify — exit `3`). Notes never
|
|
296
|
+
fail it either way.)
|
|
285
297
|
|
|
286
298
|
To gate in CI, pick the severity you want:
|
|
287
299
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fidelity tiers & answer paths
|
|
2
2
|
|
|
3
|
-
Self-contained reference. Tracks `cowork-harness 0.
|
|
3
|
+
Self-contained reference. Tracks `cowork-harness 0.31.0` (baseline `desktop-1.20186.1`).
|
|
4
4
|
|
|
5
5
|
## Fidelity tiers (`fidelity:` in the scenario)
|
|
6
6
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Scenario & session schema, assertion catalog, web_fetch, full gotchas
|
|
2
2
|
|
|
3
|
-
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.
|
|
4
|
-
(baseline `desktop-1.20186.
|
|
3
|
+
Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 0.31.0`
|
|
4
|
+
(baseline `desktop-1.20186.1`). If your checkout is newer, prefer the live `docs/scenario.md`,
|
|
5
5
|
`docs/session.md`, and `SPEC.md`.
|
|
6
6
|
|
|
7
7
|
**Minimal scenario** — `prompt` is the only required field:
|
|
@@ -262,6 +262,7 @@ same set live from the schema.
|
|
|
262
262
|
| `subagent_tool_absent: <glob>` | no sub-agent used a tool matching this glob (same rejection) |
|
|
263
263
|
| `no_vm_path_file_op: true` | **`fidelity: hostloop` only** — NO gated file tool attempted a `/sessions`(-prefixed) path (`RunResult.fileToolAttempts`) — content-class, replay-checkable without `controlOut`; any other tier FAILS "cannot verify" (`/sessions/...` is valid there). **Only `true` is valid** |
|
|
264
264
|
| `subagent_file_write: {path?, path_suffix?, tool?}` | a sub-agent-origin write attempt whose raw path equals `path` (exact) or ends with `path_suffix` has a paired non-error tool_result — the causal half of a delivery probe; requires one of `path`/`path_suffix`; `tool` defaults to Write/Edit/MultiEdit; content-class; tier-agnostic |
|
|
265
|
+
| `subagent_dispatch_healthy: {type?, delivered?, path?, path_suffix?, no_vm_paths?}` | **`fidelity: hostloop` only** — composite: selects dispatch(es) via `type` (same matching as `subagent_dispatched`; omit to require every dispatch) and, for EACH selected dispatch, checks it (not just any sub-agent) delivered a paired non-error write (`delivered`, default true — narrowed by `path`/`path_suffix`, same exact-vs-suffix precedence as `subagent_file_write`) and made no `/sessions` VM-path attempt (`no_vm_paths`, default true) — both scoped to that dispatch's OWN `parentToolUseId`, the per-dispatch correlation `subagent_file_write` (which matches ANY sub-agent write) cannot express; a `type` that matches no dispatch FAILS; content-class (`RunResult.fileToolAttempts` + `RunResult.toolResults`); any non-hostloop tier FAILS "cannot verify" |
|
|
265
266
|
| `subagent_dispatched: <regex>` | a sub-agent whose `dispatchAgentType`, binary-*resolved* `resolvedAgentType`, **or dispatch description** matches |
|
|
266
267
|
| `subagent_declared_but_unused: <Tool>` | a sub-agent declared the tool but never used **that** tool (even if it used others) |
|
|
267
268
|
| `subagent_output_contains: {match?, contains}` | a dispatched sub-agent's own output contains the substring `contains` — `match` (optional regex over `dispatchAgentType`/`resolvedAgentType`/`description`) narrows to specific dispatch(es); omitted, checks whether ANY dispatch's output contains it (existence check, not "all"); a miss against an output that was **truncated at the assert cap** reports evidence-unavailable instead of a proven absence — the substring could lie past the cut |
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Each recipe composes facts that live scattered across SKILL.md and the other references into one
|
|
4
4
|
decision path. Every one answers a question a real fleet owner had to work out the hard way. Facts track the harness version in SKILL.md's
|
|
5
|
-
front-matter (currently 0.
|
|
5
|
+
front-matter (currently 0.31.0). Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
|
|
6
6
|
Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
|
|
7
7
|
needed if your CLI meets SKILL.md's version floor.
|
|
8
8
|
|
|
@@ -39,6 +39,7 @@ system PyYAML is preferred when present.
|
|
|
39
39
|
from __future__ import annotations
|
|
40
40
|
|
|
41
41
|
import argparse
|
|
42
|
+
import functools
|
|
42
43
|
import json
|
|
43
44
|
import re
|
|
44
45
|
import sys
|
|
@@ -66,6 +67,7 @@ CONTENT_KEYS = {
|
|
|
66
67
|
"subagent_output_contains",
|
|
67
68
|
"no_vm_path_file_op",
|
|
68
69
|
"subagent_file_write",
|
|
70
|
+
"subagent_dispatch_healthy",
|
|
69
71
|
"dispatch_count_max",
|
|
70
72
|
"skill_triggered",
|
|
71
73
|
"no_skill_triggered",
|
|
@@ -741,6 +743,53 @@ def _finding_plugin_root_guarded(path, line, ctx_label):
|
|
|
741
743
|
)
|
|
742
744
|
|
|
743
745
|
|
|
746
|
+
# A self-heal `find`'s `-path '<glob>'` / `-path "<glob>"` value (same-quote char, non-greedy so a
|
|
747
|
+
# quote inside the glob — unlikely in practice — doesn't get swallowed).
|
|
748
|
+
_FIND_PATH_VALUE = re.compile(r"""-path\s+(['"])(.*?)\1""")
|
|
749
|
+
# The skill segment out of a `*/skills/<name>/...` glob.
|
|
750
|
+
_FIND_PATH_SKILLS_SEG = re.compile(r"/skills/([A-Za-z0-9_.-]+)/")
|
|
751
|
+
# The plugin segment out of a `*/plugins/<name>/...` glob.
|
|
752
|
+
_FIND_PATH_PLUGINS_SEG = re.compile(r"/plugins/([A-Za-z0-9_.-]+)/")
|
|
753
|
+
# A generic `*/<name>/scripts` glob (no `skills/`/`plugins/` literal prefix) — the segment right
|
|
754
|
+
# before a `/scripts` path component. This is what a plugin-level self-heal targeting its own
|
|
755
|
+
# `<plugin>/scripts/...` layout looks like once the real mount path (e.g. `mnt/.local-plugins/...`)
|
|
756
|
+
# is glob-abbreviated to `*/<plugin>/scripts`.
|
|
757
|
+
_FIND_PATH_SCRIPTS_SEG = re.compile(r"/([A-Za-z0-9_.-]+)/scripts\b")
|
|
758
|
+
|
|
759
|
+
|
|
760
|
+
def _extract_find_path_token(find_cmd_line):
|
|
761
|
+
"""Pull the skill/plugin-naming token out of a self-heal `find`'s `-path` glob (on the SAME
|
|
762
|
+
line as the `find`, matching how `_SELF_HEAL` itself is line-scoped), or None if there's no
|
|
763
|
+
`-path` clause or none of the recognized glob shapes match. A bare glob wildcard segment (e.g.
|
|
764
|
+
`*/scripts/*`, no name between the slashes) intentionally does not match the `[A-Za-z0-9_.-]+`
|
|
765
|
+
character class — that's a real absence of an extractable token, not a token, so it stays
|
|
766
|
+
conservative (INFO) rather than manufacturing a bogus `*` token to compare."""
|
|
767
|
+
m = _FIND_PATH_VALUE.search(find_cmd_line)
|
|
768
|
+
if not m:
|
|
769
|
+
return None
|
|
770
|
+
glob = m.group(2)
|
|
771
|
+
for pat in (_FIND_PATH_SKILLS_SEG, _FIND_PATH_PLUGINS_SEG, _FIND_PATH_SCRIPTS_SEG):
|
|
772
|
+
seg = pat.search(glob)
|
|
773
|
+
if seg:
|
|
774
|
+
return seg.group(1)
|
|
775
|
+
return None
|
|
776
|
+
|
|
777
|
+
|
|
778
|
+
def _finding_guard_pattern_mismatch(path, line, ctx_label, token, skill_name, plugin_name):
|
|
779
|
+
plugin_part = f" (plugin `{plugin_name}`)" if plugin_name else ""
|
|
780
|
+
return Finding(
|
|
781
|
+
"WARN",
|
|
782
|
+
"guard-pattern-mismatch",
|
|
783
|
+
f"`${{CLAUDE_PLUGIN_ROOT}}` used in an in-VM bash context ({ctx_label}); the block's self-heal "
|
|
784
|
+
f"`find` targets `{token}`, but this skill is `{skill_name}`{plugin_part} — the guard won't "
|
|
785
|
+
"discover THIS skill's mount (likely a copy-pasted self-heal).",
|
|
786
|
+
f"Fix the `find` pattern's `-path` to match this skill/plugin's own layout (`{skill_name}`"
|
|
787
|
+
f"{plugin_part}), not `{token}`.",
|
|
788
|
+
path,
|
|
789
|
+
line,
|
|
790
|
+
)
|
|
791
|
+
|
|
792
|
+
|
|
744
793
|
def _finding_hook_host_write(path, line, what):
|
|
745
794
|
return Finding(
|
|
746
795
|
"WARN",
|
|
@@ -778,15 +827,46 @@ def _lint_skill_text(path, raw_lines, force_json=False):
|
|
|
778
827
|
bash_token_lines = []
|
|
779
828
|
bash_block_text = []
|
|
780
829
|
|
|
830
|
+
# Self-heal find-pattern guard: identity of the skill/plugin under lint, used to check a self-heal `find -path`
|
|
831
|
+
# actually names THIS skill or its enclosing plugin (not a copy-pasted mismatch). Only
|
|
832
|
+
# meaningful when linting an actual SKILL.md (force_json=False is exactly that case here — a
|
|
833
|
+
# hooks.json body never enters the "bash" ctx below, so this is otherwise unused).
|
|
834
|
+
skill_name = None
|
|
835
|
+
plugin_name = None
|
|
836
|
+
self_plugin_tokens = set()
|
|
837
|
+
if not force_json:
|
|
838
|
+
dir_name = Path(path).resolve().parent.name
|
|
839
|
+
fm_name = _agent_name_from_frontmatter(path, _require_yaml())
|
|
840
|
+
# Prefer the frontmatter `name:` (the skill's declared identity) for display when present;
|
|
841
|
+
# the parent-dir name is always in the match set as a cross-check (both count as "this
|
|
842
|
+
# skill" — a self-heal naming either is not a mismatch).
|
|
843
|
+
skill_name = fm_name or dir_name
|
|
844
|
+
self_plugin_tokens.add(dir_name)
|
|
845
|
+
if fm_name:
|
|
846
|
+
self_plugin_tokens.add(fm_name)
|
|
847
|
+
plugin_dir = _find_enclosing_plugin_dir(path)
|
|
848
|
+
if plugin_dir is not None:
|
|
849
|
+
plugin_name = _read_plugin_name(plugin_dir)
|
|
850
|
+
if plugin_name:
|
|
851
|
+
self_plugin_tokens.add(plugin_name)
|
|
852
|
+
self_plugin_tokens.add(Path(plugin_dir).name)
|
|
853
|
+
|
|
781
854
|
def flush_bash():
|
|
782
855
|
if bash_token_lines:
|
|
783
|
-
|
|
856
|
+
self_heal_line = next((bl for bl in bash_block_text if _SELF_HEAL.search(bl)), None)
|
|
857
|
+
healed = self_heal_line is not None
|
|
858
|
+
token = _extract_find_path_token(self_heal_line) if healed else None
|
|
784
859
|
for ln in bash_token_lines:
|
|
785
|
-
|
|
786
|
-
|
|
787
|
-
|
|
788
|
-
|
|
789
|
-
|
|
860
|
+
if not healed:
|
|
861
|
+
findings.append(_finding_plugin_root(path, ln, "```bash block"))
|
|
862
|
+
elif token is not None and token not in self_plugin_tokens:
|
|
863
|
+
findings.append(
|
|
864
|
+
_finding_guard_pattern_mismatch(
|
|
865
|
+
path, ln, "```bash block", token, skill_name, plugin_name
|
|
866
|
+
)
|
|
867
|
+
)
|
|
868
|
+
else:
|
|
869
|
+
findings.append(_finding_plugin_root_guarded(path, ln, "```bash block"))
|
|
790
870
|
bash_token_lines.clear()
|
|
791
871
|
bash_block_text.clear()
|
|
792
872
|
|
|
@@ -838,6 +918,210 @@ def _lint_skill_text(path, raw_lines, force_json=False):
|
|
|
838
918
|
return findings
|
|
839
919
|
|
|
840
920
|
|
|
921
|
+
# --------------------------------------------------------------------------- #
|
|
922
|
+
# subagent_type static resolution
|
|
923
|
+
# --------------------------------------------------------------------------- #
|
|
924
|
+
#
|
|
925
|
+
# A pinned `subagent_type:` value that doesn't resolve to a real agent fails a definition lookup at
|
|
926
|
+
# dispatch time — but that's only discoverable via a live dispatch today. Resolve it statically from
|
|
927
|
+
# a plugin's own `.claude-plugin/plugin.json` (or `plugin.json`) + `agents/*.md` frontmatter instead.
|
|
928
|
+
#
|
|
929
|
+
# HONEST LIMIT: there is no harness registry of built-in agent types (the built-in set is
|
|
930
|
+
# agent-binary-version-dependent) — only `general-purpose` is harness-known. So an unresolved bare
|
|
931
|
+
# value is surfaced as INFO, never failed as WARN/ERROR; the linter can't disprove it's a real
|
|
932
|
+
# built-in. Do NOT add a committed built-in agent-type list here — that would silently go stale and
|
|
933
|
+
# either false-warn a real built-in or false-clear a typo.
|
|
934
|
+
|
|
935
|
+
_SUBAGENT_TYPE_RE = re.compile(r"subagent_type\s*[:=]\s*['\"]?([A-Za-z0-9_.:/-]+)['\"]?")
|
|
936
|
+
|
|
937
|
+
|
|
938
|
+
def _read_plugin_name(plugin_dir):
|
|
939
|
+
"""Return the `name` field from `<plugin_dir>/.claude-plugin/plugin.json` (fallback
|
|
940
|
+
`<plugin_dir>/plugin.json`), or None if neither file exists or is parsable. Never raises."""
|
|
941
|
+
p = Path(plugin_dir)
|
|
942
|
+
for candidate in (p / ".claude-plugin" / "plugin.json", p / "plugin.json"):
|
|
943
|
+
if candidate.is_file():
|
|
944
|
+
try:
|
|
945
|
+
data = json.loads(candidate.read_text(encoding="utf-8"))
|
|
946
|
+
except Exception:
|
|
947
|
+
return None
|
|
948
|
+
name = data.get("name") if isinstance(data, dict) else None
|
|
949
|
+
return name if isinstance(name, str) and name.strip() else None
|
|
950
|
+
return None
|
|
951
|
+
|
|
952
|
+
|
|
953
|
+
_AGENT_FRONTMATTER = re.compile(r"^---\s*\n(.*?\n)---\s*(?:\n|$)", re.DOTALL)
|
|
954
|
+
|
|
955
|
+
|
|
956
|
+
def _agent_name_from_frontmatter(md_path, yaml_mod):
|
|
957
|
+
"""Return a markdown file's `name:` frontmatter value, or None if there's no frontmatter, no
|
|
958
|
+
`name:` field, or it fails to parse. Originally for `agents/*.md` (caller falls back to the
|
|
959
|
+
filename stem there); also reused by the self-heal find-pattern guard for a SKILL.md, whose frontmatter has the same
|
|
960
|
+
`---\\nname: ...\\n---` shape — the parser itself is generic, only the name is agent-specific."""
|
|
961
|
+
try:
|
|
962
|
+
text = Path(md_path).read_text(encoding="utf-8")
|
|
963
|
+
except Exception:
|
|
964
|
+
return None
|
|
965
|
+
m = _AGENT_FRONTMATTER.match(text)
|
|
966
|
+
if not m:
|
|
967
|
+
return None
|
|
968
|
+
try:
|
|
969
|
+
data = yaml_mod.safe_load(m.group(1))
|
|
970
|
+
except Exception:
|
|
971
|
+
return None
|
|
972
|
+
if isinstance(data, dict):
|
|
973
|
+
name = data.get("name")
|
|
974
|
+
if isinstance(name, str) and name.strip():
|
|
975
|
+
return name.strip()
|
|
976
|
+
return None
|
|
977
|
+
|
|
978
|
+
|
|
979
|
+
def _resolve_plugin_agents(plugin_dir):
|
|
980
|
+
"""Resolve in-plugin agent types: return the set of valid `<plugin>:<agent>` subagent types defined WITHIN plugin_dir.
|
|
981
|
+
Reads the plugin name from plugin.json and each agents/*.md's `name:` frontmatter (filename stem
|
|
982
|
+
fallback). Returns an empty set (never crashes) when no plugin.json is found — a bare SKILL.md
|
|
983
|
+
dir with no plugin manifest has nothing to resolve against."""
|
|
984
|
+
plugin_name = _read_plugin_name(plugin_dir)
|
|
985
|
+
if not plugin_name:
|
|
986
|
+
return set()
|
|
987
|
+
agents_dir = Path(plugin_dir) / "agents"
|
|
988
|
+
if not agents_dir.is_dir():
|
|
989
|
+
return set()
|
|
990
|
+
yaml = _require_yaml()
|
|
991
|
+
types = set()
|
|
992
|
+
for md in sorted(agents_dir.glob("*.md")):
|
|
993
|
+
agent_name = _agent_name_from_frontmatter(md, yaml) or md.stem
|
|
994
|
+
types.add(f"{plugin_name}:{agent_name}")
|
|
995
|
+
return types
|
|
996
|
+
|
|
997
|
+
|
|
998
|
+
def cmd_resolve_agent_types(args):
|
|
999
|
+
types = sorted(_resolve_plugin_agents(args.plugin_dir))
|
|
1000
|
+
if args.json:
|
|
1001
|
+
print(json.dumps(types))
|
|
1002
|
+
else:
|
|
1003
|
+
for t in types:
|
|
1004
|
+
print(t)
|
|
1005
|
+
return 0
|
|
1006
|
+
|
|
1007
|
+
|
|
1008
|
+
def _find_enclosing_plugin_dir(skill_md_path):
|
|
1009
|
+
"""Resolve the enclosing plugin: walk up from a SKILL.md to the nearest ancestor dir containing
|
|
1010
|
+
`.claude-plugin/plugin.json` or `plugin.json` — that's the enclosing plugin. None if no ancestor
|
|
1011
|
+
has one (a bare SKILL.md dir with no plugin manifest anywhere above it)."""
|
|
1012
|
+
start = Path(skill_md_path).resolve().parent
|
|
1013
|
+
for anc in [start, *start.parents]:
|
|
1014
|
+
if (anc / ".claude-plugin" / "plugin.json").is_file() or (anc / "plugin.json").is_file():
|
|
1015
|
+
return anc
|
|
1016
|
+
return None
|
|
1017
|
+
|
|
1018
|
+
|
|
1019
|
+
def _finding_subagent_unresolvable(path, line, value):
|
|
1020
|
+
return Finding(
|
|
1021
|
+
"INFO",
|
|
1022
|
+
"subagent-type-unresolvable",
|
|
1023
|
+
f"pinned type `{value}` belongs to another plugin — can't confirm it resolves from here",
|
|
1024
|
+
"Verify it resolves in that plugin's own agents/ dir (e.g. `scenario.py resolve-agent-types "
|
|
1025
|
+
"<that-plugin-dir>`), or dispatch without pinning a cross-plugin type.",
|
|
1026
|
+
path,
|
|
1027
|
+
line,
|
|
1028
|
+
)
|
|
1029
|
+
|
|
1030
|
+
|
|
1031
|
+
def _finding_subagent_unknown(path, line, value):
|
|
1032
|
+
return Finding(
|
|
1033
|
+
"INFO",
|
|
1034
|
+
"subagent-type-unknown",
|
|
1035
|
+
f"pinned type `{value}` is not defined in this plugin and is not the `general-purpose` "
|
|
1036
|
+
"built-in — can't confirm statically (may be an agent-binary built-in)",
|
|
1037
|
+
"If it's meant to be an in-plugin agent, add `agents/<name>.md` with a `name:` frontmatter "
|
|
1038
|
+
"matching the pinned value (or rely on the filename-stem fallback). If it's a real built-in "
|
|
1039
|
+
"agent type, this INFO is expected — the linter has no built-in registry to check it against.",
|
|
1040
|
+
path,
|
|
1041
|
+
line,
|
|
1042
|
+
)
|
|
1043
|
+
|
|
1044
|
+
|
|
1045
|
+
def _finding_subagent_not_found_in_plugin(path, line, value, plugin_name, agent, sorted_agents):
|
|
1046
|
+
"""A pinned `<this-plugin>:<agent>` value whose prefix names the RESOLVED plugin (this SKILL.md's
|
|
1047
|
+
own enclosing plugin) but whose agent isn't in its enumerated agents/*.md set. Unlike
|
|
1048
|
+
`subagent-type-unknown`, this can never be another binary's built-in — the namespace prefix
|
|
1049
|
+
already commits it to this plugin, and the plugin's agent set was fully enumerable — so it's a
|
|
1050
|
+
provable typo, not an unconfirmable unknown."""
|
|
1051
|
+
return Finding(
|
|
1052
|
+
"INFO",
|
|
1053
|
+
"subagent-type-not-found-in-plugin",
|
|
1054
|
+
f"pinned type `{value}` names this plugin (`{plugin_name}`) but `{agent}` is not among its "
|
|
1055
|
+
f"agents [{', '.join(sorted_agents)}] — likely a typo.",
|
|
1056
|
+
"Fix the agent name to match one of the listed agents, or add `agents/<agent>.md` with a "
|
|
1057
|
+
"`name:` frontmatter matching the pinned value (or rely on the filename-stem fallback) if it "
|
|
1058
|
+
"was meant to exist. Check `scenario.py resolve-agent-types <plugin-dir>` to confirm.",
|
|
1059
|
+
path,
|
|
1060
|
+
line,
|
|
1061
|
+
)
|
|
1062
|
+
|
|
1063
|
+
|
|
1064
|
+
def _classify_subagent_type(value, plugin_name, plugin_agent_types):
|
|
1065
|
+
"""subagent_type severity ladder. Returns a Finding-builder (path, line, value) -> Finding, or None if
|
|
1066
|
+
clean. Severity is ALWAYS INFO — see the HONEST LIMIT note above this section; never WARN/ERROR.
|
|
1067
|
+
|
|
1068
|
+
Precedence:
|
|
1069
|
+
1. Resolves in-plugin, or is `general-purpose` -> clean (None).
|
|
1070
|
+
2. Has a `<prefix>:<agent>` shape whose prefix EQUALS the resolved plugin's own name AND the
|
|
1071
|
+
plugin's agent set was non-empty (i.e. enumerable) -> `subagent-type-not-found-in-plugin`.
|
|
1072
|
+
A value namespaced under this plugin's own name can never be another binary's built-in, so
|
|
1073
|
+
once the plugin is fully enumerated a miss here is a provable typo, not an unconfirmable
|
|
1074
|
+
unknown.
|
|
1075
|
+
3. Has a `<prefix>:<agent>` shape whose prefix does NOT equal the resolved plugin's name (or
|
|
1076
|
+
there's no resolved plugin at all) -> `subagent-type-unresolvable` (belongs to another
|
|
1077
|
+
plugin, or an empty plugin_name means we truly can't tell whose namespace it is).
|
|
1078
|
+
4. Otherwise (no colon, or same-plugin prefix but the plugin set was empty/unenumerable) ->
|
|
1079
|
+
`subagent-type-unknown` — genuinely can't confirm statically.
|
|
1080
|
+
"""
|
|
1081
|
+
if value == "general-purpose" or value in plugin_agent_types:
|
|
1082
|
+
return None
|
|
1083
|
+
if ":" in value:
|
|
1084
|
+
prefix, agent = value.split(":", 1)
|
|
1085
|
+
if plugin_name is not None and prefix == plugin_name:
|
|
1086
|
+
if plugin_agent_types:
|
|
1087
|
+
sorted_agents = sorted(
|
|
1088
|
+
t.split(":", 1)[1] for t in plugin_agent_types if t.split(":", 1)[0] == plugin_name
|
|
1089
|
+
)
|
|
1090
|
+
return functools.partial(
|
|
1091
|
+
_finding_subagent_not_found_in_plugin,
|
|
1092
|
+
plugin_name=plugin_name,
|
|
1093
|
+
agent=agent,
|
|
1094
|
+
sorted_agents=sorted_agents,
|
|
1095
|
+
)
|
|
1096
|
+
# Plugin set is empty — couldn't enumerate (e.g. no agents/ dir) — genuinely can't confirm.
|
|
1097
|
+
return _finding_subagent_unknown
|
|
1098
|
+
return _finding_subagent_unresolvable
|
|
1099
|
+
return _finding_subagent_unknown
|
|
1100
|
+
|
|
1101
|
+
|
|
1102
|
+
def _lint_subagent_types(path, raw_lines):
|
|
1103
|
+
"""Scan a SKILL.md's raw text (not limited to fenced blocks — a pinned `subagent_type`
|
|
1104
|
+
can appear in prose or YAML frontmatter) for pinned `subagent_type` values and classify each
|
|
1105
|
+
against the enclosing plugin's in-plugin agent set."""
|
|
1106
|
+
matches = []
|
|
1107
|
+
for i, line in enumerate(raw_lines, start=1):
|
|
1108
|
+
for m in _SUBAGENT_TYPE_RE.finditer(line):
|
|
1109
|
+
matches.append((i, m.group(1)))
|
|
1110
|
+
if not matches:
|
|
1111
|
+
return []
|
|
1112
|
+
|
|
1113
|
+
plugin_dir = _find_enclosing_plugin_dir(path)
|
|
1114
|
+
plugin_name = _read_plugin_name(plugin_dir) if plugin_dir else None
|
|
1115
|
+
plugin_agent_types = _resolve_plugin_agents(plugin_dir) if plugin_dir else set()
|
|
1116
|
+
|
|
1117
|
+
findings = []
|
|
1118
|
+
for line_no, value in matches:
|
|
1119
|
+
builder = _classify_subagent_type(value, plugin_name, plugin_agent_types)
|
|
1120
|
+
if builder is not None:
|
|
1121
|
+
findings.append(builder(path, line_no, value))
|
|
1122
|
+
return findings
|
|
1123
|
+
|
|
1124
|
+
|
|
841
1125
|
def _resolve_skill_targets(arg):
|
|
842
1126
|
"""Return (skill_md_path_or_None, [hooks.json paths]) for a directory or file arg."""
|
|
843
1127
|
p = Path(arg)
|
|
@@ -872,7 +1156,9 @@ def cmd_lint_skill(args):
|
|
|
872
1156
|
continue
|
|
873
1157
|
if md is not None:
|
|
874
1158
|
n_files += 1
|
|
875
|
-
|
|
1159
|
+
md_lines = Path(md).read_text(encoding="utf-8").splitlines()
|
|
1160
|
+
all_findings.extend(_lint_skill_text(md, md_lines))
|
|
1161
|
+
all_findings.extend(_lint_subagent_types(md, md_lines))
|
|
876
1162
|
for hp in hooks:
|
|
877
1163
|
n_files += 1
|
|
878
1164
|
all_findings.extend(
|
|
@@ -883,7 +1169,13 @@ def cmd_lint_skill(args):
|
|
|
883
1169
|
else:
|
|
884
1170
|
_print_findings(all_findings, n_files, kind="skill file", clean_suffix=" — no Cowork host-loop footguns.")
|
|
885
1171
|
has_error = any(x.severity == "ERROR" for x in all_findings)
|
|
886
|
-
|
|
1172
|
+
# --strict fails on WARN too, per its own --help text ("exit non-zero on WARN too, not just ERROR")
|
|
1173
|
+
# — but NEVER on INFO. The subagent_type ladder (subagent-type-unresolvable /
|
|
1174
|
+
# -not-found-in-plugin / -unknown) and plugin-root-guarded are always INFO by design (there is no
|
|
1175
|
+
# harness registry to disprove an unknown value against — see the subparser help above), so they
|
|
1176
|
+
# must always be surfaced, never failed, even under --strict.
|
|
1177
|
+
has_warn = any(x.severity == "WARN" for x in all_findings)
|
|
1178
|
+
if has_error or (args.strict and has_warn):
|
|
887
1179
|
return 1
|
|
888
1180
|
return 0
|
|
889
1181
|
|
|
@@ -1057,7 +1349,7 @@ def main(argv=None):
|
|
|
1057
1349
|
|
|
1058
1350
|
lsp = sub.add_parser(
|
|
1059
1351
|
"lint-skill",
|
|
1060
|
-
help="lint SKILL.md bodies for
|
|
1352
|
+
help="lint SKILL.md bodies for Cowork host-loop footguns + static subagent_type resolution",
|
|
1061
1353
|
description=(
|
|
1062
1354
|
"Inspect skill bodies (SKILL.md + any sibling hooks.json) for two antipatterns a paid "
|
|
1063
1355
|
"Cowork host-loop run would expose:\n"
|
|
@@ -1068,7 +1360,15 @@ def main(argv=None):
|
|
|
1068
1360
|
"ONLY a fenced ```bash/```sh/```shell block, a hooks-config JSON \"command\" value, or a "
|
|
1069
1361
|
"Bash(...) directive. Host-side prose and Read/Grep directives (the correct way to read a "
|
|
1070
1362
|
"reference via ${CLAUDE_PLUGIN_ROOT}/...) are left alone. False negatives are expected: a token "
|
|
1071
|
-
"in an indented/unfenced shell snippet won't be caught
|
|
1363
|
+
"in an indented/unfenced shell snippet won't be caught.\n\n"
|
|
1364
|
+
"Also statically resolves any pinned `subagent_type` value in the SKILL.md against the "
|
|
1365
|
+
"enclosing plugin's `agents/*.md` (see `resolve-agent-types`): a value that resolves in-plugin "
|
|
1366
|
+
"or is `general-purpose` is clean; a `<this-plugin>:<agent>` whose agent is missing from an "
|
|
1367
|
+
"enumerable plugin is a provable typo, reported as `subagent-type-not-found-in-plugin`; a "
|
|
1368
|
+
"`<other-plugin>:<agent>` is reported as `subagent-type-unresolvable`; any other unresolved "
|
|
1369
|
+
"value (including a same-plugin prefix the linter couldn't enumerate) is "
|
|
1370
|
+
"`subagent-type-unknown` — all three are always INFO, never WARN, since there is no harness "
|
|
1371
|
+
"registry of built-in agent types to disprove an unknown value against."
|
|
1072
1372
|
),
|
|
1073
1373
|
formatter_class=argparse.RawDescriptionHelpFormatter,
|
|
1074
1374
|
)
|
|
@@ -1077,6 +1377,24 @@ def main(argv=None):
|
|
|
1077
1377
|
lsp.add_argument("--strict", action="store_true", help="exit non-zero on WARN too, not just ERROR")
|
|
1078
1378
|
lsp.set_defaults(func=cmd_lint_skill)
|
|
1079
1379
|
|
|
1380
|
+
rap = sub.add_parser(
|
|
1381
|
+
"resolve-agent-types",
|
|
1382
|
+
help="print a plugin's valid <plugin>:<agent> subagent types (from plugin.json + agents/*.md)",
|
|
1383
|
+
description=(
|
|
1384
|
+
"Statically resolve the set of `<plugin>:<agent>` subagent types defined WITHIN a plugin "
|
|
1385
|
+
"dir: the plugin name comes from `.claude-plugin/plugin.json` (fallback `plugin.json`), "
|
|
1386
|
+
"each agent name comes from `agents/*.md`'s `name:` frontmatter (filename-stem fallback "
|
|
1387
|
+
"when a file has no `name:`). Prints an empty result (exit 0) for a dir with no "
|
|
1388
|
+
"plugin.json — nothing to resolve against. This is the token-free 'does "
|
|
1389
|
+
"`<plugin>:<agent>` resolve within this plugin?' answer that backs the `subagent_type` "
|
|
1390
|
+
"check folded into `lint-skill`."
|
|
1391
|
+
),
|
|
1392
|
+
formatter_class=argparse.RawDescriptionHelpFormatter,
|
|
1393
|
+
)
|
|
1394
|
+
rap.add_argument("plugin_dir", help="plugin directory (containing .claude-plugin/plugin.json or plugin.json)")
|
|
1395
|
+
rap.add_argument("--json", action="store_true", help="emit the resolved types as a JSON array instead of one per line")
|
|
1396
|
+
rap.set_defaults(func=cmd_resolve_agent_types)
|
|
1397
|
+
|
|
1080
1398
|
sp = sub.add_parser("scaffold", help="emit a valid scenario skeleton (self-linted)")
|
|
1081
1399
|
sp.add_argument("--name", default="my-scenario", help="scenario name (default: my-scenario)")
|
|
1082
1400
|
sp.add_argument("--prompt", help="the user turn (the prompt: block)")
|