cowork-harness 3.6.0 → 3.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/.claude/skills/cowork-harness/SKILL.md +4 -4
  2. package/.claude/skills/cowork-harness/references/ci-recipe.md +5 -5
  3. package/.claude/skills/cowork-harness/references/critique.md +22 -9
  4. package/.claude/skills/cowork-harness/references/fidelity-and-answers.md +8 -1
  5. package/.claude/skills/cowork-harness/references/scenario-schema.md +1 -1
  6. package/.claude/skills/cowork-harness/references/task-recipes.md +1 -1
  7. package/.claude/skills/cowork-harness/scripts/scenario.py +352 -23
  8. package/CHANGELOG.md +232 -1
  9. package/CODE_OF_CONDUCT.md +143 -0
  10. package/CONTRIBUTING.md +15 -4
  11. package/DESIGN.md +30 -3
  12. package/README.md +33 -6
  13. package/RELEASING.md +8 -0
  14. package/dist/cli.js +4 -4
  15. package/dist/critique/armor.js +10 -1
  16. package/dist/critique/command.js +77 -23
  17. package/dist/critique/corpus-walk.js +80 -0
  18. package/dist/critique/evaluator.js +17 -8
  19. package/dist/critique/limitations.js +1 -1
  20. package/dist/critique/package-evidence.js +172 -120
  21. package/dist/critique/resolve-agents.js +175 -0
  22. package/dist/critique/resolve-references.js +261 -0
  23. package/dist/run/analyze-skill.js +1 -1
  24. package/docs/README.md +3 -1
  25. package/docs/boundary.md +11 -0
  26. package/docs/cassette.md +17 -0
  27. package/docs/ci.md +1 -1
  28. package/docs/cli.md +54 -18
  29. package/docs/companion-skill.md +8 -18
  30. package/docs/critique.md +102 -18
  31. package/docs/decisions/README.md +2 -0
  32. package/docs/fidelity-gaps.md +27 -3
  33. package/docs/protocol.md +20 -0
  34. package/docs/scenario.md +11 -0
  35. package/docs/subagents.md +16 -0
  36. package/examples/README.md +2 -2
  37. package/examples/replays/README.md +1 -1
  38. package/llms.txt +1 -1
  39. package/package.json +2 -1
  40. package/python/test_scenario_lint.py +233 -0
  41. package/schema/critique-report.json +32 -2
  42. package/scripts/gen-surface.ts +14 -3
@@ -3,8 +3,8 @@ name: cowork-harness
3
3
  description: Test or debug a Claude Code skill/plugin under Claude Cowork's runtime — sandboxed agent, default-deny egress, the can_use_tool permission/question protocol — using the cowork-harness CLI. Use when validating or regression-testing a skill, authoring or debugging a scenario YAML (prompt + scripted answers + assert:), choosing a fidelity tier, scripting AskUserQuestion / tool-permission answers, or asserting artifacts, egress, or sub-agent dispatch. Especially when a harness run no-ops an assertion, fails on an unanswered gate, false-greens, a steered answer never reaches the model, or a web_fetch is unexpectedly denied or gated. Also when iterating or hardening a skill across fixes, or grounding a skill's self-critique against its own run evidence — including a document-analysis skill (cap table, deck, financial model, transcript) that needs an uploaded file attached to be critiqued at all. NOT for generic unit testing (pytest/vitest of your own scripts) or non-Cowork CI. Covers the skill / run / chat / record / replay / trace / decide / assertions / scaffold commands and the session-vs-scenario split.
4
4
  metadata:
5
5
  author: cowork-harness
6
- version: 3.6.0
7
- tracks-harness: cowork-harness 3.6.0 (baseline desktop-2.2553.1)
6
+ version: 3.7.0
7
+ tracks-harness: cowork-harness 3.7.0 (baseline desktop-2.2553.1)
8
8
  ---
9
9
 
10
10
  # cowork-harness
@@ -25,7 +25,7 @@ flagged with a loud `::warning::`, not silent — auto-answer a gate, observe an
25
25
  allowlist). This skill exists mostly to keep you out of those traps — the Gotchas section below is
26
26
  the highest-value part. Read it.
27
27
 
28
- > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.6.0` (baseline
28
+ > **Version note:** the facts and `file:line` pointers here track `cowork-harness 3.7.0` (baseline
29
29
  > `desktop-2.2553.1`). If your checkout is newer, prefer the live `--help` and — in a repo checkout —
30
30
  > `SPEC.md` / `docs/*.md` over this snapshot, and re-run the bundled linter.
31
31
 
@@ -42,7 +42,7 @@ Before the first command, confirm the CLI is reachable and **fail loud** (never
42
42
 
43
43
  - **One-shot check.** Run `cowork-harness doctor [--tier <tier>]` first — a read-only prerequisite check that inspects Docker, the staged agent, the token, and the baseline in one pass. The bullets below explain each thing it checks (and how to fix it).
44
44
  - **Replay-only? Skip `doctor`.** Replaying committed cassettes needs no Docker, no staged agent, and no token — and every tier's `doctor` validates the auth token (the live tiers also Docker + the staged agent), so a ✗ there is expected, not a blocker. Go straight to `cowork-harness replay <cassette>`.
45
- - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.6.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.6.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.6.0"`. **Pin `@^3.6.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
45
+ - **CLI on PATH, recent enough?** Run `cowork-harness --version` — this skill needs **≥ 3.7.0**. If it's missing or older, prefix every command with the version floor `npx "cowork-harness@^3.7.0" <cmd>` (Node ≥ 22), or install once with `npm i -g "cowork-harness@^3.7.0"`. **Pin `@^3.7.0`, never `@latest`** — `@latest` can silently fetch an older CLI and the new commands fail as "unknown command", whereas the floor **fails loud** if no compatible version is published.
46
46
 
47
47
  This skill documents the CURRENT surface, not release history. If `cowork-harness --version` is
48
48
  OLDER than the floor, the per-release record of what you are missing is [CHANGELOG.md](https://github.com/yaniv-golan/cowork-harness/blob/main/CHANGELOG.md)
@@ -1,6 +1,6 @@
1
1
  # CI recipe — replay vs live lanes
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`).
4
4
 
5
5
  **Fastest path: the packaged Action.** One step gets you `replay`/`lint`/`verify-cassettes` plus a PR
6
6
  job-summary reporter (verdict table, staleness findings, cost/turns when available):
@@ -17,7 +17,7 @@ job-summary reporter (verdict table, staleness findings, cost/turns when availab
17
17
  CLI major reaches your workflow the moment it is promoted even though your `uses:` ref never changed — so a
18
18
  copy-pasted recipe that omits the input takes a major bump with no say in it. `^2` holds the major, needs no
19
19
  patch number to remember, and only wants a human decision at the next major. Pin an exact version
20
- (e.g. `version: "3.6.0"`) instead when you want byte-reproducible CI.
20
+ (e.g. `version: "3.7.0"`) instead when you want byte-reproducible CI.
21
21
 
22
22
  Reach for the manual multi-step form below only when you need per-step control the Action's inputs don't
23
23
  cover (a custom flag combination, a different runner matrix per step, or `lint`/`verify-cassettes` gated
@@ -73,7 +73,7 @@ sha256-*checked* but not hard-blocking on mismatch — it's advisory for an inte
73
73
  GitHub-hosted runners, no token/Docker/agent:
74
74
 
75
75
  ```yaml
76
- - run: npm i -g "cowork-harness@^3.6.0"
76
+ - run: npm i -g "cowork-harness@^3.7.0"
77
77
  - run: cowork-harness lint scenarios/*.yaml --strict --min-severity WARN
78
78
  # no silent false-greens. WITHOUT --strict this
79
79
  # step cannot fail on a WARN-class rule (e.g.
@@ -350,7 +350,7 @@ jobs:
350
350
  with: { node-version: '24' }
351
351
  - uses: actions/setup-python@v5
352
352
  with: { python-version: '3.x' } # python3 only — PyYAML is bundled with the linter
353
- - run: npm i -g "cowork-harness@^3.6.0"
353
+ - run: npm i -g "cowork-harness@^3.7.0"
354
354
  - run: cowork-harness lint scenarios/*.yaml # no-silent-false-green (needs python3; PyYAML bundled)
355
355
  - run: cowork-harness verify-cassettes cassettes/ --output-format json # privacy + staleness gate
356
356
  - run: cowork-harness replay cassettes/ --output-format json # token-free content/structure
@@ -379,7 +379,7 @@ jobs:
379
379
  echo "live=true" >> "$GITHUB_OUTPUT"
380
380
  fi
381
381
  - if: steps.guard.outputs.live == 'true'
382
- run: npm i -g "cowork-harness@^3.6.0"
382
+ run: npm i -g "cowork-harness@^3.7.0"
383
383
  - if: steps.guard.outputs.live == 'true'
384
384
  run: cowork-harness run scenarios/ --output-format json
385
385
  env:
@@ -1,6 +1,6 @@
1
1
  # Critique — the facts a plugin install can't otherwise reach
2
2
 
3
- Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`). This is **not** a trim of the full
3
+ Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`). This is **not** a trim of the full
4
4
  [`docs/critique.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/critique.md) (repo-only —
5
5
  flags, cost, reproduction discipline, known limitations all live there). This file covers exactly what a
6
6
  plugin install cannot otherwise discover: the run-dir artifact a harvester actually reads, the report's
@@ -103,19 +103,30 @@ fingerprints differently).
103
103
 
104
104
  ## What the evaluator was actually shown — `evidenceBudget`
105
105
 
106
- Skill-authored content (`SKILL.md`, every `references/**` file, `agents/<skill>.md`) ships **WHOLE, not
107
- rationed** — up to a **512 KiB combined ceiling** across all three together. A breach cuts **loudly**: the
108
- named file and byte counts are reported, never silent, never refused. The **transcript** is bounded
106
+ Skill-authored content ships **WHOLE, not rationed**: `SKILL.md`, every file under the skill's own
107
+ `references/**`, every dispatchable `agents/**.md`, and — for a multi-skill plugin — the shared
108
+ plugin-root `references/` files that the skill's own text, a packaged agent body, or the graded agent's
109
+ own read of it actually points at, up to a **512 KiB combined ceiling** across all four together. A
110
+ breach cuts **loudly**: the named file and byte counts are reported, never silent, never refused. The
111
+ plugin-root class is narrow by design — packaging the WHOLE shared tree instead was measured and
112
+ rejected: on a real 6-skill plugin it pushed one skill's corpus to 107% of the ceiling and cost that
113
+ skill's own `SKILL.md` 37,295 B, and `already-covered` (below) judges by presence with no notion of
114
+ authorship, so another skill's shared docs would silently excuse a real gap — a root file the rule leaves
115
+ out is reported in `corpusOmitted` below, never dropped silently. The **transcript** is bounded
109
116
  separately at **128 KiB**, cut **head+tail with an elided middle**, so a run's setup and its conclusion
110
117
  both survive a cut instead of just one end.
111
118
 
112
119
  You do not need a paid run to find out where you stand: **`cowork-harness lint-skill <skill-dir>` sizes
113
120
  your corpus against the same ceiling**, reporting `skill-corpus-near-evidence-ceiling` (INFO) from 80%
114
- and `skill-corpus-over-evidence-ceiling` (WARN, so it fails `--strict`) past it. It counts the three
115
- classes the ceiling governs — `SKILL.md`, every file under `references/` (**any extension**: the packager
116
- applies no extension filter, so JSON schemas and rule packs count toward your total), and a plugin
117
- skill's `agents/<name>.md`. It does not apply staging's git-tracked filter, so an untracked reference
118
- inflates the figure; `corpusCuts` below stays the authority.
121
+ and `skill-corpus-over-evidence-ceiling` (WARN, so it fails `--strict`) past it. It counts only the
122
+ skill-local classes — `SKILL.md`, every file under `references/` (**any extension**: the packager
123
+ applies no extension filter, so JSON schemas and rule packs count toward your total), and every
124
+ dispatchable `agents/**.md` a plugin skill can dispatch, and every plugin-root `references/` file the
125
+ skill links — the same four classes the packager ships. The one clause it cannot mirror is the
126
+ run-dependent one (a root reference included only because the graded agent READ it), since a static lint
127
+ has no run to read, so its number can under-report there. It also does not apply staging's git-tracked filter,
128
+ so an untracked skill-local reference inflates the figure the other way; `corpusCuts`/`corpusOmitted`
129
+ below stay the authority.
119
130
 
120
131
  The report's `evidenceBudget` object says exactly what was shown — read it instead of inferring budgets
121
132
  from `dist/` source:
@@ -124,7 +135,9 @@ from `dist/` source:
124
135
  |---|---|
125
136
  | `corpusBytes` | total skill-content bytes found, BEFORE any cut |
126
137
  | `corpusCeiling` | the 512 KiB combined ceiling |
138
+ | `corpusPackaged` | every corpus file whose CONTENT shipped into the corpus sections, by the same key `corpusCuts`/`corpusOmitted` use — so a reader can see exactly which sub-agent bodies and shared references the grade rests on. A file the ceiling zeroed is NOT listed; a partially cut one is, with its loss in `corpusCuts`; a placeholder for an unreadable file is never a corpus entry at all |
127
139
  | `corpusCuts` | per-file cut record — empty on every real skill; non-empty only once the ceiling is actually breached |
140
+ | `corpusOmitted` | plugin-root `references/` files present on the HOST under `<plugin>/references/` but NOT packaged (a raw walk, so an untracked file staging would not deliver is listed too — `alsoUntracked: true` says so, and the property is ABSENT rather than `false` when trackedness could not be evaluated), and why: `not-linked` (nothing in the skill's authored text, a packaged agent body, or the graded agent's own read points at it), `not-utf8` (fails to decode as clean UTF-8 — the skill's **own** `references/**` has no such filter), or `ambiguous-read` (the graded agent read a path that exists under both the skill's own `references/` and the plugin root's, so which tree it read cannot be attributed) |
128
141
  | `corpusExcluded` | skill files present on the host but never delivered to the agent (see below) |
129
142
  | `trimRecord` | which section the transcript trim shaved, and by how much |
130
143
  | `packageTruncated` | `true` if ANY section was cut — check this, not `corpusCuts`, for "was anything trimmed": a transcript-only cut leaves `corpusCuts` empty and would otherwise read as "nothing cut" |
@@ -1,6 +1,13 @@
1
1
  # Fidelity tiers & answer paths
2
2
 
3
- Self-contained reference. Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`).
3
+ Self-contained reference. Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`).
4
+
5
+ > **This page vs. the repo docs.** This is the **offline snapshot** that ships inside the installed
6
+ > plugin — it is self-contained on purpose. The repo carries four other fidelity views, each answering a
7
+ > different question: *which tier do I pick?* ([README → Fidelity tiers](https://github.com/yaniv-golan/cowork-harness/blob/main/README.md#fidelity-tiers-pick-per-scenario--per-ci-job)),
8
+ > *what does a tier enforce?* ([docs/boundary.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/boundary.md)), *what does it NOT reproduce?*
9
+ > ([docs/fidelity-gaps.md](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/fidelity-gaps.md)), and *why is it built this way?*
10
+ > ([DESIGN.md § 2](https://github.com/yaniv-golan/cowork-harness/blob/main/DESIGN.md#2-parity-matrix-per-tier)).
4
11
 
5
12
  ## Fidelity tiers (`fidelity:` in the scenario)
6
13
 
@@ -1,6 +1,6 @@
1
1
  # Scenario & session schema, assertion catalog, web_fetch, authoring gotchas
2
2
 
3
- Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.6.0`
3
+ Self-contained reference for authoring `cowork-harness` scenarios. Tracks `cowork-harness 3.7.0`
4
4
  (baseline `desktop-2.2553.1`). If your checkout is newer, prefer the live [`docs/scenario.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/scenario.md),
5
5
  [`docs/session.md`](https://github.com/yaniv-golan/cowork-harness/blob/main/docs/session.md), and `SPEC.md`.
6
6
 
@@ -2,7 +2,7 @@
2
2
 
3
3
  Each recipe composes facts that live scattered across SKILL.md and the other references into one
4
4
  decision path. Every one answers a question a real fleet owner had to work out the hard way.
5
- Tracks `cowork-harness 3.6.0` (baseline `desktop-2.2553.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
5
+ Tracks `cowork-harness 3.7.0` (baseline `desktop-2.2553.1`), same as SKILL.md's front-matter. Recipe 2's `resolved-tier`/`unverifiable-tier` staleness classes and
6
6
  Recipe 3's `init-redact` shipped in 0.24.0 and are part of the current feature set — no version gate
7
7
  needed if your CLI meets SKILL.md's version floor.
8
8
 
@@ -1762,23 +1762,342 @@ def _agent_name_from_frontmatter(md_path, yaml_mod):
1762
1762
  return None
1763
1763
 
1764
1764
 
1765
+ def _enumerate_plugin_agents(plugin_dir):
1766
+ """Every agent markdown under `<plugin_dir>/agents/`, RECURSIVELY, as (declared_name, Path) pairs.
1767
+
1768
+ Recursive on purpose. Claude Code discovers `agents/sub/x.md` (see src/run/analyze-skill.ts:981-983)
1769
+ and skill-hash.ts's `agentSkillName` already attributes both the flat and nested shapes, so a nested
1770
+ agent is dispatchable, hash-attributed and analyze-scanned. The old non-recursive `glob("*.md")` made
1771
+ it invisible to BOTH this linter's subagent_type resolution and the critique corpus.
1772
+
1773
+ This CAN change a severity in both directions, verified, not assumed:
1774
+ - a literal naming a nested agent stops being a `subagent-type-not-found-in-plugin` WARN (it really
1775
+ is dispatchable, so that WARN was a false positive); and
1776
+ - for a plugin whose `agents/` holds ONLY subdirectories, the agent set was previously EMPTY, which
1777
+ sent every same-plugin literal down the `plugin_agent_types` falsy branch in
1778
+ `_classify_subagent_type` to `subagent-type-unknown` (INFO). The set is now non-empty, so a
1779
+ genuinely typo'd literal is reported as the WARN it always was -- a true positive that used to be
1780
+ suppressed, but a NEW `--strict` failure on an unchanged tree."""
1781
+ agents_dir = Path(plugin_dir) / "agents"
1782
+ if not agents_dir.is_dir():
1783
+ return []
1784
+ yaml = _require_yaml()
1785
+ out = []
1786
+ for md in sorted(agents_dir.rglob("*.md")):
1787
+ if not md.is_file():
1788
+ continue
1789
+ out.append((_agent_name_from_frontmatter(md, yaml) or md.stem, md))
1790
+ return out
1791
+
1792
+
1765
1793
  def _resolve_plugin_agents(plugin_dir):
1766
1794
  """Resolve in-plugin agent types: return the set of valid `<plugin>:<agent>` subagent types defined WITHIN plugin_dir.
1767
- Reads the plugin name from plugin.json and each agents/*.md's `name:` frontmatter (filename stem
1795
+ Reads the plugin name from plugin.json and each agents/**.md's `name:` frontmatter (filename stem
1768
1796
  fallback). Returns an empty set (never crashes) when no plugin.json is found — a bare SKILL.md
1769
1797
  dir with no plugin manifest has nothing to resolve against."""
1770
1798
  plugin_name = _read_plugin_name(plugin_dir)
1771
1799
  if not plugin_name:
1772
1800
  return set()
1773
- agents_dir = Path(plugin_dir) / "agents"
1774
- if not agents_dir.is_dir():
1775
- return set()
1776
- yaml = _require_yaml()
1777
- types = set()
1778
- for md in sorted(agents_dir.glob("*.md")):
1779
- agent_name = _agent_name_from_frontmatter(md, yaml) or md.stem
1780
- types.add(f"{plugin_name}:{agent_name}")
1781
- return types
1801
+ return {f"{plugin_name}:{name}" for name, _ in _enumerate_plugin_agents(plugin_dir)}
1802
+
1803
+
1804
+ def _literal_to_agent_name(value, plugin_name):
1805
+ """Map a pinned `subagent_type` literal to a declared agent name WITHIN this plugin, or None.
1806
+
1807
+ A colon-bearing literal is namespaced (`<plugin>:<agent>`) and its prefix MUST equal this plugin's
1808
+ name, or a cross-plugin `other:market-sizing` would match THIS plugin's `market-sizing`. A bare
1809
+ literal matches a declared name directly. MIRRORED in src/critique/resolve-agents.ts; the two are
1810
+ pinned by test/fixtures/dispatchable-agents.json, which both sides execute."""
1811
+ if ":" not in value:
1812
+ return value or None
1813
+ if not plugin_name:
1814
+ return None
1815
+ prefix, _, suffix = value.partition(":")
1816
+ return suffix if prefix == plugin_name and suffix else None
1817
+
1818
+
1819
+ def _resolve_corpus_agents(skill_dir):
1820
+ """Every agent file the critique packager will put in THIS skill's evidence corpus, as a sorted list
1821
+ of Paths. Mirrors `resolveDispatchableAgents` in src/critique/resolve-agents.ts -- union of:
1822
+ 1. `agents/<skillname>.md` by FILENAME (what the packager did through 3.6.0),
1823
+ 2. every in-plugin agent a pinned `subagent_type` literal in SKILL.md / references/** resolves to,
1824
+ 3. every agent whose DECLARED name equals the skill name,
1825
+ then a transitive closure over the resolved agent bodies (an agent that dispatches another agent).
1826
+
1827
+ The union is the SAFETY property: this can never resolve to less than the single agents/<skill>.md
1828
+ the old sizing counted, so a skill's reported corpus never shrinks under this change."""
1829
+ skill_dir = Path(skill_dir)
1830
+ # Walk UP for the manifest, exactly as the TypeScript side does (`findEnclosingPluginDir`). The
1831
+ # old `parent.name == "skills"` LAYOUT rule disagreed with the packager outside the
1832
+ # `<root>/skills/<name>` shape: a skill at `<root>/foo/SKILL.md` with a manifest at `<root>` made the
1833
+ # packager size a corpus this sized as ZERO, which is the packager-vs-linter divergence 0003e38
1834
+ # closed in one shape and left open in another.
1835
+ # Nearest MANIFEST first (matching the TypeScript `findEnclosingPluginDir`), falling back to the
1836
+ # `<root>/skills/<name>` LAYOUT. Manifest-only diverged from the packager for a manifest-less plugin,
1837
+ # where the TS side still enumerates agents because it is handed the root explicitly; layout-only
1838
+ # diverged for a skill that is not under `skills/` but does have a manifest above it, which the
1839
+ # packager sizes and this sized as ZERO. Neither rule alone matches; the union does.
1840
+ plugin_dir = _find_enclosing_plugin_dir(skill_dir / "SKILL.md") or (
1841
+ skill_dir.parent.parent if skill_dir.parent.name == "skills" else None
1842
+ )
1843
+ if plugin_dir is None or not (plugin_dir / "agents").is_dir():
1844
+ return []
1845
+ all_agents = _enumerate_plugin_agents(plugin_dir)
1846
+ if not all_agents:
1847
+ return []
1848
+ plugin_name = _read_plugin_name(plugin_dir)
1849
+ skill_name = skill_dir.name
1850
+ picked = {}
1851
+ sources = [] # declared before `add`: every picked agent is itself a dispatch SOURCE
1852
+
1853
+ def add(path):
1854
+ """Record an agent AND queue its body as a dispatch source.
1855
+
1856
+ Clauses 1 and 3 used to record without queueing, so a skill's own primary agent -- the most
1857
+ likely dispatcher of a second agent -- was never scanned and the transitive closure did not run
1858
+ for the dominant real shape. `scanned` below still guarantees termination."""
1859
+ if str(path) in picked:
1860
+ return
1861
+ picked[str(path)] = path
1862
+ sources.append(path)
1863
+
1864
+ for name, md in all_agents:
1865
+ # clause 1 (filename) and clause 3 (declared name)
1866
+ if md == plugin_dir / "agents" / f"{skill_name}.md" or name == skill_name:
1867
+ add(md)
1868
+
1869
+ skill_md = skill_dir / "SKILL.md"
1870
+ if skill_md.is_file():
1871
+ sources.append(skill_md)
1872
+ refs = skill_dir / "references"
1873
+ if refs.is_dir():
1874
+ sources.extend(p for p in sorted(refs.rglob("*")) if p.is_file())
1875
+ scanned = set()
1876
+ while sources:
1877
+ src = sources.pop(0)
1878
+ if str(src) in scanned:
1879
+ continue
1880
+ scanned.add(str(src))
1881
+ try:
1882
+ text = src.read_text(encoding="utf-8")
1883
+ except Exception:
1884
+ continue
1885
+ for m in _SUBAGENT_TYPE_RE.finditer(text):
1886
+ agent_name = _literal_to_agent_name(m.group(1), plugin_name)
1887
+ if not agent_name:
1888
+ continue
1889
+ for name, md in all_agents:
1890
+ if name == agent_name:
1891
+ add(md) # queues the body too -- transitive: this agent may dispatch another
1892
+ return [picked[k] for k in sorted(picked)]
1893
+
1894
+
1895
+ # --------------------------------------------------------------------------- #
1896
+ # plugin-root references/ resolution, for corpus sizing only
1897
+ #
1898
+ # Mirrors `resolveRootReferences` in src/critique/resolve-references.ts (clauses 1 + 2 only -- see the
1899
+ # docstring on `_resolve_corpus_root_references` below for why clause 3 cannot be mirrored here). The TS
1900
+ # packager now puts a multi-skill plugin's SHARED plugin-root `references/` files into the evaluator
1901
+ # corpus when the graded skill's own text (or a sub-agent it can dispatch) points at them, so
1902
+ # `_lint_skill_corpus_size` must size the same files or the ceiling warning under-reports exactly the
1903
+ # plugins this feature targets.
1904
+ # --------------------------------------------------------------------------- #
1905
+
1906
+ # Trailing punctuation a prose or markdown token picks up, stripped as a RUN and WITHOUT requiring
1907
+ # balance -- byte-identical rule to TRAILING_PUNCT in resolve-references.ts.
1908
+ _TRAILING_PUNCT_RE = re.compile(r"""[)\]}>.,;:!?'"]+$""")
1909
+ _LEADING_QUOTE_RE = re.compile(r"""^["'<([]+""")
1910
+ _HASH_QUERY_RE = re.compile(r"[#?].*$")
1911
+ _BACKTICK_SPAN_RE = re.compile(r"`([^`]+)`")
1912
+ _MD_LINK_TARGET_RE = re.compile(r"\]\(([^)\s]+)")
1913
+ _HREF_TARGET_RE = re.compile(r"""href=["']([^"']+)["']""")
1914
+
1915
+
1916
+ def _line_tokens(line):
1917
+ """Candidate tokens on one line: backticked spans, markdown/href link targets, and bare whitespace-
1918
+ delimited runs. Only those containing `/` are resolved as paths (`pathish`); the bare ones matter only
1919
+ for the ARMED basename pass. Mirrors `lineTokens()` in resolve-references.ts."""
1920
+ pathish = []
1921
+ bare = []
1922
+ raw = (
1923
+ _BACKTICK_SPAN_RE.findall(line)
1924
+ + _MD_LINK_TARGET_RE.findall(line)
1925
+ + _HREF_TARGET_RE.findall(line)
1926
+ + line.split()
1927
+ )
1928
+ for t0 in raw:
1929
+ t = _HASH_QUERY_RE.sub("", t0)
1930
+ t = _LEADING_QUOTE_RE.sub("", t)
1931
+ t = _TRAILING_PUNCT_RE.sub("", t)
1932
+ t = t.replace("\\", "/")
1933
+ if not t:
1934
+ continue
1935
+ if "/" in t:
1936
+ pathish.append(t)
1937
+ else:
1938
+ bare.append(t)
1939
+ return pathish, bare
1940
+
1941
+
1942
+ def _token_to_abs(token, plugin_root, plugin_name, file_dir):
1943
+ """Resolve one path-ish token to an absolute (not yet realpath'd) host path. `${CLAUDE_PLUGIN_ROOT}`
1944
+ and a leading `<pluginName>/` both mean the plugin root; everything else is relative to the
1945
+ directory of the file the token was found in. Mirrors `tokenToAbs()`."""
1946
+ if "${CLAUDE_PLUGIN_ROOT}" in token:
1947
+ return os.path.normpath(token.replace("${CLAUDE_PLUGIN_ROOT}", plugin_root))
1948
+ if plugin_name and token.startswith(plugin_name + "/"):
1949
+ return os.path.normpath(os.path.join(plugin_root, token[len(plugin_name) + 1 :]))
1950
+ if os.path.isabs(token):
1951
+ return os.path.normpath(token)
1952
+ return os.path.normpath(os.path.join(file_dir, token))
1953
+
1954
+
1955
+ def _realpath_strict(path):
1956
+ """Mirrors node's `realpathSync`: resolves symlinks and raises if any path component doesn't exist
1957
+ (unlike `os.path.realpath`, which resolves lexically even against a nonexistent path)."""
1958
+ return str(Path(path).resolve(strict=True))
1959
+
1960
+
1961
+ def _links_in(text, file_dir, plugin_root, plugin_name, by_basename):
1962
+ """Every plugin-root reference `text` (from a file at `file_dir`) points at, mapped to the 1-based
1963
+ line it was pointed at from. Mirrors `linksIn()` -- including the ARMING rule: a token resolving to
1964
+ the plugin-root `references/` DIRECTORY ITSELF arms bare-basename matching for THAT LINE ONLY; it
1965
+ never recurses and never carries to the next line."""
1966
+ try:
1967
+ refs_dir = _realpath_strict(os.path.join(plugin_root, "references"))
1968
+ except OSError:
1969
+ return {}
1970
+ hits = {}
1971
+ for i, line in enumerate(text.splitlines()):
1972
+ pathish, bare = _line_tokens(line)
1973
+ armed = False
1974
+ for t in pathish:
1975
+ abs_path = _token_to_abs(t, plugin_root, plugin_name, file_dir)
1976
+ try:
1977
+ rp = _realpath_strict(abs_path)
1978
+ except OSError:
1979
+ continue # most prose tokens are not paths at all
1980
+ if rp == refs_dir:
1981
+ armed = True
1982
+ continue
1983
+ if rp.startswith(refs_dir + os.sep):
1984
+ rel = "references/" + rp[len(refs_dir) + 1 :].replace(os.sep, "/")
1985
+ if rel not in hits:
1986
+ hits[rel] = i + 1
1987
+ if not armed:
1988
+ continue
1989
+ for b in bare:
1990
+ rel = by_basename.get(b)
1991
+ if rel is not None and rel not in hits:
1992
+ hits[rel] = i + 1
1993
+ return hits
1994
+
1995
+
1996
+ def _is_clean_utf8(path):
1997
+ """Valid UTF-8, decided by the decoder rather than by scanning the decoded string -- mirrors
1998
+ `isCleanUtf8()` (a `TextDecoder("utf-8", { fatal: true })` decode). `bytes.decode("utf-8")` is
1999
+ strict by default, so a decode failure (not a scan of the result) is what disqualifies a file."""
2000
+ try:
2001
+ path.read_bytes().decode("utf-8")
2002
+ return True
2003
+ except (OSError, UnicodeDecodeError):
2004
+ return False
2005
+
2006
+
2007
+ def _resolve_root_references(skill_dir, plugin_dir, agents):
2008
+ """Low-level worker, taking an explicit `plugin_dir` so the same-directory guard below is directly
2009
+ unit-testable regardless of how a caller derives `plugin_dir`. Mirrors `resolveRootReferences()`,
2010
+ clauses 1 + 2 only:
2011
+ 1. the skill's own authored text (SKILL.md + its own `references/**`), and
2012
+ 2. every sub-agent body already in the corpus (`agents`, from `_resolve_corpus_agents`).
2013
+
2014
+ Clause 3 in the TS packager (files the graded AGENT actually read during a live run) is
2015
+ run-dependent -- it needs a recorded turn's access log -- and CANNOT be mirrored by a static lint
2016
+ pass with no run to inspect, so it is deliberately NOT reproduced here. This makes this function's
2017
+ count a floor, never an exact match, relative to what a real critique run would package; that is the
2018
+ same "warn early, the packager's report is the authority" posture the rest of this linter already
2019
+ has for the ceiling.
2020
+
2021
+ Two further known, deliberate divergences (not fixed here -- out of scope for this change):
2022
+ - byte counting: the packager measures UTF-8-DECODED string length while `_lint_skill_corpus_size`
2023
+ (like its pre-existing SKILL.md/references/agents counting) sums `stat().st_size` -- raw bytes
2024
+ on disk. These agree for pure single-byte UTF-8 text and diverge slightly otherwise.
2025
+ - cleanliness filtering: this module's existing `references/` walks (`rglob("*")`, used here and
2026
+ for the skill's own references/** in `_lint_skill_corpus_size`) have no git-tracked-set or
2027
+ symlink-containment filter, unlike `listSkillFilesRecursive` in corpus-walk.ts. This function
2028
+ reuses that same untared posture for consistency with the rest of this linter, not because it is
2029
+ provably safe against a hostile references/ tree."""
2030
+ skill_dir = Path(skill_dir)
2031
+ plugin_dir = Path(plugin_dir)
2032
+ # SKIP (not "dedupe") the standalone-skill shape: those files are already packaged as skill-local,
2033
+ # and running this pass too would double-count them under two different display keys.
2034
+ try:
2035
+ same = skill_dir.resolve(strict=True) == plugin_dir.resolve(strict=True)
2036
+ except OSError:
2037
+ same = skill_dir.resolve() == plugin_dir.resolve()
2038
+ if same:
2039
+ return []
2040
+ refs_root = plugin_dir / "references"
2041
+ if not refs_root.is_dir():
2042
+ return []
2043
+ all_files = sorted(p for p in refs_root.rglob("*") if p.is_file())
2044
+ if not all_files:
2045
+ return []
2046
+ plugin_name = _read_plugin_name(plugin_dir) or plugin_dir.name
2047
+ # Later (sorted) entries win on a real basename collision -- matches the TS `Map` construction,
2048
+ # where each duplicate key overwrites the previous.
2049
+ by_basename = {}
2050
+ for p in all_files:
2051
+ rel = "references/" + p.relative_to(refs_root).as_posix()
2052
+ by_basename[p.name] = rel
2053
+
2054
+ linked = set()
2055
+ sources = [skill_dir / "SKILL.md"]
2056
+ local_refs = skill_dir / "references"
2057
+ if local_refs.is_dir():
2058
+ sources.extend(p for p in sorted(local_refs.rglob("*")) if p.is_file())
2059
+ sources.extend(Path(a) for a in agents)
2060
+ for src in sources:
2061
+ try:
2062
+ text = src.read_text(encoding="utf-8")
2063
+ except (OSError, UnicodeDecodeError):
2064
+ continue
2065
+ for rel in _links_in(text, str(src.parent), str(plugin_dir), plugin_name, by_basename):
2066
+ linked.add(rel)
2067
+
2068
+ packaged = []
2069
+ for p in all_files:
2070
+ rel = "references/" + p.relative_to(refs_root).as_posix()
2071
+ if rel not in linked:
2072
+ continue
2073
+ if not _is_clean_utf8(p):
2074
+ continue
2075
+ packaged.append(p)
2076
+ return packaged
2077
+
2078
+
2079
+ def _resolve_corpus_root_references(skill_dir, agents):
2080
+ """Every plugin-root `references/` file the critique packager will put in THIS skill's evidence
2081
+ corpus, as a sorted list of Paths. Derives `plugin_dir` the same way `_resolve_corpus_agents` does
2082
+ (the `skills/<name>` multi-skill-plugin shape); a standalone skill (no `skills/` parent) has no
2083
+ plugin-root references pass and is unaffected."""
2084
+ skill_dir = Path(skill_dir)
2085
+ # Walk UP for the manifest, exactly as the TypeScript side does (`findEnclosingPluginDir`). The
2086
+ # old `parent.name == "skills"` LAYOUT rule disagreed with the packager outside the
2087
+ # `<root>/skills/<name>` shape: a skill at `<root>/foo/SKILL.md` with a manifest at `<root>` made the
2088
+ # packager size a corpus this sized as ZERO, which is the packager-vs-linter divergence 0003e38
2089
+ # closed in one shape and left open in another.
2090
+ # Nearest MANIFEST first (matching the TypeScript `findEnclosingPluginDir`), falling back to the
2091
+ # `<root>/skills/<name>` LAYOUT. Manifest-only diverged from the packager for a manifest-less plugin,
2092
+ # where the TS side still enumerates agents because it is handed the root explicitly; layout-only
2093
+ # diverged for a skill that is not under `skills/` but does have a manifest above it, which the
2094
+ # packager sizes and this sized as ZERO. Neither rule alone matches; the union does.
2095
+ plugin_dir = _find_enclosing_plugin_dir(skill_dir / "SKILL.md") or (
2096
+ skill_dir.parent.parent if skill_dir.parent.name == "skills" else None
2097
+ )
2098
+ if plugin_dir is None:
2099
+ return []
2100
+ return _resolve_root_references(skill_dir, plugin_dir, agents)
1782
2101
 
1783
2102
 
1784
2103
  def cmd_resolve_agent_types(args):
@@ -1938,9 +2257,12 @@ _EVIDENCE_CORPUS_NOTICE_RATIO = 0.8
1938
2257
  def _lint_skill_corpus_size(md_path):
1939
2258
  """Total skill-authored bytes against the critique evidence ceiling.
1940
2259
 
1941
- Counts the SAME THREE CLASSES the ceiling governs: SKILL.md, every file under references/ (any
1942
- extension -- the packager applies no extension filter, so JSON schemas and rule packs count), and,
1943
- for a skill inside a multi-skill plugin, the invoked skill's <root>/agents/<name>.md.
2260
+ Counts the SAME FOUR CLASSES the ceiling governs: SKILL.md, every file under references/ (any
2261
+ extension -- the packager applies no extension filter, so JSON schemas and rule packs count), for a
2262
+ skill inside a multi-skill plugin every <root>/agents/**.md the skill can dispatch, and -- also for a
2263
+ multi-skill plugin -- every <root>/references/** file the skill's authored text (or a dispatchable
2264
+ agent) links (see `_resolve_corpus_root_references`; the run-dependent fourth TS clause, files the
2265
+ agent actually read live, is out of scope for a static lint).
1944
2266
 
1945
2267
  Omitting the agents md is not a rounding error: a plugin whose SKILL.md + references sit in the INFO
1946
2268
  band while the agents md carries the corpus past the ceiling reported INFO and PASSED --strict on
@@ -1956,13 +2278,18 @@ def _lint_skill_corpus_size(md_path):
1956
2278
  refs = skill_dir / "references"
1957
2279
  if refs.is_dir():
1958
2280
  files.extend(p for p in sorted(refs.rglob("*")) if p.is_file())
1959
- # Multi-skill plugin layout: skillDir is <root>/skills/<name> and the invoked skill's sub-agent
1960
- # system prompt is <root>/agents/<name>.md -- the same resolution the critique command performs.
1961
- # A standalone skill (no `skills/` parent) has no agents md and is unaffected.
1962
- if skill_dir.parent.name == "skills":
1963
- agents_md = skill_dir.parent.parent / "agents" / f"{skill_dir.name}.md"
1964
- if agents_md.is_file():
1965
- files.append(agents_md)
2281
+ # Multi-skill plugin layout: skillDir is <root>/skills/<name> and the skill's dispatchable sub-agent
2282
+ # prompts live under <root>/agents/ -- the same resolution the critique command performs, via the
2283
+ # shared `_resolve_corpus_agents`. This counted exactly ONE file (agents/<name>.md) while the packager
2284
+ # shipped N, which under-reported the ceiling for precisely the multi-agent plugins that need the
2285
+ # warning most. A standalone skill (no `skills/` parent) has no agents and is unaffected.
2286
+ agents = _resolve_corpus_agents(skill_dir)
2287
+ files.extend(agents)
2288
+ # The packager also puts a multi-skill plugin's SHARED plugin-root `references/` files into the
2289
+ # evidence corpus when the graded skill's own text (or a dispatchable agent) points at them -- see
2290
+ # `_resolve_corpus_root_references`. Omitting this would under-report the ceiling for exactly the
2291
+ # plugins that feature targets, the same failure mode the agents fix above was written to close.
2292
+ files.extend(_resolve_corpus_root_references(skill_dir, agents))
1966
2293
  for p in files:
1967
2294
  try:
1968
2295
  total += p.stat().st_size
@@ -1977,9 +2304,11 @@ def _lint_skill_corpus_size(md_path):
1977
2304
  f"skill content is {total:,} B ({pct:.0f}% of the {_EVIDENCE_CORPUS_CEILING:,} B critique "
1978
2305
  f"evidence ceiling) — a critique will cut it before grading.",
1979
2306
  "Split or trim the largest references/ files. This counts SKILL.md + references/** + "
1980
- "agents/<skill>.md, the same three classes the ceiling governs; it does not apply "
1981
- "staging's git-tracked filter, so an untracked reference inflates it. The critique "
1982
- "report's corpusCuts names exactly which files lose bytes.",
2307
+ "agents/** (every agent the skill can dispatch) + linked plugin-root references/** "
2308
+ "(the shared files the skill's own text or a dispatchable agent points at), the same "
2309
+ "classes the ceiling governs; it does not apply staging's git-tracked filter, so an "
2310
+ "untracked reference inflates it. The critique report's corpusCuts names exactly which "
2311
+ "files lose bytes.",
1983
2312
  str(skill_dir),
1984
2313
  )
1985
2314
  ]