@ssheleg/make-skill 0.25.1 → 0.25.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,105 @@
1
1
  # Changelog
2
2
 
3
+ ## v0.25.3 — the runner we said did not exist, and the date a claim about someone else now carries
4
+
5
+ - **B-122 closed: `references/authoring.md` told authors for four weeks that
6
+ there was no built-in eval runner, and prescribed a home-rolled suite on that
7
+ basis.** `claude plugin eval` had shipped. The section now names it, its case
8
+ format (`evals/<case>/case.yaml`, or `prompt.md` beside `graders/*.md`), and
9
+ the fact that `--ablation with-without` performs steps 3 and 5 of the very
10
+ procedure the paragraph prescribes by hand — plus the detail that it publishes
11
+ its HTML report to claude.ai unless `--no-publish` is passed, which is the
12
+ wrong default for an unreleased skill. **The replacement is a measurement, not
13
+ a correction of one absence into another**: on 2026-08-31, Claude Code 2.1.236,
14
+ every path (`eval`, `eval init`, `eval init --bare`, `eval <target>`) prints
15
+ ``plugin eval` is currently in early access`, writes nothing and exits 1, with
16
+ no settings key, environment variable or feature flag turning it on. That is
17
+ why `test/evals/` remains what runs here, and the paragraph says which day it
18
+ should be migrated. Two derived statements went with it: `distribution.md`'s
19
+ layout comment called the suite *"run by a human"*, and `retrofit.md`'s item 8
20
+ asked only that evaluations *exist* — the state MSK-01 had just spent a whole
21
+ run closing.
22
+ - **The class, not the sentence: `THIRD_PARTY_CLAIMS` in `test/validate.py`.**
23
+ A load-bearing claim about someone else's tool is registered beside the command
24
+ that re-checks it, and the guard fails unless the claim, that command and an
25
+ ISO date still sit within eight lines of each other — and fails equally when a
26
+ registered claim has been deleted from the doctrine but left in the registry,
27
+ so the registry cannot outlive what it describes. The skill's own rule *"No
28
+ time-sensitive statements"* sits 100 lines above the sentence that broke it and
29
+ nothing enforced it; this is the enforcement.
30
+ - **A generic detector was built first, measured, and rejected with the number.**
31
+ Over the shipped doctrine a pattern for absence-claims flagged **10 lines, of
32
+ which roughly half are legitimate** — the fallback rows of
33
+ `host-capabilities.md` ("on any other host there are no subagents") are made of
34
+ that sentence shape on purpose. A guard with a ~50% false-positive rate is the
35
+ over-defense that gets switched off, so the explicit registry replaced it; the
36
+ reasoning is recorded in the code beside the table rather than lost.
37
+ - **The three plants are in CI, and the group count still computes.** Added as
38
+ cases to the existing *claims, budgets and runnable commands* step rather than
39
+ as a new step, so `grep -c 'name: Negative self-test'` stays at **9** and
40
+ `CONTRIBUTING.md`'s *"9 negative self-test groups"* — a claim this validator
41
+ compares against the workflow — remains true. Each case asserts its own
42
+ expected message. The first local run of the guard **passed its plants for the
43
+ wrong reason**: the claim wraps a line break and the check was line-anchored,
44
+ so all three plants tripped the registry-rot branch instead of their own. The
45
+ matcher is whitespace-tolerant now, and every plant in the block asserts it
46
+ changed something before the validator is asked anything.
47
+
48
+ ## v0.25.2 — the evals have been run, and the discovery check runs somewhere
49
+
50
+ - **MSK-01 closed: the evaluation suite was executed against two models, and the
51
+ numbers live in `test/evals/RESULTS.md` as a dated run (2026-08-31), not a
52
+ claim.** All 22 trigger queries probed blind per model — haiku 22/22, sonnet
53
+ 20/22 (q01 and q10 answered "none": under-triggering on positives, ZERO false
54
+ positives on the near-miss negatives across both models) — and all four
55
+ behavioural scenarios executed and scored per `expected_behavior` line (sonnet
56
+ 31/35, haiku 24/35; the s01 repositories were re-audited with the bundled
57
+ auditor rather than trusted from the agent's report). The Method section
58
+ states the harness and its limits: one probe per query instead of the
59
+ README's fired/3, Agent-tool subagents on the author's machine whose global
60
+ config carries the family routing map (a bias that makes the clean negatives
61
+ the strongest result), model aliases rather than exact ids, and two probes
62
+ counted with a stated non-compliance caveat. Headline finding: the standing
63
+ worry (turn-stealing) did not materialize; the observed failure is sonnet
64
+ under-triggering on two positive categories, and s03-haiku answered a
65
+ surfaces question without consulting the installed skill — selection and
66
+ consultation measured apart, on purpose.
67
+ - **MSK-03 closed: the discovery check the body demands now runs in CI.**
68
+ SKILL.md's stray-SKILL.md gotcha ends "verify with the skills CLI's `--list`"
69
+ — and nothing ran it anywhere: the operator machine's hygiene guard refuses
70
+ the bare skills CLI on family members by design, and CI never asked. The new
71
+ `skills-cli-discovery` job in `validate.yml` runs the listing against
72
+ `ssheleg/make-skill` and asserts it serves exactly ONE skill, named
73
+ make-skill, with ANSI stripped before asserting. Its own job, like
74
+ `claude-plugin-validate`, so a registry outage cannot mask an
75
+ offline-validator failure; it reads the repository's default branch, so the
76
+ post-merge push run and the release's `workflow_call` on the tag assert the
77
+ published tree.
78
+ - **MSK-04 closed: body headroom restored by displacement — the words moved,
79
+ none were lost.** The Cyrillic budget measurement (1.9–2.3 chars/token,
80
+ 3408→1885) moved to `authoring.md` → *Content guidelines*; the six skeleton
81
+ filenames moved into `distribution.md`'s layout tree; the five verified facts
82
+ and the machine-refresh command sequence are now referenced at their single
83
+ home in `distribution.md`'s checklists instead of being restated; the
84
+ long-file rule already lived in `authoring.md` and the body now points there.
85
+ Re-measured with the bundled auditor: body ~4742/4750 tokens (8 spare) →
86
+ ~4677/4750 (73 spare), after also paying for the next bullet. The
87
+ description is untouched at 965/970 — its headroom is spent only on trigger
88
+ changes, and this release makes none.
89
+ - **XF-10: the MCP/A2A carve-out now runs both directions at the layer an agent
90
+ reads first.** The References-table rows for `mcp.md` and `a2a.md` carry "the
91
+ protocol wire itself → `agent-interop` (agent-stack)" — the two reference
92
+ files' own headers already said it, and now the router-level table says it
93
+ too, so a reader deciding what to load learns the boundary before opening
94
+ either file.
95
+ - Also: `test/evals/README.md`'s file table said "20 realistic queries" over a
96
+ `triggers.json` holding 22 — the same drift class as MSK-02, in a wording no
97
+ guard could parse. Reworded to the checkable "22 trigger queries" form, which
98
+ puts the sentence under `test/validate.py`'s counted-claims sweep instead of
99
+ under anyone's memory. `SKILL-CARD.md`'s evaluation status and posture
100
+ sections now describe the executed run and the replication still owed,
101
+ instead of "never executed".
102
+
3
103
  ## v0.25.1 — the rule keeps its default, and the family's exception is written down
4
104
 
5
105
  - **The same-name command rule now carries its recorded exception** (operator
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@ssheleg/make-skill",
3
- "version": "0.25.1",
4
- "description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way — conformance to the Agent Skills open standard AND Anthropic's platform rules (front-matter limits, disclosure budgets, per-surface runtime limits, the Skills API, evals) plus the Claude Code plugin reference (manifest schemas, component layout, claude plugin validate --strict), marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, the review checklist for third-party skills, and MCP / A2A rules for protocol-connected skills. This package is the installer CLI.",
3
+ "version": "0.25.3",
4
+ "description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way \u2014 conformance to the Agent Skills open standard AND Anthropic's platform rules (front-matter limits, disclosure budgets, per-surface runtime limits, the Skills API, evals) plus the Claude Code plugin reference (manifest schemas, component layout, claude plugin validate --strict), marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, the review checklist for third-party skills, and MCP / A2A rules for protocol-connected skills. This package is the installer CLI.",
5
5
  "keywords": [
6
6
  "skill",
7
7
  "plugin",
@@ -3,7 +3,7 @@
3
3
  "name": "make-skill",
4
4
  "displayName": "Make Skill",
5
5
  "description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way: conformance to the Agent Skills open standard, Anthropic's platform rules (surfaces, Skills API, evals) and the Claude Code plugin reference, marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, end-to-end first publish, the review checklist for third-party skills, plus MCP / A2A references for protocol-connected skills.",
6
- "version": "0.25.1",
6
+ "version": "0.25.3",
7
7
  "author": {
8
8
  "name": "ssheleg",
9
9
  "url": "https://x.com/sshlg93"
@@ -5,7 +5,7 @@ license: MIT
5
5
  compatibility: Authoring works on any agent. The bundled scripts/ need python3. Publishing steps need git, gh, node and npm; the plugin gates need the claude CLI. Not usable on the Claude API surface, which has no network and no runtime package install.
6
6
  metadata:
7
7
  author: ssheleg
8
- version: "0.25.1"
8
+ version: "0.25.3"
9
9
  homepage: https://github.com/ssheleg/make-skill
10
10
  ---
11
11
 
@@ -27,8 +27,8 @@ orchestrator, release automation). **make-skill itself** is built to this canon.
27
27
  | `references/host-capabilities.md` | shipping a **hook, subagent, command, script or MCP dependency** — what each buys and costs, hook events and exit codes, the degradation clauses |
28
28
  | `references/claude-code-plugin.md` | anything shipping as a **Claude Code plugin/marketplace** — manifest schemas, component layout, path variables, `validate` failures |
29
29
  | `references/distribution.md` | the repo layout, releases, and all five channels — plugin, skills CLI, npx, Cursor, umbrella family repo |
30
- | `references/mcp.md` | skill vs **MCP** server, declaring the dependency, consent and untrusted-output rules |
31
- | `references/a2a.md` | the skill spans two autonomous agents (**A2A**) — choosing it, the two meanings of "skill", driving a peer safely |
30
+ | `references/mcp.md` | skill vs **MCP** server, declaring the dependency, consent and untrusted-output rules; the protocol wire itself → `agent-interop` (agent-stack) |
31
+ | `references/a2a.md` | the skill spans two autonomous agents (**A2A**) — choosing it, the two meanings of "skill", driving a peer safely; the wire itself → `agent-interop` (agent-stack) |
32
32
 
33
33
  Missing from this copy? Raw fallback:
34
34
  `raw.githubusercontent.com/ssheleg/make-skill/main/plugins/make-skill/skills/make-skill/references/<file>`
@@ -81,8 +81,8 @@ differences and the checklist: `references/agent-skills-spec.md`.
81
81
  99% of budget turns the next correction into a fight with the validator.
82
82
  Heavier material goes to `references/`, `scripts/`, `assets/` INSIDE the skill
83
83
  dir, one level deep, each linked from the body with a stated load trigger
84
- ("read X when Y") — never a bare "see references/". Reference files >100 lines
85
- open with a `## Contents` list: a partial `head` read is what agents get.
84
+ ("read X when Y") — never a bare "see references/". Long-file and layout rules:
85
+ `references/authoring.md` → *Content guidelines*.
86
86
  - Gotchas stay in `SKILL.md`: the agent can't know to open a file about a trap
87
87
  it doesn't know exists.
88
88
  - **Write for the weakest surface you claim** (`references/surfaces.md`): the
@@ -97,9 +97,8 @@ House additions on top of the spec:
97
97
  survives in four places only, where the string itself is the point: a
98
98
  **trigger phrase**, a **refusal phrase** (the operator types both — translated,
99
99
  they no longer match what was said), a **proper noun**, a **language example**
100
- (`«вы»/«ты»`). A budget rule before a style one: Russian encodes at 1.9–2.3
101
- chars/token against English's 5.0 (`cl100k`), and rewriting the eight ssheleg
102
- routers into English cut them **3408 → 1885 tokens** with no loss of meaning.
100
+ (`«вы»/«ты»`). A budget rule before a style one — the measured cost of
101
+ Cyrillic prose is in `references/authoring.md` → *Content guidelines*.
103
102
  - `description` starts "Use when …" and lists concrete trigger phrases — English
104
103
  AND Russian (user works in both). A skill nobody triggers is dead weight. Hold
105
104
  **5% headroom here too** (≤970 of 1024): a near-miss neighbour forces a "NOT
@@ -153,10 +152,9 @@ A fallback you know but did not write is not a fallback.
153
152
  Retrofit), wrapped as `bin/make-skill-audit` for Claude Code. Beside them:
154
153
  `hooks/` (PostToolUse, silent unless a `SKILL.md` was written),
155
154
  `commands/skill-audit.md` (deliberately NOT the skill's name),
156
- `agents/skill-auditor.md`, and six skeletons: `assets/SKILL.template.md`,
157
- `assets/plugin.template.json`, `assets/marketplace.template.json`,
158
- `assets/hooks.template.json`, `assets/agent.template.md`,
159
- `assets/command.template.md`.
155
+ `agents/skill-auditor.md`, and six `assets/*.template.*` skeletons — one per
156
+ component, filenames in `references/distribution.md` → *The distributable repo
157
+ layout*.
160
158
 
161
159
  ## Create (personal)
162
160
 
@@ -196,8 +194,7 @@ regardless:
196
194
  Take it ALL the way, no half-done handoffs. Only the first publish needs a human
197
195
  (npm 2FA); **arming CI publishing is part of shipping**, so the second does not.
198
196
  **The 11-step sequence is in `references/distribution.md` → *First publish*.**
199
- Done = five VERIFIED facts: repo + CI green, npm resolvable via npx, plugin
200
- installed, skills-CLI discovery working, next tag publishing without a human.
197
+ Done = the five VERIFIED facts in that sequence's step 10 — nothing assumed.
201
198
 
202
199
  ## Retrofit (bring an existing skill/repo up to standard)
203
200
 
@@ -310,9 +307,8 @@ session:
310
307
  - everything green BEFORE the tag: `python3 test/validate.py` plus BOTH
311
308
  `claude plugin validate … --strict` runs;
312
309
  - **refresh THIS machine's global installs as Definition of Done** (per global
313
- `~/.claude/CLAUDE.md`): `claude plugin marketplace update <name>` →
314
- `claude plugin update <name>@<name>` → `npx skills update <name> --global
315
- --yes && rm -f ~/.claude/skills/<name>`, then remind about the restart;
310
+ `~/.claude/CLAUDE.md`) — the exact three-command sequence is checklist step 6;
311
+ then remind about the restart;
316
312
  - **move the family pin in the SAME session.** A member released without its
317
313
  umbrella pin bumped is invisible: `list` advertises the old version and
318
314
  `update` installs it (seen here 2026-08-10);
@@ -155,6 +155,11 @@ the body.
155
155
  first 100 lines and never learn the rest exists.
156
156
  - **One level deep from `SKILL.md`.** A file reachable only through another file
157
157
  gets partially read or missed entirely. Every reference links from the body.
158
+ - **English prose is a budget rule before a style one** (the house rule in
159
+ `SKILL.md` states which four literals stay Cyrillic). The measured cost:
160
+ Russian encodes at 1.9–2.3 chars/token against English's 5.0 (`cl100k`), and
161
+ rewriting the eight ssheleg routers into English cut them **3408 → 1885
162
+ tokens** with no loss of meaning.
158
163
 
159
164
  ## Scripts — the rules that separate a script from a liability
160
165
 
@@ -237,8 +242,24 @@ skill documents imagined problems.
237
242
  4. **Write the minimum** that fixes the gaps.
238
243
  5. **Iterate** — re-run, compare to baseline, refine.
239
244
 
240
- Evaluation record — there is no built-in runner, so keep them in the repo as
241
- data and run them yourself. House layout: `test/evals/triggers.json` (~20
245
+ **The upstream runner exists — check whether it runs for you before building
246
+ around it.** `claude plugin eval <target>` reads `evals/<case>/case.yaml`, or a
247
+ `prompt.md` beside `graders/*.md`, and `--ablation with-without` runs the
248
+ no-plugin arm and reports the delta — steps 3 and 5 above, performed for you.
249
+ It defaults to publishing its HTML report to claude.ai; `--no-publish` keeps it
250
+ local, which is the right default for an unreleased skill.
251
+
252
+ **Measured 2026-08-31 on Claude Code 2.1.236: every path — `eval`, `eval init`,
253
+ `eval init --bare`, `eval <target>` — prints ``plugin eval` is currently in
254
+ early access`, writes nothing and exits 1.** No settings key, environment
255
+ variable or feature flag on that machine turned it on. That is why the layout
256
+ below exists and is what actually runs here; the day `claude plugin eval --help`
257
+ is followed by a run rather than that line, this suite is the one to migrate.
258
+ **Re-check before trusting either half of this paragraph** — an absence is the
259
+ most perishable thing a document can assert, and nothing in this repository
260
+ changes on the day it stops being true.
261
+
262
+ House layout: `test/evals/triggers.json` (~20
242
263
  queries, half near-miss negatives) and `test/evals/scenarios.json` (≥3, the
243
264
  shape below), with any input under `test/evals/fixtures/` — **never named
244
265
  `SKILL.md`**, which would ship as a real skill:
@@ -51,7 +51,11 @@ the shape from `ssheleg/super-ux`:
51
51
  │ ├── .claude-plugin/plugin.json # ONLY the manifest lives in .claude-plugin/
52
52
  │ ├── skills/<skill>/SKILL.md + references/*.md + scripts/ + assets/
53
53
  │ │ # skeletons live HERE, not at the repo root:
54
- │ │ # only the skill dir travels on every channel
54
+ │ │ # only the skill dir travels on every channel.
55
+ │ │ # make-skill ships six: SKILL.template.md,
56
+ │ │ # plugin.template.json, marketplace.template.json,
57
+ │ │ # hooks.template.json, agent.template.md,
58
+ │ │ # command.template.md
55
59
  │ ├── bin/<exe> # Claude Code puts this on the Bash PATH:
56
60
  │ │ # the only reliable way to hand the agent
57
61
  │ │ # a runnable command (no path variable)
@@ -61,7 +65,7 @@ the shape from `ssheleg/super-ux`:
61
65
  ├── cursor/rules/*.mdc # if agent-rules make sense for Cursor
62
66
  ├── bin/<name>.js + package.json # npx installer (zero-dep Node)
63
67
  ├── test/validate.py # consistency validator (stdlib only)
64
- ├── test/evals/ # triggers.json + scenarios.json (data, run by a human)
68
+ ├── test/evals/ # triggers.json + scenarios.json (driven here; `claude plugin eval` is gated, authoring.md)
65
69
  ├── .github/workflows/validate.yml # validator on push+PR (+ release.yml, off by default)
66
70
  ├── install.sh # POSIX fallback
67
71
  ├── README.md (English-first), CHANGELOG.md, LICENSE (MIT)
@@ -98,11 +98,13 @@ Report the table before changing anything, then fix.
98
98
  7. **Validator**: present, green, and able to fail — run the negative test. CI
99
99
  present, last run `success`, with `claude plugin validate --strict` as its own
100
100
  job so an upstream outage cannot mask a house failure.
101
- 8. **Evaluations** exist in `test/evals/` (`references/authoring.md`): ≥3
101
+ 8. **Evaluations** exist and have been executed (`references/authoring.md`): ≥3
102
102
  behavioral scenarios, a trigger set whose negatives are near-misses, both
103
103
  classes on both sides of the train/validation split, coexistence checked
104
104
  against the skills already installed, and a re-run on every model the skill
105
- claims support for.
105
+ claims support for. Try `claude plugin eval` first and record what it did —
106
+ authored-but-never-executed is the state this item exists to catch, and the
107
+ house layout in `test/evals/` is the fallback for when that runner is gated.
106
108
  9. **README**: badges (npm/CI/license), install + update matrix, English-first
107
109
  prose, and the bundled `references/` listed so a reader sees what ships.
108
110
  10. **Distribution live-checks** (`references/distribution.md`): `npx --yes