@ssheleg/make-skill 0.25.2 → 0.26.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +77 -0
- package/package.json +1 -1
- package/plugins/make-skill/.claude-plugin/plugin.json +1 -1
- package/plugins/make-skill/skills/make-skill/SKILL.md +1 -1
- package/plugins/make-skill/skills/make-skill/references/authoring.md +58 -2
- package/plugins/make-skill/skills/make-skill/references/distribution.md +1 -1
- package/plugins/make-skill/skills/make-skill/references/retrofit.md +4 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,82 @@
|
|
|
1
|
+
## v0.26.0 — where a rule lives decides whether it exists
|
|
2
|
+
|
|
3
|
+
`authoring.md` treated progressive disclosure as a budget question. It is also a
|
|
4
|
+
**delivery** question, and that half was measured.
|
|
5
|
+
|
|
6
|
+
**Soft channels failed completely.** A README, an `AGENTS.md` beside the code, comments
|
|
7
|
+
inside a dependency directory, a warning field in an API response — none of them steered
|
|
8
|
+
anything. Two failures inside that deserve their own names, because each looks like it
|
|
9
|
+
should work: agents **rarely open files inside dependency directories**, so a rule written
|
|
10
|
+
there is a rule nobody reads; and agents **parse the data out of an API response while
|
|
11
|
+
ignoring the warning field in the same payload** — the bytes arrived, the guidance did not.
|
|
12
|
+
|
|
13
|
+
**Hard channels worked**: the skill file itself, an error message, CLI help, an install
|
|
14
|
+
prompt. So a rule that matters belongs in `SKILL.md` or a reference the body links, never
|
|
15
|
+
in a neighbouring document that is merely nearby — **proximity is not loading**. And an
|
|
16
|
+
error message is a steering surface rather than an apology: it arrives exactly when the
|
|
17
|
+
agent is wrong, in a channel it is already reading.
|
|
18
|
+
|
|
19
|
+
The numbers, all on hard channels: restructuring for progressive disclosure bought about
|
|
20
|
+
**10%** better performance *at lower token cost*; promoting a skill from CLI login converted
|
|
21
|
+
at **30–35%** to an install; moving templates and examples inside the skill rather than into
|
|
22
|
+
the system prompt cut time-to-first-token by **18.1%**.
|
|
23
|
+
|
|
24
|
+
**And the trigger corpus's negative half now carries the number that defends it.** House
|
|
25
|
+
style already mandates *"~20 queries, half near-miss negatives"* — with no evidence that
|
|
26
|
+
omitting them costs anything, which is exactly the shape of a rule an author trims under
|
|
27
|
+
deadline. Moving to skill-based routing dropped triggering by about **20%** in targeted
|
|
28
|
+
evals before negative examples and edge-case coverage were added back; with them the same
|
|
29
|
+
skill reached **73% → 85%** routing accuracy. A corpus that is all positives measures
|
|
30
|
+
whether the skill fires and never whether it stays quiet — and staying quiet is half of what
|
|
31
|
+
routing is.
|
|
32
|
+
|
|
1
33
|
# Changelog
|
|
2
34
|
|
|
35
|
+
## v0.25.3 — the runner we said did not exist, and the date a claim about someone else now carries
|
|
36
|
+
|
|
37
|
+
- **B-122 closed: `references/authoring.md` told authors for four weeks that
|
|
38
|
+
there was no built-in eval runner, and prescribed a home-rolled suite on that
|
|
39
|
+
basis.** `claude plugin eval` had shipped. The section now names it, its case
|
|
40
|
+
format (`evals/<case>/case.yaml`, or `prompt.md` beside `graders/*.md`), and
|
|
41
|
+
the fact that `--ablation with-without` performs steps 3 and 5 of the very
|
|
42
|
+
procedure the paragraph prescribes by hand — plus the detail that it publishes
|
|
43
|
+
its HTML report to claude.ai unless `--no-publish` is passed, which is the
|
|
44
|
+
wrong default for an unreleased skill. **The replacement is a measurement, not
|
|
45
|
+
a correction of one absence into another**: on 2026-08-31, Claude Code 2.1.236,
|
|
46
|
+
every path (`eval`, `eval init`, `eval init --bare`, `eval <target>`) prints
|
|
47
|
+
``plugin eval` is currently in early access`, writes nothing and exits 1, with
|
|
48
|
+
no settings key, environment variable or feature flag turning it on. That is
|
|
49
|
+
why `test/evals/` remains what runs here, and the paragraph says which day it
|
|
50
|
+
should be migrated. Two derived statements went with it: `distribution.md`'s
|
|
51
|
+
layout comment called the suite *"run by a human"*, and `retrofit.md`'s item 8
|
|
52
|
+
asked only that evaluations *exist* — the state MSK-01 had just spent a whole
|
|
53
|
+
run closing.
|
|
54
|
+
- **The class, not the sentence: `THIRD_PARTY_CLAIMS` in `test/validate.py`.**
|
|
55
|
+
A load-bearing claim about someone else's tool is registered beside the command
|
|
56
|
+
that re-checks it, and the guard fails unless the claim, that command and an
|
|
57
|
+
ISO date still sit within eight lines of each other — and fails equally when a
|
|
58
|
+
registered claim has been deleted from the doctrine but left in the registry,
|
|
59
|
+
so the registry cannot outlive what it describes. The skill's own rule *"No
|
|
60
|
+
time-sensitive statements"* sits 100 lines above the sentence that broke it and
|
|
61
|
+
nothing enforced it; this is the enforcement.
|
|
62
|
+
- **A generic detector was built first, measured, and rejected with the number.**
|
|
63
|
+
Over the shipped doctrine a pattern for absence-claims flagged **10 lines, of
|
|
64
|
+
which roughly half are legitimate** — the fallback rows of
|
|
65
|
+
`host-capabilities.md` ("on any other host there are no subagents") are made of
|
|
66
|
+
that sentence shape on purpose. A guard with a ~50% false-positive rate is the
|
|
67
|
+
over-defense that gets switched off, so the explicit registry replaced it; the
|
|
68
|
+
reasoning is recorded in the code beside the table rather than lost.
|
|
69
|
+
- **The three plants are in CI, and the group count still computes.** Added as
|
|
70
|
+
cases to the existing *claims, budgets and runnable commands* step rather than
|
|
71
|
+
as a new step, so `grep -c 'name: Negative self-test'` stays at **9** and
|
|
72
|
+
`CONTRIBUTING.md`'s *"9 negative self-test groups"* — a claim this validator
|
|
73
|
+
compares against the workflow — remains true. Each case asserts its own
|
|
74
|
+
expected message. The first local run of the guard **passed its plants for the
|
|
75
|
+
wrong reason**: the claim wraps a line break and the check was line-anchored,
|
|
76
|
+
so all three plants tripped the registry-rot branch instead of their own. The
|
|
77
|
+
matcher is whitespace-tolerant now, and every plant in the block asserts it
|
|
78
|
+
changed something before the validator is asked anything.
|
|
79
|
+
|
|
3
80
|
## v0.25.2 — the evals have been run, and the discovery check runs somewhere
|
|
4
81
|
|
|
5
82
|
- **MSK-01 closed: the evaluation suite was executed against two models, and the
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ssheleg/make-skill",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.26.0",
|
|
4
4
|
"description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way \u2014 conformance to the Agent Skills open standard AND Anthropic's platform rules (front-matter limits, disclosure budgets, per-surface runtime limits, the Skills API, evals) plus the Claude Code plugin reference (manifest schemas, component layout, claude plugin validate --strict), marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, the review checklist for third-party skills, and MCP / A2A rules for protocol-connected skills. This package is the installer CLI.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"skill",
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"name": "make-skill",
|
|
4
4
|
"displayName": "Make Skill",
|
|
5
5
|
"description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way: conformance to the Agent Skills open standard, Anthropic's platform rules (surfaces, Skills API, evals) and the Claude Code plugin reference, marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, end-to-end first publish, the review checklist for third-party skills, plus MCP / A2A references for protocol-connected skills.",
|
|
6
|
-
"version": "0.
|
|
6
|
+
"version": "0.26.0",
|
|
7
7
|
"author": {
|
|
8
8
|
"name": "ssheleg",
|
|
9
9
|
"url": "https://x.com/sshlg93"
|
|
@@ -5,7 +5,7 @@ license: MIT
|
|
|
5
5
|
compatibility: Authoring works on any agent. The bundled scripts/ need python3. Publishing steps need git, gh, node and npm; the plugin gates need the claude CLI. Not usable on the Claude API surface, which has no network and no runtime package install.
|
|
6
6
|
metadata:
|
|
7
7
|
author: ssheleg
|
|
8
|
-
version: "0.
|
|
8
|
+
version: "0.26.0"
|
|
9
9
|
homepage: https://github.com/ssheleg/make-skill
|
|
10
10
|
---
|
|
11
11
|
|
|
@@ -14,6 +14,7 @@ and [agentskills.io](https://agentskills.io/skill-creation/best-practices).
|
|
|
14
14
|
|
|
15
15
|
- Naming — what to call the skill
|
|
16
16
|
- Description — the entire triggering budget
|
|
17
|
+
- Where a rule lives decides whether it exists
|
|
17
18
|
- Trigger eval loop — when firing is wrong
|
|
18
19
|
- Degrees of freedom — how prescriptive to be
|
|
19
20
|
- Body patterns worth copying
|
|
@@ -65,11 +66,50 @@ each of which changes whether the skill fires:
|
|
|
65
66
|
- Agents skip skills for tasks they can already do in one step. Descriptions earn
|
|
66
67
|
their keep on specialized or multi-step work.
|
|
67
68
|
|
|
69
|
+
## Where a rule lives decides whether it exists
|
|
70
|
+
|
|
71
|
+
**If your guidance was not in the loaded context, it did not happen.** That is not a
|
|
72
|
+
figure of speech about attention — it is the measured difference between two kinds of
|
|
73
|
+
channel.
|
|
74
|
+
|
|
75
|
+
| | Channel | Result |
|
|
76
|
+
|---|---|---|
|
|
77
|
+
| **Soft** | a README, an `AGENTS.md` beside the code, comments inside a dependency directory, a warning field in an API response | **failed completely** |
|
|
78
|
+
| **Hard** | the skill file itself, an error message, CLI help, an install prompt | worked |
|
|
79
|
+
|
|
80
|
+
Two failures inside that are worth naming separately, because each looks like it should
|
|
81
|
+
work. Agents **rarely open files inside dependency directories** — a rule written there is
|
|
82
|
+
a rule nobody reads. And agents **parse the data out of an API response while ignoring the
|
|
83
|
+
warning field in the same payload**: the bytes arrived, the guidance did not.
|
|
84
|
+
|
|
85
|
+
The consequences for a skill author:
|
|
86
|
+
|
|
87
|
+
- **A rule that matters belongs in `SKILL.md` or in a reference the body links**, not in a
|
|
88
|
+
neighbouring document that happens to be nearby in the repository. Proximity is not
|
|
89
|
+
loading.
|
|
90
|
+
- **An error message is a steering surface**, and often the best one — it arrives exactly
|
|
91
|
+
when the agent is wrong, in the channel it is already reading. Error-based steering
|
|
92
|
+
reliably corrected requests where a soft warning did not.
|
|
93
|
+
- **Restructuring for progressive disclosure is not only a budget move**: doing it bought
|
|
94
|
+
about **10% better performance at lower token cost**, because what remained in the body
|
|
95
|
+
was what had to be read every time.
|
|
96
|
+
|
|
97
|
+
Two figures for the other end of the funnel, both about hard channels: promoting a skill
|
|
98
|
+
from CLI login converted at **30–35%** to an install, and moving templates and examples
|
|
99
|
+
*inside* the skill rather than into the system prompt cut time-to-first-token by **18.1%**.
|
|
100
|
+
|
|
68
101
|
## Trigger eval loop — when firing is wrong
|
|
69
102
|
|
|
70
103
|
1. Write ~20 realistic queries: 8–10 `should_trigger: true`, 8–10 `false`. The
|
|
71
104
|
valuable negatives are **near-misses** that share keywords but need something
|
|
72
105
|
else.
|
|
106
|
+
|
|
107
|
+
**The negatives are the half an author deletes first, and the one with a number
|
|
108
|
+
behind it.** Moving to skill-based routing dropped triggering by about **20%** in
|
|
109
|
+
targeted evals before negative examples and edge-case coverage were added back; with
|
|
110
|
+
them the same skill reached **73% → 85%** routing accuracy. A trigger corpus that is
|
|
111
|
+
all positives measures whether the skill fires, never whether it stays quiet — and
|
|
112
|
+
staying quiet is half of what routing is.
|
|
73
113
|
2. Run each 3× against the agent with the skill installed → trigger rate;
|
|
74
114
|
pass threshold 0.5.
|
|
75
115
|
3. Split 60% train / 40% validation, fixed across iterations. Tune only on train
|
|
@@ -242,8 +282,24 @@ skill documents imagined problems.
|
|
|
242
282
|
4. **Write the minimum** that fixes the gaps.
|
|
243
283
|
5. **Iterate** — re-run, compare to baseline, refine.
|
|
244
284
|
|
|
245
|
-
|
|
246
|
-
|
|
285
|
+
**The upstream runner exists — check whether it runs for you before building
|
|
286
|
+
around it.** `claude plugin eval <target>` reads `evals/<case>/case.yaml`, or a
|
|
287
|
+
`prompt.md` beside `graders/*.md`, and `--ablation with-without` runs the
|
|
288
|
+
no-plugin arm and reports the delta — steps 3 and 5 above, performed for you.
|
|
289
|
+
It defaults to publishing its HTML report to claude.ai; `--no-publish` keeps it
|
|
290
|
+
local, which is the right default for an unreleased skill.
|
|
291
|
+
|
|
292
|
+
**Measured 2026-08-31 on Claude Code 2.1.236: every path — `eval`, `eval init`,
|
|
293
|
+
`eval init --bare`, `eval <target>` — prints ``plugin eval` is currently in
|
|
294
|
+
early access`, writes nothing and exits 1.** No settings key, environment
|
|
295
|
+
variable or feature flag on that machine turned it on. That is why the layout
|
|
296
|
+
below exists and is what actually runs here; the day `claude plugin eval --help`
|
|
297
|
+
is followed by a run rather than that line, this suite is the one to migrate.
|
|
298
|
+
**Re-check before trusting either half of this paragraph** — an absence is the
|
|
299
|
+
most perishable thing a document can assert, and nothing in this repository
|
|
300
|
+
changes on the day it stops being true.
|
|
301
|
+
|
|
302
|
+
House layout: `test/evals/triggers.json` (~20
|
|
247
303
|
queries, half near-miss negatives) and `test/evals/scenarios.json` (≥3, the
|
|
248
304
|
shape below), with any input under `test/evals/fixtures/` — **never named
|
|
249
305
|
`SKILL.md`**, which would ship as a real skill:
|
|
@@ -65,7 +65,7 @@ the shape from `ssheleg/super-ux`:
|
|
|
65
65
|
├── cursor/rules/*.mdc # if agent-rules make sense for Cursor
|
|
66
66
|
├── bin/<name>.js + package.json # npx installer (zero-dep Node)
|
|
67
67
|
├── test/validate.py # consistency validator (stdlib only)
|
|
68
|
-
├── test/evals/ # triggers.json + scenarios.json (
|
|
68
|
+
├── test/evals/ # triggers.json + scenarios.json (driven here; `claude plugin eval` is gated, authoring.md)
|
|
69
69
|
├── .github/workflows/validate.yml # validator on push+PR (+ release.yml, off by default)
|
|
70
70
|
├── install.sh # POSIX fallback
|
|
71
71
|
├── README.md (English-first), CHANGELOG.md, LICENSE (MIT)
|
|
@@ -98,11 +98,13 @@ Report the table before changing anything, then fix.
|
|
|
98
98
|
7. **Validator**: present, green, and able to fail — run the negative test. CI
|
|
99
99
|
present, last run `success`, with `claude plugin validate --strict` as its own
|
|
100
100
|
job so an upstream outage cannot mask a house failure.
|
|
101
|
-
8. **Evaluations** exist
|
|
101
|
+
8. **Evaluations** exist and have been executed (`references/authoring.md`): ≥3
|
|
102
102
|
behavioral scenarios, a trigger set whose negatives are near-misses, both
|
|
103
103
|
classes on both sides of the train/validation split, coexistence checked
|
|
104
104
|
against the skills already installed, and a re-run on every model the skill
|
|
105
|
-
claims support for.
|
|
105
|
+
claims support for. Try `claude plugin eval` first and record what it did —
|
|
106
|
+
authored-but-never-executed is the state this item exists to catch, and the
|
|
107
|
+
house layout in `test/evals/` is the fallback for when that runner is gated.
|
|
106
108
|
9. **README**: badges (npm/CI/license), install + update matrix, English-first
|
|
107
109
|
prose, and the bundled `references/` listed so a reader sees what ships.
|
|
108
110
|
10. **Distribution live-checks** (`references/distribution.md`): `npx --yes
|