@ssheleg/make-skill 0.25.3 → 0.27.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,3 +1,71 @@
1
+ ## v0.27.0 — the standard-keeper asked WHEN nine times and WHAT never
2
+
3
+ `B-139`. `audit_skill.py` carried nine `DESC_*` rules — length, headroom, person, Russian
4
+ triggers, XML, target, the `Use when` opener — and **not one asked whether the
5
+ description says what the skill DOES.** That is the half of Anthropic's guidance `B-76`
6
+ quoted directly: *a description that never says what the skill does passes*. Measured
7
+ 2026-09-03 across the family: **28 of 28 skills pass `DESC_USEWHEN`, 0 gaps** — the WHEN
8
+ half is universal and the WHAT half was unchecked.
9
+
10
+ `DESC_WHAT` measures the description with its mechanical parts removed — the `Use when`
11
+ opener, the trigger list, the `Not for` clause and the opt-out sentence — and refuses
12
+ what is left below 60 characters.
13
+
14
+ **Built on the PARSED description, which is the whole reason the previous attempt was
15
+ refused.** That prototype read raw front matter, reported a **0-character** WHAT half for
16
+ six skills and missed the opening clause of twenty, because several descriptions are YAML
17
+ block scalars (`>-`) a raw-text regex reads straight past. `parse_frontmatter` already
18
+ resolves them.
19
+
20
+ **The floor is stated with its margin.** Measured across the shipped family the smallest
21
+ honest WHAT half is **149** characters (`ux-audit`) and the largest **949**
22
+ (`seo-aeo-audit`), so 60 clears every real description by more than double while still
23
+ catching `Use when the user asks. Triggers - "делай" / "do it".`, whose WHAT half is
24
+ **13**. A case asserts that margin, so raising the floor without re-measuring fails.
25
+
26
+ **This rule finds no gap today and that is said out loud.** A standard is for the
27
+ description not yet written, and a rule that fires on nothing now is worth only what its
28
+ plants prove — so it was watched refusing a WHEN-only description and watched going
29
+ silent when disabled.
30
+
31
+ **One case was reworded after being caught claiming somebody else's work.** It asserted
32
+ the WHAT half is read from the parsed value rather than raw text — and the plant for that
33
+ property is caught by an existing case guarding `parse_frontmatter`'s folding, one layer
34
+ up. It now claims only what it holds: both spellings of one description reach the same
35
+ verdict, neither passes vacuously, and the parsed value carries no newline.
36
+
37
+ ## v0.26.0 — where a rule lives decides whether it exists
38
+
39
+ `authoring.md` treated progressive disclosure as a budget question. It is also a
40
+ **delivery** question, and that half was measured.
41
+
42
+ **Soft channels failed completely.** A README, an `AGENTS.md` beside the code, comments
43
+ inside a dependency directory, a warning field in an API response — none of them steered
44
+ anything. Two failures inside that deserve their own names, because each looks like it
45
+ should work: agents **rarely open files inside dependency directories**, so a rule written
46
+ there is a rule nobody reads; and agents **parse the data out of an API response while
47
+ ignoring the warning field in the same payload** — the bytes arrived, the guidance did not.
48
+
49
+ **Hard channels worked**: the skill file itself, an error message, CLI help, an install
50
+ prompt. So a rule that matters belongs in `SKILL.md` or a reference the body links, never
51
+ in a neighbouring document that is merely nearby — **proximity is not loading**. And an
52
+ error message is a steering surface rather than an apology: it arrives exactly when the
53
+ agent is wrong, in a channel it is already reading.
54
+
55
+ The numbers, all on hard channels: restructuring for progressive disclosure bought about
56
+ **10%** better performance *at lower token cost*; promoting a skill from CLI login converted
57
+ at **30–35%** to an install; moving templates and examples inside the skill rather than into
58
+ the system prompt cut time-to-first-token by **18.1%**.
59
+
60
+ **And the trigger corpus's negative half now carries the number that defends it.** House
61
+ style already mandates *"~20 queries, half near-miss negatives"* — with no evidence that
62
+ omitting them costs anything, which is exactly the shape of a rule an author trims under
63
+ deadline. Moving to skill-based routing dropped triggering by about **20%** in targeted
64
+ evals before negative examples and edge-case coverage were added back; with them the same
65
+ skill reached **73% → 85%** routing accuracy. A corpus that is all positives measures
66
+ whether the skill fires and never whether it stays quiet — and staying quiet is half of what
67
+ routing is.
68
+
1
69
  # Changelog
2
70
 
3
71
  ## v0.25.3 — the runner we said did not exist, and the date a claim about someone else now carries
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ssheleg/make-skill",
3
- "version": "0.25.3",
3
+ "version": "0.27.0",
4
4
  "description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way \u2014 conformance to the Agent Skills open standard AND Anthropic's platform rules (front-matter limits, disclosure budgets, per-surface runtime limits, the Skills API, evals) plus the Claude Code plugin reference (manifest schemas, component layout, claude plugin validate --strict), marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, the review checklist for third-party skills, and MCP / A2A rules for protocol-connected skills. This package is the installer CLI.",
5
5
  "keywords": [
6
6
  "skill",
@@ -3,7 +3,7 @@
3
3
  "name": "make-skill",
4
4
  "displayName": "Make Skill",
5
5
  "description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way: conformance to the Agent Skills open standard, Anthropic's platform rules (surfaces, Skills API, evals) and the Claude Code plugin reference, marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, end-to-end first publish, the review checklist for third-party skills, plus MCP / A2A references for protocol-connected skills.",
6
- "version": "0.25.3",
6
+ "version": "0.27.0",
7
7
  "author": {
8
8
  "name": "ssheleg",
9
9
  "url": "https://x.com/sshlg93"
@@ -5,7 +5,7 @@ license: MIT
5
5
  compatibility: Authoring works on any agent. The bundled scripts/ need python3. Publishing steps need git, gh, node and npm; the plugin gates need the claude CLI. Not usable on the Claude API surface, which has no network and no runtime package install.
6
6
  metadata:
7
7
  author: ssheleg
8
- version: "0.25.3"
8
+ version: "0.27.0"
9
9
  homepage: https://github.com/ssheleg/make-skill
10
10
  ---
11
11
 
@@ -14,6 +14,7 @@ and [agentskills.io](https://agentskills.io/skill-creation/best-practices).
14
14
 
15
15
  - Naming — what to call the skill
16
16
  - Description — the entire triggering budget
17
+ - Where a rule lives decides whether it exists
17
18
  - Trigger eval loop — when firing is wrong
18
19
  - Degrees of freedom — how prescriptive to be
19
20
  - Body patterns worth copying
@@ -65,11 +66,50 @@ each of which changes whether the skill fires:
65
66
  - Agents skip skills for tasks they can already do in one step. Descriptions earn
66
67
  their keep on specialized or multi-step work.
67
68
 
69
+ ## Where a rule lives decides whether it exists
70
+
71
+ **If your guidance was not in the loaded context, it did not happen.** That is not a
72
+ figure of speech about attention — it is the measured difference between two kinds of
73
+ channel.
74
+
75
+ | | Channel | Result |
76
+ |---|---|---|
77
+ | **Soft** | a README, an `AGENTS.md` beside the code, comments inside a dependency directory, a warning field in an API response | **failed completely** |
78
+ | **Hard** | the skill file itself, an error message, CLI help, an install prompt | worked |
79
+
80
+ Two failures inside that are worth naming separately, because each looks like it should
81
+ work. Agents **rarely open files inside dependency directories** — a rule written there is
82
+ a rule nobody reads. And agents **parse the data out of an API response while ignoring the
83
+ warning field in the same payload**: the bytes arrived, the guidance did not.
84
+
85
+ The consequences for a skill author:
86
+
87
+ - **A rule that matters belongs in `SKILL.md` or in a reference the body links**, not in a
88
+ neighbouring document that happens to be nearby in the repository. Proximity is not
89
+ loading.
90
+ - **An error message is a steering surface**, and often the best one — it arrives exactly
91
+ when the agent is wrong, in the channel it is already reading. Error-based steering
92
+ reliably corrected requests where a soft warning did not.
93
+ - **Restructuring for progressive disclosure is not only a budget move**: doing it bought
94
+ about **10% better performance at lower token cost**, because what remained in the body
95
+ was what had to be read every time.
96
+
97
+ Two figures for the other end of the funnel, both about hard channels: promoting a skill
98
+ from CLI login converted at **30–35%** to an install, and moving templates and examples
99
+ *inside* the skill rather than into the system prompt cut time-to-first-token by **18.1%**.
100
+
68
101
  ## Trigger eval loop — when firing is wrong
69
102
 
70
103
  1. Write ~20 realistic queries: 8–10 `should_trigger: true`, 8–10 `false`. The
71
104
  valuable negatives are **near-misses** that share keywords but need something
72
105
  else.
106
+
107
+ **The negatives are the half an author deletes first, and the one with a number
108
+ behind it.** Moving to skill-based routing dropped triggering by about **20%** in
109
+ targeted evals before negative examples and edge-case coverage were added back; with
110
+ them the same skill reached **73% → 85%** routing accuracy. A trigger corpus that is
111
+ all positives measures whether the skill fires, never whether it stays quiet — and
112
+ staying quiet is half of what routing is.
73
113
  2. Run each 3× against the agent with the skill installed → trigger rate;
74
114
  pass threshold 0.5.
75
115
  3. Split 60% train / 40% validation, fixed across iterations. Tune only on train
@@ -38,6 +38,48 @@ DESC_MAX = 1024 # spec: description is 1-1024 characters
38
38
  # budget AND the field that must grow when a near-miss skill appears ("say what
39
39
  # it is NOT for"). A description at 98% of cap cannot absorb that sentence.
40
40
  DESC_TARGET = 970
41
+ # B-139 — the nine `DESC_*` rules ask WHEN and never WHAT.
42
+ #
43
+ # Anthropic's guidance asks a description to say what the skill DOES as well as when to
44
+ # use it, and `B-76` quoted the failure directly: *a description that never says what the
45
+ # skill does passes*. Measured 2026-09-03 across the family: **28 of 28 skills pass
46
+ # `DESC_USEWHEN`, 0 gaps** — the WHEN half is universal and the WHAT half was unchecked.
47
+ #
48
+ # The WHAT half is the description with its mechanical parts removed: the `Use when`
49
+ # opener, the trigger list, the `Not for` clause and the opt-out sentence. What remains is
50
+ # the prose that names the act. A description that is an opener plus a trigger list leaves
51
+ # almost nothing, which is exactly the shape the rule refuses.
52
+ #
53
+ # The floor is 60 and it is deliberately permissive. Measured on the shipped family, the
54
+ # smallest honest WHAT half is **149** characters (`ux-audit`) and the largest 949, so 60
55
+ # clears every real description by more than double and still catches
56
+ # `Use when the user asks. Triggers - "x" / "у".`, whose WHAT half is 13.
57
+ #
58
+ # **This rule finds no gap today, and that is stated rather than hidden.** A standard is
59
+ # for the description not yet written; a rule that fires on nothing now is only worth
60
+ # anything if it was watched firing on a plant, which `test/checker_parity_test.py` does.
61
+ DESC_WHAT_MIN = 60
62
+
63
+ # Built on the PARSED description, never the raw front matter. A prototype was built
64
+ # against raw text and refused: it reported a 0-character WHAT half for six skills and
65
+ # missed the opening clause of twenty, because several descriptions are YAML block
66
+ # scalars (`>-`) that a raw-text regex reads straight past. `parse_frontmatter` already
67
+ # resolves them, and the prototype did not use it.
68
+ _WHAT_OPENER = re.compile(r"^use when\s+", re.I)
69
+ _WHAT_STRIP = (
70
+ re.compile(r"\bTriggers?\s*[-–—:].*", re.S | re.I),
71
+ re.compile(r"\bNot for\b.*", re.S | re.I),
72
+ re.compile(r"\bsay\s+['\"«].*", re.S | re.I),
73
+ )
74
+
75
+
76
+ def what_half(description):
77
+ """The prose that names the act, with the mechanical parts removed."""
78
+ core = _WHAT_OPENER.sub("", " ".join(str(description or "").split()))
79
+ for rx in _WHAT_STRIP:
80
+ core = rx.sub("", core)
81
+ return core.strip(" .,;—-")
82
+
41
83
  COMPAT_MAX = 500 # spec: compatibility is 1-500 characters
42
84
  BODY_MAX_LINES = 500 # spec + Anthropic: keep the body under 500 lines
43
85
  BODY_MAX_TOKENS = 5000 # spec + Anthropic: level-2 budget
@@ -270,6 +312,17 @@ def _check_description(a, fm, lines, rel, house):
270
312
  a.gap("DESC_RU", "description carries no Russian trigger phrases (house rule)", rel, ln)
271
313
  else:
272
314
  a.ok("DESC_RU", "description carries Russian trigger phrases (house rule)", rel, ln)
315
+ what = what_half(desc)
316
+ if len(what) < DESC_WHAT_MIN:
317
+ a.gap("DESC_WHAT", "description says WHEN and never WHAT: %d chars remain "
318
+ "after the opener, the trigger list and the refusal are removed, and "
319
+ "the floor is %d (house rule). Anthropic's guidance asks for both "
320
+ "halves, and a description that is an opener plus a trigger list "
321
+ "selects the skill without telling the model what it will do"
322
+ % (len(what), DESC_WHAT_MIN), rel, ln)
323
+ else:
324
+ a.ok("DESC_WHAT", "description states WHAT the skill does in %d chars "
325
+ "beyond its triggers (house rule)" % len(what), rel, ln)
273
326
  if DESC_TARGET < len(desc) <= DESC_MAX:
274
327
  a.gap("DESC_HEADROOM", "description is %d chars — inside the %d cap but past the "
275
328
  "%d working limit (house rule): leave room for the 'what this is NOT for' "