@ssheleg/make-skill 0.25.0 → 0.25.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +83 -0
- package/cursor/rules/make-skill.mdc +4 -1
- package/package.json +2 -2
- package/plugins/make-skill/.claude-plugin/plugin.json +1 -1
- package/plugins/make-skill/skills/make-skill/SKILL.md +13 -17
- package/plugins/make-skill/skills/make-skill/references/authoring.md +5 -0
- package/plugins/make-skill/skills/make-skill/references/distribution.md +5 -1
- package/plugins/make-skill/skills/make-skill/references/host-capabilities.md +13 -2
- package/plugins/make-skill/skills/make-skill/references/retrofit.md +5 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,88 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v0.25.2 — the evals have been run, and the discovery check runs somewhere
|
|
4
|
+
|
|
5
|
+
- **MSK-01 closed: the evaluation suite was executed against two models, and the
|
|
6
|
+
numbers live in `test/evals/RESULTS.md` as a dated run (2026-08-31), not a
|
|
7
|
+
claim.** All 22 trigger queries probed blind per model — haiku 22/22, sonnet
|
|
8
|
+
20/22 (q01 and q10 answered "none": under-triggering on positives, ZERO false
|
|
9
|
+
positives on the near-miss negatives across both models) — and all four
|
|
10
|
+
behavioural scenarios executed and scored per `expected_behavior` line (sonnet
|
|
11
|
+
31/35, haiku 24/35; the s01 repositories were re-audited with the bundled
|
|
12
|
+
auditor rather than trusted from the agent's report). The Method section
|
|
13
|
+
states the harness and its limits: one probe per query instead of the
|
|
14
|
+
README's fired/3, Agent-tool subagents on the author's machine whose global
|
|
15
|
+
config carries the family routing map (a bias that makes the clean negatives
|
|
16
|
+
the strongest result), model aliases rather than exact ids, and two probes
|
|
17
|
+
counted with a stated non-compliance caveat. Headline finding: the standing
|
|
18
|
+
worry (turn-stealing) did not materialize; the observed failure is sonnet
|
|
19
|
+
under-triggering on two positive categories, and s03-haiku answered a
|
|
20
|
+
surfaces question without consulting the installed skill — selection and
|
|
21
|
+
consultation measured apart, on purpose.
|
|
22
|
+
- **MSK-03 closed: the discovery check the body demands now runs in CI.**
|
|
23
|
+
SKILL.md's stray-SKILL.md gotcha ends "verify with the skills CLI's `--list`"
|
|
24
|
+
— and nothing ran it anywhere: the operator machine's hygiene guard refuses
|
|
25
|
+
the bare skills CLI on family members by design, and CI never asked. The new
|
|
26
|
+
`skills-cli-discovery` job in `validate.yml` runs the listing against
|
|
27
|
+
`ssheleg/make-skill` and asserts it serves exactly ONE skill, named
|
|
28
|
+
make-skill, with ANSI stripped before asserting. Its own job, like
|
|
29
|
+
`claude-plugin-validate`, so a registry outage cannot mask an
|
|
30
|
+
offline-validator failure; it reads the repository's default branch, so the
|
|
31
|
+
post-merge push run and the release's `workflow_call` on the tag assert the
|
|
32
|
+
published tree.
|
|
33
|
+
- **MSK-04 closed: body headroom restored by displacement — the words moved,
|
|
34
|
+
none were lost.** The Cyrillic budget measurement (1.9–2.3 chars/token,
|
|
35
|
+
3408→1885) moved to `authoring.md` → *Content guidelines*; the six skeleton
|
|
36
|
+
filenames moved into `distribution.md`'s layout tree; the five verified facts
|
|
37
|
+
and the machine-refresh command sequence are now referenced at their single
|
|
38
|
+
home in `distribution.md`'s checklists instead of being restated; the
|
|
39
|
+
long-file rule already lived in `authoring.md` and the body now points there.
|
|
40
|
+
Re-measured with the bundled auditor: body ~4742/4750 tokens (8 spare) →
|
|
41
|
+
~4677/4750 (73 spare), after also paying for the next bullet. The
|
|
42
|
+
description is untouched at 965/970 — its headroom is spent only on trigger
|
|
43
|
+
changes, and this release makes none.
|
|
44
|
+
- **XF-10: the MCP/A2A carve-out now runs both directions at the layer an agent
|
|
45
|
+
reads first.** The References-table rows for `mcp.md` and `a2a.md` carry "the
|
|
46
|
+
protocol wire itself → `agent-interop` (agent-stack)" — the two reference
|
|
47
|
+
files' own headers already said it, and now the router-level table says it
|
|
48
|
+
too, so a reader deciding what to load learns the boundary before opening
|
|
49
|
+
either file.
|
|
50
|
+
- Also: `test/evals/README.md`'s file table said "20 realistic queries" over a
|
|
51
|
+
`triggers.json` holding 22 — the same drift class as MSK-02, in a wording no
|
|
52
|
+
guard could parse. Reworded to the checkable "22 trigger queries" form, which
|
|
53
|
+
puts the sentence under `test/validate.py`'s counted-claims sweep instead of
|
|
54
|
+
under anyone's memory. `SKILL-CARD.md`'s evaluation status and posture
|
|
55
|
+
sections now describe the executed run and the replication still owed,
|
|
56
|
+
instead of "never executed".
|
|
57
|
+
|
|
58
|
+
## v0.25.1 — the rule keeps its default, and the family's exception is written down
|
|
59
|
+
|
|
60
|
+
- **The same-name command rule now carries its recorded exception** (operator
|
|
61
|
+
decision, 2026-08-30: no renames anywhere in the family). The default in
|
|
62
|
+
`references/host-capabilities.md` is unchanged — never name a command after a
|
|
63
|
+
skill in the same plugin — and the one recorded exception is enumerated beside
|
|
64
|
+
it: the ssheleg family ships same-named commands deliberately (`task-pipeline`,
|
|
65
|
+
`project-audit`, `seo-aeo-audit`, `sheleg-design`, `agent-sync`, and
|
|
66
|
+
`super-ux`'s `vision` / `ux-foundation` / `ux-flows` / `ux-audit`), with the
|
|
67
|
+
cost stated rather than waived: the skill wins the trigger, and each command
|
|
68
|
+
stays an always-on token cost that may be unreachable in the picker.
|
|
69
|
+
`retrofit.md` item 14 and the Cursor rule teach the auditor the same path, so
|
|
70
|
+
family audits report ASY-02 / SEO-02 / TPA-02 / SHD-06 / SUX-03 as
|
|
71
|
+
*deliberate, recorded — no change needed* instead of open gaps. A collision
|
|
72
|
+
NOT on a dated, recorded exception list is still the gap the rule names.
|
|
73
|
+
- **A hand-typed eval count drifted, exactly as this skill's own gotcha predicts**
|
|
74
|
+
(MSK-02). `test/evals/RESULTS.md` said "20 queries" and `SKILL-CARD.md` said
|
|
75
|
+
"20 trigger queries" over a `triggers.json` holding 22. Both corrected — and
|
|
76
|
+
both statements are now compared, not typed: `test/evals_validate.py` parses
|
|
77
|
+
RESULTS.md's stated counts against the artifacts and refuses a claim-free
|
|
78
|
+
RESULTS.md rather than passing vacuously, with negative self-tests planting an
|
|
79
|
+
off-by-one count and a claim-free file; `test/validate.py`'s counted-claims
|
|
80
|
+
sweep gains the `N trigger queries` / `N behavioural scenarios` patterns for
|
|
81
|
+
every other document. Both guards were watched failing against the real
|
|
82
|
+
20-vs-22 defect before the numbers were corrected.
|
|
83
|
+
- The `SKILL.md` body is untouched on purpose: at ~4742/5000 tokens it has 8
|
|
84
|
+
tokens of headroom, so the whole amendment lives in the references.
|
|
85
|
+
|
|
3
86
|
## v0.25.0 — the installer refuses the shadow it documents, loudly
|
|
4
87
|
|
|
5
88
|
- **The family audit of 2026-08-29 reproduced the shadow live:** a bare
|
|
@@ -140,7 +140,10 @@ through a quoted `"${CLAUDE_PLUGIN_ROOT}/…"` with a `timeout`; hook scripts ex
|
|
|
140
140
|
0 silently when the event is not theirs and when the interpreter is missing;
|
|
141
141
|
`PreToolUse` blocks (exit 2 sends stderr to the model), `PostToolUse` advises via
|
|
142
142
|
`systemMessage`; plugin agents may not carry `hooks`, `mcpServers` or
|
|
143
|
-
`permissionMode`; a command is never named after a skill in the same plugin
|
|
143
|
+
`permissionMode`; a command is never named after a skill in the same plugin —
|
|
144
|
+
unless the collision is on the recorded, dated exception list in the skill's
|
|
145
|
+
`references/host-capabilities.md` (the ssheleg family, operator decision
|
|
146
|
+
2026-08-30), which an audit reports as deliberate and recorded, not as a gap.
|
|
144
147
|
|
|
145
148
|
## Hard rules
|
|
146
149
|
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@ssheleg/make-skill",
|
|
3
|
-
"version": "0.25.
|
|
4
|
-
"description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way
|
|
3
|
+
"version": "0.25.2",
|
|
4
|
+
"description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way \u2014 conformance to the Agent Skills open standard AND Anthropic's platform rules (front-matter limits, disclosure budgets, per-surface runtime limits, the Skills API, evals) plus the Claude Code plugin reference (manifest schemas, component layout, claude plugin validate --strict), marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, the review checklist for third-party skills, and MCP / A2A rules for protocol-connected skills. This package is the installer CLI.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"skill",
|
|
7
7
|
"plugin",
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"name": "make-skill",
|
|
4
4
|
"displayName": "Make Skill",
|
|
5
5
|
"description": "Create, retrofit, audit, and ship agent skills & Claude Code plugins the proven ssheleg way: conformance to the Agent Skills open standard, Anthropic's platform rules (surfaces, Skills API, evals) and the Claude Code plugin reference, marketplace repo layout, version sync, validator + CI, multi-channel distribution (plugin, vercel skills CLI, npx, Cursor), npm gotchas, end-to-end first publish, the review checklist for third-party skills, plus MCP / A2A references for protocol-connected skills.",
|
|
6
|
-
"version": "0.25.
|
|
6
|
+
"version": "0.25.2",
|
|
7
7
|
"author": {
|
|
8
8
|
"name": "ssheleg",
|
|
9
9
|
"url": "https://x.com/sshlg93"
|
|
@@ -5,7 +5,7 @@ license: MIT
|
|
|
5
5
|
compatibility: Authoring works on any agent. The bundled scripts/ need python3. Publishing steps need git, gh, node and npm; the plugin gates need the claude CLI. Not usable on the Claude API surface, which has no network and no runtime package install.
|
|
6
6
|
metadata:
|
|
7
7
|
author: ssheleg
|
|
8
|
-
version: "0.25.
|
|
8
|
+
version: "0.25.2"
|
|
9
9
|
homepage: https://github.com/ssheleg/make-skill
|
|
10
10
|
---
|
|
11
11
|
|
|
@@ -27,8 +27,8 @@ orchestrator, release automation). **make-skill itself** is built to this canon.
|
|
|
27
27
|
| `references/host-capabilities.md` | shipping a **hook, subagent, command, script or MCP dependency** — what each buys and costs, hook events and exit codes, the degradation clauses |
|
|
28
28
|
| `references/claude-code-plugin.md` | anything shipping as a **Claude Code plugin/marketplace** — manifest schemas, component layout, path variables, `validate` failures |
|
|
29
29
|
| `references/distribution.md` | the repo layout, releases, and all five channels — plugin, skills CLI, npx, Cursor, umbrella family repo |
|
|
30
|
-
| `references/mcp.md` | skill vs **MCP** server, declaring the dependency, consent and untrusted-output rules |
|
|
31
|
-
| `references/a2a.md` | the skill spans two autonomous agents (**A2A**) — choosing it, the two meanings of "skill", driving a peer safely |
|
|
30
|
+
| `references/mcp.md` | skill vs **MCP** server, declaring the dependency, consent and untrusted-output rules; the protocol wire itself → `agent-interop` (agent-stack) |
|
|
31
|
+
| `references/a2a.md` | the skill spans two autonomous agents (**A2A**) — choosing it, the two meanings of "skill", driving a peer safely; the wire itself → `agent-interop` (agent-stack) |
|
|
32
32
|
|
|
33
33
|
Missing from this copy? Raw fallback:
|
|
34
34
|
`raw.githubusercontent.com/ssheleg/make-skill/main/plugins/make-skill/skills/make-skill/references/<file>`
|
|
@@ -81,8 +81,8 @@ differences and the checklist: `references/agent-skills-spec.md`.
|
|
|
81
81
|
99% of budget turns the next correction into a fight with the validator.
|
|
82
82
|
Heavier material goes to `references/`, `scripts/`, `assets/` INSIDE the skill
|
|
83
83
|
dir, one level deep, each linked from the body with a stated load trigger
|
|
84
|
-
("read X when Y") — never a bare "see references/".
|
|
85
|
-
|
|
84
|
+
("read X when Y") — never a bare "see references/". Long-file and layout rules:
|
|
85
|
+
`references/authoring.md` → *Content guidelines*.
|
|
86
86
|
- Gotchas stay in `SKILL.md`: the agent can't know to open a file about a trap
|
|
87
87
|
it doesn't know exists.
|
|
88
88
|
- **Write for the weakest surface you claim** (`references/surfaces.md`): the
|
|
@@ -97,9 +97,8 @@ House additions on top of the spec:
|
|
|
97
97
|
survives in four places only, where the string itself is the point: a
|
|
98
98
|
**trigger phrase**, a **refusal phrase** (the operator types both — translated,
|
|
99
99
|
they no longer match what was said), a **proper noun**, a **language example**
|
|
100
|
-
(`«вы»/«ты»`). A budget rule before a style one
|
|
101
|
-
|
|
102
|
-
routers into English cut them **3408 → 1885 tokens** with no loss of meaning.
|
|
100
|
+
(`«вы»/«ты»`). A budget rule before a style one — the measured cost of
|
|
101
|
+
Cyrillic prose is in `references/authoring.md` → *Content guidelines*.
|
|
103
102
|
- `description` starts "Use when …" and lists concrete trigger phrases — English
|
|
104
103
|
AND Russian (user works in both). A skill nobody triggers is dead weight. Hold
|
|
105
104
|
**5% headroom here too** (≤970 of 1024): a near-miss neighbour forces a "NOT
|
|
@@ -153,10 +152,9 @@ A fallback you know but did not write is not a fallback.
|
|
|
153
152
|
Retrofit), wrapped as `bin/make-skill-audit` for Claude Code. Beside them:
|
|
154
153
|
`hooks/` (PostToolUse, silent unless a `SKILL.md` was written),
|
|
155
154
|
`commands/skill-audit.md` (deliberately NOT the skill's name),
|
|
156
|
-
`agents/skill-auditor.md`, and six
|
|
157
|
-
`
|
|
158
|
-
|
|
159
|
-
`assets/command.template.md`.
|
|
155
|
+
`agents/skill-auditor.md`, and six `assets/*.template.*` skeletons — one per
|
|
156
|
+
component, filenames in `references/distribution.md` → *The distributable repo
|
|
157
|
+
layout*.
|
|
160
158
|
|
|
161
159
|
## Create (personal)
|
|
162
160
|
|
|
@@ -196,8 +194,7 @@ regardless:
|
|
|
196
194
|
Take it ALL the way, no half-done handoffs. Only the first publish needs a human
|
|
197
195
|
(npm 2FA); **arming CI publishing is part of shipping**, so the second does not.
|
|
198
196
|
**The 11-step sequence is in `references/distribution.md` → *First publish*.**
|
|
199
|
-
Done = five VERIFIED facts
|
|
200
|
-
installed, skills-CLI discovery working, next tag publishing without a human.
|
|
197
|
+
Done = the five VERIFIED facts in that sequence's step 10 — nothing assumed.
|
|
201
198
|
|
|
202
199
|
## Retrofit (bring an existing skill/repo up to standard)
|
|
203
200
|
|
|
@@ -310,9 +307,8 @@ session:
|
|
|
310
307
|
- everything green BEFORE the tag: `python3 test/validate.py` plus BOTH
|
|
311
308
|
`claude plugin validate … --strict` runs;
|
|
312
309
|
- **refresh THIS machine's global installs as Definition of Done** (per global
|
|
313
|
-
`~/.claude/CLAUDE.md`)
|
|
314
|
-
|
|
315
|
-
--yes && rm -f ~/.claude/skills/<name>`, then remind about the restart;
|
|
310
|
+
`~/.claude/CLAUDE.md`) — the exact three-command sequence is checklist step 6;
|
|
311
|
+
then remind about the restart;
|
|
316
312
|
- **move the family pin in the SAME session.** A member released without its
|
|
317
313
|
umbrella pin bumped is invisible: `list` advertises the old version and
|
|
318
314
|
`update` installs it (seen here 2026-08-10);
|
|
@@ -155,6 +155,11 @@ the body.
|
|
|
155
155
|
first 100 lines and never learn the rest exists.
|
|
156
156
|
- **One level deep from `SKILL.md`.** A file reachable only through another file
|
|
157
157
|
gets partially read or missed entirely. Every reference links from the body.
|
|
158
|
+
- **English prose is a budget rule before a style one** (the house rule in
|
|
159
|
+
`SKILL.md` states which four literals stay Cyrillic). The measured cost:
|
|
160
|
+
Russian encodes at 1.9–2.3 chars/token against English's 5.0 (`cl100k`), and
|
|
161
|
+
rewriting the eight ssheleg routers into English cut them **3408 → 1885
|
|
162
|
+
tokens** with no loss of meaning.
|
|
158
163
|
|
|
159
164
|
## Scripts — the rules that separate a script from a liability
|
|
160
165
|
|
|
@@ -51,7 +51,11 @@ the shape from `ssheleg/super-ux`:
|
|
|
51
51
|
│ ├── .claude-plugin/plugin.json # ONLY the manifest lives in .claude-plugin/
|
|
52
52
|
│ ├── skills/<skill>/SKILL.md + references/*.md + scripts/ + assets/
|
|
53
53
|
│ │ # skeletons live HERE, not at the repo root:
|
|
54
|
-
│ │ # only the skill dir travels on every channel
|
|
54
|
+
│ │ # only the skill dir travels on every channel.
|
|
55
|
+
│ │ # make-skill ships six: SKILL.template.md,
|
|
56
|
+
│ │ # plugin.template.json, marketplace.template.json,
|
|
57
|
+
│ │ # hooks.template.json, agent.template.md,
|
|
58
|
+
│ │ # command.template.md
|
|
55
59
|
│ ├── bin/<exe> # Claude Code puts this on the Bash PATH:
|
|
56
60
|
│ │ # the only reliable way to hand the agent
|
|
57
61
|
│ │ # a runnable command (no path variable)
|
|
@@ -143,7 +143,17 @@ body receives `$ARGUMENTS`. Two rules cost a debugging round each:
|
|
|
143
143
|
|
|
144
144
|
- **Never name a command after a skill in the same plugin.** Commands are skills
|
|
145
145
|
now: both claim `/<name>`, the skill wins, the command is unreachable always-on
|
|
146
|
-
cost visible only in `claude plugin details`.
|
|
146
|
+
cost visible only in `claude plugin details`. That is the default for every new
|
|
147
|
+
plugin. **One recorded exception (operator decision, 2026-08-30): the ssheleg
|
|
148
|
+
family ships same-named commands deliberately** — `task-pipeline`,
|
|
149
|
+
`project-audit`, `seo-aeo-audit`, `sheleg-design`, `agent-sync`, and
|
|
150
|
+
`super-ux`'s `vision`, `ux-foundation`, `ux-flows`, `ux-audit`. The cost is
|
|
151
|
+
accepted, not waived: the skill wins the trigger, and each command stays an
|
|
152
|
+
always-on token cost that may be unreachable in the picker. An audit that finds
|
|
153
|
+
a collision on this list reports it as *deliberate, recorded — no change
|
|
154
|
+
needed*; a collision NOT on a dated, recorded exception list is still the gap
|
|
155
|
+
this rule names. A rule's exception is enumerated, dated, and carries its cost
|
|
156
|
+
— never implied.
|
|
147
157
|
- **Quote `argument-hint`.** Bare `[a | b]` is a YAML flow sequence; a comma
|
|
148
158
|
inside it drops the entire frontmatter block, leaving a command with no
|
|
149
159
|
description and no warning.
|
|
@@ -236,7 +246,8 @@ what the agent reads at the exact moment something is missing.
|
|
|
236
246
|
- [ ] `PostToolUse` advises (`systemMessage`), `PreToolUse` blocks — not the reverse
|
|
237
247
|
- [ ] Hook commands quote `"${CLAUDE_PLUGIN_ROOT}"` and set a `timeout`
|
|
238
248
|
- [ ] Hook scripts are executable, have a shebang, and need no `jq`
|
|
239
|
-
- [ ] No command named after a skill
|
|
249
|
+
- [ ] No command named after a skill — or the collision is on a recorded, dated
|
|
250
|
+
exception list (see *Commands*); every `argument-hint` quoted
|
|
240
251
|
- [ ] Plugin agents carry no `hooks` / `mcpServers` / `permissionMode`
|
|
241
252
|
- [ ] Scripts are stdlib-only, inside the skill dir, invoked by a resolvable path
|
|
242
253
|
- [ ] No command the agent is told to RUN contains `${CLAUDE_PLUGIN_ROOT}` — it is
|
|
@@ -125,8 +125,11 @@ Report the table before changing anything, then fix.
|
|
|
125
125
|
MCP server: the degradation contract written in the body for all three axes
|
|
126
126
|
(not Claude Code / recommended plugin absent / tool absent); hooks that
|
|
127
127
|
exit 0 silently when the event is not theirs; `PostToolUse` advising rather
|
|
128
|
-
than blocking; commands quoted and never named after a skill
|
|
129
|
-
|
|
128
|
+
than blocking; commands quoted and never named after a skill (a collision on
|
|
129
|
+
the recorded, dated exception list in `references/host-capabilities.md` —
|
|
130
|
+
the ssheleg family's same-named commands, operator decision 2026-08-30 — is
|
|
131
|
+
reported as *deliberate, recorded — no change needed*, not as a gap); plugin
|
|
132
|
+
agents free of `hooks`, `mcpServers` and `permissionMode`.
|
|
130
133
|
|
|
131
134
|
## Personal skills — the short form
|
|
132
135
|
|