@zalom/plastic 1.0.0-alpha.10 → 1.0.0-alpha.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/PLASTIC.md +106 -475
- package/package.json +1 -1
- package/skills/auto/SKILL.md +4 -0
- package/skills/auto/references/agent-architecture.md +60 -0
- package/skills/continuing/SKILL.md +4 -0
- package/skills/continuing/references/context-management.md +32 -0
- package/skills/creating-intent/SKILL.md +5 -0
- package/skills/creating-intent/references/lifecycle.md +74 -0
- package/skills/creating-intent/references/wikilinks.md +8 -0
- package/skills/creating-project/SKILL.md +4 -0
- package/skills/creating-project/references/hubs-projects.md +55 -0
- package/skills/doctor/SKILL.md +4 -0
- package/skills/doctor/references/gates-stuck-detection.md +38 -0
- package/skills/evaluating-skills/SKILL.md +140 -0
- package/skills/evaluating-skills/assets/eval-template.json +12 -0
- package/skills/evaluating-skills/evals/evals.json +75 -0
- package/skills/evaluating-skills/references/convention-checks.md +76 -0
- package/skills/evaluating-skills/references/eval-methodology.md +154 -0
- package/skills/linking-intents/SKILL.md +4 -0
- package/skills/linking-intents/references/zettelkasten.md +33 -0
- package/skills/managing-index/SKILL.md +4 -0
- package/skills/releasing/SKILL.md +4 -0
- package/skills/releasing/references/deprecations.md +44 -0
- package/skills/savepoint/SKILL.md +4 -0
- package/skills/savepoint/references/context-management.md +32 -0
- package/skills/writing-instructions/SKILL.md +159 -0
- package/skills/writing-instructions/references/agentskills-spec.md +135 -0
|
@@ -0,0 +1,154 @@
|
|
|
1
|
+
# Evaluation Methodology
|
|
2
|
+
|
|
3
|
+
Load when choosing a grader type, interpreting pass rate results, or deciding
|
|
4
|
+
whether to graduate or retire evals.
|
|
5
|
+
|
|
6
|
+
Sources: agentskills.io, Anthropic eval engineering, Tessl eval framework,
|
|
7
|
+
Philipp Schmid skill testing guide.
|
|
8
|
+
|
|
9
|
+
## Three-Tier Grader Taxonomy
|
|
10
|
+
|
|
11
|
+
Layer graders like a Swiss cheese model — no single tier catches everything.
|
|
12
|
+
Use the simplest grader that covers the assertion. Combine tiers for coverage.
|
|
13
|
+
|
|
14
|
+
### Code-Based Graders
|
|
15
|
+
|
|
16
|
+
Best for structural and mechanical checks:
|
|
17
|
+
- File existence and correct path
|
|
18
|
+
- JSON/YAML validity
|
|
19
|
+
- Line count, character count, token budget
|
|
20
|
+
- Regex pattern matching (frontmatter fields, required sections)
|
|
21
|
+
- Exit codes from bundled scripts
|
|
22
|
+
|
|
23
|
+
Use as the default tier. Fast, deterministic, reproducible.
|
|
24
|
+
|
|
25
|
+
### LLM-as-Judge Graders
|
|
26
|
+
|
|
27
|
+
Best for subjective quality and semantic checks:
|
|
28
|
+
- "Does this description convey when to use the skill?"
|
|
29
|
+
- "Are the gotchas concrete corrections, not general advice?"
|
|
30
|
+
- "Does the output address the user's actual intent?"
|
|
31
|
+
|
|
32
|
+
Calibration protocol:
|
|
33
|
+
1. Write a natural-language rubric (not binary pass/fail)
|
|
34
|
+
2. Run the judge on 10-15 cases where you already know the correct grade
|
|
35
|
+
3. Compare judge grades to your grades — adjust rubric until >80% agreement
|
|
36
|
+
4. Spot-check with human grading periodically (every 5th eval run)
|
|
37
|
+
|
|
38
|
+
Rubric template:
|
|
39
|
+
- PASS: [specific criteria with examples]
|
|
40
|
+
- PARTIAL: [what partial credit looks like]
|
|
41
|
+
- FAIL: [specific failure modes]
|
|
42
|
+
|
|
43
|
+
### Human Graders
|
|
44
|
+
|
|
45
|
+
Best for edge cases and final calibration:
|
|
46
|
+
- Novel failure modes the other tiers miss
|
|
47
|
+
- Calibration set for LLM judges
|
|
48
|
+
- Final sign-off on graduating evals to regression
|
|
49
|
+
|
|
50
|
+
Use sparingly — human grading doesn't scale. Reserve for calibration
|
|
51
|
+
and cases where code + LLM judges disagree.
|
|
52
|
+
|
|
53
|
+
## pass@k vs pass^k
|
|
54
|
+
|
|
55
|
+
Two metrics that diverge dramatically. Always track both.
|
|
56
|
+
|
|
57
|
+
### pass@k — Capability
|
|
58
|
+
|
|
59
|
+
"Did it succeed at least once in k trials?"
|
|
60
|
+
|
|
61
|
+
Formula: pass@k = 1 - (1 - p)^k where p = single-trial pass rate
|
|
62
|
+
|
|
63
|
+
Measures: Can the skill/agent do this at all?
|
|
64
|
+
|
|
65
|
+
Example at p=0.7, k=10: pass@k = 97.2%
|
|
66
|
+
|
|
67
|
+
Use when: evaluating whether a skill enables a new capability,
|
|
68
|
+
initial development, deciding whether to invest more iteration.
|
|
69
|
+
|
|
70
|
+
### pass^k — Reliability
|
|
71
|
+
|
|
72
|
+
"Did it succeed every time in k trials?"
|
|
73
|
+
|
|
74
|
+
Formula: pass^k = p^k
|
|
75
|
+
|
|
76
|
+
Measures: Will it always do this correctly?
|
|
77
|
+
|
|
78
|
+
Example at p=0.7, k=10: pass^k = 2.8%
|
|
79
|
+
|
|
80
|
+
Use when: evaluating production readiness, regression testing,
|
|
81
|
+
deciding whether a skill is reliable enough to ship.
|
|
82
|
+
|
|
83
|
+
### Interpreting the Gap
|
|
84
|
+
|
|
85
|
+
| pass@k | pass^k | Interpretation |
|
|
86
|
+
|--------|--------|----------------|
|
|
87
|
+
| High | High | Reliable — ready for production |
|
|
88
|
+
| High | Low | Capable but flaky — needs iteration on consistency |
|
|
89
|
+
| Low | Low | Not yet capable — needs fundamental skill improvement |
|
|
90
|
+
| Low | High | Impossible (pass^k <= pass@k always) |
|
|
91
|
+
|
|
92
|
+
Run k=3 minimum for meaningful results. k=5 for production decisions.
|
|
93
|
+
|
|
94
|
+
## Capability-to-Regression Graduation
|
|
95
|
+
|
|
96
|
+
Track pass rates across iterations. When a capability eval consistently
|
|
97
|
+
hits ~100% (pass@k=1.0 for 3+ consecutive runs):
|
|
98
|
+
|
|
99
|
+
1. Graduate the eval from "capability" to "regression"
|
|
100
|
+
2. Regression evals run on every skill change — they protect against backsliding
|
|
101
|
+
3. If a regression eval starts failing, the recent change broke something
|
|
102
|
+
4. Investigate the failing regression before iterating further
|
|
103
|
+
|
|
104
|
+
Graduation is one-way. Once an eval is regression, it stays regression
|
|
105
|
+
unless the underlying requirement changes.
|
|
106
|
+
|
|
107
|
+
## Skill Retirement Detection
|
|
108
|
+
|
|
109
|
+
Monitor the with/without skill delta over time:
|
|
110
|
+
|
|
111
|
+
1. Run paired evals (with-skill vs without-skill) periodically
|
|
112
|
+
2. Compute the delta in pass rates
|
|
113
|
+
3. If delta shrinks to near-zero across 3+ consecutive runs:
|
|
114
|
+
- The model has likely internalized the skill's knowledge
|
|
115
|
+
- The skill may be ready for retirement
|
|
116
|
+
4. Before retiring: run one final full eval suite to confirm
|
|
117
|
+
5. Archive the skill (don't delete — may need to restore if model changes)
|
|
118
|
+
|
|
119
|
+
Common cause of false retirement signals: model update changed capabilities.
|
|
120
|
+
Re-test after model updates.
|
|
121
|
+
|
|
122
|
+
## Tessl Three-Layer Eval Taxonomy
|
|
123
|
+
|
|
124
|
+
Three layers of increasing realism. Each catches failures the others miss.
|
|
125
|
+
|
|
126
|
+
### Layer 1: Skill Review (Structural Lint)
|
|
127
|
+
|
|
128
|
+
Does the skill itself follow best practices?
|
|
129
|
+
- SKILL.md structure (frontmatter, sections, length)
|
|
130
|
+
- Description quality (imperative, user-intent, trigger keywords)
|
|
131
|
+
- Reference organization (conditional triggers, one level deep)
|
|
132
|
+
- Convention compliance (for Plastic skills, see convention-checks.md)
|
|
133
|
+
|
|
134
|
+
Fast, cheap, runs without executing the skill.
|
|
135
|
+
|
|
136
|
+
### Layer 2: Task Evals (Synthetic)
|
|
137
|
+
|
|
138
|
+
Does the skill improve agent output on synthetic tasks?
|
|
139
|
+
- Paired with/without comparison
|
|
140
|
+
- Controlled prompts with known-good expected outputs
|
|
141
|
+
- Measures the delta the skill adds
|
|
142
|
+
|
|
143
|
+
This is the core eval loop (Steps 2-5 in the SKILL.md procedure).
|
|
144
|
+
|
|
145
|
+
### Layer 3: Repo Evals (Real Codebase)
|
|
146
|
+
|
|
147
|
+
Does the skill work correctly in a real project context?
|
|
148
|
+
- Install the skill in a real repository
|
|
149
|
+
- Run real tasks (not synthetic prompts)
|
|
150
|
+
- Measure whether the agent uses the skill correctly in situ
|
|
151
|
+
|
|
152
|
+
Most expensive, most realistic. Catches skills that pass synthetic tests
|
|
153
|
+
but fail under real project complexity. Use for high-stakes skills
|
|
154
|
+
or before shipping to users.
|
|
@@ -70,3 +70,7 @@ Add a wikilink in the `## Links` section of **both** intents (bidirectional).
|
|
|
70
70
|
|
|
71
71
|
### 4. Update INDEX.md Clusters
|
|
72
72
|
If both intents share a topic, ensure they're in the same cluster.
|
|
73
|
+
|
|
74
|
+
## References
|
|
75
|
+
|
|
76
|
+
- Read `references/zettelkasten.md` for the three Zettelkasten structures (Folgezettel, directed graph, tags), ID encoding rules, and dual-mode (Obsidian + programmatic) design
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# Zettelkasten Structure
|
|
2
|
+
|
|
3
|
+
Plastic implements three Zettelkasten structures:
|
|
4
|
+
|
|
5
|
+
| Structure | Implementation | Purpose |
|
|
6
|
+
|---|---|---|
|
|
7
|
+
| Folgezettel (linked list) | `sources` + `chain` in frontmatter | Sequential provenance |
|
|
8
|
+
| Directed graph (web of notes) | `## Links` with wikilinks | Obsidian navigation |
|
|
9
|
+
| Tag-based taxonomy | `tags` in frontmatter | Topic grouping |
|
|
10
|
+
|
|
11
|
+
INDEX.md is a structure note (hub), not a table of contents.
|
|
12
|
+
|
|
13
|
+
## Folgezettel IDs
|
|
14
|
+
|
|
15
|
+
IDs encode lineage using Luhmann's alternating convention:
|
|
16
|
+
- Root intents: sequential numbers (`1`, `2`, `3`...)
|
|
17
|
+
- Branches alternate letters and numbers: `1` → `1a` → `1a1` → `1a1a` → ...
|
|
18
|
+
- Multiple branches from the same parent increment: `1a`, `1b`, `1c`
|
|
19
|
+
- IDs are assigned at creation time and never change
|
|
20
|
+
|
|
21
|
+
## Knowledge Graph
|
|
22
|
+
|
|
23
|
+
`sources` + `chain` form the double-linked knowledge graph:
|
|
24
|
+
- `sources` = what fed into this intent (parents, inspirations, prerequisites)
|
|
25
|
+
- `chain` = what this intent produced (children, follow-ups, spin-offs)
|
|
26
|
+
|
|
27
|
+
## Dual-Mode
|
|
28
|
+
|
|
29
|
+
This store works in two modes without modification:
|
|
30
|
+
- **Obsidian** (human, offline) — browse, link, write markdown
|
|
31
|
+
- **Programmatic** (any agent) — read/write via filesystem operations
|
|
32
|
+
|
|
33
|
+
No special tooling required for either mode.
|
|
@@ -64,3 +64,7 @@ When 3+ intents share tags but aren't in a cluster, suggest a new cluster headin
|
|
|
64
64
|
Intents with no links (empty `sources`, empty `chain`, no `## Links` entries, not in any cluster) should be flagged for curation.
|
|
65
65
|
|
|
66
66
|
REQUIRED BACKGROUND: linking-intents (for understanding connection types and Zettelkasten theory)
|
|
67
|
+
|
|
68
|
+
## References
|
|
69
|
+
|
|
70
|
+
- Read `references/zettelkasten-linking.md` for the three structural layers (Folgezettel, directed graph, tags) and how they map to INDEX.md organization
|
|
@@ -193,3 +193,7 @@ git tag -a v0.1.0 <commit-sha> -m "v0.1.0 — [description]"
|
|
|
193
193
|
```
|
|
194
194
|
|
|
195
195
|
Use `git log --oneline` to find the right commits (look for version bump commits or major feature merges).
|
|
196
|
+
|
|
197
|
+
## References
|
|
198
|
+
|
|
199
|
+
- Read `references/deprecations.md` for the full deprecation process, severity levels, deprecations.yml schema, and dismissal rules when adding or managing deprecations
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Deprecation Process
|
|
2
|
+
|
|
3
|
+
When removing a feature, changing a convention, or making a breaking change:
|
|
4
|
+
|
|
5
|
+
1. Add entry to `deprecations.yml` in the Plastic source root
|
|
6
|
+
2. Set severity: `info` (awareness), `warning` (action needed), `critical` (urgent)
|
|
7
|
+
3. Provide clear migration steps
|
|
8
|
+
4. Set `removal` version at least 2 minor versions ahead
|
|
9
|
+
5. SessionStart hook displays active deprecations automatically
|
|
10
|
+
6. Remove the feature AND the deprecation entry together
|
|
11
|
+
|
|
12
|
+
## Severity Levels
|
|
13
|
+
|
|
14
|
+
| Severity | When to use | Dismissable? |
|
|
15
|
+
|----------|-------------|--------------|
|
|
16
|
+
| `info` | Upcoming change | Yes |
|
|
17
|
+
| `warning` | Action needed | Yes (re-shown at removal) |
|
|
18
|
+
| `critical` | Urgent/security | Never |
|
|
19
|
+
|
|
20
|
+
## deprecations.yml Schema
|
|
21
|
+
|
|
22
|
+
```yaml
|
|
23
|
+
deprecations:
|
|
24
|
+
- id: unique-slug
|
|
25
|
+
severity: info | warning | critical
|
|
26
|
+
summary: "One-line description"
|
|
27
|
+
migration_steps:
|
|
28
|
+
- "Step 1"
|
|
29
|
+
- "Step 2"
|
|
30
|
+
introduced: "0.9.0"
|
|
31
|
+
removal: "1.0.0"
|
|
32
|
+
link: "optional URL"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
## Dismissal
|
|
36
|
+
|
|
37
|
+
Users dismiss by adding the id to `deprecations_dismissed` in `config.yml`:
|
|
38
|
+
|
|
39
|
+
```yaml
|
|
40
|
+
deprecations_dismissed:
|
|
41
|
+
- some-deprecated-feature
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Critical deprecations and final-version warnings ignore dismissal.
|
|
@@ -55,3 +55,7 @@ git commit -m "chore: savepoint — [active intent name]"
|
|
|
55
55
|
|
|
56
56
|
### 5. Notify User
|
|
57
57
|
Tell the user: "Context is getting large. I've saved progress to intent [ID] — [name]. Please run `/clear` and say `continue` to resume."
|
|
58
|
+
|
|
59
|
+
## References
|
|
60
|
+
|
|
61
|
+
- Read `references/context-management.md` for the full save/continue protocol when the save flow needs debugging or you need to understand the full resume sequence
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# Context Management (Start-Save-Continue)
|
|
2
|
+
|
|
3
|
+
## Save Point
|
|
4
|
+
Triggered by PreCompact hook or manually:
|
|
5
|
+
1. Find active intent(s) from `~/.plastic/INDEX.md`
|
|
6
|
+
2. Update active intent's `checklist.md` (check off completed items)
|
|
7
|
+
3. Update active intent's `savepoint.md` (in-progress, next steps, blockers, discoveries)
|
|
8
|
+
4. Add observations to `## Insights`
|
|
9
|
+
5. Update INDEX.md
|
|
10
|
+
6. Commit: `cd ~/.plastic && git add . && git commit -m "chore: savepoint — [intent name]"`
|
|
11
|
+
7. Notify user to `/clear`
|
|
12
|
+
|
|
13
|
+
## Continue
|
|
14
|
+
Triggered by UserPromptSubmit hook when user says "continue". Priority order:
|
|
15
|
+
|
|
16
|
+
**1. Active intents first (resume work):**
|
|
17
|
+
1. Read INDEX.md → find active intent(s)
|
|
18
|
+
2. Read active intent's `intent.md` → what and why
|
|
19
|
+
3. Read active intent's `savepoint.md` → where we left off
|
|
20
|
+
4. Read active intent's `checklist.md` → what's next
|
|
21
|
+
5. Announce: intent name, current state, next step, blockers
|
|
22
|
+
6. Resume
|
|
23
|
+
|
|
24
|
+
**2. No active intents → offer future intents:**
|
|
25
|
+
1. List all future intents from INDEX.md
|
|
26
|
+
2. Present them as options
|
|
27
|
+
3. When user picks one, move to Active in INDEX.md
|
|
28
|
+
|
|
29
|
+
**3. Stale future intents (untouched 3+ days) → triage:**
|
|
30
|
+
- **activate** — start working on it now
|
|
31
|
+
- **abandon** — mark as abandoned
|
|
32
|
+
- **defer to agent** — implement, research, or ideate
|
|
@@ -0,0 +1,159 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: writing-instructions
|
|
3
|
+
description: >
|
|
4
|
+
Write or restructure agent instructions, conventions files, and SKILL.md
|
|
5
|
+
content using progressive disclosure and the agentskills.io specification.
|
|
6
|
+
Use when creating PLASTIC.md, rewriting convention docs, authoring new
|
|
7
|
+
skills, restructuring large instruction files to fit context budgets,
|
|
8
|
+
or when instructions feel too long, too vague, or agents aren't following
|
|
9
|
+
them. Also use when the user says "progressive disclosure", "restructure
|
|
10
|
+
instructions", or "the instructions are too big".
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# Writing Agent Instructions
|
|
14
|
+
|
|
15
|
+
Based on the [agentskills.io specification](https://agentskills.io).
|
|
16
|
+
|
|
17
|
+
## Progressive Disclosure Architecture
|
|
18
|
+
|
|
19
|
+
All agent instructions follow three loading stages. Design for this — it's the
|
|
20
|
+
architecture, not a nice-to-have.
|
|
21
|
+
|
|
22
|
+
| Stage | What loads | Budget | Design for |
|
|
23
|
+
|-------|-----------|--------|------------|
|
|
24
|
+
| **Discovery** | name + description | ~100 tokens | Trigger accuracy |
|
|
25
|
+
| **Activation** | Full instruction body | <5000 tokens / <500 lines | Core procedures |
|
|
26
|
+
| **Execution** | references/, scripts/, assets/ | As needed | Deep detail |
|
|
27
|
+
|
|
28
|
+
**The description carries the entire burden of triggering.** If agents aren't
|
|
29
|
+
activating your skill, the description is the problem.
|
|
30
|
+
|
|
31
|
+
## Procedure
|
|
32
|
+
|
|
33
|
+
### Step 1: Audit the content
|
|
34
|
+
|
|
35
|
+
Before writing or restructuring, classify every piece of content:
|
|
36
|
+
|
|
37
|
+
| Classification | Where it goes | Example |
|
|
38
|
+
|---------------|---------------|---------|
|
|
39
|
+
| **Trigger context** | description field | "Use when...", keywords |
|
|
40
|
+
| **Core procedure** | SKILL.md body | Step-by-step workflows |
|
|
41
|
+
| **Gotchas** | SKILL.md body (early) | Facts that defy assumptions |
|
|
42
|
+
| **Deep reference** | references/ | API details, full schemas |
|
|
43
|
+
| **Templates** | assets/ or inline | Output format examples |
|
|
44
|
+
| **Executable logic** | scripts/ | Validation, data processing |
|
|
45
|
+
|
|
46
|
+
Ask for each piece: "Would the agent get this wrong without this?"
|
|
47
|
+
If no — cut it. If unsure — test it.
|
|
48
|
+
|
|
49
|
+
### Step 2: Write the description
|
|
50
|
+
|
|
51
|
+
The description must convey WHEN to use, not just WHAT it does.
|
|
52
|
+
|
|
53
|
+
**Rules:**
|
|
54
|
+
- Imperative phrasing: "Use this skill when..." not "This skill does..."
|
|
55
|
+
- Focus on user intent, not implementation
|
|
56
|
+
- Err on the side of being pushy — list contexts explicitly
|
|
57
|
+
- Include cases where user doesn't name the domain directly
|
|
58
|
+
- Under 1024 characters (hard limit)
|
|
59
|
+
- Include specific trigger keywords
|
|
60
|
+
|
|
61
|
+
**Template:**
|
|
62
|
+
```yaml
|
|
63
|
+
description: >
|
|
64
|
+
[One sentence: what it does]. Use when [primary trigger context],
|
|
65
|
+
[secondary trigger], or when [indirect trigger where user doesn't
|
|
66
|
+
name the domain]. Also use when [edge case trigger].
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
### Step 3: Write the body
|
|
70
|
+
|
|
71
|
+
**Principles (in priority order):**
|
|
72
|
+
|
|
73
|
+
1. **Procedures over declarations** — Teach HOW to approach a class of problems,
|
|
74
|
+
not WHAT to produce for a specific instance
|
|
75
|
+
2. **Defaults, not menus** — Pick one approach, mention alternatives briefly.
|
|
76
|
+
"Use X. For Y cases, use Z instead." Never present options as equals.
|
|
77
|
+
3. **Calibrate control to fragility** — Prescriptive for fragile/sequential tasks,
|
|
78
|
+
flexible when multiple approaches are valid. Most skills have a mix.
|
|
79
|
+
4. **Reasoning over rigid directives** — "Do X because Y tends to cause Z" beats
|
|
80
|
+
"ALWAYS do X, NEVER do Y"
|
|
81
|
+
5. **Add what the agent lacks, omit what it knows** — No explaining HTTP, PDFs,
|
|
82
|
+
or what a migration is. Jump to what's non-obvious.
|
|
83
|
+
|
|
84
|
+
**Structure:**
|
|
85
|
+
```markdown
|
|
86
|
+
# Title
|
|
87
|
+
|
|
88
|
+
[1-2 sentence purpose statement]
|
|
89
|
+
|
|
90
|
+
## Gotchas
|
|
91
|
+
- [Highest-value content first — facts that defy assumptions]
|
|
92
|
+
- [Each gotcha is a concrete correction, not general advice]
|
|
93
|
+
|
|
94
|
+
## Procedure
|
|
95
|
+
### Step 1: ...
|
|
96
|
+
### Step 2: ...
|
|
97
|
+
|
|
98
|
+
## Patterns
|
|
99
|
+
[Only if the skill covers multiple approaches to similar problems]
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
### Step 4: Extract to references/
|
|
103
|
+
|
|
104
|
+
Anything over 500 lines or 5000 tokens goes to `references/`. But tell the
|
|
105
|
+
agent WHEN to load each file — conditional references, not a generic "see
|
|
106
|
+
references/ for details."
|
|
107
|
+
|
|
108
|
+
**Good:** "Read `references/api-errors.md` if the API returns a non-200 status"
|
|
109
|
+
**Bad:** "See references/ for more information"
|
|
110
|
+
|
|
111
|
+
### Step 5: Validate
|
|
112
|
+
|
|
113
|
+
Run through this checklist:
|
|
114
|
+
|
|
115
|
+
- [ ] Description under 1024 chars
|
|
116
|
+
- [ ] Body under 500 lines / 5000 tokens
|
|
117
|
+
- [ ] Every instruction passes "would the agent get this wrong without it?"
|
|
118
|
+
- [ ] Gotchas are concrete corrections, not general advice
|
|
119
|
+
- [ ] Defaults chosen, not menus presented
|
|
120
|
+
- [ ] Control calibrated: prescriptive where fragile, flexible where tolerant
|
|
121
|
+
- [ ] References have conditional load triggers
|
|
122
|
+
- [ ] No explaining what the agent already knows
|
|
123
|
+
|
|
124
|
+
For structured evaluation beyond this checklist (paired evals, pass rate
|
|
125
|
+
tracking, regression testing), use `plastic:evaluating-skills`.
|
|
126
|
+
|
|
127
|
+
## Gotchas
|
|
128
|
+
|
|
129
|
+
- Hook `additionalContext` truncates at **10,000 characters**. If instructions
|
|
130
|
+
must load via hooks (not skills), they must fit this budget.
|
|
131
|
+
- `~/.claude/rules/*.md` files load fully with no truncation — use for
|
|
132
|
+
always-loaded conventions that don't fit in a skill.
|
|
133
|
+
- CLAUDE.md supports `@path/to/file` imports (max depth 4) — another way
|
|
134
|
+
to load large instruction sets without truncation.
|
|
135
|
+
- Agents only consult skills for tasks beyond what they can handle alone.
|
|
136
|
+
Simple one-step requests may not trigger even with a perfect description.
|
|
137
|
+
- Content from real domain expertise (runbooks, incident reports, code review)
|
|
138
|
+
dramatically outperforms LLM-generated instructions without project context.
|
|
139
|
+
- The most common cause of agents not following instructions: the instruction
|
|
140
|
+
was too vague, didn't apply to the current task, or presented too many
|
|
141
|
+
options without a clear default. Read execution traces to diagnose.
|
|
142
|
+
|
|
143
|
+
## For Convention Files (PLASTIC.md, CLAUDE.md)
|
|
144
|
+
|
|
145
|
+
Convention files aren't skills — they load differently. Apply progressive
|
|
146
|
+
disclosure by splitting:
|
|
147
|
+
|
|
148
|
+
| Layer | Mechanism | Budget | Content |
|
|
149
|
+
|-------|-----------|--------|---------|
|
|
150
|
+
| **Always-on** | `~/.claude/rules/` or hook | <10K chars | Core identity, gotchas, critical procedures |
|
|
151
|
+
| **On-demand** | Skills (SKILL.md) | <5K tokens each | Specific workflows activated by task |
|
|
152
|
+
| **Deep reference** | `references/` in skills | Unlimited | Full specs, schemas, examples |
|
|
153
|
+
|
|
154
|
+
**The always-on layer should answer: "What does this agent need to know about
|
|
155
|
+
every single task?"** Everything else activates on demand.
|
|
156
|
+
|
|
157
|
+
## References
|
|
158
|
+
|
|
159
|
+
- Read `references/agentskills-spec.md` for the complete agentskills.io specification including frontmatter fields, description optimization methodology, evaluation framework, and script design requirements
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
# agentskills.io Full Reference
|
|
2
|
+
|
|
3
|
+
Source: https://agentskills.io (all sections, verified June 2026)
|
|
4
|
+
|
|
5
|
+
## Specification Details
|
|
6
|
+
|
|
7
|
+
### Frontmatter Fields
|
|
8
|
+
|
|
9
|
+
| Field | Required | Constraints |
|
|
10
|
+
|-------|----------|-------------|
|
|
11
|
+
| name | Yes | 1-64 chars. Lowercase alphanumeric + hyphens. No leading/trailing/consecutive hyphens. Must match directory name. |
|
|
12
|
+
| description | Yes | 1-1024 chars. Non-empty. What + when. |
|
|
13
|
+
| license | No | Short — name or filename reference |
|
|
14
|
+
| compatibility | No | 1-500 chars. Environment requirements only when needed. |
|
|
15
|
+
| metadata | No | String→string map. Use unique key names. |
|
|
16
|
+
| allowed-tools | No | Space-separated. Experimental. |
|
|
17
|
+
|
|
18
|
+
### Progressive Disclosure Token Budgets
|
|
19
|
+
|
|
20
|
+
- Discovery: ~100 tokens per skill (name + description only)
|
|
21
|
+
- Activation: <5000 tokens / <500 lines recommended for SKILL.md body
|
|
22
|
+
- Execution: Unbounded — files in scripts/, references/, assets/ load as needed
|
|
23
|
+
|
|
24
|
+
### File References
|
|
25
|
+
|
|
26
|
+
- Use relative paths from skill root
|
|
27
|
+
- Keep one level deep from SKILL.md
|
|
28
|
+
- Agent resolves paths automatically
|
|
29
|
+
|
|
30
|
+
## Description Optimization
|
|
31
|
+
|
|
32
|
+
### Evaluation Methodology
|
|
33
|
+
|
|
34
|
+
1. Create ~20 eval queries (8-10 should-trigger, 8-10 should-not)
|
|
35
|
+
2. Split 60/40 train/validation (proportional mix in each)
|
|
36
|
+
3. Run each query 3 times, compute trigger rate
|
|
37
|
+
4. Pass threshold: 0.5
|
|
38
|
+
5. Near-miss negatives are most valuable (share keywords, need different thing)
|
|
39
|
+
6. Iterate on train set only, validate on held-out set
|
|
40
|
+
7. Select best by validation pass rate, not last iteration
|
|
41
|
+
8. 5 iterations usually sufficient
|
|
42
|
+
|
|
43
|
+
### Description Anti-patterns
|
|
44
|
+
|
|
45
|
+
- "Helps with PDFs" — too vague, no trigger context
|
|
46
|
+
- "Process CSV files" — no when/why, no user-intent focus
|
|
47
|
+
- Implementation details instead of user intent
|
|
48
|
+
- Missing edge case triggers (user doesn't name the domain)
|
|
49
|
+
|
|
50
|
+
## Instruction Best Practices
|
|
51
|
+
|
|
52
|
+
### Gotchas — Highest-Value Content
|
|
53
|
+
|
|
54
|
+
Concrete corrections, not general advice:
|
|
55
|
+
|
|
56
|
+
```markdown
|
|
57
|
+
## Gotchas
|
|
58
|
+
- The `users` table uses soft deletes. Queries must include
|
|
59
|
+
`WHERE deleted_at IS NULL`.
|
|
60
|
+
- User ID is `user_id` in DB, `uid` in auth, `accountId` in billing.
|
|
61
|
+
All three refer to the same value.
|
|
62
|
+
- The `/health` endpoint returns 200 even if DB is down. Use `/ready`.
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
### Calibrating Control
|
|
66
|
+
|
|
67
|
+
Prescriptive when:
|
|
68
|
+
- Operations are fragile
|
|
69
|
+
- Consistency matters
|
|
70
|
+
- Specific sequence must be followed
|
|
71
|
+
|
|
72
|
+
Flexible when:
|
|
73
|
+
- Multiple approaches are valid
|
|
74
|
+
- Task tolerates variation
|
|
75
|
+
- Explaining WHY is more effective than rigid rules
|
|
76
|
+
|
|
77
|
+
### Instruction Patterns
|
|
78
|
+
|
|
79
|
+
1. **Validation loops**: Do work → validate → fix → repeat
|
|
80
|
+
2. **Plan-validate-execute**: Create plan → validate vs source of truth → execute
|
|
81
|
+
3. **Checklists**: Track progress, enforce dependencies, validation gates
|
|
82
|
+
4. **Bundled scripts**: If agent reinvents same logic each run, bundle it
|
|
83
|
+
5. **Templates**: Concrete output structures > prose descriptions
|
|
84
|
+
|
|
85
|
+
## Script Design
|
|
86
|
+
|
|
87
|
+
### Hard Requirements
|
|
88
|
+
- No interactive prompts (hard requirement — agents hang indefinitely)
|
|
89
|
+
- All input via flags, env vars, or stdin
|
|
90
|
+
|
|
91
|
+
### Agent-Friendly Design
|
|
92
|
+
- --help as primary interface documentation
|
|
93
|
+
- Helpful error messages: what wrong + what expected + what to try
|
|
94
|
+
- Structured output (JSON/CSV/TSV), data on stdout, diagnostics on stderr
|
|
95
|
+
- Idempotent operations (agents may retry)
|
|
96
|
+
- Dry-run for destructive operations
|
|
97
|
+
- Meaningful exit codes documented in --help
|
|
98
|
+
- Output size control: default to summaries, support --offset pagination
|
|
99
|
+
- Agent harnesses truncate at 10-30K characters
|
|
100
|
+
|
|
101
|
+
## Evaluation Framework
|
|
102
|
+
|
|
103
|
+
### Test Case Structure
|
|
104
|
+
```json
|
|
105
|
+
{
|
|
106
|
+
"skill_name": "name",
|
|
107
|
+
"evals": [{
|
|
108
|
+
"id": 1,
|
|
109
|
+
"prompt": "realistic user message",
|
|
110
|
+
"expected_output": "what success looks like",
|
|
111
|
+
"files": ["evals/files/input.csv"],
|
|
112
|
+
"assertions": ["specific, verifiable checks"]
|
|
113
|
+
}]
|
|
114
|
+
}
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
### Running Evals
|
|
118
|
+
- With-skill vs without-skill (or previous version) comparison
|
|
119
|
+
- Clean context per run (subagents or separate sessions)
|
|
120
|
+
- Capture timing: total_tokens, duration_ms
|
|
121
|
+
- Start with 2-3 test cases, expand after first results
|
|
122
|
+
|
|
123
|
+
### Assertion Quality
|
|
124
|
+
Good: Programmatically verifiable, specific, countable
|
|
125
|
+
Weak: Vague ("the output is good")
|
|
126
|
+
Brittle: Exact phrase matching
|
|
127
|
+
|
|
128
|
+
Principle: Require concrete evidence for PASS. No benefit of the doubt.
|
|
129
|
+
|
|
130
|
+
### Iteration Loop
|
|
131
|
+
1. Run evals → grade assertions → aggregate benchmarks
|
|
132
|
+
2. Identify failures (assertions, human feedback, execution transcripts)
|
|
133
|
+
3. Feed all three + SKILL.md to LLM for proposed changes
|
|
134
|
+
4. Apply changes → re-run → compare
|
|
135
|
+
5. Stop when consistently empty feedback or no meaningful improvement
|