@zalom/plastic 1.0.0-alpha.10 → 1.0.0-alpha.11
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/PLASTIC.md +106 -475
- package/package.json +1 -1
- package/skills/auto/SKILL.md +4 -0
- package/skills/auto/references/agent-architecture.md +60 -0
- package/skills/continuing/SKILL.md +4 -0
- package/skills/continuing/references/context-management.md +32 -0
- package/skills/creating-intent/SKILL.md +5 -0
- package/skills/creating-intent/references/lifecycle.md +74 -0
- package/skills/creating-intent/references/wikilinks.md +8 -0
- package/skills/creating-project/SKILL.md +4 -0
- package/skills/creating-project/references/hubs-projects.md +55 -0
- package/skills/doctor/SKILL.md +4 -0
- package/skills/doctor/references/gates-stuck-detection.md +38 -0
- package/skills/evaluating-skills/SKILL.md +140 -0
- package/skills/evaluating-skills/assets/eval-template.json +12 -0
- package/skills/evaluating-skills/evals/evals.json +75 -0
- package/skills/evaluating-skills/references/convention-checks.md +76 -0
- package/skills/evaluating-skills/references/eval-methodology.md +154 -0
- package/skills/linking-intents/SKILL.md +4 -0
- package/skills/linking-intents/references/zettelkasten.md +33 -0
- package/skills/managing-index/SKILL.md +4 -0
- package/skills/releasing/SKILL.md +4 -0
- package/skills/releasing/references/deprecations.md +44 -0
- package/skills/savepoint/SKILL.md +4 -0
- package/skills/savepoint/references/context-management.md +32 -0
- package/skills/writing-instructions/SKILL.md +159 -0
- package/skills/writing-instructions/references/agentskills-spec.md +135 -0
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Agent Architecture
|
|
2
|
+
|
|
3
|
+
## Main Orchestrator
|
|
4
|
+
|
|
5
|
+
The Main Orchestrator manages the global store (Main Knowledge Base). It:
|
|
6
|
+
- Recognizes, creates, updates, and groups intents
|
|
7
|
+
- Spawns Project Orchestrators for registered projects
|
|
8
|
+
- Receives contributions back from Project Orchestrators
|
|
9
|
+
- Is the only agent that runs in a loop (continuous Build→Observe→Repeat)
|
|
10
|
+
|
|
11
|
+
## Project Orchestrators
|
|
12
|
+
|
|
13
|
+
Project Orchestrators manage project stores (Project Knowledge Bases). They:
|
|
14
|
+
- Care about intents and execution within their project
|
|
15
|
+
- Spawn teams to develop and execute intents
|
|
16
|
+
- Contribute back to the Main Orchestrator when new intents are born
|
|
17
|
+
that could enrich the Main Knowledge Base
|
|
18
|
+
|
|
19
|
+
Rules:
|
|
20
|
+
- 1 Main Orchestrator : 1 Global Store (`~/.plastic/`)
|
|
21
|
+
- 1 Main Orchestrator : N Project Orchestrators
|
|
22
|
+
- 1 Project Orchestrator : 1 Project Store
|
|
23
|
+
- 1 Agent : 1 Intent (exclusive assignment)
|
|
24
|
+
- 1 Agent : N Sub-agents (for parallel Actions within an intent)
|
|
25
|
+
|
|
26
|
+
## Two Modes
|
|
27
|
+
|
|
28
|
+
- **Human-driven:** Human chats with Main Orchestrator, creates intents,
|
|
29
|
+
brainstorms, then Main Orchestrator dispatches Project Orchestrators and
|
|
30
|
+
Agents for execution.
|
|
31
|
+
- **Autonomous:** Human gives Main Orchestrator a starting intent with defined
|
|
32
|
+
outcomes. Main Orchestrator runs the full cycle — Agents do the lifecycle
|
|
33
|
+
(What→Why→How→Exec), Main Orchestrator reviews Insights, spawns next intents,
|
|
34
|
+
dispatches again.
|
|
35
|
+
|
|
36
|
+
## Autonomous Delivery
|
|
37
|
+
|
|
38
|
+
Human owns What and Why for human-initiated intents. Agent assists (research,
|
|
39
|
+
exploration) but human drives until handoff. When Why is complete — or human
|
|
40
|
+
triggers `plastic:auto` — the agent takes over How and Exec autonomously.
|
|
41
|
+
|
|
42
|
+
- **Safe-by-default:** Agent always prefers non-destructive routes (rename vs
|
|
43
|
+
delete, additive migrations, backups before changes). Destructive actions on
|
|
44
|
+
existing projects require human approval unless `--skip-permissions` is set.
|
|
45
|
+
- **One agent per intent.** Agent follows the full W→W→H→E lifecycle.
|
|
46
|
+
- **Notification only on:** finish or hard stop (blocked on destructive action,
|
|
47
|
+
unresolvable error). No progress reports — `## Insights` tracks everything.
|
|
48
|
+
- **Greenfield autonomy:** During initial project creation, all decisions are
|
|
49
|
+
non-destructive (nothing to destroy). Agent has full autonomy for greenfield choices.
|
|
50
|
+
- **Autonomous decisions** are logged in `## Insights` with `(autonomous)` marker.
|
|
51
|
+
|
|
52
|
+
## Coordinator Loop
|
|
53
|
+
|
|
54
|
+
When "work on Project X":
|
|
55
|
+
1. Read `projects.yml` → find project path
|
|
56
|
+
2. Load global config (defaults)
|
|
57
|
+
3. Load project config (overrides)
|
|
58
|
+
4. Load global INDEX.md → find hub intents tagged `project-<name>`
|
|
59
|
+
5. Load project INDEX.md → tactical intents
|
|
60
|
+
6. Coordinator has full picture, dispatches Agent teams
|
|
@@ -102,3 +102,7 @@ Auto-commit all triage changes.
|
|
|
102
102
|
2. **Project context** — if in a registered project, show governing intent + tactical intents
|
|
103
103
|
3. **Stale future intents** — surface for triage
|
|
104
104
|
4. **Fresh future intents** — offer as next work
|
|
105
|
+
|
|
106
|
+
## References
|
|
107
|
+
|
|
108
|
+
- Read `references/context-management.md` for the full save/continue protocol when resuming from a savepoint or when the resume flow needs debugging
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
# Context Management (Start-Save-Continue)
|
|
2
|
+
|
|
3
|
+
## Save Point
|
|
4
|
+
Triggered by PreCompact hook or manually:
|
|
5
|
+
1. Find active intent(s) from `~/.plastic/INDEX.md`
|
|
6
|
+
2. Update active intent's `checklist.md` (check off completed items)
|
|
7
|
+
3. Update active intent's `savepoint.md` (in-progress, next steps, blockers, discoveries)
|
|
8
|
+
4. Add observations to `## Insights`
|
|
9
|
+
5. Update INDEX.md
|
|
10
|
+
6. Commit: `cd ~/.plastic && git add . && git commit -m "chore: savepoint — [intent name]"`
|
|
11
|
+
7. Notify user to `/clear`
|
|
12
|
+
|
|
13
|
+
## Continue
|
|
14
|
+
Triggered by UserPromptSubmit hook when user says "continue". Priority order:
|
|
15
|
+
|
|
16
|
+
**1. Active intents first (resume work):**
|
|
17
|
+
1. Read INDEX.md → find active intent(s)
|
|
18
|
+
2. Read active intent's `intent.md` → what and why
|
|
19
|
+
3. Read active intent's `savepoint.md` → where we left off
|
|
20
|
+
4. Read active intent's `checklist.md` → what's next
|
|
21
|
+
5. Announce: intent name, current state, next step, blockers
|
|
22
|
+
6. Resume
|
|
23
|
+
|
|
24
|
+
**2. No active intents → offer future intents:**
|
|
25
|
+
1. List all future intents from INDEX.md
|
|
26
|
+
2. Present them as options
|
|
27
|
+
3. When user picks one, move to Active in INDEX.md
|
|
28
|
+
|
|
29
|
+
**3. Stale future intents (untouched 3+ days) → triage:**
|
|
30
|
+
- **activate** — start working on it now
|
|
31
|
+
- **abandon** — mark as abandoned
|
|
32
|
+
- **defer to agent** — implement, research, or ideate
|
|
@@ -120,3 +120,8 @@ cd <store-root> && git add . && git commit -m "feat: create intent ID — [name]
|
|
|
120
120
|
### 8. Announce
|
|
121
121
|
|
|
122
122
|
"Created intent ID — [name]. Placed in: [Active|Future]. Store: [global|project:<slug>|local]."
|
|
123
|
+
|
|
124
|
+
## References
|
|
125
|
+
|
|
126
|
+
- Read `references/lifecycle.md` for the full What→Why→How→Exec stage detail, filesystem-as-schema conventions, and creating-intent step-by-step
|
|
127
|
+
- Read `references/wikilinks.md` for the wikilink syntax table when adding `## Links` to intents
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Building an Intent — Full Lifecycle Detail
|
|
2
|
+
|
|
3
|
+
## What → `## Intent` section
|
|
4
|
+
|
|
5
|
+
The desire. One paragraph. What the human or agent wants.
|
|
6
|
+
This exists from the moment the intent is created.
|
|
7
|
+
|
|
8
|
+
**Deliverable:** `{ID}--{slug}.md`
|
|
9
|
+
|
|
10
|
+
## Why → `## Context` + `### Decisions` sections
|
|
11
|
+
|
|
12
|
+
Why this intent exists. Grows over time through brainstorming and exploration.
|
|
13
|
+
|
|
14
|
+
- **Context** — what we knew going in + what we decided along the way
|
|
15
|
+
- **Decisions** — main premises derived from Context plus decisions from brainstorming/grilling
|
|
16
|
+
- Decisions are Why-level: "status belongs on actions because multiple workstreams", not How-level: "use ACTION_N.md files"
|
|
17
|
+
|
|
18
|
+
**Deliverable:** `spec.md` (consolidated specification from Context + Decisions + brainstorming)
|
|
19
|
+
|
|
20
|
+
## How → Planning and preparation
|
|
21
|
+
|
|
22
|
+
Research decisions, create the implementation plan, define actions.
|
|
23
|
+
|
|
24
|
+
**Deliverable:** `plan.md` + `actions/` + `checklist.md` (execution registry with checkboxes covering all actions)
|
|
25
|
+
|
|
26
|
+
## Exec → Execute actions
|
|
27
|
+
|
|
28
|
+
Execute actions from the plan, track progress via checklist.
|
|
29
|
+
|
|
30
|
+
**Deliverable:** `outcome.md` (detailed result). `## Outcome` in intent.md = short summary written as last step.
|
|
31
|
+
|
|
32
|
+
## `## Insights` — Append-only work log
|
|
33
|
+
|
|
34
|
+
Captured throughout ALL stages. One-liner bullet points.
|
|
35
|
+
Never modified, only appended.
|
|
36
|
+
|
|
37
|
+
Tracks: stage transitions, decisions, shifts, blocks, cancellations, material for future intents.
|
|
38
|
+
This is how execution is tracked. When this intent completes, Insights
|
|
39
|
+
is where to look for what comes next. New intents spawned from Insights
|
|
40
|
+
appear in the `chain` field.
|
|
41
|
+
|
|
42
|
+
## `## Links`
|
|
43
|
+
|
|
44
|
+
Wikilinks for Obsidian graph navigation. Human-facing counterpart to the
|
|
45
|
+
frontmatter knowledge graph.
|
|
46
|
+
|
|
47
|
+
## Conventions — Filesystem as Schema
|
|
48
|
+
|
|
49
|
+
State is derived from what exists, not from what's declared.
|
|
50
|
+
|
|
51
|
+
| Convention | Signal |
|
|
52
|
+
|---|---|
|
|
53
|
+
| No `## Context` | Intent is fleeting (quick capture, non-actionable) |
|
|
54
|
+
| `## Context` has content | Intent is permanent (developed, actionable) |
|
|
55
|
+
| `## Outcome` has content | Intent is done |
|
|
56
|
+
| `## Insights` has `(autonomous)` entries | Intent is/was being delivered autonomously |
|
|
57
|
+
|
|
58
|
+
### Transitions
|
|
59
|
+
|
|
60
|
+
- Fleeting → permanent: add `## Context` (one-way, also makes it actionable)
|
|
61
|
+
- There is no separate "non-actionable → actionable" transition — permanence implies actionability
|
|
62
|
+
- Even research intents are actionable: the research itself is the action, the conclusion is the outcome
|
|
63
|
+
|
|
64
|
+
## Creating an Intent — Full Steps
|
|
65
|
+
|
|
66
|
+
1. Determine the target store: `~/.plastic/store/` for global intents (default), `~/.plastic/projects/{slug}/store/` for project intents
|
|
67
|
+
2. Determine the Folgezettel ID: If root (no parent), find highest root number +1. If branch, run `"${CLAUDE_PLUGIN_ROOT}/scripts/folgezettel-id" <parent_id> <store_path>`
|
|
68
|
+
3. Create the intent directory in the chosen store (e.g., `ID--three-to-five-words`)
|
|
69
|
+
4. Create `{ID}--{slug}.md` with frontmatter (id, intent, sources, chain, created, author, tags)
|
|
70
|
+
5. Write `## Intent` — the What
|
|
71
|
+
6. Add remaining sections: `## Context`, `## Outcome`, `## Insights`, `## Links`
|
|
72
|
+
7. Update the appropriate `INDEX.md` — add to Active section and appropriate cluster
|
|
73
|
+
|
|
74
|
+
A fleeting intent can skip `## Context` — just `## Intent` and empty sections.
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
# Wikilink Conventions
|
|
2
|
+
|
|
3
|
+
| Syntax | Meaning |
|
|
4
|
+
|--------|---------|
|
|
5
|
+
| `[[ID]]` | Link to intent in same store (e.g., `[[1a1]]`) |
|
|
6
|
+
| `[[ID\|display text]]` | Link with human-readable label |
|
|
7
|
+
| `[[global:ID]]` | Link to intent in `~/.plastic/store/` |
|
|
8
|
+
| `[[project-slug:ID]]` | Link to intent in `~/.plastic/projects/{slug}/store/` |
|
|
@@ -164,3 +164,7 @@ Log in `## Insights` of each founding intent:
|
|
|
164
164
|
|
|
165
165
|
Announce to user:
|
|
166
166
|
> "Project `<slug>` created at `<path>`. AGENTS.md populated with [N] decisions from [founding intent IDs]. Tactical mirror `1` is now the active intent in the project store."
|
|
167
|
+
|
|
168
|
+
## References
|
|
169
|
+
|
|
170
|
+
- Read `references/hubs-projects.md` for the full hub/project relationship model, project creation flow, and cross-linking conventions
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Hubs and Projects
|
|
2
|
+
|
|
3
|
+
## Hubs
|
|
4
|
+
|
|
5
|
+
A Hub is a cloud of intents around related topics. Hubs emerge naturally from
|
|
6
|
+
Folgezettel branching — intents that spawn in the same direction cluster.
|
|
7
|
+
|
|
8
|
+
- A Hub can spawn a Project. The Hub holds the founding ideas.
|
|
9
|
+
- A single intent can also spawn a Project.
|
|
10
|
+
- Hub-spawned projects revolve around different ideas around related topics.
|
|
11
|
+
- Intent-spawned projects revolve around the single founding intent.
|
|
12
|
+
- A Project is the deliverable outcome of one or more intents.
|
|
13
|
+
|
|
14
|
+
Hubs are represented as clusters in INDEX.md.
|
|
15
|
+
|
|
16
|
+
## Projects — Full Detail
|
|
17
|
+
|
|
18
|
+
A Project is a deliverable grouping of intents. Projects have two stores:
|
|
19
|
+
|
|
20
|
+
- **Global store** (`~/.plastic/store/`): strategic intents
|
|
21
|
+
- **Project store** (`~/.plastic/projects/{slug}/store/`): tactical intents
|
|
22
|
+
|
|
23
|
+
`projects.yml` maps project slugs to codebase paths:
|
|
24
|
+
```yaml
|
|
25
|
+
projects:
|
|
26
|
+
plastic:
|
|
27
|
+
path: "/path/to/plastic"
|
|
28
|
+
remote: "git@github.com:org/plastic.git"
|
|
29
|
+
registered: '2026-05-26'
|
|
30
|
+
status: active
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
Config resolution: `~/.plastic/projects/{slug}/config.yml` overrides `~/.plastic/config.yml`.
|
|
34
|
+
|
|
35
|
+
Cross-linking: project intents reference global intents via `[[global:ID]]`. Global intents reference project intents via `[[project-slug:ID]]`.
|
|
36
|
+
|
|
37
|
+
## Privacy and Collaboration
|
|
38
|
+
|
|
39
|
+
**Plastic is personal.** All intent data lives under `~/.plastic/` — one location,
|
|
40
|
+
one git repo, never pushed. Each person has their own intent store.
|
|
41
|
+
|
|
42
|
+
Collaboration happens through pull requests and project conventions, not shared intents.
|
|
43
|
+
When an intent delivers something that changes how a project works, the decision gets
|
|
44
|
+
written into the project's shared files (README, docs, config). The intents themselves
|
|
45
|
+
are private working memory.
|
|
46
|
+
|
|
47
|
+
## Project Creation Flow
|
|
48
|
+
|
|
49
|
+
When an implementation intent spawns a project:
|
|
50
|
+
1. Determine project path from config `project_roots` or intent context
|
|
51
|
+
2. `gh repo create --private` (agent-created repos are always private by default)
|
|
52
|
+
3. Set up project directory, git init, AGENTS.md with founding intent decisions
|
|
53
|
+
4. Register in `projects.yml`
|
|
54
|
+
5. Create tactical mirror in project store
|
|
55
|
+
6. The global intent completes; the tactical mirror becomes the active intent
|
package/skills/doctor/SKILL.md
CHANGED
|
@@ -110,3 +110,7 @@ This keeps the update flow clean when nothing is wrong.
|
|
|
110
110
|
- The script outputs JSON to stdout. Any diagnostic errors go to stderr.
|
|
111
111
|
- Non-zero exit codes mean "issues found", not "script crashed".
|
|
112
112
|
Always parse stdout regardless of exit code.
|
|
113
|
+
|
|
114
|
+
## References
|
|
115
|
+
|
|
116
|
+
- Read `references/gates-stuck-detection.md` for the full gate enforcement table, bridge file pattern, and stuck detection thresholds when diagnosing gate failures or stuck agents
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
# Gate Enforcement and Stuck Detection
|
|
2
|
+
|
|
3
|
+
## Bridge File Pattern
|
|
4
|
+
|
|
5
|
+
`/tmp/plastic-{session}.json` is the hot cache; the filesystem is the authority.
|
|
6
|
+
SessionStart rebuilds the bridge file from the filesystem — the bridge is always disposable.
|
|
7
|
+
|
|
8
|
+
## Gate Taxonomy
|
|
9
|
+
|
|
10
|
+
| Gate type | Purpose |
|
|
11
|
+
|---|---|
|
|
12
|
+
| **Pre-flight** | Prerequisites exist? |
|
|
13
|
+
| **Revision** | New info invalidated prior work? |
|
|
14
|
+
| **Escalation** | Blocked items surface to user |
|
|
15
|
+
| **Abort** | Inconsistent state detected |
|
|
16
|
+
|
|
17
|
+
## Full Gate Enforcement
|
|
18
|
+
|
|
19
|
+
| Gate | Blocked action | Required prerequisite |
|
|
20
|
+
|---|---|---|
|
|
21
|
+
| Pre-flight | Cannot write `plan.md` | `spec.md` must exist |
|
|
22
|
+
| Pre-flight | Cannot create `actions/` | `spec.md` must exist |
|
|
23
|
+
| Pre-flight | Cannot write `outcome.md` | `checklist.md` must exist with all items checked |
|
|
24
|
+
| Revision | Cannot proceed to Exec | Context changed after spec.md — re-derive spec |
|
|
25
|
+
| Abort | Cannot complete intent | `outcome.md` exists but checklist has unchecked items |
|
|
26
|
+
|
|
27
|
+
Hard blocking — hooks exit with code 2 when gates fail.
|
|
28
|
+
|
|
29
|
+
## Stuck Detection
|
|
30
|
+
|
|
31
|
+
| Condition | Threshold | Action |
|
|
32
|
+
|---|---|---|
|
|
33
|
+
| Consecutive gate failures | 3+ | Warning |
|
|
34
|
+
| Consecutive gate failures | 5+ | Force savepoint + escalate |
|
|
35
|
+
| No activity | 5+ min | Warning |
|
|
36
|
+
| No activity | 10+ min | Force savepoint + escalate |
|
|
37
|
+
| Context pressure | 80% | Warning |
|
|
38
|
+
| Context pressure | 90% | Force savepoint |
|
|
@@ -0,0 +1,140 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: evaluating-skills
|
|
3
|
+
description: >
|
|
4
|
+
Evaluate Plastic skills for correctness, convention compliance, and
|
|
5
|
+
progressive disclosure. Use when testing whether a skill produces good
|
|
6
|
+
outputs, verifying convention compliance after changes, running evals
|
|
7
|
+
against skills or instructions, creating evals for a new or updated
|
|
8
|
+
skill, checking if a description triggers correctly, or assessing
|
|
9
|
+
whether a skill is still needed. Also use when the user says "evaluate",
|
|
10
|
+
"test the skill", "run evals", "check conventions", or "write evals".
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# Evaluating Skills
|
|
14
|
+
|
|
15
|
+
Eval methodology for Plastic skills, based on
|
|
16
|
+
[agentskills.io](https://agentskills.io/skill-creation/evaluating-skills)
|
|
17
|
+
and [Anthropic's eval guide](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents).
|
|
18
|
+
|
|
19
|
+
## Gotchas
|
|
20
|
+
|
|
21
|
+
- Assertions written before observing output are almost always wrong — run
|
|
22
|
+
the eval first, observe actual output, THEN write assertions
|
|
23
|
+
- Near-miss negative test cases are the most valuable — prompts that share
|
|
24
|
+
keywords with should-trigger cases but need a different skill entirely
|
|
25
|
+
- Select the best skill iteration by validation pass rate, not the last one
|
|
26
|
+
- Grade outcomes, not execution paths — if the agent solved the task via an
|
|
27
|
+
unexpected route but produced correct output, that is a pass
|
|
28
|
+
- Same skill can behave differently across agent frameworks — test on each
|
|
29
|
+
target agent (Claude Code, Hermes, OpenClaw, Codex)
|
|
30
|
+
|
|
31
|
+
## Procedure
|
|
32
|
+
|
|
33
|
+
### Step 1: Choose eval scope
|
|
34
|
+
|
|
35
|
+
Determine what you are evaluating:
|
|
36
|
+
|
|
37
|
+
- **Description triggering** — does the agent activate the right skill for
|
|
38
|
+
a given prompt? Tests the description field effectiveness.
|
|
39
|
+
- **Output quality** — does the skill produce correct results when activated?
|
|
40
|
+
Tests the skill body and references.
|
|
41
|
+
- **Convention compliance** — does the output follow Plastic conventions?
|
|
42
|
+
Read `references/convention-checks.md` for the full assertion library.
|
|
43
|
+
|
|
44
|
+
Multiple scopes can apply to the same skill. Start with the scope that
|
|
45
|
+
addresses your immediate concern, add others as needed.
|
|
46
|
+
|
|
47
|
+
### Step 2: Design test cases
|
|
48
|
+
|
|
49
|
+
Create `evals/evals.json` in the skill being evaluated. Copy the starter
|
|
50
|
+
template from `assets/eval-template.json` in this skill.
|
|
51
|
+
|
|
52
|
+
**For description triggering:**
|
|
53
|
+
- Write ~20 queries: 8-10 should-trigger, 8-10 should-not-trigger
|
|
54
|
+
- Split 60/40 into train and validation sets (proportional mix in each)
|
|
55
|
+
- Include near-miss negatives that share keywords but need a different skill
|
|
56
|
+
- In `expected_output`, describe whether the skill should or should not activate
|
|
57
|
+
and why
|
|
58
|
+
|
|
59
|
+
**For output quality:**
|
|
60
|
+
- Start with 2-3 test cases, expand after first results
|
|
61
|
+
- Use realistic user prompts with varied phrasing, detail level, and formality
|
|
62
|
+
- In `expected_output`, describe what correct output looks like — not exact text
|
|
63
|
+
- Use `files` array for any input files the test needs
|
|
64
|
+
|
|
65
|
+
**For convention compliance:**
|
|
66
|
+
- Start with 2-3 test cases targeting specific convention areas
|
|
67
|
+
- In `expected_output`, describe which conventions must be met
|
|
68
|
+
- Read `references/convention-checks.md` for the full assertion library
|
|
69
|
+
|
|
70
|
+
Leave `assertions` arrays empty. They are populated after Step 4.
|
|
71
|
+
|
|
72
|
+
### Step 3: Run paired evals
|
|
73
|
+
|
|
74
|
+
Dispatch a subagent per test case to ensure clean context — no leakage
|
|
75
|
+
between test runs. Run each case twice:
|
|
76
|
+
|
|
77
|
+
1. **With skill** — the skill is available and loaded
|
|
78
|
+
2. **Without skill** — the skill is not available (baseline comparison)
|
|
79
|
+
|
|
80
|
+
The delta between with-skill and without-skill measures what the skill adds.
|
|
81
|
+
If the delta is negligible, the skill may not be adding value for that case.
|
|
82
|
+
|
|
83
|
+
Grade outcomes, not paths. An unexpected tool-call sequence that produces
|
|
84
|
+
correct output is still a pass.
|
|
85
|
+
|
|
86
|
+
### Step 4: Write assertions after observing
|
|
87
|
+
|
|
88
|
+
Review actual outputs from Step 3. Write specific, verifiable assertions
|
|
89
|
+
based on what you observed — not what you expected beforehand.
|
|
90
|
+
|
|
91
|
+
Choose the grader type that fits each assertion:
|
|
92
|
+
- **Code-based** — structure, file existence, format validity, counts
|
|
93
|
+
- **LLM-as-judge** — quality, completeness, tone, semantic correctness
|
|
94
|
+
- **Human** — edge cases, calibration, judgment calls
|
|
95
|
+
|
|
96
|
+
Read `references/eval-methodology.md` for the full grader taxonomy and
|
|
97
|
+
LLM-judge calibration protocol.
|
|
98
|
+
|
|
99
|
+
Good assertions are specific, verifiable, and countable:
|
|
100
|
+
"Output includes at least 3 concrete recommendations."
|
|
101
|
+
|
|
102
|
+
Weak assertions are vague: "Output is good."
|
|
103
|
+
Brittle assertions use exact phrase matching.
|
|
104
|
+
|
|
105
|
+
Require concrete evidence for PASS. No benefit of the doubt.
|
|
106
|
+
|
|
107
|
+
### Step 5: Grade and iterate
|
|
108
|
+
|
|
109
|
+
Compute pass rates per test case and aggregate across the eval suite.
|
|
110
|
+
|
|
111
|
+
Track two metrics separately:
|
|
112
|
+
- **pass@k** — succeeded at least once in k trials (capability)
|
|
113
|
+
- **pass^k** — succeeded every time in k trials (reliability)
|
|
114
|
+
|
|
115
|
+
Read `references/eval-methodology.md` for formulas and interpretation.
|
|
116
|
+
|
|
117
|
+
Three signal sources for skill improvement:
|
|
118
|
+
1. **Failed assertions** — specific gaps in the skill
|
|
119
|
+
2. **Human feedback** — broader quality issues not captured by assertions
|
|
120
|
+
3. **Execution transcripts** — reveals WHY things went wrong
|
|
121
|
+
|
|
122
|
+
Feed all three plus the current SKILL.md to propose targeted changes.
|
|
123
|
+
Iterate on the train set only. Check the validation set for generalization.
|
|
124
|
+
5 iterations is usually enough. If not improving after 5, the test cases
|
|
125
|
+
themselves may be the problem — revisit Step 2.
|
|
126
|
+
|
|
127
|
+
### Step 6: Graduate and monitor
|
|
128
|
+
|
|
129
|
+
Once a skill hits ~100% on capability evals (pass@k = 1.0 for 3+
|
|
130
|
+
consecutive runs):
|
|
131
|
+
|
|
132
|
+
1. **Graduate** capability evals into regression tests — run them on every
|
|
133
|
+
skill change to protect against backsliding
|
|
134
|
+
2. **Monitor** the with/without delta over time — if it shrinks to zero
|
|
135
|
+
across 3+ runs, the model may have internalized the skill
|
|
136
|
+
3. **Retire** cautiously — archive the skill, do not delete. Re-test after
|
|
137
|
+
model updates in case capabilities regress.
|
|
138
|
+
|
|
139
|
+
Read `references/eval-methodology.md` for graduation criteria, retirement
|
|
140
|
+
detection, and the three-layer eval taxonomy.
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
{
|
|
2
|
+
"skill_name": "evaluating-skills",
|
|
3
|
+
"evals": [
|
|
4
|
+
{
|
|
5
|
+
"id": 1,
|
|
6
|
+
"prompt": "I want to evaluate whether my creating-intent skill follows Plastic conventions",
|
|
7
|
+
"expected_output": "The skill should activate and guide the user through convention compliance evaluation: choose eval scope, design test cases using convention-checks reference, run paired evals, write assertions after observing.",
|
|
8
|
+
"files": [],
|
|
9
|
+
"assertions": []
|
|
10
|
+
},
|
|
11
|
+
{
|
|
12
|
+
"id": 2,
|
|
13
|
+
"prompt": "Run evals on the brainstorming skill to see if the description triggers correctly",
|
|
14
|
+
"expected_output": "The skill should activate and guide through description triggering evaluation: design ~20 queries with should-trigger and near-miss negatives, 60/40 train/validation split, compute trigger rates.",
|
|
15
|
+
"files": [],
|
|
16
|
+
"assertions": []
|
|
17
|
+
},
|
|
18
|
+
{
|
|
19
|
+
"id": 3,
|
|
20
|
+
"prompt": "Create evals for a new skill I just wrote for database migrations",
|
|
21
|
+
"expected_output": "The skill should activate, help create evals/evals.json in the migration skill directory using the template, guide through choosing eval scope (likely output quality), and design initial test cases.",
|
|
22
|
+
"files": [],
|
|
23
|
+
"assertions": []
|
|
24
|
+
},
|
|
25
|
+
{
|
|
26
|
+
"id": 4,
|
|
27
|
+
"prompt": "Check if my SKILL.md is under the token budget and follows progressive disclosure",
|
|
28
|
+
"expected_output": "The skill should activate and guide through convention compliance evaluation focused on progressive disclosure checks: body under 500 lines/5000 tokens, references have conditional triggers, description under 1024 chars.",
|
|
29
|
+
"files": [],
|
|
30
|
+
"assertions": []
|
|
31
|
+
},
|
|
32
|
+
{
|
|
33
|
+
"id": 5,
|
|
34
|
+
"prompt": "My skill's description isn't triggering on the right prompts, how do I fix it?",
|
|
35
|
+
"expected_output": "The skill should activate and guide through description triggering evaluation: create should-trigger and should-not-trigger queries, run paired evals, measure trigger rate, iterate on description wording.",
|
|
36
|
+
"files": [],
|
|
37
|
+
"assertions": []
|
|
38
|
+
},
|
|
39
|
+
{
|
|
40
|
+
"id": 6,
|
|
41
|
+
"prompt": "Write unit tests for my Ruby model that validates email addresses",
|
|
42
|
+
"expected_output": "The skill should NOT trigger. This is a code testing task, not a skill evaluation task. Near-miss negative — shares 'test' keyword but needs a different skill.",
|
|
43
|
+
"files": [],
|
|
44
|
+
"assertions": []
|
|
45
|
+
},
|
|
46
|
+
{
|
|
47
|
+
"id": 7,
|
|
48
|
+
"prompt": "Fix the bug in my login controller where sessions aren't persisting",
|
|
49
|
+
"expected_output": "The skill should NOT trigger. This is a debugging task with no relation to skill evaluation.",
|
|
50
|
+
"files": [],
|
|
51
|
+
"assertions": []
|
|
52
|
+
},
|
|
53
|
+
{
|
|
54
|
+
"id": 8,
|
|
55
|
+
"prompt": "Review this pull request for code quality issues",
|
|
56
|
+
"expected_output": "The skill should NOT trigger. Code review is different from skill evaluation. Near-miss negative — shares 'review/evaluate' concept.",
|
|
57
|
+
"files": [],
|
|
58
|
+
"assertions": []
|
|
59
|
+
},
|
|
60
|
+
{
|
|
61
|
+
"id": 9,
|
|
62
|
+
"prompt": "I updated my skill and want to make sure it still works correctly",
|
|
63
|
+
"expected_output": "The skill should activate and guide through regression evaluation: run existing evals against the updated skill, compare pass rates with previous version, check for backsliding.",
|
|
64
|
+
"files": [],
|
|
65
|
+
"assertions": []
|
|
66
|
+
},
|
|
67
|
+
{
|
|
68
|
+
"id": 10,
|
|
69
|
+
"prompt": "How do I know if a skill is still needed or if the model has learned the behavior?",
|
|
70
|
+
"expected_output": "The skill should activate and guide through skill retirement detection: monitor with/without delta over time, if delta approaches zero the model has internalized the skill.",
|
|
71
|
+
"files": [],
|
|
72
|
+
"assertions": []
|
|
73
|
+
}
|
|
74
|
+
]
|
|
75
|
+
}
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Plastic Convention Checks
|
|
2
|
+
|
|
3
|
+
Assertion library for convention compliance evals on Plastic skills.
|
|
4
|
+
Load when running convention compliance evals on a Plastic skill.
|
|
5
|
+
|
|
6
|
+
## Directory Structure
|
|
7
|
+
|
|
8
|
+
| Check | Pass criteria |
|
|
9
|
+
|-------|--------------|
|
|
10
|
+
| SKILL.md exists | File present at skill root |
|
|
11
|
+
| Name matches directory | `name` in frontmatter equals directory name |
|
|
12
|
+
| Standard directories only | Only scripts/, references/, assets/, evals/ at root level (all optional) |
|
|
13
|
+
| References one level deep | No nested directories inside references/ |
|
|
14
|
+
| No orphan files | Every file in the skill dir is referenced from SKILL.md or another skill file |
|
|
15
|
+
| evals/ has evals.json | If evals/ exists, it contains evals.json |
|
|
16
|
+
|
|
17
|
+
## Frontmatter
|
|
18
|
+
|
|
19
|
+
| Check | Pass criteria |
|
|
20
|
+
|-------|--------------|
|
|
21
|
+
| `name` present | Non-empty, 1-64 chars |
|
|
22
|
+
| `name` format | Lowercase alphanumeric + hyphens, no leading/trailing/consecutive hyphens |
|
|
23
|
+
| `name` matches dir | Exact match with skill directory name |
|
|
24
|
+
| `description` present | Non-empty, 1-1024 chars |
|
|
25
|
+
| `description` phrasing | Starts with imperative verb or "Use when" pattern |
|
|
26
|
+
| `description` triggers | Mentions at least one trigger context ("Use when...") |
|
|
27
|
+
| `description` edge cases | Includes at least one indirect trigger (user doesn't name the domain) |
|
|
28
|
+
| No unknown required fields | Only uses fields from agentskills.io spec: name, description, license, compatibility, metadata, allowed-tools |
|
|
29
|
+
|
|
30
|
+
## Progressive Disclosure
|
|
31
|
+
|
|
32
|
+
| Check | Pass criteria |
|
|
33
|
+
|-------|--------------|
|
|
34
|
+
| Body line count | SKILL.md body (excluding frontmatter) under 500 lines |
|
|
35
|
+
| Body token count | SKILL.md body under 5000 tokens (estimate: lines * 10) |
|
|
36
|
+
| Deep detail in references | Content exceeding activation budget lives in references/ |
|
|
37
|
+
| Conditional reference triggers | Every reference file mentioned in SKILL.md has a "when" condition |
|
|
38
|
+
| No generic references | No "see references/ for more" — each reference has specific load trigger |
|
|
39
|
+
| Description token budget | Description stays under ~100 tokens for discovery stage |
|
|
40
|
+
|
|
41
|
+
## Content Quality
|
|
42
|
+
|
|
43
|
+
| Check | Pass criteria |
|
|
44
|
+
|-------|--------------|
|
|
45
|
+
| Gotchas are concrete | Each gotcha is a specific correction, not general advice |
|
|
46
|
+
| Gotchas positioned early | Gotchas section appears before or near the top of procedures |
|
|
47
|
+
| Defaults, not menus | Skill picks one approach; alternatives mentioned briefly if at all |
|
|
48
|
+
| Procedures over declarations | Instructions teach HOW to approach, not WHAT to produce |
|
|
49
|
+
| Reasoning over rigid directives | Uses "Do X because Y" pattern, not "ALWAYS/NEVER" without rationale |
|
|
50
|
+
| No redundant knowledge | Every instruction passes "would the agent get this wrong without it?" |
|
|
51
|
+
| No explaining basics | Does not explain HTTP, JSON, what a migration is, etc. |
|
|
52
|
+
|
|
53
|
+
## Eval Quality (when evals/ exists)
|
|
54
|
+
|
|
55
|
+
| Check | Pass criteria |
|
|
56
|
+
|-------|--------------|
|
|
57
|
+
| Standard format | evals.json follows agentskills.io eval structure |
|
|
58
|
+
| Prompts are realistic | Each prompt reads like a real user message |
|
|
59
|
+
| Varied phrasing | Prompts use different wording, detail levels, formality |
|
|
60
|
+
| Near-miss negatives | At least 2 should-not-trigger cases that share keywords |
|
|
61
|
+
| Expected output descriptive | expected_output describes success, not exact text |
|
|
62
|
+
| Assertions (if populated) | Each assertion is specific, verifiable, and countable |
|
|
63
|
+
|
|
64
|
+
## Intent Structure (when evaluating intent compliance)
|
|
65
|
+
|
|
66
|
+
| Check | Pass criteria |
|
|
67
|
+
|-------|--------------|
|
|
68
|
+
| Intent file exists | `{ID}--{slug}.md` present in intent directory |
|
|
69
|
+
| Frontmatter complete | id, intent, sources, chain, created, author, tags present |
|
|
70
|
+
| ID format correct | Follows Luhmann alternating: digits and letters alternate |
|
|
71
|
+
| Directory name matches | `{ID}--{slug}` format, 3-5 word slug |
|
|
72
|
+
| Lifecycle artifacts | Present artifacts match the intent's lifecycle stage |
|
|
73
|
+
| spec.md gate | plan.md only exists if spec.md exists |
|
|
74
|
+
| plan.md gate | checklist.md only exists if plan.md exists |
|
|
75
|
+
| outcome.md gate | outcome.md only exists if all checklist items checked |
|
|
76
|
+
| Insights append-only | `## Insights` section only grows, never shrinks |
|