@supersuit/superskill 0.5.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +35 -0
- package/README.md +22 -12
- package/SPEC.md +109 -25
- package/package.json +1 -1
- package/src/commands/doctor.mjs +29 -3
- package/src/commands/miss.mjs +7 -4
- package/src/goldens.mjs +9 -2
- package/src/misses.mjs +22 -13
- package/src/realruns.mjs +122 -0
- package/src/rules/skill.mjs +71 -0
- package/src/rules/superskill.mjs +81 -43
- package/src/run/claude.mjs +29 -0
- package/src/run/codex.mjs +14 -0
- package/src/run/fake.mjs +7 -2
- package/src/run/index.mjs +54 -12
- package/src/snippet.md +3 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,40 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.6.0 (2026-10-06)
|
|
4
|
+
|
|
5
|
+
**A skill's embodiment is what it has absorbed, not one run of it.** 0.5.0 made a golden from a
|
|
6
|
+
real run the top-level requirement. Its first real test reversed it: nine harvested candidates,
|
|
7
|
+
each a real accepted run, and none was the standard the skill should be held to. The tests and
|
|
8
|
+
fixes a skill has absorbed are. **Breaking for anyone at superskill today:** the top level now
|
|
9
|
+
needs a real-run record and a 0.6.0 `--run` (one that records its `skill_sha` and the triggers).
|
|
10
|
+
|
|
11
|
+
- **Superskill is now** every miss closed with a regression eval (`misses-closed`, which replaces
|
|
12
|
+
`no-stale-open-miss`: an open miss fails however recent it is), and `evals-pass`: the last
|
|
13
|
+
`--run` was against the current `SKILL.md`, the suite passed (default 90% of runs), every fixed
|
|
14
|
+
miss's regression eval passed every run, and the trigger evals were right (default 90%).
|
|
15
|
+
`run-evidence` and `run-fresh` stand.
|
|
16
|
+
- **`real-runs`:** a real-run record where one model+harness pair, alone, has at least 5 real runs
|
|
17
|
+
at 80% one-shot (no correction). Every pair is reported; pairs are never pooled, and a pair that
|
|
18
|
+
says `unknown` never counts. The record is a harness-neutral file, `superskill-real-runs/1`, read
|
|
19
|
+
from `--real-runs <dir>` / `SUPERSKILL_REAL_RUNS`, then `evals/real-runs.json`, then Freedom's
|
|
20
|
+
skill ledger live. It replaces `real-use`.
|
|
21
|
+
- **Configurable bar:** `--min-real-runs`, `--min-one-shot`, `--min-pass-rate`,
|
|
22
|
+
`--min-trigger-rate` (and `SUPERSKILL_MIN_*`). Never frontmatter: a skill cannot lower its own bar.
|
|
23
|
+
- **Goldens are optional evidence.** `golden-approved` never fails. A golden is an eval graded on
|
|
24
|
+
`goldens/<id>/expectations.json`, a checklist of expected behavior (grade outcomes, not paths);
|
|
25
|
+
`output.md` is the reference that proves the task solvable. Provenance rules still say which
|
|
26
|
+
goldens are real runs.
|
|
27
|
+
- **`doctor --run` runs the trigger evals** (did the harness load the skill exactly when it
|
|
28
|
+
should) and writes `skill_sha` and `triggers` into `latest.json`.
|
|
29
|
+
- **`history-in-skill` (warn, level skill):** a dated incident story in `SKILL.md` ("Earned
|
|
30
|
+
2026-09-08", "(Gary, 2026-09-16: ...)", "on 2026-09-13 a session ...") belongs in `MISSES.md`,
|
|
31
|
+
which now holds a story's full entry: id, date, what happened, fix, eval, optional `Quote:`, with
|
|
32
|
+
indented continuation lines. Tuned on Freedom's 103 shipped skills. Still no line-count rule.
|
|
33
|
+
- `superskill miss` takes `--quote` and `--date`.
|
|
34
|
+
- Tests: 17 new; each guard broken on purpose and seen red (a stale-sha pass counting, pooled pairs
|
|
35
|
+
counting, an unknown pair counting, an open miss passing, a half-passing regression eval passing,
|
|
36
|
+
examples and provenance stamps flagged as history).
|
|
37
|
+
|
|
3
38
|
## 0.5.0 (2026-10-06)
|
|
4
39
|
|
|
5
40
|
**Superskill needs a golden from a real run.** Before, any golden a person approved counted, so a
|
package/README.md
CHANGED
|
@@ -4,8 +4,9 @@ Score any agent skill folder as **skill**, **tested**, or **superskill**, and ge
|
|
|
4
4
|
to-do list for the next level. Works on skills for Claude Code, Codex, or any harness that
|
|
5
5
|
reads the [Agent Skills](https://agentskills.io) format. Zero dependencies, Node 20 or later.
|
|
6
6
|
|
|
7
|
-
A superskill
|
|
8
|
-
|
|
7
|
+
A superskill is fixed every time it gets something wrong, with a test that keeps each fix fixed;
|
|
8
|
+
passes its whole eval suite on a current model; and does the job for real people without a
|
|
9
|
+
correction, often enough to count. [What that means](https://supersuit.wiki/concepts/superskill);
|
|
9
10
|
[the standard](SPEC.md).
|
|
10
11
|
|
|
11
12
|
## 30 seconds
|
|
@@ -34,12 +35,17 @@ file error. `--json` prints one JSON document and nothing else.
|
|
|
34
35
|
|
|
35
36
|
1. **skill**: a valid `SKILL.md` (name matches the folder, description of 1024 characters or
|
|
36
37
|
fewer that says when to use it, hard rules above the compaction fold, headings in long bodies, nothing said twice), references one level deep, no
|
|
37
|
-
hard-coded machine paths, nothing that reads like a prompt injection.
|
|
38
|
+
hard-coded machine paths, nothing that reads like a prompt injection. It warns on dated
|
|
39
|
+
incident stories in `SKILL.md`, which belong in `MISSES.md`.
|
|
38
40
|
2. **tested**: at least three task evals with checks a machine can verify, and a trigger set of
|
|
39
41
|
at least ten requests, some that should load the skill and some near-misses that should not.
|
|
40
|
-
3. **superskill**:
|
|
41
|
-
|
|
42
|
-
|
|
42
|
+
3. **superskill**: no open miss, and every fixed miss guarded by a regression eval; a recent
|
|
43
|
+
`--run` against the `SKILL.md` that is there now, where the suite passes (90%), every
|
|
44
|
+
regression eval passes every run, the trigger set is 90% right, and the skill beats the same
|
|
45
|
+
task done without it; and a real-run record where one model+harness has at least 5 real runs
|
|
46
|
+
at 80% one-shot (no correction). Pairs are never pooled. "Recent" follows `metadata.cadence`
|
|
47
|
+
(a weekly skill's proof lasts 30 days). Goldens are optional evidence. Every threshold is a
|
|
48
|
+
flag (`--min-real-runs`, `--min-one-shot`, `--min-pass-rate`, `--min-trigger-rate`).
|
|
43
49
|
|
|
44
50
|
Every rule and threshold is in [SPEC.md](SPEC.md).
|
|
45
51
|
|
|
@@ -52,7 +58,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
|
|
|
52
58
|
| `superskill doctor <skill> --run [--harness claude\|codex] [--repeat 3] [--yes]` | Run the evals for real, with and without the skill, and write `evals/results/latest.json`. **The only command that spends model calls**; it prints an estimate and asks first. |
|
|
53
59
|
| `superskill init <skill>` | Add missing `evals/`, `goldens/`, `MISSES.md`. Never overwrites. |
|
|
54
60
|
| `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
|
|
55
|
-
| `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
|
|
61
|
+
| `superskill miss <skill> "<what happened>" [--expected "..."] [--quote "..."] [--date YYYY-MM-DD]` | Log a time the skill got it wrong, or the story behind a rule it learned. |
|
|
56
62
|
| `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
|
|
57
63
|
| `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). At a terminal it asks for your name; from your phone, an agent records your tap with `--approved-by` and `--via`. Approvals accumulate. |
|
|
58
64
|
| `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
|
|
@@ -70,9 +76,10 @@ my-skill/
|
|
|
70
76
|
SKILL.md
|
|
71
77
|
evals/evals.json task evals (Anthropic skill-creator format)
|
|
72
78
|
evals/triggers.json should / should-not load (skill-creator format)
|
|
73
|
-
goldens/<id>/ input.md, output.md, PROVENANCE.json, APPROVAL.json
|
|
74
|
-
MISSES.md every miss, open or fixed with its eval
|
|
75
|
-
evals/results/latest.json the last --run
|
|
79
|
+
goldens/<id>/ optional: input.md, expectations.json, output.md, PROVENANCE.json, APPROVAL.json
|
|
80
|
+
MISSES.md every miss and its story, open or fixed with its eval
|
|
81
|
+
evals/results/latest.json the last --run (which SKILL.md it proved, suite and triggers)
|
|
82
|
+
evals/real-runs.json real uses per model+harness, one-shot or not (or --real-runs <dir>)
|
|
76
83
|
```
|
|
77
84
|
|
|
78
85
|
Harnesses ignore folders they do not know, so none of this changes how the skill loads.
|
|
@@ -84,13 +91,16 @@ Harnesses ignore folders they do not know, so none of this changes how the skill
|
|
|
84
91
|
- Codex: the skill is linked at `.agents/skills/<name>`. Codex cannot switch skills off, so a
|
|
85
92
|
copy installed in `~/.agents/skills` can leak into the baseline; move it aside while proving.
|
|
86
93
|
- Machine checks (`contains:`, `regex:`, `file_exists:`) are free. Each plain-language expectation
|
|
87
|
-
|
|
94
|
+
costs one grader call per run.
|
|
95
|
+
- Trigger evals run each query in `evals/triggers.json` with the skill installed and record
|
|
96
|
+
whether the harness loaded it (Claude Code: a `Skill` call or a read of its `SKILL.md`; Codex:
|
|
97
|
+
a read of its `SKILL.md`).
|
|
88
98
|
|
|
89
99
|
## Freedom
|
|
90
100
|
|
|
91
101
|
Nothing here needs [Freedom](https://getfreedom.wiki). If a skill has Freedom's `HDSOP.md`, the
|
|
92
102
|
doctor shows it as a bonus; `miss import --freedom-ledger` reads Freedom's run ledger as plain
|
|
93
|
-
files. `superskill snippet` gives any agent the same habits with no Freedom installed.
|
|
103
|
+
files, and with no `evals/real-runs.json` the doctor reads the real-run record from it too. `superskill snippet` gives any agent the same habits with no Freedom installed.
|
|
94
104
|
|
|
95
105
|
## Releasing
|
|
96
106
|
|
package/SPEC.md
CHANGED
|
@@ -1,13 +1,21 @@
|
|
|
1
1
|
# The superskill standard
|
|
2
2
|
|
|
3
|
-
**Version 0.
|
|
3
|
+
**Version 0.6.0** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
|
|
4
4
|
document and the checker disagree, the checker has a bug.
|
|
5
5
|
|
|
6
|
-
A **superskill** runs on frontier intelligence, is
|
|
7
|
-
|
|
6
|
+
A **superskill** runs on frontier intelligence, is fixed every time it gets something wrong, and
|
|
7
|
+
does the job for real people without a correction
|
|
8
8
|
([definition](https://supersuit.wiki/concepts/superskill)). This standard turns each clause into
|
|
9
9
|
a file in the skill's own folder, so the evidence travels with the skill wherever it is copied.
|
|
10
10
|
|
|
11
|
+
**What a skill has absorbed is its embodiment, not one run of it (0.6.0).** Until 0.5.0 the top
|
|
12
|
+
level required a golden: one real run a person approved as the standard. A random accepted run
|
|
13
|
+
is not the ultimate embodiment of a skill; the misses it was fixed for, and the evals that keep
|
|
14
|
+
each fix fixed, are. This follows Anthropic's guidance on agent evals: build the suite from real
|
|
15
|
+
failures, hold regression evals near 100%, grade outcomes rather than paths, and use a reference
|
|
16
|
+
solution to prove a task is solvable, not as the answer to match. Goldens remain, as optional
|
|
17
|
+
evidence.
|
|
18
|
+
|
|
11
19
|
## Contents
|
|
12
20
|
|
|
13
21
|
- [The clauses and their evidence](#the-clauses-and-their-evidence)
|
|
@@ -23,9 +31,10 @@ a file in the skill's own folder, so the evidence travels with the skill whereve
|
|
|
23
31
|
| The clause | What proves it | Where it lives |
|
|
24
32
|
|---|---|---|
|
|
25
33
|
| A skill at all | A valid `SKILL.md` under the Agent Skills spec, plus the hygiene rules below | `SKILL.md` |
|
|
26
|
-
|
|
|
27
|
-
|
|
|
28
|
-
|
|
|
34
|
+
| Fixed every time it gets something wrong | A miss log where every miss is fixed, each with a regression eval that would catch it again, and the story of each fix | `MISSES.md` + `evals/evals.json` |
|
|
35
|
+
| Runs on frontier intelligence | Its whole suite and trigger set last passed against the current `SKILL.md`, on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
|
|
36
|
+
| Does the job for real | Enough real runs on one model+harness, enough of them one-shot (no correction) | `evals/real-runs.json` (or an operator's export, see below) |
|
|
37
|
+
| Optional: an example a person approved | A golden: a real input, the checklist a right answer meets, a reference output, where it came from | `goldens/<id>/` (or a private folder) |
|
|
29
38
|
|
|
30
39
|
```
|
|
31
40
|
my-skill/
|
|
@@ -33,12 +42,14 @@ my-skill/
|
|
|
33
42
|
references/ scripts/ ... (as the Agent Skills spec allows)
|
|
34
43
|
evals/evals.json task evals (level: tested)
|
|
35
44
|
evals/triggers.json trigger evals (level: tested)
|
|
36
|
-
|
|
45
|
+
MISSES.md miss log + history (level: superskill)
|
|
46
|
+
evals/results/latest.json last --run (level: superskill)
|
|
47
|
+
evals/real-runs.json real-run record (level: superskill)
|
|
48
|
+
goldens/<id>/input.md optional evidence
|
|
49
|
+
goldens/<id>/expectations.json
|
|
37
50
|
goldens/<id>/output.md
|
|
38
51
|
goldens/<id>/PROVENANCE.json
|
|
39
52
|
goldens/<id>/APPROVAL.json
|
|
40
|
-
MISSES.md miss log (level: superskill)
|
|
41
|
-
evals/results/latest.json last --run (level: superskill)
|
|
42
53
|
```
|
|
43
54
|
|
|
44
55
|
## Compatibility
|
|
@@ -78,6 +89,7 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
|
|
|
78
89
|
| `body-size` | info | reports lines and estimated tokens (characters / 4) once the body passes about 5000 tokens. Length alone is never a defect. |
|
|
79
90
|
| `rules-above-the-fold` | warn | in a body past about 5000 tokens, every hard rule (a shouted NEVER, ALWAYS, MUST, DO NOT, REFUSE, or a bolded **Never ...** command, outside code fences) appears in the first 5000 tokens, either there or restated there. After compaction Claude Code keeps only that much of each invoked skill. |
|
|
80
91
|
| `reference-says-when` | warn | every link from SKILL.md to a markdown file sits on a line that says when to read it (before, when, if, for, read ...). Step files (`steps/<step>.md`) are the recommended way to keep a long skill's detail out of the always-loaded body: they are read fresh when the step comes up, so compaction does not lose them. |
|
|
92
|
+
| `history-in-skill` | warn | no dated incident story in the body outside code: "Earned 2026-09-08", "(Gary, 2026-09-16: ...)", "Wilson, 2026-09-05: \*"..."\*", "on 2026-09-13 a session ...", "until 2026-09-21 the flag ...", "measured 2026-09-20", "(2026-09-30, #324)", "- 2026-09-15 (#147): ...". Matched per paragraph, since a line wrap can split a name from its date. Not flagged: a line with "e.g." or "example", a date in a code span or fence, a heading, an HTML comment (a generator's provenance stamp), "as of <date>". The story goes in `MISSES.md` (below); `SKILL.md` keeps the rule and a one-line why. There is deliberately no line-count rule (retired 2026-09-28): length alone is never a defect. |
|
|
81
93
|
| `navigable` | warn | a body over 300 lines has no run of more than 150 lines without a heading |
|
|
82
94
|
| `no-repeated-paragraphs` | warn | no paragraph of 100+ characters appears twice |
|
|
83
95
|
| `references-one-deep` | fail | a markdown file linked from `SKILL.md` links on to no further local file |
|
|
@@ -99,13 +111,20 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
|
|
|
99
111
|
|
|
100
112
|
| Rule | Severity | Threshold |
|
|
101
113
|
|---|---|---|
|
|
102
|
-
| `golden-approved` | fail | at least one golden **from a real run** (its `PROVENANCE.json` says `source: real-run`, names the run, and names who accepted it and when; an anonymized twin also carries `ANONYMIZED.json`) has an approval with non-empty `approved_by` and a valid `approved_at`. An approved golden without that provenance is named and does not count. Info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
|
|
103
|
-
| `real-use` | info | when the run ledger records what the person's next message made of each run (Freedom's `next_turn`), the share of those runs in the last 30 days they accepted; sandbox runs are never counted |
|
|
104
114
|
| `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
|
|
105
|
-
| `
|
|
115
|
+
| `misses-closed` | fail | no miss is open. A miss is a known way the skill fails; a skill that still fails a known way is not at the top level, however recent the miss |
|
|
106
116
|
| `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
|
|
117
|
+
| `evals-pass` | fail | the last `--run` was against the current `SKILL.md` (`skill_sha` equals the sha256 of `SKILL.md`); its suite passed at least the minimum share of runs with the skill (default 90%); every fixed miss's regression eval passed every run; and it ran the trigger evals and got at least the minimum share right (default 90%) |
|
|
107
118
|
| `run-evidence` | fail | `evals/results/latest.json` exists and `with_skill.pass_rate` > `without_skill.pass_rate` |
|
|
108
119
|
| `run-fresh` | fail / warn | the last run is younger than the cadence window: `daily` or `weekly` 30 days, `monthly` 60, `quarterly` 120, `yearly` 365, none declared 30 (info). A `yearly` skill always warns to `--run` before its next real use. An unknown cadence warns |
|
|
120
|
+
| `real-runs` | fail | a real-run record (below) in which at least one model+harness pair, on its own, has at least the minimum runs (default 5) and one-shot share (default 80%). Each pair is reported as info; **pairs are never pooled**, since a clean record on one model in one harness proves nothing about another. A pair whose model or harness is `unknown` is reported and never counts |
|
|
121
|
+
| `golden-approved` | info / warn | never fails. Reports each golden that is approved and from a real run (with its approval weight), names approved goldens that are not from a real run and why, names a golden with no `expectations.json`, and warns on an `APPROVAL.json` or `PROVENANCE.json` that does not parse |
|
|
122
|
+
|
|
123
|
+
**The thresholds are configurable, never by the skill itself.** A skill cannot lower its own bar,
|
|
124
|
+
so they are not frontmatter. `doctor` takes `--min-real-runs <n>`, `--min-one-shot <0..1>`,
|
|
125
|
+
`--min-pass-rate <0..1>` and `--min-trigger-rate <0..1>`, or the environment variables
|
|
126
|
+
`SUPERSKILL_MIN_REAL_RUNS`, `SUPERSKILL_MIN_ONE_SHOT`, `SUPERSKILL_MIN_PASS_RATE` and
|
|
127
|
+
`SUPERSKILL_MIN_TRIGGER_RATE`. A report scored with non-default thresholds should say so.
|
|
109
128
|
|
|
110
129
|
`metadata.cadence` in `SKILL.md` frontmatter declares how often the skill really runs:
|
|
111
130
|
|
|
@@ -155,11 +174,18 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
|
|
|
155
174
|
]
|
|
156
175
|
```
|
|
157
176
|
|
|
158
|
-
### `goldens/<id>/`
|
|
177
|
+
### `goldens/<id>/` (optional evidence)
|
|
178
|
+
|
|
179
|
+
**A golden is an eval whose grading is a checklist of expected behavior (0.6.0).** It is never
|
|
180
|
+
required for any level. Grade the outcome, not the path: the checklist says what a right answer
|
|
181
|
+
does, and the reference output proves the task is solvable and calibrates the grader. A golden
|
|
182
|
+
that exists is still held to the provenance rules below, because "a person accepted this when
|
|
183
|
+
it ran" is a claim that has to be true.
|
|
159
184
|
|
|
160
185
|
- `input.md`: the real request.
|
|
161
|
-
- `
|
|
162
|
-
|
|
186
|
+
- `expectations.json`: an array of expectations (the same forms as `evals.json`), graded by
|
|
187
|
+
`--run`. Without it, `--run` grades by likeness to the reference output.
|
|
188
|
+
- `output.md` (or any other file that is not `input.*` or a metadata file): the reference output.
|
|
163
189
|
- `APPROVAL.json`, written only by `superskill approve`: at an interactive terminal, or relayed by an agent with `--approved-by` and `--via` after the person approved with a tap (the `via` field records where):
|
|
164
190
|
|
|
165
191
|
```json
|
|
@@ -184,11 +210,9 @@ approval whose `note` is its rationale.
|
|
|
184
210
|
A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
|
|
185
211
|
output.
|
|
186
212
|
|
|
187
|
-
**`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run
|
|
188
|
-
|
|
189
|
-
right" on something no person's work produced
|
|
190
|
-
author's fiction. Invented cases still belong in `evals/evals.json`, where they hold a skill at
|
|
191
|
-
`tested`.
|
|
213
|
+
**`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run is
|
|
214
|
+
reported as real evidence. An invented input with an invented output puts "a person said this
|
|
215
|
+
was right" on something no person's work produced. Invented cases belong in `evals/evals.json`.
|
|
192
216
|
|
|
193
217
|
```json
|
|
194
218
|
{
|
|
@@ -208,13 +232,12 @@ author's fiction. Invented cases still belong in `evals/evals.json`, where they
|
|
|
208
232
|
anonymizer's receipt (`checker`, `checked_at`, `counts` by kind, `fingerprint`, never the
|
|
209
233
|
mapping). Without the receipt the twin does not count.
|
|
210
234
|
- `superskill init --from-session` writes `source: real-run` with `accepted` empty, so the golden
|
|
211
|
-
|
|
235
|
+
is not reported as real until someone records who accepted it.
|
|
212
236
|
|
|
213
237
|
**Private goldens.** A real run's input and output are usually about real people and should not
|
|
214
238
|
travel with a skill that is shared. `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`)
|
|
215
239
|
makes `doctor`, `approve` and `doctor --run` also read `<dir>/<skill-name>/<id>/`, laid out exactly
|
|
216
|
-
like `goldens/<id>/`. A
|
|
217
|
-
by an eval or golden that ships with the skill.
|
|
240
|
+
like `goldens/<id>/`. A miss is closed only by an eval or golden that ships with the skill.
|
|
218
241
|
|
|
219
242
|
### `MISSES.md`
|
|
220
243
|
|
|
@@ -226,11 +249,29 @@ by an eval or golden that ships with the skill.
|
|
|
226
249
|
- Should have: Merged duplicates into one line.
|
|
227
250
|
- Fix: a1b2c3d
|
|
228
251
|
- Eval: m1
|
|
252
|
+
- Quote: "why is Atlas in here twice" (optional)
|
|
229
253
|
- Source: freedom-ledger inv_2026-09-20T10-00-00Z_ab12 (optional)
|
|
230
254
|
```
|
|
231
255
|
|
|
232
256
|
Headings are `## <id> · <YYYY-MM-DD> · <open|fixed>`; `|` or `-` also separate. Ids are `m1`,
|
|
233
|
-
`m2`, ... Other headings are ignored.
|
|
257
|
+
`m2`, ... Other headings are ignored. A field runs on over indented lines that follow it.
|
|
258
|
+
|
|
259
|
+
**`MISSES.md` is where a skill's history lives (0.6.0).** `SKILL.md` is instructions, loaded on
|
|
260
|
+
every run; the story of the incident that earned a rule is history, and Anthropic's authoring
|
|
261
|
+
guidance is to keep time-sensitive content out of instructions. So a story becomes a miss entry:
|
|
262
|
+
|
|
263
|
+
| Field | Holds |
|
|
264
|
+
|---|---|
|
|
265
|
+
| id, date | `m<N>`, and the date the incident happened (not the date it was written down) |
|
|
266
|
+
| What happened | the incident, in full: what the skill did, what it cost, what was noticed |
|
|
267
|
+
| Should have | what the skill should have done |
|
|
268
|
+
| Fix | the commit, or the rule now in `SKILL.md` that the incident earned |
|
|
269
|
+
| Eval | the regression eval that would catch it again. Empty until one exists, and the doctor says so (`fixed-miss-has-eval`): a story with no eval is a fix nothing guards yet |
|
|
270
|
+
| Quote | optional: what the person said, verbatim |
|
|
271
|
+
| Source | optional: where it was recorded (a ledger id, an issue) |
|
|
272
|
+
|
|
273
|
+
`SKILL.md` keeps the rule and a one-line why. `superskill miss <skill> "<what>" --date <d> --quote
|
|
274
|
+
"<q>"` writes one.
|
|
234
275
|
|
|
235
276
|
### `evals/results/latest.json`
|
|
236
277
|
|
|
@@ -241,15 +282,53 @@ Written only by `superskill doctor --run`:
|
|
|
241
282
|
"run_at": "2026-09-20T10:00:00.000Z",
|
|
242
283
|
"harness": "claude",
|
|
243
284
|
"model": "<model id the harness reported>",
|
|
285
|
+
"skill_sha": "<sha256 of the SKILL.md the run proved>",
|
|
244
286
|
"cases": 4,
|
|
245
287
|
"repeat": 3,
|
|
246
288
|
"with_skill": { "pass_rate": 1.0, "mean_ms": 21000, "mean_tokens": 4100 },
|
|
247
289
|
"without_skill": { "pass_rate": 0.33, "mean_ms": 18000, "mean_tokens": 3900 },
|
|
248
|
-
"per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }]
|
|
290
|
+
"per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }],
|
|
291
|
+
"triggers": { "cases": 12, "runs": 36, "passes": 35, "pass_rate": 0.972,
|
|
292
|
+
"per_query": [{ "query": "write my weekly status", "should_trigger": true, "runs": 3, "passes": 3 }] }
|
|
249
293
|
}
|
|
250
294
|
```
|
|
251
295
|
|
|
252
296
|
A run passes when every expectation of its case passes; `pass_rate` is passing runs over runs.
|
|
297
|
+
A trigger run passes when the harness loaded the skill exactly when `should_trigger` says
|
|
298
|
+
(Claude Code: a `Skill` tool call naming it or a read of its `SKILL.md`; Codex: a read of its
|
|
299
|
+
`SKILL.md`). A golden with `expectations.json` is graded on that checklist; one without is graded
|
|
300
|
+
by likeness to its reference output.
|
|
301
|
+
|
|
302
|
+
### `evals/real-runs.json` (the real-run record)
|
|
303
|
+
|
|
304
|
+
Harness-neutral: any harness, ledger or script may write it, and the doctor only reads it.
|
|
305
|
+
|
|
306
|
+
```json
|
|
307
|
+
{
|
|
308
|
+
"format": "superskill-real-runs/1",
|
|
309
|
+
"skill": "weekly-status",
|
|
310
|
+
"generated_at": "2026-10-06T12:00:00Z",
|
|
311
|
+
"source": "freedom-skill-ledger",
|
|
312
|
+
"pairs": [
|
|
313
|
+
{ "model": "claude-opus-5-5", "harness": "claude-code", "runs": 12, "one_shot": 11,
|
|
314
|
+
"first_at": "2026-09-01T09:00:00Z", "last_at": "2026-10-05T18:00:00Z" }
|
|
315
|
+
]
|
|
316
|
+
}
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
- A **run** is one real use of the skill by a person. Never a sandbox run (`doctor --run`), never
|
|
320
|
+
a test.
|
|
321
|
+
- It is **one-shot** when the person needed no correction, rescue or redirect, it did not fail or
|
|
322
|
+
get abandoned, and it was not corrected after it handed back. A taste note is the person's
|
|
323
|
+
preference, not the skill's defect, and does not break one-shot.
|
|
324
|
+
- `model` is the model id the harness ran (as the transcript or environment reports it);
|
|
325
|
+
`harness` is `claude-code`, `codex`, or another harness's name. A writer that cannot tell
|
|
326
|
+
writes `unknown`, never a guess. Counts only: the file carries no input, output or names, so it
|
|
327
|
+
can ship with the skill.
|
|
328
|
+
|
|
329
|
+
**Where the doctor reads it**, first found wins: `<dir>/<skill-name>.json` where `<dir>` is
|
|
330
|
+
`--real-runs` or `SUPERSKILL_REAL_RUNS` (an operator's own export, kept out of the skill); then
|
|
331
|
+
`evals/real-runs.json` in the skill; then Freedom's skill ledger, read live (below).
|
|
253
332
|
|
|
254
333
|
## Collections and plugins
|
|
255
334
|
|
|
@@ -285,6 +364,11 @@ exist they are read as plain files:
|
|
|
285
364
|
sandbox run) is skipped. A record corrected after it handed back (`corrected_after`, with
|
|
286
365
|
`next_turn_ref`) gets a "Should have" that points at the person's correction: the session and
|
|
287
366
|
the time, never their words.
|
|
367
|
+
- **The real-run record.** Freedom's skill ledger records, on every run, the `model` id and the
|
|
368
|
+
`harness` (`claude-code`, `codex`, ...), read from the session transcript or the environment and
|
|
369
|
+
`unknown` when neither says. Freedom exports it in the `superskill-real-runs/1` format above,
|
|
370
|
+
one file per skill, for `--real-runs`. With no export and no `evals/real-runs.json`, the doctor
|
|
371
|
+
reads the same ledger files live and counts them the same way (sandbox runs skipped).
|
|
288
372
|
- **`doctor --run` is never recorded as a use.** Every harness spawns its child with
|
|
289
373
|
`FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`, in a folder named `superskill-run-*`.
|
|
290
374
|
|
package/package.json
CHANGED
package/src/commands/doctor.mjs
CHANGED
|
@@ -9,8 +9,9 @@ export const help = `superskill doctor <path...> [options]
|
|
|
9
9
|
|
|
10
10
|
Score one skill, a folder of skills, or a plugin (skills/*/SKILL.md).
|
|
11
11
|
Levels: skill (spec-valid, hygienic) < tested (real evals + triggers) < superskill
|
|
12
|
-
(
|
|
13
|
-
--run that beats no-skill
|
|
12
|
+
(every miss fixed with a regression eval; the suite and triggers pass against this SKILL.md
|
|
13
|
+
in a fresh --run that beats no-skill; and a real-run record: at least 5 real runs, 80%
|
|
14
|
+
one-shot, on one model+harness). Goldens are optional evidence.
|
|
14
15
|
|
|
15
16
|
Options:
|
|
16
17
|
--level <skill|tested|superskill> target level for the exit code (default skill)
|
|
@@ -23,6 +24,13 @@ Options:
|
|
|
23
24
|
model calls; see "superskill doctor --run --help")
|
|
24
25
|
--private-goldens <dir> also read goldens from <dir>/<skill-name>/<id>/ (or set
|
|
25
26
|
SUPERSKILL_PRIVATE_GOLDENS): real runs kept out of the skill
|
|
27
|
+
--real-runs <dir> read the real-run record from <dir>/<skill-name>.json (or
|
|
28
|
+
set SUPERSKILL_REAL_RUNS); else evals/real-runs.json, else
|
|
29
|
+
Freedom's skill ledger
|
|
30
|
+
--min-real-runs <n> real runs one model+harness needs (default 5)
|
|
31
|
+
--min-one-shot <0..1> one-shot share that pair needs (default 0.8)
|
|
32
|
+
--min-pass-rate <0..1> suite pass rate the last --run needs (default 0.9)
|
|
33
|
+
--min-trigger-rate <0..1> trigger evals right in the last --run (default 0.9)
|
|
26
34
|
--now <iso date> evaluate dates as of this moment
|
|
27
35
|
--help this text
|
|
28
36
|
|
|
@@ -37,7 +45,7 @@ export async function run(argv) {
|
|
|
37
45
|
const level = a.flags.level || "skill";
|
|
38
46
|
if (!LEVELS.includes(level)) throw new UsageError(`--level must be one of ${LEVELS.join(", ")}`);
|
|
39
47
|
if (!a._.length) throw new UsageError("doctor needs a path");
|
|
40
|
-
const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null };
|
|
48
|
+
const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null, ...thresholds(a.flags) };
|
|
41
49
|
if (a.flags.changed) {
|
|
42
50
|
const all = a._.flatMap((p) => findSkills(p));
|
|
43
51
|
const files = a._.flatMap((p) => changedFiles(p, a.flags.base));
|
|
@@ -64,3 +72,21 @@ export async function run(argv) {
|
|
|
64
72
|
}
|
|
65
73
|
return result.ok ? 0 : 1;
|
|
66
74
|
}
|
|
75
|
+
|
|
76
|
+
/** The superskill bar's thresholds, from flags then environment. Unset means the documented default. */
|
|
77
|
+
export function thresholds(flags, env = process.env) {
|
|
78
|
+
const pick = (flag, envName, lo, hi) => {
|
|
79
|
+
const raw = flags[flag] ?? env[envName];
|
|
80
|
+
if (raw === undefined || raw === "") return undefined;
|
|
81
|
+
const n = Number(raw);
|
|
82
|
+
if (!Number.isFinite(n) || n < lo || (hi !== null && n > hi)) throw new UsageError(`--${flag} must be a number${hi === null ? ` of at least ${lo}` : ` from ${lo} to ${hi}`}`);
|
|
83
|
+
return n;
|
|
84
|
+
};
|
|
85
|
+
return {
|
|
86
|
+
realRuns: flags["real-runs"] || env.SUPERSKILL_REAL_RUNS || null,
|
|
87
|
+
minRealRuns: pick("min-real-runs", "SUPERSKILL_MIN_REAL_RUNS", 1, null),
|
|
88
|
+
minOneShot: pick("min-one-shot", "SUPERSKILL_MIN_ONE_SHOT", 0, 1),
|
|
89
|
+
minPassRate: pick("min-pass-rate", "SUPERSKILL_MIN_PASS_RATE", 0, 1),
|
|
90
|
+
minTriggerRate: pick("min-trigger-rate", "SUPERSKILL_MIN_TRIGGER_RATE", 0, 1),
|
|
91
|
+
};
|
|
92
|
+
}
|
package/src/commands/miss.mjs
CHANGED
|
@@ -2,11 +2,12 @@ import { parseArgs, clock, UsageError } from "../args.mjs";
|
|
|
2
2
|
import { skillDir, today } from "./common.mjs";
|
|
3
3
|
import { readMisses, appendMisses, nextMissId } from "../misses.mjs";
|
|
4
4
|
|
|
5
|
-
export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"]
|
|
5
|
+
export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"] [--quote "<what they said>"] [--date YYYY-MM-DD]
|
|
6
6
|
superskill miss import <skill> --freedom-ledger [--ledger <file>]
|
|
7
7
|
|
|
8
|
-
Log a time the skill got something wrong. The entry opens today
|
|
9
|
-
|
|
8
|
+
Log a time the skill got something wrong. The entry opens today (or on --date, for a story
|
|
9
|
+
from the past); an open miss blocks the superskill level. Close it with \`superskill fix\`.
|
|
10
|
+
This is also where the story behind a rule goes: SKILL.md keeps the rule and a one-line why.
|
|
10
11
|
`;
|
|
11
12
|
|
|
12
13
|
export async function run(argv) {
|
|
@@ -18,7 +19,9 @@ export async function run(argv) {
|
|
|
18
19
|
if (!what) throw new UsageError('miss needs a description: superskill miss <skill> "<what happened>"');
|
|
19
20
|
const misses = readMisses(dir) || [];
|
|
20
21
|
const id = nextMissId(misses);
|
|
21
|
-
|
|
22
|
+
const date = a.flags.date || today(clock(a.flags));
|
|
23
|
+
if (!/^\d{4}-\d{2}-\d{2}$/.test(date)) throw new UsageError("--date must be YYYY-MM-DD");
|
|
24
|
+
appendMisses(dir, [{ id, date, status: "open", what, expected: a.flags.expected || "", quote: a.flags.quote || "" }]);
|
|
22
25
|
process.stdout.write(`${id} logged (open). Close it with: superskill fix ${a._[0]} ${id} --eval <id>\n`);
|
|
23
26
|
return 0;
|
|
24
27
|
}
|
package/src/goldens.mjs
CHANGED
|
@@ -1,4 +1,5 @@
|
|
|
1
|
-
// goldens/<id>/: input.md,
|
|
1
|
+
// goldens/<id>/: input.md, expectations.json (0.6.0: the checklist a right answer meets,
|
|
2
|
+
// graded by --run), the reference output (output.md, or any other non-input file),
|
|
2
3
|
// APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
|
|
3
4
|
// rationale, basis, evidence}] } (or the single-approval shape from before 0.3.0),
|
|
4
5
|
// PROVENANCE.json saying where the example came from (0.5.0), and, for an anonymized twin of a
|
|
@@ -11,7 +12,7 @@ import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
|
|
|
11
12
|
import { join } from "node:path";
|
|
12
13
|
|
|
13
14
|
/** Files in a golden folder that describe it rather than being its input or output. */
|
|
14
|
-
export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json"]);
|
|
15
|
+
export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json", "expectations.json"]);
|
|
15
16
|
|
|
16
17
|
/** Where goldens are read from: the skill's own goldens/, then the private folder for its name. */
|
|
17
18
|
export function goldenRoots(dir, { privateGoldens = null, name = null } = {}) {
|
|
@@ -40,6 +41,8 @@ export function readGoldens(dir, opts = {}) {
|
|
|
40
41
|
const prov = readJsonFile(join(gdir, "PROVENANCE.json"));
|
|
41
42
|
const anon = readJsonFile(join(gdir, "ANONYMIZED.json"));
|
|
42
43
|
const provenance = prov.value;
|
|
44
|
+
const exp = readJsonFile(join(gdir, "expectations.json"));
|
|
45
|
+
const expList = Array.isArray(exp.value) ? exp.value : Array.isArray(exp.value?.expectations) ? exp.value.expectations : [];
|
|
43
46
|
out.push({
|
|
44
47
|
id,
|
|
45
48
|
dir: gdir,
|
|
@@ -54,6 +57,10 @@ export function readGoldens(dir, opts = {}) {
|
|
|
54
57
|
provenanceError: prov.error,
|
|
55
58
|
anonymized: anon.value,
|
|
56
59
|
origin: originOf(provenance, anon.value),
|
|
60
|
+
// 0.6.0: what a right answer does, graded by --run. output.md is the reference that
|
|
61
|
+
// proves the task is solvable; it is no longer what the output has to look like.
|
|
62
|
+
expectations: expList.filter((x) => typeof x === "string" && x.trim()),
|
|
63
|
+
expectationsError: exp.error,
|
|
57
64
|
});
|
|
58
65
|
}
|
|
59
66
|
}
|
package/src/misses.mjs
CHANGED
|
@@ -1,41 +1,49 @@
|
|
|
1
|
-
// MISSES.md: one dated entry per time the skill got something wrong.
|
|
1
|
+
// MISSES.md: one dated entry per time the skill got something wrong. It is also where a
|
|
2
|
+
// skill's history lives (0.6.0): the story of the incident that earned a rule goes here, and
|
|
3
|
+
// SKILL.md keeps only the rule and a one-line why.
|
|
2
4
|
//
|
|
3
5
|
// ## m1 · 2026-09-20 · fixed
|
|
4
|
-
// - What happened: ...
|
|
6
|
+
// - What happened: ... (an indented line under a field continues it)
|
|
5
7
|
// - Should have: ...
|
|
6
|
-
// - Fix: <commit sha or note>
|
|
8
|
+
// - Fix: <commit sha or note: the rule now in SKILL.md>
|
|
7
9
|
// - Eval: m1
|
|
10
|
+
// - Quote: "what the person said" (optional)
|
|
8
11
|
// - Source: freedom-ledger <id> (optional)
|
|
9
12
|
import { readFileSync, existsSync, writeFileSync } from "node:fs";
|
|
10
13
|
import { join } from "node:path";
|
|
11
14
|
|
|
12
15
|
export const MISSES_HEADER = `# Misses
|
|
13
16
|
|
|
14
|
-
Every time this skill got something wrong
|
|
15
|
-
superskill level; a fixed miss must name the eval that would catch it again.
|
|
17
|
+
Every time this skill got something wrong, and the story behind each rule it learned. An open
|
|
18
|
+
miss blocks the superskill level; a fixed miss must name the eval that would catch it again.
|
|
16
19
|
Written by \`superskill miss\` and \`superskill fix\`, and readable by hand.
|
|
17
20
|
`;
|
|
18
21
|
|
|
19
22
|
const HEAD_RE = /^##\s+(m\d+)\s*[·|\-–]\s*(\d{4}-\d{2}-\d{2})\s*[·|\-–]\s*(open|fixed)\s*$/i;
|
|
20
|
-
const FIELDS = { "what happened": "what", "should have": "expected", fix: "fix", eval: "eval", source: "source" };
|
|
23
|
+
const FIELDS = { "what happened": "what", "should have": "expected", fix: "fix", eval: "eval", quote: "quote", source: "source" };
|
|
21
24
|
|
|
22
25
|
export function parseMisses(text) {
|
|
23
26
|
const out = [];
|
|
24
27
|
let cur = null;
|
|
28
|
+
let last = null;
|
|
25
29
|
const lines = String(text).replace(/\r\n?/g, "\n").split("\n");
|
|
26
30
|
lines.forEach((line, i) => {
|
|
27
31
|
const h = line.match(HEAD_RE);
|
|
28
32
|
if (h) {
|
|
29
|
-
cur = { id: h[1].toLowerCase(), date: h[2], status: h[3].toLowerCase(), what: "", expected: "", fix: "", eval: "", source: "", line: i + 1 };
|
|
33
|
+
cur = { id: h[1].toLowerCase(), date: h[2], status: h[3].toLowerCase(), what: "", expected: "", fix: "", eval: "", quote: "", source: "", line: i + 1 };
|
|
34
|
+
last = null;
|
|
30
35
|
out.push(cur);
|
|
31
36
|
return;
|
|
32
37
|
}
|
|
33
|
-
if (/^##\s/.test(line)) { cur = null; return; }
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
38
|
+
if (/^##\s/.test(line)) { cur = null; last = null; return; }
|
|
39
|
+
if (!cur) return;
|
|
40
|
+
const f = line.match(/^\s*[-*]\s*([A-Za-z ]+?)\s*:\s*(.*)$/);
|
|
41
|
+
const key = f && FIELDS[f[1].toLowerCase()];
|
|
42
|
+
if (key) { cur[key] = f[2].trim(); last = key; return; }
|
|
43
|
+
// A story rarely fits on one line: an indented line continues the field above it.
|
|
44
|
+
if (last && /^\s{2,}\S/.test(line)) { cur[last] = `${cur[last]} ${line.trim()}`.trim(); return; }
|
|
45
|
+
if (!line.trim()) return;
|
|
46
|
+
last = null;
|
|
39
47
|
});
|
|
40
48
|
return out;
|
|
41
49
|
}
|
|
@@ -50,6 +58,7 @@ export function formatMiss(m) {
|
|
|
50
58
|
const lines = [`## ${m.id} · ${m.date} · ${m.status}`, `- What happened: ${m.what}`];
|
|
51
59
|
if (m.expected) lines.push(`- Should have: ${m.expected}`);
|
|
52
60
|
lines.push(`- Fix: ${m.fix || ""}`, `- Eval: ${m.eval || ""}`);
|
|
61
|
+
if (m.quote) lines.push(`- Quote: ${m.quote}`);
|
|
53
62
|
if (m.source) lines.push(`- Source: ${m.source}`);
|
|
54
63
|
return lines.join("\n") + "\n";
|
|
55
64
|
}
|
package/src/realruns.mjs
ADDED
|
@@ -0,0 +1,122 @@
|
|
|
1
|
+
// The real-run record (0.6.0): how often the skill did the job for a real person with no
|
|
2
|
+
// correction, counted separately for every model and harness it ran on.
|
|
3
|
+
//
|
|
4
|
+
// Harness-neutral on purpose. Any harness, ledger or script can write the file; the doctor only
|
|
5
|
+
// reads it. Format `superskill-real-runs/1`:
|
|
6
|
+
//
|
|
7
|
+
// {
|
|
8
|
+
// "format": "superskill-real-runs/1",
|
|
9
|
+
// "skill": "weekly-status",
|
|
10
|
+
// "generated_at": "2026-10-06T12:00:00Z",
|
|
11
|
+
// "source": "freedom-skill-ledger",
|
|
12
|
+
// "pairs": [
|
|
13
|
+
// { "model": "claude-opus-5-5", "harness": "claude-code", "runs": 12, "one_shot": 11,
|
|
14
|
+
// "first_at": "2026-09-01T09:00:00Z", "last_at": "2026-10-05T18:00:00Z" }
|
|
15
|
+
// ]
|
|
16
|
+
// }
|
|
17
|
+
//
|
|
18
|
+
// A run is one real use: never a sandbox run (`doctor --run`), never a test. It is one-shot when
|
|
19
|
+
// the person needed no correction, rescue or redirect, it did not fail or get abandoned, and it
|
|
20
|
+
// was not corrected after it handed back. A taste note is the person's preference, not a defect,
|
|
21
|
+
// and does not break one-shot.
|
|
22
|
+
//
|
|
23
|
+
// Where the doctor reads it, first found wins:
|
|
24
|
+
// 1. <dir>/<skill-name>.json, where <dir> is --real-runs or SUPERSKILL_REAL_RUNS (an
|
|
25
|
+
// operator's own export, kept out of the skill);
|
|
26
|
+
// 2. <skill>/evals/real-runs.json (counts only, so it can travel with the skill);
|
|
27
|
+
// 3. Freedom's skill ledger, read as plain files, when neither exists.
|
|
28
|
+
import { readFileSync, existsSync } from "node:fs";
|
|
29
|
+
import { join } from "node:path";
|
|
30
|
+
import { ledgerPaths } from "./ledger.mjs";
|
|
31
|
+
|
|
32
|
+
export const FORMAT = "superskill-real-runs/1";
|
|
33
|
+
export const UNKNOWN = "unknown";
|
|
34
|
+
/** The defaults the real-runs rule holds a skill to. `--min-real-runs` / `--min-one-shot` change them. */
|
|
35
|
+
export const MIN_REAL_RUNS = 5;
|
|
36
|
+
export const MIN_ONE_SHOT = 0.8;
|
|
37
|
+
|
|
38
|
+
const MISS_KINDS = new Set(["redirect", "correction", "rescue"]);
|
|
39
|
+
|
|
40
|
+
/** Is this a known model+harness pair? A record that could not say proves nothing about either. */
|
|
41
|
+
export const knownPair = (p) => Boolean(p.model && p.harness && p.model !== UNKNOWN && p.harness !== UNKNOWN);
|
|
42
|
+
|
|
43
|
+
function normalize(doc, where) {
|
|
44
|
+
if (!doc || typeof doc !== "object") return { error: `${where} is not a JSON object` };
|
|
45
|
+
if (doc.format && doc.format !== FORMAT) return { error: `${where} has format ${JSON.stringify(doc.format)}, not ${FORMAT}` };
|
|
46
|
+
if (!Array.isArray(doc.pairs)) return { error: `${where} has no pairs array` };
|
|
47
|
+
const pairs = doc.pairs
|
|
48
|
+
.filter((p) => p && typeof p === "object")
|
|
49
|
+
.map((p) => ({
|
|
50
|
+
model: String(p.model || UNKNOWN),
|
|
51
|
+
harness: String(p.harness || UNKNOWN),
|
|
52
|
+
runs: Math.max(0, Number(p.runs) || 0),
|
|
53
|
+
one_shot: Math.max(0, Number(p.one_shot) || 0),
|
|
54
|
+
first_at: p.first_at || null,
|
|
55
|
+
last_at: p.last_at || null,
|
|
56
|
+
}))
|
|
57
|
+
.map((p) => ({ ...p, one_shot: Math.min(p.one_shot, p.runs) }));
|
|
58
|
+
return { pairs, generated_at: doc.generated_at || null, source: doc.source || null, where };
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
function readFile(p, where) {
|
|
62
|
+
try { return normalize(JSON.parse(readFileSync(p, "utf8")), where); }
|
|
63
|
+
catch (e) { return { error: `${where} is not valid JSON: ${e.message}` }; }
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
/** Is a ledger record a one-shot run? Exported so a writer can count the same way the reader does. */
|
|
67
|
+
export function isOneShot(rec) {
|
|
68
|
+
const kinds = (Array.isArray(rec?.interventions) ? rec.interventions : []).filter((i) => i && MISS_KINDS.has(i.kind));
|
|
69
|
+
return !kinds.length && !["failed", "abandoned"].includes(rec?.outcome) && !rec?.corrected_after;
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
/** Summarize ledger records (Freedom's shape, or any with model/harness/started) per model+harness. */
|
|
73
|
+
export function pairsFromRecords(records, skillName = null) {
|
|
74
|
+
const by = new Map();
|
|
75
|
+
for (const rec of records) {
|
|
76
|
+
if (!rec || rec.synthetic) continue;
|
|
77
|
+
if (skillName && rec.skill && bare(rec.skill) !== bare(skillName)) continue;
|
|
78
|
+
const model = String(rec.model || UNKNOWN), harness = String(rec.harness || UNKNOWN);
|
|
79
|
+
const k = `${model}\u0000${harness}`;
|
|
80
|
+
const p = by.get(k) || { model, harness, runs: 0, one_shot: 0, first_at: null, last_at: null };
|
|
81
|
+
p.runs++;
|
|
82
|
+
if (isOneShot(rec)) p.one_shot++;
|
|
83
|
+
const t = typeof rec.started === "string" ? rec.started : null;
|
|
84
|
+
if (t && (!p.first_at || t < p.first_at)) p.first_at = t;
|
|
85
|
+
if (t && (!p.last_at || t > p.last_at)) p.last_at = t;
|
|
86
|
+
by.set(k, p);
|
|
87
|
+
}
|
|
88
|
+
return [...by.values()].sort((a, b) => b.runs - a.runs || a.model.localeCompare(b.model));
|
|
89
|
+
}
|
|
90
|
+
|
|
91
|
+
const bare = (name) => String(name || "").split(":").pop();
|
|
92
|
+
|
|
93
|
+
function fromLedger(dir, name) {
|
|
94
|
+
const paths = ledgerPaths(dir, name);
|
|
95
|
+
if (!paths.length) return null;
|
|
96
|
+
const recs = [];
|
|
97
|
+
for (const p of paths) {
|
|
98
|
+
let text = "";
|
|
99
|
+
try { text = readFileSync(p, "utf8"); } catch { continue; }
|
|
100
|
+
for (const line of text.split("\n")) {
|
|
101
|
+
if (!line.trim()) continue;
|
|
102
|
+
try { recs.push(JSON.parse(line)); } catch {}
|
|
103
|
+
}
|
|
104
|
+
}
|
|
105
|
+
return { pairs: pairsFromRecords(recs, name), generated_at: null, source: "freedom-skill-ledger (read live)", where: "Freedom's skill ledger" };
|
|
106
|
+
}
|
|
107
|
+
|
|
108
|
+
/** The real-run record for one skill, or null when nothing records any. */
|
|
109
|
+
export function readRealRuns(dir, { name, realRunsDir = null } = {}) {
|
|
110
|
+
if (realRunsDir && name) {
|
|
111
|
+
const p = join(realRunsDir, `${name}.json`);
|
|
112
|
+
if (existsSync(p)) return readFile(p, `${realRunsDir}/${name}.json`);
|
|
113
|
+
}
|
|
114
|
+
const local = join(dir, "evals", "real-runs.json");
|
|
115
|
+
if (existsSync(local)) return readFile(local, "evals/real-runs.json");
|
|
116
|
+
return fromLedger(dir, name);
|
|
117
|
+
}
|
|
118
|
+
|
|
119
|
+
/** Does any single pair clear the bar? Never pooled: each pair stands or falls on its own runs. */
|
|
120
|
+
export function meetsBar(pairs, { minRuns = MIN_REAL_RUNS, minOneShot = MIN_ONE_SHOT } = {}) {
|
|
121
|
+
return pairs.filter(knownPair).filter((p) => p.runs >= minRuns && p.one_shot / p.runs >= minOneShot);
|
|
122
|
+
}
|
package/src/rules/skill.mjs
CHANGED
|
@@ -26,6 +26,63 @@ function proseLines(body) {
|
|
|
26
26
|
return { line, fenced };
|
|
27
27
|
});
|
|
28
28
|
}
|
|
29
|
+
// A dated incident story, as written in real skills: "Earned 2026-09-08", "(Gary, 2026-09-16:
|
|
30
|
+
// ...)", "Wilson, live, 2026-09-05: *"...", "on 2026-09-13 the in-process version sat...",
|
|
31
|
+
// "until 2026-09-21 the flag...", "measured on 2026-09-20", "(2026-09-30, #324)",
|
|
32
|
+
// "- 2026-09-15 (#147): ...". Matched per paragraph, because hard-wrapped prose puts the name on
|
|
33
|
+
// one line and its date on the next. Tuned against Freedom's 103 shipped skills (2026-10-06): a
|
|
34
|
+
// date in an example, a template, a code span, a fence, or a provenance line ending at the date
|
|
35
|
+
// ("vendored ... on 2026-09-19.") is not a story.
|
|
36
|
+
const D = "20\\d\\d-\\d\\d-\\d\\d";
|
|
37
|
+
const HISTORY = [
|
|
38
|
+
new RegExp(`\\bearned\\b[^.]{0,24}?\\b(${D})`, "gi"),
|
|
39
|
+
new RegExp(`\\([^()\\n]{0,80}?\\b(${D})\\s*[:),.;—]`, "g"),
|
|
40
|
+
new RegExp(`(?:\\b[A-Z][\\w'-]+|\\b(?:operator|owner|maintainer|client|teammate)),\\s*(?:[a-z]+,\\s*)?(${D})\\s*[,:]`, "g"),
|
|
41
|
+
new RegExp(`\\b(?:on|until|since|before|after)\\s+(${D})\\b`, "gi"),
|
|
42
|
+
new RegExp(`\\bfrom\\s+(${D})\\s+\\(`, "gi"),
|
|
43
|
+
new RegExp(`\\b(?:measured|reported|found|caught|hit|corrected|retired|aligned|renamed|refused|broke|failed|observed|verified|watched|fixed|softened|flipped|restored|removed|changed|added|introduced)\\s+(?:[a-z]+\\s+)?(?:on\\s+|in\\s+)?(${D})`, "gi"),
|
|
44
|
+
new RegExp(`\\b(?:rule|ruling|default|correction|incident|reversal)\\s+(?:of|from)\\s+(${D})`, "gi"),
|
|
45
|
+
new RegExp(`\\bthe\\s+(${D})\\b`, "gi"),
|
|
46
|
+
new RegExp(`(${D})'s\\b`, "g"),
|
|
47
|
+
new RegExp(`^\\s*[-*]\\s+[*_]*(${D})[*_]*\\s*[(:]`, "gm"),
|
|
48
|
+
];
|
|
49
|
+
const EXAMPLE = /\b(e\.g\.|example|for instance|such as)\b/i;
|
|
50
|
+
|
|
51
|
+
/**
|
|
52
|
+
* Body lines that narrate a dated incident, outside code fences and inline code: one entry per
|
|
53
|
+
* line that carries the date of a story.
|
|
54
|
+
*/
|
|
55
|
+
export function historyLines(body) {
|
|
56
|
+
const lines = proseLines(body);
|
|
57
|
+
const hit = new Map();
|
|
58
|
+
let para = [];
|
|
59
|
+
const flush = () => {
|
|
60
|
+
if (!para.length) return;
|
|
61
|
+
// Join the paragraph, remembering where each line starts, so a match maps back to its line.
|
|
62
|
+
let text = "";
|
|
63
|
+
const starts = [];
|
|
64
|
+
for (const { i, line } of para) { starts.push({ i, at: text.length }); text += line.replace(/`[^`]*`/g, "``") + "\n"; }
|
|
65
|
+
for (const re of HISTORY) {
|
|
66
|
+
re.lastIndex = 0;
|
|
67
|
+
for (const m of text.matchAll(re)) {
|
|
68
|
+
const at = m.index + m[0].lastIndexOf(m[1]);
|
|
69
|
+
const owner = [...starts].reverse().find((s) => s.at <= at);
|
|
70
|
+
const raw = lines[owner.i].line;
|
|
71
|
+
if (EXAMPLE.test(raw)) continue;
|
|
72
|
+
if (!hit.has(owner.i)) hit.set(owner.i, raw.trim());
|
|
73
|
+
}
|
|
74
|
+
}
|
|
75
|
+
para = [];
|
|
76
|
+
};
|
|
77
|
+
lines.forEach(({ line, fenced }, i) => {
|
|
78
|
+
// A heading, an HTML comment (a generator's provenance stamp) and a fence are not prose.
|
|
79
|
+
if (fenced || !line.trim() || /^\s*(#|<!--)/.test(line)) { flush(); return; }
|
|
80
|
+
para.push({ i, line });
|
|
81
|
+
});
|
|
82
|
+
flush();
|
|
83
|
+
return [...hit.entries()].sort((a, b) => a[0] - b[0]).map(([i, text]) => ({ lineNo: i + 1, text }));
|
|
84
|
+
}
|
|
85
|
+
|
|
29
86
|
const normRule = (line) => line.toLowerCase().replace(/[*_`>#-]/g, "").replace(/\s+/g, " ").trim();
|
|
30
87
|
const broken = (ctx) => Boolean(ctx.error || ctx.parseError);
|
|
31
88
|
const str = (v) => (typeof v === "string" ? v : "");
|
|
@@ -194,6 +251,20 @@ export const skillRules = defineRules([
|
|
|
194
251
|
return [f("warn", `${late.length} hard rule(s) sit past the first ~${FOLD_TOKENS} tokens and would not survive compaction (${shown})`, "Restate them in a short Rules section near the top, or move them up. Length is fine; the rules just need to be above the fold. Step-specific detail can move into step files (steps/<step>.md) that are read fresh when the step comes up.")];
|
|
195
252
|
},
|
|
196
253
|
},
|
|
254
|
+
{
|
|
255
|
+
// 0.6.0: SKILL.md is instructions, read on every run. The story of the incident that earned
|
|
256
|
+
// a rule is history: it belongs in MISSES.md, under the miss it records, and SKILL.md keeps
|
|
257
|
+
// the rule and a one-line why. Anthropic's authoring guidance: avoid time-sensitive content.
|
|
258
|
+
id: "history-in-skill",
|
|
259
|
+
level: "skill",
|
|
260
|
+
check(ctx) {
|
|
261
|
+
if (broken(ctx)) return [];
|
|
262
|
+
const hits = historyLines(ctx.body);
|
|
263
|
+
if (!hits.length) return [];
|
|
264
|
+
const shown = hits.slice(0, 3).map((h) => `line ${h.lineNo + ctx.bodyStartLine - 1}: ${h.text.slice(0, 70)}`).join("; ");
|
|
265
|
+
return [f("warn", `${hits.length} dated incident stor${hits.length === 1 ? "y" : "ies"} in SKILL.md (${shown})`, "Move each story into MISSES.md as a miss entry (id, date, what happened, fix, eval, optional quote) and leave the rule plus a one-line why in SKILL.md.")];
|
|
266
|
+
},
|
|
267
|
+
},
|
|
197
268
|
{
|
|
198
269
|
id: "navigable",
|
|
199
270
|
level: "skill",
|
package/src/rules/superskill.mjs
CHANGED
|
@@ -1,16 +1,19 @@
|
|
|
1
|
-
// Level 3, "superskill":
|
|
2
|
-
//
|
|
1
|
+
// Level 3, "superskill" (0.6.0): fixed every time it got something wrong, with a regression eval
|
|
2
|
+
// for each fix; its whole suite and trigger set pass against the SKILL.md that is here now, on a
|
|
3
|
+
// current model and better than no skill; and it did the job for real people, one-shot, often
|
|
4
|
+
// enough on at least one model+harness. Goldens are optional evidence.
|
|
3
5
|
import { createHash } from "node:crypto";
|
|
4
6
|
import { defineRules } from "./define.mjs";
|
|
5
7
|
import { readGoldens, isApproved, isRealApproved, weightOf, goldenOpts } from "../goldens.mjs";
|
|
6
|
-
import {
|
|
8
|
+
import { readRealRuns, meetsBar, knownPair, MIN_REAL_RUNS, MIN_ONE_SHOT } from "../realruns.mjs";
|
|
7
9
|
import { readMisses } from "../misses.mjs";
|
|
8
10
|
import { readEvals, readLatestRun } from "../evals.mjs";
|
|
9
11
|
|
|
10
12
|
const f = (severity, message, fix) => ({ severity, message, fix });
|
|
11
13
|
const DAY = 86400000;
|
|
12
|
-
|
|
13
|
-
export const
|
|
14
|
+
/** Defaults for evals-pass; `--min-pass-rate` / `--min-trigger-rate` change them. */
|
|
15
|
+
export const MIN_PASS_RATE = 0.9;
|
|
16
|
+
export const MIN_TRIGGER_RATE = 0.9;
|
|
14
17
|
/** How long a --run stays fresh, by metadata.cadence. */
|
|
15
18
|
export const FRESH_DAYS = { daily: 30, weekly: 30, monthly: 60, quarterly: 120, yearly: 365 };
|
|
16
19
|
const DEFAULT_FRESH = 30;
|
|
@@ -19,55 +22,57 @@ const days = (now, date) => Math.floor((now.getTime() - new Date(date).getTime()
|
|
|
19
22
|
|
|
20
23
|
export const superskillRules = defineRules([
|
|
21
24
|
{
|
|
25
|
+
// 0.6.0: a golden is optional evidence, never the bar. A random accepted run is not the
|
|
26
|
+
// embodiment of a skill; the misses it absorbed and the evals that hold them are. A golden
|
|
27
|
+
// that exists is still an eval (its checklist is graded by --run), so its files are checked.
|
|
22
28
|
id: "golden-approved",
|
|
23
29
|
level: "superskill",
|
|
24
30
|
check(ctx, opts = {}) {
|
|
25
31
|
if (ctx.error) return [];
|
|
26
32
|
const goldens = readGoldens(ctx.dir, goldenOpts(ctx, opts));
|
|
27
|
-
|
|
28
|
-
// 0.5.0: only a golden from a real run a person accepted counts. An invented example, however
|
|
29
|
-
// careful, puts "a person said this was right" on something no person's work produced, and a
|
|
30
|
-
// skill could then reach the top level on its author's fiction.
|
|
31
|
-
const approved = goldens.filter(isRealApproved);
|
|
33
|
+
if (!goldens.length) return [];
|
|
32
34
|
const out = [];
|
|
33
35
|
const where = (g) => (g.private ? `private golden ${g.id}` : `goldens/${g.id}`);
|
|
34
36
|
for (const g of goldens.filter((x) => x.approvalError))
|
|
35
|
-
out.push(f("
|
|
37
|
+
out.push(f("warn", `${where(g)}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
|
|
36
38
|
for (const g of goldens.filter((x) => x.provenanceError))
|
|
37
|
-
out.push(f("
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
return out;
|
|
46
|
-
}
|
|
39
|
+
out.push(f("warn", `${where(g)}/PROVENANCE.json is not valid JSON`, "Rewrite it: source, run, accepted (see SPEC.md, goldens)."));
|
|
40
|
+
for (const g of goldens.filter((x) => !x.expectations.length))
|
|
41
|
+
out.push(f("info", `${where(g)} has no expectations.json, so --run grades it by likeness to output.md`, "Write goldens/<id>/expectations.json: the behaviors a right answer shows (grade the outcome, not the path). output.md then stays as the reference that proves the task is solvable."));
|
|
42
|
+
const approved = goldens.filter(isApproved);
|
|
43
|
+
const real = goldens.filter(isRealApproved);
|
|
44
|
+
const unreal = approved.filter((g) => !real.includes(g));
|
|
45
|
+
if (unreal.length) out.push(f("info", `approved golden${unreal.length === 1 ? "" : "s"} not from a real run a person accepted: ${unreal.map((g) => `${g.id} ${g.origin.why}`).join("; ")}`, "Only a golden whose PROVENANCE.json names a real run and who accepted it is evidence of real use; an invented one is an ordinary eval."));
|
|
46
|
+
if (!real.length) return out;
|
|
47
47
|
const sha = createHash("sha256").update(ctx.raw).digest("hex");
|
|
48
|
-
const stale =
|
|
48
|
+
const stale = real.filter((g) => g.approval.skill_sha && g.approval.skill_sha !== sha).map((g) => g.id);
|
|
49
49
|
if (stale.length) out.push(f("info", `golden${stale.length === 1 ? "" : "s"} ${stale.join(", ")} approved against an earlier SKILL.md`, "Re-run the golden and re-approve if the output still holds."));
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
const w = approved.map((g) => ({ id: g.id, ...weightOf(g) }));
|
|
53
|
-
const proven = w.filter((x) => x.outcome > 0);
|
|
54
|
-
const summary = w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ");
|
|
55
|
-
if (!proven.length) out.push(f("info", `approved on judgment only, no outcome recorded yet (${summary})`, "When a golden produces a real result, record it: `superskill approve <skill> <golden> --basis outcome --evidence \"<what happened, where to check>\"`."));
|
|
56
|
-
else out.push(f("info", `approval weight: ${summary}`, ""));
|
|
50
|
+
const w = real.map((g) => ({ id: g.id, ...weightOf(g) }));
|
|
51
|
+
out.push(f("info", `real goldens (optional evidence): ${w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ")}`, ""));
|
|
57
52
|
return out;
|
|
58
53
|
},
|
|
59
54
|
},
|
|
60
55
|
{
|
|
61
|
-
// The real
|
|
62
|
-
//
|
|
63
|
-
id: "real-
|
|
56
|
+
// The record from real use, per model and harness. A clean record on one model in one harness
|
|
57
|
+
// proves nothing about another, so pairs are never pooled: one pair has to clear the bar alone.
|
|
58
|
+
id: "real-runs",
|
|
64
59
|
level: "superskill",
|
|
65
|
-
check(ctx,
|
|
60
|
+
check(ctx, opts = {}) {
|
|
66
61
|
if (ctx.error) return [];
|
|
67
62
|
const name = (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName;
|
|
68
|
-
const
|
|
69
|
-
|
|
70
|
-
|
|
63
|
+
const minRuns = num(opts.minRealRuns, MIN_REAL_RUNS), minOneShot = num(opts.minOneShot, MIN_ONE_SHOT);
|
|
64
|
+
const bar = `at least ${minRuns} real runs and ${pct(minOneShot)} one-shot on one model+harness`;
|
|
65
|
+
const realRunsDir = opts.realRuns ?? ctx.realRuns ?? process.env.SUPERSKILL_REAL_RUNS ?? null;
|
|
66
|
+
const rec = readRealRuns(ctx.dir, { name, realRunsDir });
|
|
67
|
+
const how = "Export the run record (format superskill-real-runs/1, SPEC.md) to evals/real-runs.json or --real-runs <dir>; Freedom's skill ledger exports it.";
|
|
68
|
+
if (!rec) return [f("fail", `no real-run record (needs ${bar})`, how)];
|
|
69
|
+
if (rec.error) return [f("fail", rec.error, how)];
|
|
70
|
+
const out = rec.pairs.map((p) => f("info", `real runs, ${p.model} / ${p.harness}: ${p.one_shot} of ${p.runs} one-shot${p.runs ? ` (${pct(p.one_shot / p.runs)})` : ""}${knownPair(p) ? "" : ", not counted: the record does not say which model or harness"}`, ""));
|
|
71
|
+
if (!meetsBar(rec.pairs, { minRuns, minOneShot }).length) {
|
|
72
|
+
const best = rec.pairs.filter(knownPair)[0];
|
|
73
|
+
out.push(f("fail", `real-run record below the bar (${bar}; ${rec.where})${best ? `: best is ${best.model} / ${best.harness}, ${best.one_shot} of ${best.runs}` : ": no run names its model and harness"}`, "Use the skill for real and fix what it gets wrong; every correction is a miss to close with an eval."));
|
|
74
|
+
}
|
|
75
|
+
return out;
|
|
71
76
|
},
|
|
72
77
|
},
|
|
73
78
|
{
|
|
@@ -79,18 +84,17 @@ export const superskillRules = defineRules([
|
|
|
79
84
|
},
|
|
80
85
|
},
|
|
81
86
|
{
|
|
82
|
-
|
|
87
|
+
// 0.6.0: every miss is closed. An open miss is a known way the skill fails today, and a skill
|
|
88
|
+
// that still fails a known way is not at the top level, however recent the miss is.
|
|
89
|
+
id: "misses-closed",
|
|
83
90
|
level: "superskill",
|
|
84
91
|
check(ctx, { now = new Date() } = {}) {
|
|
85
92
|
if (ctx.error) return [];
|
|
86
93
|
const misses = readMisses(ctx.dir) || [];
|
|
87
|
-
|
|
88
|
-
for (const m of misses.filter((x) => x.status === "open")) {
|
|
94
|
+
return misses.filter((x) => x.status === "open").map((m) => {
|
|
89
95
|
const age = days(now, m.date);
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
}
|
|
93
|
-
return out;
|
|
96
|
+
return f("fail", `miss ${m.id} is open (${age} day${age === 1 ? "" : "s"}): ${m.what.slice(0, 80)}`, `Fix it, add an eval that catches it, then \`superskill fix <skill> ${m.id} --eval <id>\`.`);
|
|
97
|
+
});
|
|
94
98
|
},
|
|
95
99
|
},
|
|
96
100
|
{
|
|
@@ -108,6 +112,39 @@ export const superskillRules = defineRules([
|
|
|
108
112
|
.map((m) => f("fail", m.eval ? `miss ${m.id} names eval "${m.eval}", which is not in evals.json or goldens/` : `miss ${m.id} is fixed with no regression eval`, `Add a case to evals/evals.json that would catch ${m.id} again, and name its id on the Eval line.`));
|
|
109
113
|
},
|
|
110
114
|
},
|
|
115
|
+
{
|
|
116
|
+
// The suite and the trigger set pass against the SKILL.md that is here now. A pass recorded
|
|
117
|
+
// against an earlier SKILL.md proves the earlier skill.
|
|
118
|
+
id: "evals-pass",
|
|
119
|
+
level: "superskill",
|
|
120
|
+
check(ctx, opts = {}) {
|
|
121
|
+
if (ctx.error) return [];
|
|
122
|
+
const r = readLatestRun(ctx.dir);
|
|
123
|
+
if (!r.run) return [];
|
|
124
|
+
const run = r.run;
|
|
125
|
+
const out = [];
|
|
126
|
+
const sha = createHash("sha256").update(ctx.raw).digest("hex");
|
|
127
|
+
if (!run.skill_sha) out.push(f("fail", "latest.json does not say which SKILL.md it ran against (no skill_sha)", "Re-run `superskill doctor <skill> --run` with superskill 0.6.0 or later."));
|
|
128
|
+
else if (run.skill_sha !== sha) out.push(f("fail", "the last --run was against an earlier SKILL.md", "Re-run `superskill doctor <skill> --run`; a change to the skill needs its evals re-run."));
|
|
129
|
+
const minPass = num(opts.minPassRate, MIN_PASS_RATE), minTrig = num(opts.minTriggerRate, MIN_TRIGGER_RATE);
|
|
130
|
+
const w = Number(run.with_skill?.pass_rate);
|
|
131
|
+
if (Number.isFinite(w) && w < minPass) out.push(f("fail", `the suite passed ${pct(w)} of runs with the skill (needs ${pct(minPass)})`, "Read the failures in latest.json per_case, fix the skill, re-run."));
|
|
132
|
+
// Every fixed miss's regression eval passes every run: a regression that comes back half the
|
|
133
|
+
// time has come back.
|
|
134
|
+
const fixed = (readMisses(ctx.dir) || []).filter((m) => m.status === "fixed" && m.eval);
|
|
135
|
+
const cases = new Map((Array.isArray(run.per_case) ? run.per_case : []).map((c) => [String(c.id), c]));
|
|
136
|
+
for (const m of fixed) {
|
|
137
|
+
const c = cases.get(String(m.eval)) || cases.get(`golden:${m.eval}`);
|
|
138
|
+
if (!c) { if (run.skill_sha) out.push(f("fail", `regression eval ${m.eval} (miss ${m.id}) is not in the last --run`, "Re-run `superskill doctor <skill> --run`.")); continue; }
|
|
139
|
+
const ws = c.with_skill || {};
|
|
140
|
+
if (!(ws.runs > 0 && ws.passes === ws.runs)) out.push(f("fail", `regression eval ${m.eval} (miss ${m.id}) passed ${ws.passes || 0} of ${ws.runs || 0} runs`, `Miss ${m.id} is back: fix the skill until its eval passes every run.`));
|
|
141
|
+
}
|
|
142
|
+
const t = run.triggers;
|
|
143
|
+
if (!t || !Number.isFinite(Number(t.pass_rate))) out.push(f("fail", "the last --run did not run the trigger evals", "Re-run `superskill doctor <skill> --run` with superskill 0.6.0 or later; it runs evals/triggers.json too."));
|
|
144
|
+
else if (Number(t.pass_rate) < minTrig) out.push(f("fail", `trigger evals ${pct(Number(t.pass_rate))} right (needs ${pct(minTrig)})`, "Read triggers.per_query in latest.json: sharpen the description so it loads when it should and not otherwise."));
|
|
145
|
+
return out;
|
|
146
|
+
},
|
|
147
|
+
},
|
|
111
148
|
{
|
|
112
149
|
id: "run-evidence",
|
|
113
150
|
level: "superskill",
|
|
@@ -146,3 +183,4 @@ export const superskillRules = defineRules([
|
|
|
146
183
|
]);
|
|
147
184
|
|
|
148
185
|
const pct = (x) => `${Math.round(x * 100)}%`;
|
|
186
|
+
const num = (v, d) => (v === undefined || v === null || v === "" || !Number.isFinite(Number(v)) ? d : Number(v));
|
package/src/run/claude.mjs
CHANGED
|
@@ -35,3 +35,32 @@ export function ask(prompt, { model } = {}) {
|
|
|
35
35
|
if (model) args.push("--model", model);
|
|
36
36
|
return call(args, cwd).output;
|
|
37
37
|
}
|
|
38
|
+
|
|
39
|
+
/**
|
|
40
|
+
* Did Claude Code load the skill for this query? Run headless with the skill installed and read
|
|
41
|
+
* the stream: a Skill tool call naming it, or a Read of its SKILL.md, is a load.
|
|
42
|
+
*/
|
|
43
|
+
export function loadedSkill(stream, skillName) {
|
|
44
|
+
for (const line of String(stream).split("\n")) {
|
|
45
|
+
if (!line.trim().startsWith("{")) continue;
|
|
46
|
+
let ev; try { ev = JSON.parse(line); } catch { continue; }
|
|
47
|
+
const content = ev?.message?.content;
|
|
48
|
+
if (!Array.isArray(content)) continue;
|
|
49
|
+
for (const c of content) {
|
|
50
|
+
if (c?.type !== "tool_use") continue;
|
|
51
|
+
const input = c.input || {};
|
|
52
|
+
if (c.name === "Skill" && [input.skill, input.command, input.name].some((v) => typeof v === "string" && v.split(":").pop() === skillName)) return true;
|
|
53
|
+
if (typeof input.file_path === "string" && input.file_path.endsWith(`/${skillName}/SKILL.md`)) return true;
|
|
54
|
+
}
|
|
55
|
+
}
|
|
56
|
+
return false;
|
|
57
|
+
}
|
|
58
|
+
|
|
59
|
+
export function triggerCase({ skillDir, skillName, query, model }) {
|
|
60
|
+
const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".claude/skills" });
|
|
61
|
+
const args = ["-p", query, "--output-format", "stream-json", "--verbose", "--add-dir", cwd];
|
|
62
|
+
if (model) args.push("--model", model);
|
|
63
|
+
const r = spawnSync("claude", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
64
|
+
if (r.error) throw Object.assign(new Error(`claude failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
|
|
65
|
+
return { triggered: loadedSkill(r.stdout, skillName), failed: r.status !== 0 && !r.stdout, cwd };
|
|
66
|
+
}
|
package/src/run/codex.mjs
CHANGED
|
@@ -31,3 +31,17 @@ export function ask(prompt, { model } = {}) {
|
|
|
31
31
|
const cwd = prepareWorkspace({ skillDir: null, skillName: null });
|
|
32
32
|
return call(prompt, cwd, model).output;
|
|
33
33
|
}
|
|
34
|
+
|
|
35
|
+
/**
|
|
36
|
+
* Did Codex load the skill for this query? Codex reads a skill's SKILL.md when it uses it, so the
|
|
37
|
+
* JSON event stream mentioning <name>/SKILL.md is a load.
|
|
38
|
+
*/
|
|
39
|
+
export function triggerCase({ skillDir, skillName, query, model }) {
|
|
40
|
+
const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".agents/skills" });
|
|
41
|
+
const args = ["exec", "--json", "--skip-git-repo-check", "-C", cwd];
|
|
42
|
+
if (model) args.push("-m", model);
|
|
43
|
+
args.push(query);
|
|
44
|
+
const r = spawnSync("codex", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
45
|
+
if (r.error) throw Object.assign(new Error(`codex failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
|
|
46
|
+
return { triggered: String(r.stdout).includes(`${skillName}/SKILL.md`), failed: r.status !== 0 && !r.stdout, cwd };
|
|
47
|
+
}
|
package/src/run/fake.mjs
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
// A stand-in harness for tests: SUPERSKILL_FAKE_HARNESS names a node script that reads one
|
|
2
|
-
// JSON request on stdin ({mode: "case"|"ask", prompt, withSkill, cwd}) and
|
|
3
|
-
// {output, tokens?, model?}. Never calls a model.
|
|
2
|
+
// JSON request on stdin ({mode: "case"|"ask"|"trigger", prompt|query, withSkill, cwd}) and
|
|
3
|
+
// prints {output, tokens?, model?} (or {triggered} for a trigger). Never calls a model.
|
|
4
4
|
import { spawnSync } from "node:child_process";
|
|
5
5
|
import { prepareWorkspace } from "./workspace.mjs";
|
|
6
6
|
import { sandboxEnv } from "./env.mjs";
|
|
@@ -22,6 +22,11 @@ export function create(script) {
|
|
|
22
22
|
const cwd = prepareWorkspace({ skillDir, skillName, files, linkAt: withSkill ? ".claude/skills" : null });
|
|
23
23
|
return { ...call(script, { mode: "case", prompt, withSkill, cwd }), cwd };
|
|
24
24
|
},
|
|
25
|
+
triggerCase({ skillDir, skillName, query }) {
|
|
26
|
+
const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".claude/skills" });
|
|
27
|
+
const doc = JSON.parse(spawnSync(process.execPath, [script], { input: JSON.stringify({ mode: "trigger", query, cwd }), env: sandboxEnv(), encoding: "utf8" }).stdout || "{}");
|
|
28
|
+
return { triggered: Boolean(doc.triggered), failed: false, cwd };
|
|
29
|
+
},
|
|
25
30
|
ask(prompt) {
|
|
26
31
|
return call(script, { mode: "ask", prompt }).output;
|
|
27
32
|
},
|
package/src/run/index.mjs
CHANGED
|
@@ -8,7 +8,8 @@ import { createInterface } from "node:readline/promises";
|
|
|
8
8
|
import { clock, UsageError } from "../args.mjs";
|
|
9
9
|
import { findSkills } from "../doctor.mjs";
|
|
10
10
|
import { parseSkillFile } from "../frontmatter.mjs";
|
|
11
|
-
import { readEvals } from "../evals.mjs";
|
|
11
|
+
import { readEvals, readTriggers } from "../evals.mjs";
|
|
12
|
+
import { createHash } from "node:crypto";
|
|
12
13
|
import { readGoldens } from "../goldens.mjs";
|
|
13
14
|
import { grade, isMachineCheck } from "./grade.mjs";
|
|
14
15
|
import * as claude from "./claude.mjs";
|
|
@@ -18,7 +19,9 @@ import { create as createFake } from "./fake.mjs";
|
|
|
18
19
|
export const help = `superskill doctor <skill> --run [--harness claude|codex] [--repeat 3] [--model <id>] [--yes]
|
|
19
20
|
|
|
20
21
|
Run the skill's evals for real: every case in evals/evals.json and every golden, --repeat
|
|
21
|
-
times with the skill and --repeat times without it,
|
|
22
|
+
times with the skill and --repeat times without it, and every query in evals/triggers.json
|
|
23
|
+
--repeat times with the skill installed (did it load when it should, and not otherwise),
|
|
24
|
+
through a headless harness. Machine
|
|
22
25
|
checks (contains:, regex:, file_exists:) are free; each plain-language expectation costs
|
|
23
26
|
one grader call per run. Prints the estimated number of model calls first, and asks
|
|
24
27
|
before starting (or needs --yes when there is no terminal). Writes
|
|
@@ -45,15 +48,51 @@ export function collectCases(skillDir, { privateGoldens = process.env.SUPERSKILL
|
|
|
45
48
|
if (e.error) throw new UsageError(e.error);
|
|
46
49
|
const cases = e.cases.filter((c) => c.prompt.trim()).map((c) => ({ id: String(c.id), prompt: c.prompt, files: c.files, assertions: c.assertions.length ? c.assertions : [c.expected_output].filter(Boolean) }));
|
|
47
50
|
for (const g of readGoldens(skillDir, { privateGoldens, name })) {
|
|
48
|
-
if (!g.input || !g.output
|
|
49
|
-
|
|
51
|
+
if (!g.input || !((g.output && g.output.trim()) || g.expectations?.length)) continue;
|
|
52
|
+
// 0.6.0: grade the outcome, not the path. A golden with a checklist is graded on it; one
|
|
53
|
+
// without falls back to likeness with its reference output.
|
|
54
|
+
const assertions = g.expectations?.length ? g.expectations : [`The output matches this approved output in substance (same facts, same shape; wording may differ):\n${g.output}`];
|
|
55
|
+
cases.push({ id: `golden:${g.id}`, prompt: g.input, files: [], assertions });
|
|
50
56
|
}
|
|
51
57
|
return cases;
|
|
52
58
|
}
|
|
53
59
|
|
|
54
|
-
export function estimateCalls(cases, repeat) {
|
|
60
|
+
export function estimateCalls(cases, repeat, triggers = 0) {
|
|
55
61
|
const graded = cases.reduce((n, c) => n + c.assertions.filter((a) => !isMachineCheck(a)).length, 0);
|
|
56
|
-
|
|
62
|
+
const runs = cases.length * repeat * 2 + triggers * repeat;
|
|
63
|
+
return { runs, grader: graded * repeat * 2, triggers: triggers * repeat, total: runs + graded * repeat * 2 };
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
/** The trigger queries the run will check, from evals/triggers.json. */
|
|
67
|
+
export function collectTriggers(skillDir) {
|
|
68
|
+
const t = readTriggers(skillDir);
|
|
69
|
+
if (t.error) throw new UsageError(t.error);
|
|
70
|
+
return t.triggers;
|
|
71
|
+
}
|
|
72
|
+
|
|
73
|
+
/**
|
|
74
|
+
* Run every trigger query `repeat` times with the skill installed and record whether the harness
|
|
75
|
+
* loaded it. A query passes a run when loading matched should_trigger.
|
|
76
|
+
*/
|
|
77
|
+
export function runTriggers(skillDir, { harness, repeat, model, log = () => {} }) {
|
|
78
|
+
const skillName = parseSkillFile(readFileSync(join(skillDir, "SKILL.md"), "utf8")).data.name || "";
|
|
79
|
+
const queries = collectTriggers(skillDir);
|
|
80
|
+
if (!queries.length) return null;
|
|
81
|
+
if (typeof harness.triggerCase !== "function") throw new UsageError(`harness ${harness.name} cannot run trigger evals`);
|
|
82
|
+
let runs = 0, passes = 0;
|
|
83
|
+
const per_query = [];
|
|
84
|
+
for (const q of queries) {
|
|
85
|
+
const row = { query: q.query, should_trigger: q.should_trigger, runs: 0, passes: 0 };
|
|
86
|
+
for (let i = 0; i < repeat; i++) {
|
|
87
|
+
log(`trigger "${q.query.slice(0, 40)}", run ${i + 1}/${repeat}`);
|
|
88
|
+
const r = harness.triggerCase({ skillDir, skillName, query: q.query, model });
|
|
89
|
+
if (r.cwd) rmSync(r.cwd, { recursive: true, force: true });
|
|
90
|
+
const ok = !r.failed && Boolean(r.triggered) === q.should_trigger;
|
|
91
|
+
row.runs++; row.passes += ok ? 1 : 0; runs++; passes += ok ? 1 : 0;
|
|
92
|
+
}
|
|
93
|
+
per_query.push(row);
|
|
94
|
+
}
|
|
95
|
+
return { cases: queries.length, runs, passes, pass_rate: runs ? round(passes / runs) : 0, per_query };
|
|
57
96
|
}
|
|
58
97
|
|
|
59
98
|
export function runEvals(skillDir, { harness, repeat = 3, now = new Date(), model, log = () => {} }) {
|
|
@@ -84,7 +123,10 @@ export function runEvals(skillDir, { harness, repeat = 3, now = new Date(), mode
|
|
|
84
123
|
per_case.push(row);
|
|
85
124
|
}
|
|
86
125
|
const summary = (t) => ({ pass_rate: t.runs ? round(t.passes / t.runs) : 0, mean_ms: t.runs ? Math.round(t.ms / t.runs) : 0, mean_tokens: t.tokenRuns ? Math.round(t.tokens / t.tokenRuns) : null });
|
|
87
|
-
const
|
|
126
|
+
const triggers = runTriggers(skillDir, { harness, repeat, model, log });
|
|
127
|
+
// 0.6.0: the run says which SKILL.md it proved, so a later edit cannot ride on an old pass.
|
|
128
|
+
const skill_sha = createHash("sha256").update(readFileSync(join(skillDir, "SKILL.md"), "utf8")).digest("hex");
|
|
129
|
+
const result = { run_at: now.toISOString(), harness: harness.name, model: seenModel, skill_sha, cases: cases.length, repeat, with_skill: summary(totals.with_skill), without_skill: summary(totals.without_skill), per_case, ...(triggers ? { triggers } : {}) };
|
|
88
130
|
mkdirSync(join(skillDir, "evals", "results"), { recursive: true });
|
|
89
131
|
writeFileSync(join(skillDir, "evals", "results", "latest.json"), JSON.stringify(result, null, 2) + "\n");
|
|
90
132
|
return result;
|
|
@@ -100,11 +142,11 @@ export async function runCommand(a) {
|
|
|
100
142
|
const harness = pickHarness(a.flags.harness);
|
|
101
143
|
const dirs = a._.flatMap((p) => findSkills(p));
|
|
102
144
|
if (!dirs.length) throw new UsageError(`no skills found under ${a._.join(", ")}`);
|
|
103
|
-
const plan = dirs.map((d) => ({ dir: d, cases: collectCases(d) }));
|
|
104
|
-
const calls = plan.reduce((n, p) => n + estimateCalls(p.cases, repeat).total, 0);
|
|
145
|
+
const plan = dirs.map((d) => ({ dir: d, cases: collectCases(d), triggers: collectTriggers(d).length }));
|
|
146
|
+
const calls = plan.reduce((n, p) => n + estimateCalls(p.cases, repeat, p.triggers).total, 0);
|
|
105
147
|
for (const p of plan) {
|
|
106
|
-
const e = estimateCalls(p.cases, repeat);
|
|
107
|
-
process.stderr.write(`${p.dir}: ${p.cases.length} cases x ${repeat} x 2 = ${e.runs} runs + ${e.grader} grader calls\n`);
|
|
148
|
+
const e = estimateCalls(p.cases, repeat, p.triggers);
|
|
149
|
+
process.stderr.write(`${p.dir}: ${p.cases.length} cases x ${repeat} x 2 + ${p.triggers} triggers x ${repeat} = ${e.runs} runs + ${e.grader} grader calls\n`);
|
|
108
150
|
}
|
|
109
151
|
process.stderr.write(`estimated model calls: ${calls} through ${harness.name}\n`);
|
|
110
152
|
if (!a.flags.yes) {
|
|
@@ -122,7 +164,7 @@ export async function runCommand(a) {
|
|
|
122
164
|
for (const p of plan) {
|
|
123
165
|
const r = runEvals(p.dir, { harness, repeat, now, model: a.flags.model, log: (m) => process.stderr.write(` ${m}\n`) });
|
|
124
166
|
results.push({ path: p.dir, ...r });
|
|
125
|
-
if (!a.flags.json) process.stdout.write(`${p.dir}\n with skill ${pct(r.with_skill.pass_rate)} without ${pct(r.without_skill.pass_rate)} (${r.cases} cases x ${repeat})\n wrote evals/results/latest.json\n`);
|
|
167
|
+
if (!a.flags.json) process.stdout.write(`${p.dir}\n with skill ${pct(r.with_skill.pass_rate)} without ${pct(r.without_skill.pass_rate)} (${r.cases} cases x ${repeat})${r.triggers ? `\n triggers ${pct(r.triggers.pass_rate)} right (${r.triggers.cases} queries x ${repeat})` : ""}\n wrote evals/results/latest.json\n`);
|
|
126
168
|
}
|
|
127
169
|
if (a.flags.json) process.stdout.write(JSON.stringify({ results }, null, 2) + "\n");
|
|
128
170
|
return results.every((r) => r.with_skill.pass_rate > r.without_skill.pass_rate) ? 0 : 1;
|
package/src/snippet.md
CHANGED
|
@@ -5,12 +5,14 @@ Skills in this project are held to the superskill standard: https://supersuit.wi
|
|
|
5
5
|
- **After doing a job by hand once, propose a skill for it.** If the person agrees, write the
|
|
6
6
|
skill, then seed its evidence from this session:
|
|
7
7
|
`npx @supersuit/superskill init <skill-folder> --from-session <this session's transcript>`
|
|
8
|
-
The golden it creates waits for the person
|
|
8
|
+
The golden it creates is optional evidence and waits for the person; never approve it yourself.
|
|
9
9
|
- **Whenever a skill needed correcting** (the person redirected you, fixed its output, or you
|
|
10
10
|
worked around it), log it before moving on:
|
|
11
11
|
`npx @supersuit/superskill miss <skill-folder> "<what happened>" --expected "<what should have>"`
|
|
12
12
|
When you fix it, add an eval that would catch it again and close it with
|
|
13
13
|
`npx @supersuit/superskill fix <skill-folder> <miss-id> --eval <eval-id>`.
|
|
14
|
+
- **The story goes in MISSES.md, never in SKILL.md.** SKILL.md holds the rule and a one-line
|
|
15
|
+
why; what happened, when, and what the person said go on the miss entry (`--quote`).
|
|
14
16
|
- **Before calling a skill done**, run
|
|
15
17
|
`npx @supersuit/superskill doctor <skill-folder>`
|
|
16
18
|
and fix everything it lists for the level the skill is meant to reach.
|