@supersuit/superskill 0.5.0 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,48 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.6.1 (2026-10-06)
4
+
5
+ - **`real-runs` reading Freedom's ledger live no longer counts a rescued run as one-shot.** Freedom
6
+ derives `succeeded_with_rescues` from the operator's turns during a run even when nobody
7
+ classified them, so those records carry no intervention, and 0.6.0 read them as clean. Measured
8
+ on a real ledger the same day: a skill showed 54 of 54 one-shot that way. The exported
9
+ `superskill-real-runs/1` file was already right; only the live fallback was affected.
10
+
11
+ ## 0.6.0 (2026-10-06)
12
+
13
+ **A skill's embodiment is what it has absorbed, not one run of it.** 0.5.0 made a golden from a
14
+ real run the top-level requirement. Its first real test reversed it: nine harvested candidates,
15
+ each a real accepted run, and none was the standard the skill should be held to. The tests and
16
+ fixes a skill has absorbed are. **Breaking for anyone at superskill today:** the top level now
17
+ needs a real-run record and a 0.6.0 `--run` (one that records its `skill_sha` and the triggers).
18
+
19
+ - **Superskill is now** every miss closed with a regression eval (`misses-closed`, which replaces
20
+ `no-stale-open-miss`: an open miss fails however recent it is), and `evals-pass`: the last
21
+ `--run` was against the current `SKILL.md`, the suite passed (default 90% of runs), every fixed
22
+ miss's regression eval passed every run, and the trigger evals were right (default 90%).
23
+ `run-evidence` and `run-fresh` stand.
24
+ - **`real-runs`:** a real-run record where one model+harness pair, alone, has at least 5 real runs
25
+ at 80% one-shot (no correction). Every pair is reported; pairs are never pooled, and a pair that
26
+ says `unknown` never counts. The record is a harness-neutral file, `superskill-real-runs/1`, read
27
+ from `--real-runs <dir>` / `SUPERSKILL_REAL_RUNS`, then `evals/real-runs.json`, then Freedom's
28
+ skill ledger live. It replaces `real-use`.
29
+ - **Configurable bar:** `--min-real-runs`, `--min-one-shot`, `--min-pass-rate`,
30
+ `--min-trigger-rate` (and `SUPERSKILL_MIN_*`). Never frontmatter: a skill cannot lower its own bar.
31
+ - **Goldens are optional evidence.** `golden-approved` never fails. A golden is an eval graded on
32
+ `goldens/<id>/expectations.json`, a checklist of expected behavior (grade outcomes, not paths);
33
+ `output.md` is the reference that proves the task solvable. Provenance rules still say which
34
+ goldens are real runs.
35
+ - **`doctor --run` runs the trigger evals** (did the harness load the skill exactly when it
36
+ should) and writes `skill_sha` and `triggers` into `latest.json`.
37
+ - **`history-in-skill` (warn, level skill):** a dated incident story in `SKILL.md` ("Earned
38
+ 2026-09-08", "(Gary, 2026-09-16: ...)", "on 2026-09-13 a session ...") belongs in `MISSES.md`,
39
+ which now holds a story's full entry: id, date, what happened, fix, eval, optional `Quote:`, with
40
+ indented continuation lines. Tuned on Freedom's 103 shipped skills. Still no line-count rule.
41
+ - `superskill miss` takes `--quote` and `--date`.
42
+ - Tests: 17 new; each guard broken on purpose and seen red (a stale-sha pass counting, pooled pairs
43
+ counting, an unknown pair counting, an open miss passing, a half-passing regression eval passing,
44
+ examples and provenance stamps flagged as history).
45
+
3
46
  ## 0.5.0 (2026-10-06)
4
47
 
5
48
  **Superskill needs a golden from a real run.** Before, any golden a person approved counted, so a
package/README.md CHANGED
@@ -4,8 +4,9 @@ Score any agent skill folder as **skill**, **tested**, or **superskill**, and ge
4
4
  to-do list for the next level. Works on skills for Claude Code, Codex, or any harness that
5
5
  reads the [Agent Skills](https://agentskills.io) format. Zero dependencies, Node 20 or later.
6
6
 
7
- A superskill runs on frontier intelligence, is checked against examples you approved, and is
8
- fixed every time it gets something wrong. [What that means](https://supersuit.wiki/concepts/superskill);
7
+ A superskill is fixed every time it gets something wrong, with a test that keeps each fix fixed;
8
+ passes its whole eval suite on a current model; and does the job for real people without a
9
+ correction, often enough to count. [What that means](https://supersuit.wiki/concepts/superskill);
9
10
  [the standard](SPEC.md).
10
11
 
11
12
  ## 30 seconds
@@ -34,12 +35,17 @@ file error. `--json` prints one JSON document and nothing else.
34
35
 
35
36
  1. **skill**: a valid `SKILL.md` (name matches the folder, description of 1024 characters or
36
37
  fewer that says when to use it, hard rules above the compaction fold, headings in long bodies, nothing said twice), references one level deep, no
37
- hard-coded machine paths, nothing that reads like a prompt injection.
38
+ hard-coded machine paths, nothing that reads like a prompt injection. It warns on dated
39
+ incident stories in `SKILL.md`, which belong in `MISSES.md`.
38
40
  2. **tested**: at least three task evals with checks a machine can verify, and a trigger set of
39
41
  at least ten requests, some that should load the skill and some near-misses that should not.
40
- 3. **superskill**: a golden from a real run, accepted when it ran and approved by a person; no miss open longer than 14 days and every fixed
41
- miss guarded by an eval; a recent `--run` on file where the skill beats the same task done
42
- without it. "Recent" follows `metadata.cadence` (a weekly skill's proof lasts 30 days).
42
+ 3. **superskill**: no open miss, and every fixed miss guarded by a regression eval; a recent
43
+ `--run` against the `SKILL.md` that is there now, where the suite passes (90%), every
44
+ regression eval passes every run, the trigger set is 90% right, and the skill beats the same
45
+ task done without it; and a real-run record where one model+harness has at least 5 real runs
46
+ at 80% one-shot (no correction). Pairs are never pooled. "Recent" follows `metadata.cadence`
47
+ (a weekly skill's proof lasts 30 days). Goldens are optional evidence. Every threshold is a
48
+ flag (`--min-real-runs`, `--min-one-shot`, `--min-pass-rate`, `--min-trigger-rate`).
43
49
 
44
50
  Every rule and threshold is in [SPEC.md](SPEC.md).
45
51
 
@@ -52,7 +58,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
52
58
  | `superskill doctor <skill> --run [--harness claude\|codex] [--repeat 3] [--yes]` | Run the evals for real, with and without the skill, and write `evals/results/latest.json`. **The only command that spends model calls**; it prints an estimate and asks first. |
53
59
  | `superskill init <skill>` | Add missing `evals/`, `goldens/`, `MISSES.md`. Never overwrites. |
54
60
  | `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
55
- | `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
61
+ | `superskill miss <skill> "<what happened>" [--expected "..."] [--quote "..."] [--date YYYY-MM-DD]` | Log a time the skill got it wrong, or the story behind a rule it learned. |
56
62
  | `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
57
63
  | `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). At a terminal it asks for your name; from your phone, an agent records your tap with `--approved-by` and `--via`. Approvals accumulate. |
58
64
  | `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
@@ -70,9 +76,10 @@ my-skill/
70
76
  SKILL.md
71
77
  evals/evals.json task evals (Anthropic skill-creator format)
72
78
  evals/triggers.json should / should-not load (skill-creator format)
73
- goldens/<id>/ input.md, output.md, PROVENANCE.json, APPROVAL.json
74
- MISSES.md every miss, open or fixed with its eval
75
- evals/results/latest.json the last --run
79
+ goldens/<id>/ optional: input.md, expectations.json, output.md, PROVENANCE.json, APPROVAL.json
80
+ MISSES.md every miss and its story, open or fixed with its eval
81
+ evals/results/latest.json the last --run (which SKILL.md it proved, suite and triggers)
82
+ evals/real-runs.json real uses per model+harness, one-shot or not (or --real-runs <dir>)
76
83
  ```
77
84
 
78
85
  Harnesses ignore folders they do not know, so none of this changes how the skill loads.
@@ -84,13 +91,16 @@ Harnesses ignore folders they do not know, so none of this changes how the skill
84
91
  - Codex: the skill is linked at `.agents/skills/<name>`. Codex cannot switch skills off, so a
85
92
  copy installed in `~/.agents/skills` can leak into the baseline; move it aside while proving.
86
93
  - Machine checks (`contains:`, `regex:`, `file_exists:`) are free. Each plain-language expectation
87
- and each golden costs one grader call per run.
94
+ costs one grader call per run.
95
+ - Trigger evals run each query in `evals/triggers.json` with the skill installed and record
96
+ whether the harness loaded it (Claude Code: a `Skill` call or a read of its `SKILL.md`; Codex:
97
+ a read of its `SKILL.md`).
88
98
 
89
99
  ## Freedom
90
100
 
91
101
  Nothing here needs [Freedom](https://getfreedom.wiki). If a skill has Freedom's `HDSOP.md`, the
92
102
  doctor shows it as a bonus; `miss import --freedom-ledger` reads Freedom's run ledger as plain
93
- files. `superskill snippet` gives any agent the same habits with no Freedom installed.
103
+ files, and with no `evals/real-runs.json` the doctor reads the real-run record from it too. `superskill snippet` gives any agent the same habits with no Freedom installed.
94
104
 
95
105
  ## Releasing
96
106
 
package/SPEC.md CHANGED
@@ -1,13 +1,21 @@
1
1
  # The superskill standard
2
2
 
3
- **Version 0.5.0** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
3
+ **Version 0.6.1** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
4
4
  document and the checker disagree, the checker has a bug.
5
5
 
6
- A **superskill** runs on frontier intelligence, is checked against examples a person approved,
7
- and is fixed every time it gets something wrong
6
+ A **superskill** runs on frontier intelligence, is fixed every time it gets something wrong, and
7
+ does the job for real people without a correction
8
8
  ([definition](https://supersuit.wiki/concepts/superskill)). This standard turns each clause into
9
9
  a file in the skill's own folder, so the evidence travels with the skill wherever it is copied.
10
10
 
11
+ **What a skill has absorbed is its embodiment, not one run of it (0.6.0).** Until 0.5.0 the top
12
+ level required a golden: one real run a person approved as the standard. A random accepted run
13
+ is not the ultimate embodiment of a skill; the misses it was fixed for, and the evals that keep
14
+ each fix fixed, are. This follows Anthropic's guidance on agent evals: build the suite from real
15
+ failures, hold regression evals near 100%, grade outcomes rather than paths, and use a reference
16
+ solution to prove a task is solvable, not as the answer to match. Goldens remain, as optional
17
+ evidence.
18
+
11
19
  ## Contents
12
20
 
13
21
  - [The clauses and their evidence](#the-clauses-and-their-evidence)
@@ -23,9 +31,10 @@ a file in the skill's own folder, so the evidence travels with the skill whereve
23
31
  | The clause | What proves it | Where it lives |
24
32
  |---|---|---|
25
33
  | A skill at all | A valid `SKILL.md` under the Agent Skills spec, plus the hygiene rules below | `SKILL.md` |
26
- | Checked against examples you approved | At least one golden from a real run: the real input, the output the person accepted when it ran, where it came from, and a record of who approved it as a golden and when | `goldens/<id>/` (or a private folder, see below) |
27
- | Fixed every time it gets something wrong | A miss log where every miss is open (recently) or fixed with a regression eval that would catch it again | `MISSES.md` + `evals/evals.json` |
28
- | Runs on frontier intelligence | Its evals last passed on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
34
+ | Fixed every time it gets something wrong | A miss log where every miss is fixed, each with a regression eval that would catch it again, and the story of each fix | `MISSES.md` + `evals/evals.json` |
35
+ | Runs on frontier intelligence | Its whole suite and trigger set last passed against the current `SKILL.md`, on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
36
+ | Does the job for real | Enough real runs on one model+harness, enough of them one-shot (no correction) | `evals/real-runs.json` (or an operator's export, see below) |
37
+ | Optional: an example a person approved | A golden: a real input, the checklist a right answer meets, a reference output, where it came from | `goldens/<id>/` (or a private folder) |
29
38
 
30
39
  ```
31
40
  my-skill/
@@ -33,12 +42,14 @@ my-skill/
33
42
  references/ scripts/ ... (as the Agent Skills spec allows)
34
43
  evals/evals.json task evals (level: tested)
35
44
  evals/triggers.json trigger evals (level: tested)
36
- goldens/<id>/input.md approved examples (level: superskill)
45
+ MISSES.md miss log + history (level: superskill)
46
+ evals/results/latest.json last --run (level: superskill)
47
+ evals/real-runs.json real-run record (level: superskill)
48
+ goldens/<id>/input.md optional evidence
49
+ goldens/<id>/expectations.json
37
50
  goldens/<id>/output.md
38
51
  goldens/<id>/PROVENANCE.json
39
52
  goldens/<id>/APPROVAL.json
40
- MISSES.md miss log (level: superskill)
41
- evals/results/latest.json last --run (level: superskill)
42
53
  ```
43
54
 
44
55
  ## Compatibility
@@ -78,6 +89,7 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
78
89
  | `body-size` | info | reports lines and estimated tokens (characters / 4) once the body passes about 5000 tokens. Length alone is never a defect. |
79
90
  | `rules-above-the-fold` | warn | in a body past about 5000 tokens, every hard rule (a shouted NEVER, ALWAYS, MUST, DO NOT, REFUSE, or a bolded **Never ...** command, outside code fences) appears in the first 5000 tokens, either there or restated there. After compaction Claude Code keeps only that much of each invoked skill. |
80
91
  | `reference-says-when` | warn | every link from SKILL.md to a markdown file sits on a line that says when to read it (before, when, if, for, read ...). Step files (`steps/<step>.md`) are the recommended way to keep a long skill's detail out of the always-loaded body: they are read fresh when the step comes up, so compaction does not lose them. |
92
+ | `history-in-skill` | warn | no dated incident story in the body outside code: "Earned 2026-09-08", "(Gary, 2026-09-16: ...)", "Wilson, 2026-09-05: \*"..."\*", "on 2026-09-13 a session ...", "until 2026-09-21 the flag ...", "measured 2026-09-20", "(2026-09-30, #324)", "- 2026-09-15 (#147): ...". Matched per paragraph, since a line wrap can split a name from its date. Not flagged: a line with "e.g." or "example", a date in a code span or fence, a heading, an HTML comment (a generator's provenance stamp), "as of <date>". The story goes in `MISSES.md` (below); `SKILL.md` keeps the rule and a one-line why. There is deliberately no line-count rule (retired 2026-09-28): length alone is never a defect. |
81
93
  | `navigable` | warn | a body over 300 lines has no run of more than 150 lines without a heading |
82
94
  | `no-repeated-paragraphs` | warn | no paragraph of 100+ characters appears twice |
83
95
  | `references-one-deep` | fail | a markdown file linked from `SKILL.md` links on to no further local file |
@@ -99,13 +111,20 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
99
111
 
100
112
  | Rule | Severity | Threshold |
101
113
  |---|---|---|
102
- | `golden-approved` | fail | at least one golden **from a real run** (its `PROVENANCE.json` says `source: real-run`, names the run, and names who accepted it and when; an anonymized twin also carries `ANONYMIZED.json`) has an approval with non-empty `approved_by` and a valid `approved_at`. An approved golden without that provenance is named and does not count. Info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
103
- | `real-use` | info | when the run ledger records what the person's next message made of each run (Freedom's `next_turn`), the share of those runs in the last 30 days they accepted; sandbox runs are never counted |
104
114
  | `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
105
- | `no-stale-open-miss` | fail | no miss has been open more than 14 days |
115
+ | `misses-closed` | fail | no miss is open. A miss is a known way the skill fails; a skill that still fails a known way is not at the top level, however recent the miss |
106
116
  | `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
117
+ | `evals-pass` | fail | the last `--run` was against the current `SKILL.md` (`skill_sha` equals the sha256 of `SKILL.md`); its suite passed at least the minimum share of runs with the skill (default 90%); every fixed miss's regression eval passed every run; and it ran the trigger evals and got at least the minimum share right (default 90%) |
107
118
  | `run-evidence` | fail | `evals/results/latest.json` exists and `with_skill.pass_rate` > `without_skill.pass_rate` |
108
119
  | `run-fresh` | fail / warn | the last run is younger than the cadence window: `daily` or `weekly` 30 days, `monthly` 60, `quarterly` 120, `yearly` 365, none declared 30 (info). A `yearly` skill always warns to `--run` before its next real use. An unknown cadence warns |
120
+ | `real-runs` | fail | a real-run record (below) in which at least one model+harness pair, on its own, has at least the minimum runs (default 5) and one-shot share (default 80%). Each pair is reported as info; **pairs are never pooled**, since a clean record on one model in one harness proves nothing about another. A pair whose model or harness is `unknown` is reported and never counts |
121
+ | `golden-approved` | info / warn | never fails. Reports each golden that is approved and from a real run (with its approval weight), names approved goldens that are not from a real run and why, names a golden with no `expectations.json`, and warns on an `APPROVAL.json` or `PROVENANCE.json` that does not parse |
122
+
123
+ **The thresholds are configurable, never by the skill itself.** A skill cannot lower its own bar,
124
+ so they are not frontmatter. `doctor` takes `--min-real-runs <n>`, `--min-one-shot <0..1>`,
125
+ `--min-pass-rate <0..1>` and `--min-trigger-rate <0..1>`, or the environment variables
126
+ `SUPERSKILL_MIN_REAL_RUNS`, `SUPERSKILL_MIN_ONE_SHOT`, `SUPERSKILL_MIN_PASS_RATE` and
127
+ `SUPERSKILL_MIN_TRIGGER_RATE`. A report scored with non-default thresholds should say so.
109
128
 
110
129
  `metadata.cadence` in `SKILL.md` frontmatter declares how often the skill really runs:
111
130
 
@@ -155,11 +174,18 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
155
174
  ]
156
175
  ```
157
176
 
158
- ### `goldens/<id>/`
177
+ ### `goldens/<id>/` (optional evidence)
178
+
179
+ **A golden is an eval whose grading is a checklist of expected behavior (0.6.0).** It is never
180
+ required for any level. Grade the outcome, not the path: the checklist says what a right answer
181
+ does, and the reference output proves the task is solvable and calibrates the grader. A golden
182
+ that exists is still held to the provenance rules below, because "a person accepted this when
183
+ it ran" is a claim that has to be true.
159
184
 
160
185
  - `input.md`: the real request.
161
- - `output.md` (or any other file that is not `input.*` or `APPROVAL.json`): the output a person
162
- said was right.
186
+ - `expectations.json`: an array of expectations (the same forms as `evals.json`), graded by
187
+ `--run`. Without it, `--run` grades by likeness to the reference output.
188
+ - `output.md` (or any other file that is not `input.*` or a metadata file): the reference output.
163
189
  - `APPROVAL.json`, written only by `superskill approve`: at an interactive terminal, or relayed by an agent with `--approved-by` and `--via` after the person approved with a tap (the `via` field records where):
164
190
 
165
191
  ```json
@@ -184,11 +210,9 @@ approval whose `note` is its rationale.
184
210
  A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
185
211
  output.
186
212
 
187
- **`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run
188
- counts toward superskill. An invented input with an invented output puts "a person said this was
189
- right" on something no person's work produced, and a skill could then reach the top level on its
190
- author's fiction. Invented cases still belong in `evals/evals.json`, where they hold a skill at
191
- `tested`.
213
+ **`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run is
214
+ reported as real evidence. An invented input with an invented output puts "a person said this
215
+ was right" on something no person's work produced. Invented cases belong in `evals/evals.json`.
192
216
 
193
217
  ```json
194
218
  {
@@ -208,13 +232,12 @@ author's fiction. Invented cases still belong in `evals/evals.json`, where they
208
232
  anonymizer's receipt (`checker`, `checked_at`, `counts` by kind, `fingerprint`, never the
209
233
  mapping). Without the receipt the twin does not count.
210
234
  - `superskill init --from-session` writes `source: real-run` with `accepted` empty, so the golden
211
- cannot count until someone records who accepted it.
235
+ is not reported as real until someone records who accepted it.
212
236
 
213
237
  **Private goldens.** A real run's input and output are usually about real people and should not
214
238
  travel with a skill that is shared. `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`)
215
239
  makes `doctor`, `approve` and `doctor --run` also read `<dir>/<skill-name>/<id>/`, laid out exactly
216
- like `goldens/<id>/`. A private golden counts for the operator who holds it. A miss is closed only
217
- by an eval or golden that ships with the skill.
240
+ like `goldens/<id>/`. A miss is closed only by an eval or golden that ships with the skill.
218
241
 
219
242
  ### `MISSES.md`
220
243
 
@@ -226,11 +249,29 @@ by an eval or golden that ships with the skill.
226
249
  - Should have: Merged duplicates into one line.
227
250
  - Fix: a1b2c3d
228
251
  - Eval: m1
252
+ - Quote: "why is Atlas in here twice" (optional)
229
253
  - Source: freedom-ledger inv_2026-09-20T10-00-00Z_ab12 (optional)
230
254
  ```
231
255
 
232
256
  Headings are `## <id> · <YYYY-MM-DD> · <open|fixed>`; `|` or `-` also separate. Ids are `m1`,
233
- `m2`, ... Other headings are ignored.
257
+ `m2`, ... Other headings are ignored. A field runs on over indented lines that follow it.
258
+
259
+ **`MISSES.md` is where a skill's history lives (0.6.0).** `SKILL.md` is instructions, loaded on
260
+ every run; the story of the incident that earned a rule is history, and Anthropic's authoring
261
+ guidance is to keep time-sensitive content out of instructions. So a story becomes a miss entry:
262
+
263
+ | Field | Holds |
264
+ |---|---|
265
+ | id, date | `m<N>`, and the date the incident happened (not the date it was written down) |
266
+ | What happened | the incident, in full: what the skill did, what it cost, what was noticed |
267
+ | Should have | what the skill should have done |
268
+ | Fix | the commit, or the rule now in `SKILL.md` that the incident earned |
269
+ | Eval | the regression eval that would catch it again. Empty until one exists, and the doctor says so (`fixed-miss-has-eval`): a story with no eval is a fix nothing guards yet |
270
+ | Quote | optional: what the person said, verbatim |
271
+ | Source | optional: where it was recorded (a ledger id, an issue) |
272
+
273
+ `SKILL.md` keeps the rule and a one-line why. `superskill miss <skill> "<what>" --date <d> --quote
274
+ "<q>"` writes one.
234
275
 
235
276
  ### `evals/results/latest.json`
236
277
 
@@ -241,15 +282,55 @@ Written only by `superskill doctor --run`:
241
282
  "run_at": "2026-09-20T10:00:00.000Z",
242
283
  "harness": "claude",
243
284
  "model": "<model id the harness reported>",
285
+ "skill_sha": "<sha256 of the SKILL.md the run proved>",
244
286
  "cases": 4,
245
287
  "repeat": 3,
246
288
  "with_skill": { "pass_rate": 1.0, "mean_ms": 21000, "mean_tokens": 4100 },
247
289
  "without_skill": { "pass_rate": 0.33, "mean_ms": 18000, "mean_tokens": 3900 },
248
- "per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }]
290
+ "per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }],
291
+ "triggers": { "cases": 12, "runs": 36, "passes": 35, "pass_rate": 0.972,
292
+ "per_query": [{ "query": "write my weekly status", "should_trigger": true, "runs": 3, "passes": 3 }] }
249
293
  }
250
294
  ```
251
295
 
252
296
  A run passes when every expectation of its case passes; `pass_rate` is passing runs over runs.
297
+ A trigger run passes when the harness loaded the skill exactly when `should_trigger` says
298
+ (Claude Code: a `Skill` tool call naming it or a read of its `SKILL.md`; Codex: a read of its
299
+ `SKILL.md`). A golden with `expectations.json` is graded on that checklist; one without is graded
300
+ by likeness to its reference output.
301
+
302
+ ### `evals/real-runs.json` (the real-run record)
303
+
304
+ Harness-neutral: any harness, ledger or script may write it, and the doctor only reads it.
305
+
306
+ ```json
307
+ {
308
+ "format": "superskill-real-runs/1",
309
+ "skill": "weekly-status",
310
+ "generated_at": "2026-10-06T12:00:00Z",
311
+ "source": "freedom-skill-ledger",
312
+ "pairs": [
313
+ { "model": "claude-opus-5-5", "harness": "claude-code", "runs": 12, "one_shot": 11,
314
+ "first_at": "2026-09-01T09:00:00Z", "last_at": "2026-10-05T18:00:00Z" }
315
+ ]
316
+ }
317
+ ```
318
+
319
+ - A **run** is one real use of the skill by a person. Never a sandbox run (`doctor --run`), never
320
+ a test.
321
+ - It is **one-shot** when the person needed no correction, rescue or redirect, it did not fail or
322
+ get abandoned, and it was not corrected after it handed back. A ledger outcome that already
323
+ says the person rescued the run (Freedom's `succeeded_with_rescues`) is not one-shot either,
324
+ whether or not the rescue was classified. A taste note is the person's preference, not the
325
+ skill's defect, and does not break one-shot.
326
+ - `model` is the model id the harness ran (as the transcript or environment reports it);
327
+ `harness` is `claude-code`, `codex`, or another harness's name. A writer that cannot tell
328
+ writes `unknown`, never a guess. Counts only: the file carries no input, output or names, so it
329
+ can ship with the skill.
330
+
331
+ **Where the doctor reads it**, first found wins: `<dir>/<skill-name>.json` where `<dir>` is
332
+ `--real-runs` or `SUPERSKILL_REAL_RUNS` (an operator's own export, kept out of the skill); then
333
+ `evals/real-runs.json` in the skill; then Freedom's skill ledger, read live (below).
253
334
 
254
335
  ## Collections and plugins
255
336
 
@@ -285,6 +366,11 @@ exist they are read as plain files:
285
366
  sandbox run) is skipped. A record corrected after it handed back (`corrected_after`, with
286
367
  `next_turn_ref`) gets a "Should have" that points at the person's correction: the session and
287
368
  the time, never their words.
369
+ - **The real-run record.** Freedom's skill ledger records, on every run, the `model` id and the
370
+ `harness` (`claude-code`, `codex`, ...), read from the session transcript or the environment and
371
+ `unknown` when neither says. Freedom exports it in the `superskill-real-runs/1` format above,
372
+ one file per skill, for `--real-runs`. With no export and no `evals/real-runs.json`, the doctor
373
+ reads the same ledger files live and counts them the same way (sandbox runs skipped).
288
374
  - **`doctor --run` is never recorded as a use.** Every harness spawns its child with
289
375
  `FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`, in a folder named `superskill-run-*`.
290
376
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@supersuit/superskill",
3
- "version": "0.5.0",
3
+ "version": "0.6.1",
4
4
  "description": "Score any agent skill folder as skill, tested, or superskill. An open standard and a zero-dependency CLI.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -9,8 +9,9 @@ export const help = `superskill doctor <path...> [options]
9
9
 
10
10
  Score one skill, a folder of skills, or a plugin (skills/*/SKILL.md).
11
11
  Levels: skill (spec-valid, hygienic) < tested (real evals + triggers) < superskill
12
- (an approved golden from a real run a person accepted, misses fixed with evals, a fresh
13
- --run that beats no-skill).
12
+ (every miss fixed with a regression eval; the suite and triggers pass against this SKILL.md
13
+ in a fresh --run that beats no-skill; and a real-run record: at least 5 real runs, 80%
14
+ one-shot, on one model+harness). Goldens are optional evidence.
14
15
 
15
16
  Options:
16
17
  --level <skill|tested|superskill> target level for the exit code (default skill)
@@ -23,6 +24,13 @@ Options:
23
24
  model calls; see "superskill doctor --run --help")
24
25
  --private-goldens <dir> also read goldens from <dir>/<skill-name>/<id>/ (or set
25
26
  SUPERSKILL_PRIVATE_GOLDENS): real runs kept out of the skill
27
+ --real-runs <dir> read the real-run record from <dir>/<skill-name>.json (or
28
+ set SUPERSKILL_REAL_RUNS); else evals/real-runs.json, else
29
+ Freedom's skill ledger
30
+ --min-real-runs <n> real runs one model+harness needs (default 5)
31
+ --min-one-shot <0..1> one-shot share that pair needs (default 0.8)
32
+ --min-pass-rate <0..1> suite pass rate the last --run needs (default 0.9)
33
+ --min-trigger-rate <0..1> trigger evals right in the last --run (default 0.9)
26
34
  --now <iso date> evaluate dates as of this moment
27
35
  --help this text
28
36
 
@@ -37,7 +45,7 @@ export async function run(argv) {
37
45
  const level = a.flags.level || "skill";
38
46
  if (!LEVELS.includes(level)) throw new UsageError(`--level must be one of ${LEVELS.join(", ")}`);
39
47
  if (!a._.length) throw new UsageError("doctor needs a path");
40
- const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null };
48
+ const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null, ...thresholds(a.flags) };
41
49
  if (a.flags.changed) {
42
50
  const all = a._.flatMap((p) => findSkills(p));
43
51
  const files = a._.flatMap((p) => changedFiles(p, a.flags.base));
@@ -64,3 +72,21 @@ export async function run(argv) {
64
72
  }
65
73
  return result.ok ? 0 : 1;
66
74
  }
75
+
76
+ /** The superskill bar's thresholds, from flags then environment. Unset means the documented default. */
77
+ export function thresholds(flags, env = process.env) {
78
+ const pick = (flag, envName, lo, hi) => {
79
+ const raw = flags[flag] ?? env[envName];
80
+ if (raw === undefined || raw === "") return undefined;
81
+ const n = Number(raw);
82
+ if (!Number.isFinite(n) || n < lo || (hi !== null && n > hi)) throw new UsageError(`--${flag} must be a number${hi === null ? ` of at least ${lo}` : ` from ${lo} to ${hi}`}`);
83
+ return n;
84
+ };
85
+ return {
86
+ realRuns: flags["real-runs"] || env.SUPERSKILL_REAL_RUNS || null,
87
+ minRealRuns: pick("min-real-runs", "SUPERSKILL_MIN_REAL_RUNS", 1, null),
88
+ minOneShot: pick("min-one-shot", "SUPERSKILL_MIN_ONE_SHOT", 0, 1),
89
+ minPassRate: pick("min-pass-rate", "SUPERSKILL_MIN_PASS_RATE", 0, 1),
90
+ minTriggerRate: pick("min-trigger-rate", "SUPERSKILL_MIN_TRIGGER_RATE", 0, 1),
91
+ };
92
+ }
@@ -2,11 +2,12 @@ import { parseArgs, clock, UsageError } from "../args.mjs";
2
2
  import { skillDir, today } from "./common.mjs";
3
3
  import { readMisses, appendMisses, nextMissId } from "../misses.mjs";
4
4
 
5
- export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"]
5
+ export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"] [--quote "<what they said>"] [--date YYYY-MM-DD]
6
6
  superskill miss import <skill> --freedom-ledger [--ledger <file>]
7
7
 
8
- Log a time the skill got something wrong. The entry opens today; an open miss older than
9
- 14 days blocks the superskill level. Close it with \`superskill fix\`.
8
+ Log a time the skill got something wrong. The entry opens today (or on --date, for a story
9
+ from the past); an open miss blocks the superskill level. Close it with \`superskill fix\`.
10
+ This is also where the story behind a rule goes: SKILL.md keeps the rule and a one-line why.
10
11
  `;
11
12
 
12
13
  export async function run(argv) {
@@ -18,7 +19,9 @@ export async function run(argv) {
18
19
  if (!what) throw new UsageError('miss needs a description: superskill miss <skill> "<what happened>"');
19
20
  const misses = readMisses(dir) || [];
20
21
  const id = nextMissId(misses);
21
- appendMisses(dir, [{ id, date: today(clock(a.flags)), status: "open", what, expected: a.flags.expected || "" }]);
22
+ const date = a.flags.date || today(clock(a.flags));
23
+ if (!/^\d{4}-\d{2}-\d{2}$/.test(date)) throw new UsageError("--date must be YYYY-MM-DD");
24
+ appendMisses(dir, [{ id, date, status: "open", what, expected: a.flags.expected || "", quote: a.flags.quote || "" }]);
22
25
  process.stdout.write(`${id} logged (open). Close it with: superskill fix ${a._[0]} ${id} --eval <id>\n`);
23
26
  return 0;
24
27
  }
package/src/goldens.mjs CHANGED
@@ -1,4 +1,5 @@
1
- // goldens/<id>/: input.md, the approved output (output.md, or any other non-input file),
1
+ // goldens/<id>/: input.md, expectations.json (0.6.0: the checklist a right answer meets,
2
+ // graded by --run), the reference output (output.md, or any other non-input file),
2
3
  // APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
3
4
  // rationale, basis, evidence}] } (or the single-approval shape from before 0.3.0),
4
5
  // PROVENANCE.json saying where the example came from (0.5.0), and, for an anonymized twin of a
@@ -11,7 +12,7 @@ import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
11
12
  import { join } from "node:path";
12
13
 
13
14
  /** Files in a golden folder that describe it rather than being its input or output. */
14
- export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json"]);
15
+ export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json", "expectations.json"]);
15
16
 
16
17
  /** Where goldens are read from: the skill's own goldens/, then the private folder for its name. */
17
18
  export function goldenRoots(dir, { privateGoldens = null, name = null } = {}) {
@@ -40,6 +41,8 @@ export function readGoldens(dir, opts = {}) {
40
41
  const prov = readJsonFile(join(gdir, "PROVENANCE.json"));
41
42
  const anon = readJsonFile(join(gdir, "ANONYMIZED.json"));
42
43
  const provenance = prov.value;
44
+ const exp = readJsonFile(join(gdir, "expectations.json"));
45
+ const expList = Array.isArray(exp.value) ? exp.value : Array.isArray(exp.value?.expectations) ? exp.value.expectations : [];
43
46
  out.push({
44
47
  id,
45
48
  dir: gdir,
@@ -54,6 +57,10 @@ export function readGoldens(dir, opts = {}) {
54
57
  provenanceError: prov.error,
55
58
  anonymized: anon.value,
56
59
  origin: originOf(provenance, anon.value),
60
+ // 0.6.0: what a right answer does, graded by --run. output.md is the reference that
61
+ // proves the task is solvable; it is no longer what the output has to look like.
62
+ expectations: expList.filter((x) => typeof x === "string" && x.trim()),
63
+ expectationsError: exp.error,
57
64
  });
58
65
  }
59
66
  }
package/src/misses.mjs CHANGED
@@ -1,41 +1,49 @@
1
- // MISSES.md: one dated entry per time the skill got something wrong.
1
+ // MISSES.md: one dated entry per time the skill got something wrong. It is also where a
2
+ // skill's history lives (0.6.0): the story of the incident that earned a rule goes here, and
3
+ // SKILL.md keeps only the rule and a one-line why.
2
4
  //
3
5
  // ## m1 · 2026-09-20 · fixed
4
- // - What happened: ...
6
+ // - What happened: ... (an indented line under a field continues it)
5
7
  // - Should have: ...
6
- // - Fix: <commit sha or note>
8
+ // - Fix: <commit sha or note: the rule now in SKILL.md>
7
9
  // - Eval: m1
10
+ // - Quote: "what the person said" (optional)
8
11
  // - Source: freedom-ledger <id> (optional)
9
12
  import { readFileSync, existsSync, writeFileSync } from "node:fs";
10
13
  import { join } from "node:path";
11
14
 
12
15
  export const MISSES_HEADER = `# Misses
13
16
 
14
- Every time this skill got something wrong. An open miss older than 14 days blocks the
15
- superskill level; a fixed miss must name the eval that would catch it again.
17
+ Every time this skill got something wrong, and the story behind each rule it learned. An open
18
+ miss blocks the superskill level; a fixed miss must name the eval that would catch it again.
16
19
  Written by \`superskill miss\` and \`superskill fix\`, and readable by hand.
17
20
  `;
18
21
 
19
22
  const HEAD_RE = /^##\s+(m\d+)\s*[·|\-–]\s*(\d{4}-\d{2}-\d{2})\s*[·|\-–]\s*(open|fixed)\s*$/i;
20
- const FIELDS = { "what happened": "what", "should have": "expected", fix: "fix", eval: "eval", source: "source" };
23
+ const FIELDS = { "what happened": "what", "should have": "expected", fix: "fix", eval: "eval", quote: "quote", source: "source" };
21
24
 
22
25
  export function parseMisses(text) {
23
26
  const out = [];
24
27
  let cur = null;
28
+ let last = null;
25
29
  const lines = String(text).replace(/\r\n?/g, "\n").split("\n");
26
30
  lines.forEach((line, i) => {
27
31
  const h = line.match(HEAD_RE);
28
32
  if (h) {
29
- cur = { id: h[1].toLowerCase(), date: h[2], status: h[3].toLowerCase(), what: "", expected: "", fix: "", eval: "", source: "", line: i + 1 };
33
+ cur = { id: h[1].toLowerCase(), date: h[2], status: h[3].toLowerCase(), what: "", expected: "", fix: "", eval: "", quote: "", source: "", line: i + 1 };
34
+ last = null;
30
35
  out.push(cur);
31
36
  return;
32
37
  }
33
- if (/^##\s/.test(line)) { cur = null; return; }
34
- const f = cur && line.match(/^\s*[-*]\s*([A-Za-z ]+?)\s*:\s*(.*)$/);
35
- if (f) {
36
- const key = FIELDS[f[1].toLowerCase()];
37
- if (key) cur[key] = f[2].trim();
38
- }
38
+ if (/^##\s/.test(line)) { cur = null; last = null; return; }
39
+ if (!cur) return;
40
+ const f = line.match(/^\s*[-*]\s*([A-Za-z ]+?)\s*:\s*(.*)$/);
41
+ const key = f && FIELDS[f[1].toLowerCase()];
42
+ if (key) { cur[key] = f[2].trim(); last = key; return; }
43
+ // A story rarely fits on one line: an indented line continues the field above it.
44
+ if (last && /^\s{2,}\S/.test(line)) { cur[last] = `${cur[last]} ${line.trim()}`.trim(); return; }
45
+ if (!line.trim()) return;
46
+ last = null;
39
47
  });
40
48
  return out;
41
49
  }
@@ -50,6 +58,7 @@ export function formatMiss(m) {
50
58
  const lines = [`## ${m.id} · ${m.date} · ${m.status}`, `- What happened: ${m.what}`];
51
59
  if (m.expected) lines.push(`- Should have: ${m.expected}`);
52
60
  lines.push(`- Fix: ${m.fix || ""}`, `- Eval: ${m.eval || ""}`);
61
+ if (m.quote) lines.push(`- Quote: ${m.quote}`);
53
62
  if (m.source) lines.push(`- Source: ${m.source}`);
54
63
  return lines.join("\n") + "\n";
55
64
  }
@@ -0,0 +1,124 @@
1
+ // The real-run record (0.6.0): how often the skill did the job for a real person with no
2
+ // correction, counted separately for every model and harness it ran on.
3
+ //
4
+ // Harness-neutral on purpose. Any harness, ledger or script can write the file; the doctor only
5
+ // reads it. Format `superskill-real-runs/1`:
6
+ //
7
+ // {
8
+ // "format": "superskill-real-runs/1",
9
+ // "skill": "weekly-status",
10
+ // "generated_at": "2026-10-06T12:00:00Z",
11
+ // "source": "freedom-skill-ledger",
12
+ // "pairs": [
13
+ // { "model": "claude-opus-5-5", "harness": "claude-code", "runs": 12, "one_shot": 11,
14
+ // "first_at": "2026-09-01T09:00:00Z", "last_at": "2026-10-05T18:00:00Z" }
15
+ // ]
16
+ // }
17
+ //
18
+ // A run is one real use: never a sandbox run (`doctor --run`), never a test. It is one-shot when
19
+ // the person needed no correction, rescue or redirect, it did not fail or get abandoned, and it
20
+ // was not corrected after it handed back. A taste note is the person's preference, not a defect,
21
+ // and does not break one-shot.
22
+ //
23
+ // Where the doctor reads it, first found wins:
24
+ // 1. <dir>/<skill-name>.json, where <dir> is --real-runs or SUPERSKILL_REAL_RUNS (an
25
+ // operator's own export, kept out of the skill);
26
+ // 2. <skill>/evals/real-runs.json (counts only, so it can travel with the skill);
27
+ // 3. Freedom's skill ledger, read as plain files, when neither exists.
28
+ import { readFileSync, existsSync } from "node:fs";
29
+ import { join } from "node:path";
30
+ import { ledgerPaths } from "./ledger.mjs";
31
+
32
+ export const FORMAT = "superskill-real-runs/1";
33
+ export const UNKNOWN = "unknown";
34
+ /** The defaults the real-runs rule holds a skill to. `--min-real-runs` / `--min-one-shot` change them. */
35
+ export const MIN_REAL_RUNS = 5;
36
+ export const MIN_ONE_SHOT = 0.8;
37
+
38
+ const MISS_KINDS = new Set(["redirect", "correction", "rescue"]);
39
+
40
+ /** Is this a known model+harness pair? A record that could not say proves nothing about either. */
41
+ export const knownPair = (p) => Boolean(p.model && p.harness && p.model !== UNKNOWN && p.harness !== UNKNOWN);
42
+
43
+ function normalize(doc, where) {
44
+ if (!doc || typeof doc !== "object") return { error: `${where} is not a JSON object` };
45
+ if (doc.format && doc.format !== FORMAT) return { error: `${where} has format ${JSON.stringify(doc.format)}, not ${FORMAT}` };
46
+ if (!Array.isArray(doc.pairs)) return { error: `${where} has no pairs array` };
47
+ const pairs = doc.pairs
48
+ .filter((p) => p && typeof p === "object")
49
+ .map((p) => ({
50
+ model: String(p.model || UNKNOWN),
51
+ harness: String(p.harness || UNKNOWN),
52
+ runs: Math.max(0, Number(p.runs) || 0),
53
+ one_shot: Math.max(0, Number(p.one_shot) || 0),
54
+ first_at: p.first_at || null,
55
+ last_at: p.last_at || null,
56
+ }))
57
+ .map((p) => ({ ...p, one_shot: Math.min(p.one_shot, p.runs) }));
58
+ return { pairs, generated_at: doc.generated_at || null, source: doc.source || null, where };
59
+ }
60
+
61
+ function readFile(p, where) {
62
+ try { return normalize(JSON.parse(readFileSync(p, "utf8")), where); }
63
+ catch (e) { return { error: `${where} is not valid JSON: ${e.message}` }; }
64
+ }
65
+
66
+ /** Is a ledger record a one-shot run? Exported so a writer can count the same way the reader does. */
67
+ export function isOneShot(rec) {
68
+ const kinds = (Array.isArray(rec?.interventions) ? rec.interventions : []).filter((i) => i && MISS_KINDS.has(i.kind));
69
+ // Freedom derives `succeeded_with_rescues` from operator turns during the run even when nobody
70
+ // classified them, so an empty interventions list next to that outcome is still a rescue.
71
+ return !kinds.length && !["failed", "abandoned", "succeeded_with_rescues"].includes(rec?.outcome) && !rec?.corrected_after;
72
+ }
73
+
74
+ /** Summarize ledger records (Freedom's shape, or any with model/harness/started) per model+harness. */
75
+ export function pairsFromRecords(records, skillName = null) {
76
+ const by = new Map();
77
+ for (const rec of records) {
78
+ if (!rec || rec.synthetic) continue;
79
+ if (skillName && rec.skill && bare(rec.skill) !== bare(skillName)) continue;
80
+ const model = String(rec.model || UNKNOWN), harness = String(rec.harness || UNKNOWN);
81
+ const k = `${model}\u0000${harness}`;
82
+ const p = by.get(k) || { model, harness, runs: 0, one_shot: 0, first_at: null, last_at: null };
83
+ p.runs++;
84
+ if (isOneShot(rec)) p.one_shot++;
85
+ const t = typeof rec.started === "string" ? rec.started : null;
86
+ if (t && (!p.first_at || t < p.first_at)) p.first_at = t;
87
+ if (t && (!p.last_at || t > p.last_at)) p.last_at = t;
88
+ by.set(k, p);
89
+ }
90
+ return [...by.values()].sort((a, b) => b.runs - a.runs || a.model.localeCompare(b.model));
91
+ }
92
+
93
+ const bare = (name) => String(name || "").split(":").pop();
94
+
95
+ function fromLedger(dir, name) {
96
+ const paths = ledgerPaths(dir, name);
97
+ if (!paths.length) return null;
98
+ const recs = [];
99
+ for (const p of paths) {
100
+ let text = "";
101
+ try { text = readFileSync(p, "utf8"); } catch { continue; }
102
+ for (const line of text.split("\n")) {
103
+ if (!line.trim()) continue;
104
+ try { recs.push(JSON.parse(line)); } catch {}
105
+ }
106
+ }
107
+ return { pairs: pairsFromRecords(recs, name), generated_at: null, source: "freedom-skill-ledger (read live)", where: "Freedom's skill ledger" };
108
+ }
109
+
110
+ /** The real-run record for one skill, or null when nothing records any. */
111
+ export function readRealRuns(dir, { name, realRunsDir = null } = {}) {
112
+ if (realRunsDir && name) {
113
+ const p = join(realRunsDir, `${name}.json`);
114
+ if (existsSync(p)) return readFile(p, `${realRunsDir}/${name}.json`);
115
+ }
116
+ const local = join(dir, "evals", "real-runs.json");
117
+ if (existsSync(local)) return readFile(local, "evals/real-runs.json");
118
+ return fromLedger(dir, name);
119
+ }
120
+
121
+ /** Does any single pair clear the bar? Never pooled: each pair stands or falls on its own runs. */
122
+ export function meetsBar(pairs, { minRuns = MIN_REAL_RUNS, minOneShot = MIN_ONE_SHOT } = {}) {
123
+ return pairs.filter(knownPair).filter((p) => p.runs >= minRuns && p.one_shot / p.runs >= minOneShot);
124
+ }
@@ -26,6 +26,63 @@ function proseLines(body) {
26
26
  return { line, fenced };
27
27
  });
28
28
  }
29
+ // A dated incident story, as written in real skills: "Earned 2026-09-08", "(Gary, 2026-09-16:
30
+ // ...)", "Wilson, live, 2026-09-05: *"...", "on 2026-09-13 the in-process version sat...",
31
+ // "until 2026-09-21 the flag...", "measured on 2026-09-20", "(2026-09-30, #324)",
32
+ // "- 2026-09-15 (#147): ...". Matched per paragraph, because hard-wrapped prose puts the name on
33
+ // one line and its date on the next. Tuned against Freedom's 103 shipped skills (2026-10-06): a
34
+ // date in an example, a template, a code span, a fence, or a provenance line ending at the date
35
+ // ("vendored ... on 2026-09-19.") is not a story.
36
+ const D = "20\\d\\d-\\d\\d-\\d\\d";
37
+ const HISTORY = [
38
+ new RegExp(`\\bearned\\b[^.]{0,24}?\\b(${D})`, "gi"),
39
+ new RegExp(`\\([^()\\n]{0,80}?\\b(${D})\\s*[:),.;—]`, "g"),
40
+ new RegExp(`(?:\\b[A-Z][\\w'-]+|\\b(?:operator|owner|maintainer|client|teammate)),\\s*(?:[a-z]+,\\s*)?(${D})\\s*[,:]`, "g"),
41
+ new RegExp(`\\b(?:on|until|since|before|after)\\s+(${D})\\b`, "gi"),
42
+ new RegExp(`\\bfrom\\s+(${D})\\s+\\(`, "gi"),
43
+ new RegExp(`\\b(?:measured|reported|found|caught|hit|corrected|retired|aligned|renamed|refused|broke|failed|observed|verified|watched|fixed|softened|flipped|restored|removed|changed|added|introduced)\\s+(?:[a-z]+\\s+)?(?:on\\s+|in\\s+)?(${D})`, "gi"),
44
+ new RegExp(`\\b(?:rule|ruling|default|correction|incident|reversal)\\s+(?:of|from)\\s+(${D})`, "gi"),
45
+ new RegExp(`\\bthe\\s+(${D})\\b`, "gi"),
46
+ new RegExp(`(${D})'s\\b`, "g"),
47
+ new RegExp(`^\\s*[-*]\\s+[*_]*(${D})[*_]*\\s*[(:]`, "gm"),
48
+ ];
49
+ const EXAMPLE = /\b(e\.g\.|example|for instance|such as)\b/i;
50
+
51
+ /**
52
+ * Body lines that narrate a dated incident, outside code fences and inline code: one entry per
53
+ * line that carries the date of a story.
54
+ */
55
+ export function historyLines(body) {
56
+ const lines = proseLines(body);
57
+ const hit = new Map();
58
+ let para = [];
59
+ const flush = () => {
60
+ if (!para.length) return;
61
+ // Join the paragraph, remembering where each line starts, so a match maps back to its line.
62
+ let text = "";
63
+ const starts = [];
64
+ for (const { i, line } of para) { starts.push({ i, at: text.length }); text += line.replace(/`[^`]*`/g, "``") + "\n"; }
65
+ for (const re of HISTORY) {
66
+ re.lastIndex = 0;
67
+ for (const m of text.matchAll(re)) {
68
+ const at = m.index + m[0].lastIndexOf(m[1]);
69
+ const owner = [...starts].reverse().find((s) => s.at <= at);
70
+ const raw = lines[owner.i].line;
71
+ if (EXAMPLE.test(raw)) continue;
72
+ if (!hit.has(owner.i)) hit.set(owner.i, raw.trim());
73
+ }
74
+ }
75
+ para = [];
76
+ };
77
+ lines.forEach(({ line, fenced }, i) => {
78
+ // A heading, an HTML comment (a generator's provenance stamp) and a fence are not prose.
79
+ if (fenced || !line.trim() || /^\s*(#|<!--)/.test(line)) { flush(); return; }
80
+ para.push({ i, line });
81
+ });
82
+ flush();
83
+ return [...hit.entries()].sort((a, b) => a[0] - b[0]).map(([i, text]) => ({ lineNo: i + 1, text }));
84
+ }
85
+
29
86
  const normRule = (line) => line.toLowerCase().replace(/[*_`>#-]/g, "").replace(/\s+/g, " ").trim();
30
87
  const broken = (ctx) => Boolean(ctx.error || ctx.parseError);
31
88
  const str = (v) => (typeof v === "string" ? v : "");
@@ -194,6 +251,20 @@ export const skillRules = defineRules([
194
251
  return [f("warn", `${late.length} hard rule(s) sit past the first ~${FOLD_TOKENS} tokens and would not survive compaction (${shown})`, "Restate them in a short Rules section near the top, or move them up. Length is fine; the rules just need to be above the fold. Step-specific detail can move into step files (steps/<step>.md) that are read fresh when the step comes up.")];
195
252
  },
196
253
  },
254
+ {
255
+ // 0.6.0: SKILL.md is instructions, read on every run. The story of the incident that earned
256
+ // a rule is history: it belongs in MISSES.md, under the miss it records, and SKILL.md keeps
257
+ // the rule and a one-line why. Anthropic's authoring guidance: avoid time-sensitive content.
258
+ id: "history-in-skill",
259
+ level: "skill",
260
+ check(ctx) {
261
+ if (broken(ctx)) return [];
262
+ const hits = historyLines(ctx.body);
263
+ if (!hits.length) return [];
264
+ const shown = hits.slice(0, 3).map((h) => `line ${h.lineNo + ctx.bodyStartLine - 1}: ${h.text.slice(0, 70)}`).join("; ");
265
+ return [f("warn", `${hits.length} dated incident stor${hits.length === 1 ? "y" : "ies"} in SKILL.md (${shown})`, "Move each story into MISSES.md as a miss entry (id, date, what happened, fix, eval, optional quote) and leave the rule plus a one-line why in SKILL.md.")];
266
+ },
267
+ },
197
268
  {
198
269
  id: "navigable",
199
270
  level: "skill",
@@ -1,16 +1,19 @@
1
- // Level 3, "superskill": checked against examples a person approved, fixed every time it
2
- // got something wrong, and proven recently on a current model against the no-skill baseline.
1
+ // Level 3, "superskill" (0.6.0): fixed every time it got something wrong, with a regression eval
2
+ // for each fix; its whole suite and trigger set pass against the SKILL.md that is here now, on a
3
+ // current model and better than no skill; and it did the job for real people, one-shot, often
4
+ // enough on at least one model+harness. Goldens are optional evidence.
3
5
  import { createHash } from "node:crypto";
4
6
  import { defineRules } from "./define.mjs";
5
7
  import { readGoldens, isApproved, isRealApproved, weightOf, goldenOpts } from "../goldens.mjs";
6
- import { ledgerPaths, acceptedRate } from "../ledger.mjs";
8
+ import { readRealRuns, meetsBar, knownPair, MIN_REAL_RUNS, MIN_ONE_SHOT } from "../realruns.mjs";
7
9
  import { readMisses } from "../misses.mjs";
8
10
  import { readEvals, readLatestRun } from "../evals.mjs";
9
11
 
10
12
  const f = (severity, message, fix) => ({ severity, message, fix });
11
13
  const DAY = 86400000;
12
- export const OPEN_MISS_DAYS = 14;
13
- export const REAL_USE_DAYS = 30;
14
+ /** Defaults for evals-pass; `--min-pass-rate` / `--min-trigger-rate` change them. */
15
+ export const MIN_PASS_RATE = 0.9;
16
+ export const MIN_TRIGGER_RATE = 0.9;
14
17
  /** How long a --run stays fresh, by metadata.cadence. */
15
18
  export const FRESH_DAYS = { daily: 30, weekly: 30, monthly: 60, quarterly: 120, yearly: 365 };
16
19
  const DEFAULT_FRESH = 30;
@@ -19,55 +22,57 @@ const days = (now, date) => Math.floor((now.getTime() - new Date(date).getTime()
19
22
 
20
23
  export const superskillRules = defineRules([
21
24
  {
25
+ // 0.6.0: a golden is optional evidence, never the bar. A random accepted run is not the
26
+ // embodiment of a skill; the misses it absorbed and the evals that hold them are. A golden
27
+ // that exists is still an eval (its checklist is graded by --run), so its files are checked.
22
28
  id: "golden-approved",
23
29
  level: "superskill",
24
30
  check(ctx, opts = {}) {
25
31
  if (ctx.error) return [];
26
32
  const goldens = readGoldens(ctx.dir, goldenOpts(ctx, opts));
27
- const approvedAny = goldens.filter(isApproved);
28
- // 0.5.0: only a golden from a real run a person accepted counts. An invented example, however
29
- // careful, puts "a person said this was right" on something no person's work produced, and a
30
- // skill could then reach the top level on its author's fiction.
31
- const approved = goldens.filter(isRealApproved);
33
+ if (!goldens.length) return [];
32
34
  const out = [];
33
35
  const where = (g) => (g.private ? `private golden ${g.id}` : `goldens/${g.id}`);
34
36
  for (const g of goldens.filter((x) => x.approvalError))
35
- out.push(f("fail", `${where(g)}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
37
+ out.push(f("warn", `${where(g)}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
36
38
  for (const g of goldens.filter((x) => x.provenanceError))
37
- out.push(f("fail", `${where(g)}/PROVENANCE.json is not valid JSON`, "Rewrite it: source, run, accepted (see SPEC.md, goldens)."));
38
- if (!approved.length) {
39
- if (approvedAny.length) {
40
- const why = approvedAny.map((g) => `${g.id} ${g.origin.why}`).join("; ");
41
- out.push(f("fail", `${approvedAny.length} approved golden${approvedAny.length === 1 ? "" : "s"}, none from a real run a person accepted (${why})`, "A golden counts toward superskill only when goldens/<id>/PROVENANCE.json says source: real-run, names the run (session, commit, ledger id, or derived_from for an anonymized twin) and who accepted it and when. Invented examples belong in evals/evals.json, where they hold the skill at tested."));
42
- } else {
43
- out.push(f("fail", goldens.length ? `${goldens.length} golden${goldens.length === 1 ? "" : "s"}, none approved by a person` : "no goldens", goldens.length ? "A person runs `superskill approve <skill> <golden>` at a terminal (or relays a tap with --approved-by and --via) after checking the output." : "Save a real run's input and the output a person accepted under goldens/<id>/ with PROVENANCE.json, then `superskill approve`."));
44
- }
45
- return out;
46
- }
39
+ out.push(f("warn", `${where(g)}/PROVENANCE.json is not valid JSON`, "Rewrite it: source, run, accepted (see SPEC.md, goldens)."));
40
+ for (const g of goldens.filter((x) => !x.expectations.length))
41
+ out.push(f("info", `${where(g)} has no expectations.json, so --run grades it by likeness to output.md`, "Write goldens/<id>/expectations.json: the behaviors a right answer shows (grade the outcome, not the path). output.md then stays as the reference that proves the task is solvable."));
42
+ const approved = goldens.filter(isApproved);
43
+ const real = goldens.filter(isRealApproved);
44
+ const unreal = approved.filter((g) => !real.includes(g));
45
+ if (unreal.length) out.push(f("info", `approved golden${unreal.length === 1 ? "" : "s"} not from a real run a person accepted: ${unreal.map((g) => `${g.id} ${g.origin.why}`).join("; ")}`, "Only a golden whose PROVENANCE.json names a real run and who accepted it is evidence of real use; an invented one is an ordinary eval."));
46
+ if (!real.length) return out;
47
47
  const sha = createHash("sha256").update(ctx.raw).digest("hex");
48
- const stale = approved.filter((g) => g.approval.skill_sha && g.approval.skill_sha !== sha).map((g) => g.id);
48
+ const stale = real.filter((g) => g.approval.skill_sha && g.approval.skill_sha !== sha).map((g) => g.id);
49
49
  if (stale.length) out.push(f("info", `golden${stale.length === 1 ? "" : "s"} ${stale.join(", ")} approved against an earlier SKILL.md`, "Re-run the golden and re-approve if the output still holds."));
50
- // Liked is not proven. Say which weight the approvals carry, so "a person approved it" is
51
- // never read as "it worked in the world".
52
- const w = approved.map((g) => ({ id: g.id, ...weightOf(g) }));
53
- const proven = w.filter((x) => x.outcome > 0);
54
- const summary = w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ");
55
- if (!proven.length) out.push(f("info", `approved on judgment only, no outcome recorded yet (${summary})`, "When a golden produces a real result, record it: `superskill approve <skill> <golden> --basis outcome --evidence \"<what happened, where to check>\"`."));
56
- else out.push(f("info", `approval weight: ${summary}`, ""));
50
+ const w = real.map((g) => ({ id: g.id, ...weightOf(g) }));
51
+ out.push(f("info", `real goldens (optional evidence): ${w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ")}`, ""));
57
52
  return out;
58
53
  },
59
54
  },
60
55
  {
61
- // The real number beside the level: of the runs whose next message is recorded, how many did
62
- // the person accept. Info only. A level is evidence about examples; this is evidence about use.
63
- id: "real-use",
56
+ // The record from real use, per model and harness. A clean record on one model in one harness
57
+ // proves nothing about another, so pairs are never pooled: one pair has to clear the bar alone.
58
+ id: "real-runs",
64
59
  level: "superskill",
65
- check(ctx, { now = new Date() } = {}) {
60
+ check(ctx, opts = {}) {
66
61
  if (ctx.error) return [];
67
62
  const name = (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName;
68
- const r = acceptedRate(ledgerPaths(ctx.dir, name), { now, days: REAL_USE_DAYS });
69
- if (!r.judged) return [];
70
- return [f("info", `real use, last ${REAL_USE_DAYS} days: ${r.accepted} of ${r.judged} judged runs accepted (${pct(r.accepted / r.judged)})${r.synthetic ? `; ${r.synthetic} sandbox run${r.synthetic === 1 ? "" : "s"} not counted` : ""}`, "")];
63
+ const minRuns = num(opts.minRealRuns, MIN_REAL_RUNS), minOneShot = num(opts.minOneShot, MIN_ONE_SHOT);
64
+ const bar = `at least ${minRuns} real runs and ${pct(minOneShot)} one-shot on one model+harness`;
65
+ const realRunsDir = opts.realRuns ?? ctx.realRuns ?? process.env.SUPERSKILL_REAL_RUNS ?? null;
66
+ const rec = readRealRuns(ctx.dir, { name, realRunsDir });
67
+ const how = "Export the run record (format superskill-real-runs/1, SPEC.md) to evals/real-runs.json or --real-runs <dir>; Freedom's skill ledger exports it.";
68
+ if (!rec) return [f("fail", `no real-run record (needs ${bar})`, how)];
69
+ if (rec.error) return [f("fail", rec.error, how)];
70
+ const out = rec.pairs.map((p) => f("info", `real runs, ${p.model} / ${p.harness}: ${p.one_shot} of ${p.runs} one-shot${p.runs ? ` (${pct(p.one_shot / p.runs)})` : ""}${knownPair(p) ? "" : ", not counted: the record does not say which model or harness"}`, ""));
71
+ if (!meetsBar(rec.pairs, { minRuns, minOneShot }).length) {
72
+ const best = rec.pairs.filter(knownPair)[0];
73
+ out.push(f("fail", `real-run record below the bar (${bar}; ${rec.where})${best ? `: best is ${best.model} / ${best.harness}, ${best.one_shot} of ${best.runs}` : ": no run names its model and harness"}`, "Use the skill for real and fix what it gets wrong; every correction is a miss to close with an eval."));
74
+ }
75
+ return out;
71
76
  },
72
77
  },
73
78
  {
@@ -79,18 +84,17 @@ export const superskillRules = defineRules([
79
84
  },
80
85
  },
81
86
  {
82
- id: "no-stale-open-miss",
87
+ // 0.6.0: every miss is closed. An open miss is a known way the skill fails today, and a skill
88
+ // that still fails a known way is not at the top level, however recent the miss is.
89
+ id: "misses-closed",
83
90
  level: "superskill",
84
91
  check(ctx, { now = new Date() } = {}) {
85
92
  if (ctx.error) return [];
86
93
  const misses = readMisses(ctx.dir) || [];
87
- const out = [];
88
- for (const m of misses.filter((x) => x.status === "open")) {
94
+ return misses.filter((x) => x.status === "open").map((m) => {
89
95
  const age = days(now, m.date);
90
- if (age > OPEN_MISS_DAYS) out.push(f("fail", `miss ${m.id} open ${age} days (limit ${OPEN_MISS_DAYS}): ${m.what.slice(0, 80)}`, `Fix it, add an eval that catches it, then \`superskill fix <skill> ${m.id} --eval <id>\`.`));
91
- else out.push(f("info", `miss ${m.id} open ${age} day${age === 1 ? "" : "s"}`, `Fix within ${OPEN_MISS_DAYS} days of ${m.date}.`));
92
- }
93
- return out;
96
+ return f("fail", `miss ${m.id} is open (${age} day${age === 1 ? "" : "s"}): ${m.what.slice(0, 80)}`, `Fix it, add an eval that catches it, then \`superskill fix <skill> ${m.id} --eval <id>\`.`);
97
+ });
94
98
  },
95
99
  },
96
100
  {
@@ -108,6 +112,39 @@ export const superskillRules = defineRules([
108
112
  .map((m) => f("fail", m.eval ? `miss ${m.id} names eval "${m.eval}", which is not in evals.json or goldens/` : `miss ${m.id} is fixed with no regression eval`, `Add a case to evals/evals.json that would catch ${m.id} again, and name its id on the Eval line.`));
109
113
  },
110
114
  },
115
+ {
116
+ // The suite and the trigger set pass against the SKILL.md that is here now. A pass recorded
117
+ // against an earlier SKILL.md proves the earlier skill.
118
+ id: "evals-pass",
119
+ level: "superskill",
120
+ check(ctx, opts = {}) {
121
+ if (ctx.error) return [];
122
+ const r = readLatestRun(ctx.dir);
123
+ if (!r.run) return [];
124
+ const run = r.run;
125
+ const out = [];
126
+ const sha = createHash("sha256").update(ctx.raw).digest("hex");
127
+ if (!run.skill_sha) out.push(f("fail", "latest.json does not say which SKILL.md it ran against (no skill_sha)", "Re-run `superskill doctor <skill> --run` with superskill 0.6.0 or later."));
128
+ else if (run.skill_sha !== sha) out.push(f("fail", "the last --run was against an earlier SKILL.md", "Re-run `superskill doctor <skill> --run`; a change to the skill needs its evals re-run."));
129
+ const minPass = num(opts.minPassRate, MIN_PASS_RATE), minTrig = num(opts.minTriggerRate, MIN_TRIGGER_RATE);
130
+ const w = Number(run.with_skill?.pass_rate);
131
+ if (Number.isFinite(w) && w < minPass) out.push(f("fail", `the suite passed ${pct(w)} of runs with the skill (needs ${pct(minPass)})`, "Read the failures in latest.json per_case, fix the skill, re-run."));
132
+ // Every fixed miss's regression eval passes every run: a regression that comes back half the
133
+ // time has come back.
134
+ const fixed = (readMisses(ctx.dir) || []).filter((m) => m.status === "fixed" && m.eval);
135
+ const cases = new Map((Array.isArray(run.per_case) ? run.per_case : []).map((c) => [String(c.id), c]));
136
+ for (const m of fixed) {
137
+ const c = cases.get(String(m.eval)) || cases.get(`golden:${m.eval}`);
138
+ if (!c) { if (run.skill_sha) out.push(f("fail", `regression eval ${m.eval} (miss ${m.id}) is not in the last --run`, "Re-run `superskill doctor <skill> --run`.")); continue; }
139
+ const ws = c.with_skill || {};
140
+ if (!(ws.runs > 0 && ws.passes === ws.runs)) out.push(f("fail", `regression eval ${m.eval} (miss ${m.id}) passed ${ws.passes || 0} of ${ws.runs || 0} runs`, `Miss ${m.id} is back: fix the skill until its eval passes every run.`));
141
+ }
142
+ const t = run.triggers;
143
+ if (!t || !Number.isFinite(Number(t.pass_rate))) out.push(f("fail", "the last --run did not run the trigger evals", "Re-run `superskill doctor <skill> --run` with superskill 0.6.0 or later; it runs evals/triggers.json too."));
144
+ else if (Number(t.pass_rate) < minTrig) out.push(f("fail", `trigger evals ${pct(Number(t.pass_rate))} right (needs ${pct(minTrig)})`, "Read triggers.per_query in latest.json: sharpen the description so it loads when it should and not otherwise."));
145
+ return out;
146
+ },
147
+ },
111
148
  {
112
149
  id: "run-evidence",
113
150
  level: "superskill",
@@ -146,3 +183,4 @@ export const superskillRules = defineRules([
146
183
  ]);
147
184
 
148
185
  const pct = (x) => `${Math.round(x * 100)}%`;
186
+ const num = (v, d) => (v === undefined || v === null || v === "" || !Number.isFinite(Number(v)) ? d : Number(v));
@@ -35,3 +35,32 @@ export function ask(prompt, { model } = {}) {
35
35
  if (model) args.push("--model", model);
36
36
  return call(args, cwd).output;
37
37
  }
38
+
39
+ /**
40
+ * Did Claude Code load the skill for this query? Run headless with the skill installed and read
41
+ * the stream: a Skill tool call naming it, or a Read of its SKILL.md, is a load.
42
+ */
43
+ export function loadedSkill(stream, skillName) {
44
+ for (const line of String(stream).split("\n")) {
45
+ if (!line.trim().startsWith("{")) continue;
46
+ let ev; try { ev = JSON.parse(line); } catch { continue; }
47
+ const content = ev?.message?.content;
48
+ if (!Array.isArray(content)) continue;
49
+ for (const c of content) {
50
+ if (c?.type !== "tool_use") continue;
51
+ const input = c.input || {};
52
+ if (c.name === "Skill" && [input.skill, input.command, input.name].some((v) => typeof v === "string" && v.split(":").pop() === skillName)) return true;
53
+ if (typeof input.file_path === "string" && input.file_path.endsWith(`/${skillName}/SKILL.md`)) return true;
54
+ }
55
+ }
56
+ return false;
57
+ }
58
+
59
+ export function triggerCase({ skillDir, skillName, query, model }) {
60
+ const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".claude/skills" });
61
+ const args = ["-p", query, "--output-format", "stream-json", "--verbose", "--add-dir", cwd];
62
+ if (model) args.push("--model", model);
63
+ const r = spawnSync("claude", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
64
+ if (r.error) throw Object.assign(new Error(`claude failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
65
+ return { triggered: loadedSkill(r.stdout, skillName), failed: r.status !== 0 && !r.stdout, cwd };
66
+ }
package/src/run/codex.mjs CHANGED
@@ -31,3 +31,17 @@ export function ask(prompt, { model } = {}) {
31
31
  const cwd = prepareWorkspace({ skillDir: null, skillName: null });
32
32
  return call(prompt, cwd, model).output;
33
33
  }
34
+
35
+ /**
36
+ * Did Codex load the skill for this query? Codex reads a skill's SKILL.md when it uses it, so the
37
+ * JSON event stream mentioning <name>/SKILL.md is a load.
38
+ */
39
+ export function triggerCase({ skillDir, skillName, query, model }) {
40
+ const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".agents/skills" });
41
+ const args = ["exec", "--json", "--skip-git-repo-check", "-C", cwd];
42
+ if (model) args.push("-m", model);
43
+ args.push(query);
44
+ const r = spawnSync("codex", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
45
+ if (r.error) throw Object.assign(new Error(`codex failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
46
+ return { triggered: String(r.stdout).includes(`${skillName}/SKILL.md`), failed: r.status !== 0 && !r.stdout, cwd };
47
+ }
package/src/run/fake.mjs CHANGED
@@ -1,6 +1,6 @@
1
1
  // A stand-in harness for tests: SUPERSKILL_FAKE_HARNESS names a node script that reads one
2
- // JSON request on stdin ({mode: "case"|"ask", prompt, withSkill, cwd}) and prints
3
- // {output, tokens?, model?}. Never calls a model.
2
+ // JSON request on stdin ({mode: "case"|"ask"|"trigger", prompt|query, withSkill, cwd}) and
3
+ // prints {output, tokens?, model?} (or {triggered} for a trigger). Never calls a model.
4
4
  import { spawnSync } from "node:child_process";
5
5
  import { prepareWorkspace } from "./workspace.mjs";
6
6
  import { sandboxEnv } from "./env.mjs";
@@ -22,6 +22,11 @@ export function create(script) {
22
22
  const cwd = prepareWorkspace({ skillDir, skillName, files, linkAt: withSkill ? ".claude/skills" : null });
23
23
  return { ...call(script, { mode: "case", prompt, withSkill, cwd }), cwd };
24
24
  },
25
+ triggerCase({ skillDir, skillName, query }) {
26
+ const cwd = prepareWorkspace({ skillDir, skillName, linkAt: ".claude/skills" });
27
+ const doc = JSON.parse(spawnSync(process.execPath, [script], { input: JSON.stringify({ mode: "trigger", query, cwd }), env: sandboxEnv(), encoding: "utf8" }).stdout || "{}");
28
+ return { triggered: Boolean(doc.triggered), failed: false, cwd };
29
+ },
25
30
  ask(prompt) {
26
31
  return call(script, { mode: "ask", prompt }).output;
27
32
  },
package/src/run/index.mjs CHANGED
@@ -8,7 +8,8 @@ import { createInterface } from "node:readline/promises";
8
8
  import { clock, UsageError } from "../args.mjs";
9
9
  import { findSkills } from "../doctor.mjs";
10
10
  import { parseSkillFile } from "../frontmatter.mjs";
11
- import { readEvals } from "../evals.mjs";
11
+ import { readEvals, readTriggers } from "../evals.mjs";
12
+ import { createHash } from "node:crypto";
12
13
  import { readGoldens } from "../goldens.mjs";
13
14
  import { grade, isMachineCheck } from "./grade.mjs";
14
15
  import * as claude from "./claude.mjs";
@@ -18,7 +19,9 @@ import { create as createFake } from "./fake.mjs";
18
19
  export const help = `superskill doctor <skill> --run [--harness claude|codex] [--repeat 3] [--model <id>] [--yes]
19
20
 
20
21
  Run the skill's evals for real: every case in evals/evals.json and every golden, --repeat
21
- times with the skill and --repeat times without it, through a headless harness. Machine
22
+ times with the skill and --repeat times without it, and every query in evals/triggers.json
23
+ --repeat times with the skill installed (did it load when it should, and not otherwise),
24
+ through a headless harness. Machine
22
25
  checks (contains:, regex:, file_exists:) are free; each plain-language expectation costs
23
26
  one grader call per run. Prints the estimated number of model calls first, and asks
24
27
  before starting (or needs --yes when there is no terminal). Writes
@@ -45,15 +48,51 @@ export function collectCases(skillDir, { privateGoldens = process.env.SUPERSKILL
45
48
  if (e.error) throw new UsageError(e.error);
46
49
  const cases = e.cases.filter((c) => c.prompt.trim()).map((c) => ({ id: String(c.id), prompt: c.prompt, files: c.files, assertions: c.assertions.length ? c.assertions : [c.expected_output].filter(Boolean) }));
47
50
  for (const g of readGoldens(skillDir, { privateGoldens, name })) {
48
- if (!g.input || !g.output || !g.output.trim()) continue;
49
- cases.push({ id: `golden:${g.id}`, prompt: g.input, files: [], assertions: [`The output matches this approved output in substance (same facts, same shape; wording may differ):\n${g.output}`] });
51
+ if (!g.input || !((g.output && g.output.trim()) || g.expectations?.length)) continue;
52
+ // 0.6.0: grade the outcome, not the path. A golden with a checklist is graded on it; one
53
+ // without falls back to likeness with its reference output.
54
+ const assertions = g.expectations?.length ? g.expectations : [`The output matches this approved output in substance (same facts, same shape; wording may differ):\n${g.output}`];
55
+ cases.push({ id: `golden:${g.id}`, prompt: g.input, files: [], assertions });
50
56
  }
51
57
  return cases;
52
58
  }
53
59
 
54
- export function estimateCalls(cases, repeat) {
60
+ export function estimateCalls(cases, repeat, triggers = 0) {
55
61
  const graded = cases.reduce((n, c) => n + c.assertions.filter((a) => !isMachineCheck(a)).length, 0);
56
- return { runs: cases.length * repeat * 2, grader: graded * repeat * 2, total: cases.length * repeat * 2 + graded * repeat * 2 };
62
+ const runs = cases.length * repeat * 2 + triggers * repeat;
63
+ return { runs, grader: graded * repeat * 2, triggers: triggers * repeat, total: runs + graded * repeat * 2 };
64
+ }
65
+
66
+ /** The trigger queries the run will check, from evals/triggers.json. */
67
+ export function collectTriggers(skillDir) {
68
+ const t = readTriggers(skillDir);
69
+ if (t.error) throw new UsageError(t.error);
70
+ return t.triggers;
71
+ }
72
+
73
+ /**
74
+ * Run every trigger query `repeat` times with the skill installed and record whether the harness
75
+ * loaded it. A query passes a run when loading matched should_trigger.
76
+ */
77
+ export function runTriggers(skillDir, { harness, repeat, model, log = () => {} }) {
78
+ const skillName = parseSkillFile(readFileSync(join(skillDir, "SKILL.md"), "utf8")).data.name || "";
79
+ const queries = collectTriggers(skillDir);
80
+ if (!queries.length) return null;
81
+ if (typeof harness.triggerCase !== "function") throw new UsageError(`harness ${harness.name} cannot run trigger evals`);
82
+ let runs = 0, passes = 0;
83
+ const per_query = [];
84
+ for (const q of queries) {
85
+ const row = { query: q.query, should_trigger: q.should_trigger, runs: 0, passes: 0 };
86
+ for (let i = 0; i < repeat; i++) {
87
+ log(`trigger "${q.query.slice(0, 40)}", run ${i + 1}/${repeat}`);
88
+ const r = harness.triggerCase({ skillDir, skillName, query: q.query, model });
89
+ if (r.cwd) rmSync(r.cwd, { recursive: true, force: true });
90
+ const ok = !r.failed && Boolean(r.triggered) === q.should_trigger;
91
+ row.runs++; row.passes += ok ? 1 : 0; runs++; passes += ok ? 1 : 0;
92
+ }
93
+ per_query.push(row);
94
+ }
95
+ return { cases: queries.length, runs, passes, pass_rate: runs ? round(passes / runs) : 0, per_query };
57
96
  }
58
97
 
59
98
  export function runEvals(skillDir, { harness, repeat = 3, now = new Date(), model, log = () => {} }) {
@@ -84,7 +123,10 @@ export function runEvals(skillDir, { harness, repeat = 3, now = new Date(), mode
84
123
  per_case.push(row);
85
124
  }
86
125
  const summary = (t) => ({ pass_rate: t.runs ? round(t.passes / t.runs) : 0, mean_ms: t.runs ? Math.round(t.ms / t.runs) : 0, mean_tokens: t.tokenRuns ? Math.round(t.tokens / t.tokenRuns) : null });
87
- const result = { run_at: now.toISOString(), harness: harness.name, model: seenModel, cases: cases.length, repeat, with_skill: summary(totals.with_skill), without_skill: summary(totals.without_skill), per_case };
126
+ const triggers = runTriggers(skillDir, { harness, repeat, model, log });
127
+ // 0.6.0: the run says which SKILL.md it proved, so a later edit cannot ride on an old pass.
128
+ const skill_sha = createHash("sha256").update(readFileSync(join(skillDir, "SKILL.md"), "utf8")).digest("hex");
129
+ const result = { run_at: now.toISOString(), harness: harness.name, model: seenModel, skill_sha, cases: cases.length, repeat, with_skill: summary(totals.with_skill), without_skill: summary(totals.without_skill), per_case, ...(triggers ? { triggers } : {}) };
88
130
  mkdirSync(join(skillDir, "evals", "results"), { recursive: true });
89
131
  writeFileSync(join(skillDir, "evals", "results", "latest.json"), JSON.stringify(result, null, 2) + "\n");
90
132
  return result;
@@ -100,11 +142,11 @@ export async function runCommand(a) {
100
142
  const harness = pickHarness(a.flags.harness);
101
143
  const dirs = a._.flatMap((p) => findSkills(p));
102
144
  if (!dirs.length) throw new UsageError(`no skills found under ${a._.join(", ")}`);
103
- const plan = dirs.map((d) => ({ dir: d, cases: collectCases(d) }));
104
- const calls = plan.reduce((n, p) => n + estimateCalls(p.cases, repeat).total, 0);
145
+ const plan = dirs.map((d) => ({ dir: d, cases: collectCases(d), triggers: collectTriggers(d).length }));
146
+ const calls = plan.reduce((n, p) => n + estimateCalls(p.cases, repeat, p.triggers).total, 0);
105
147
  for (const p of plan) {
106
- const e = estimateCalls(p.cases, repeat);
107
- process.stderr.write(`${p.dir}: ${p.cases.length} cases x ${repeat} x 2 = ${e.runs} runs + ${e.grader} grader calls\n`);
148
+ const e = estimateCalls(p.cases, repeat, p.triggers);
149
+ process.stderr.write(`${p.dir}: ${p.cases.length} cases x ${repeat} x 2 + ${p.triggers} triggers x ${repeat} = ${e.runs} runs + ${e.grader} grader calls\n`);
108
150
  }
109
151
  process.stderr.write(`estimated model calls: ${calls} through ${harness.name}\n`);
110
152
  if (!a.flags.yes) {
@@ -122,7 +164,7 @@ export async function runCommand(a) {
122
164
  for (const p of plan) {
123
165
  const r = runEvals(p.dir, { harness, repeat, now, model: a.flags.model, log: (m) => process.stderr.write(` ${m}\n`) });
124
166
  results.push({ path: p.dir, ...r });
125
- if (!a.flags.json) process.stdout.write(`${p.dir}\n with skill ${pct(r.with_skill.pass_rate)} without ${pct(r.without_skill.pass_rate)} (${r.cases} cases x ${repeat})\n wrote evals/results/latest.json\n`);
167
+ if (!a.flags.json) process.stdout.write(`${p.dir}\n with skill ${pct(r.with_skill.pass_rate)} without ${pct(r.without_skill.pass_rate)} (${r.cases} cases x ${repeat})${r.triggers ? `\n triggers ${pct(r.triggers.pass_rate)} right (${r.triggers.cases} queries x ${repeat})` : ""}\n wrote evals/results/latest.json\n`);
126
168
  }
127
169
  if (a.flags.json) process.stdout.write(JSON.stringify({ results }, null, 2) + "\n");
128
170
  return results.every((r) => r.with_skill.pass_rate > r.without_skill.pass_rate) ? 0 : 1;
package/src/snippet.md CHANGED
@@ -5,12 +5,14 @@ Skills in this project are held to the superskill standard: https://supersuit.wi
5
5
  - **After doing a job by hand once, propose a skill for it.** If the person agrees, write the
6
6
  skill, then seed its evidence from this session:
7
7
  `npx @supersuit/superskill init <skill-folder> --from-session <this session's transcript>`
8
- The golden it creates waits for the person to approve it; never approve it yourself.
8
+ The golden it creates is optional evidence and waits for the person; never approve it yourself.
9
9
  - **Whenever a skill needed correcting** (the person redirected you, fixed its output, or you
10
10
  worked around it), log it before moving on:
11
11
  `npx @supersuit/superskill miss <skill-folder> "<what happened>" --expected "<what should have>"`
12
12
  When you fix it, add an eval that would catch it again and close it with
13
13
  `npx @supersuit/superskill fix <skill-folder> <miss-id> --eval <eval-id>`.
14
+ - **The story goes in MISSES.md, never in SKILL.md.** SKILL.md holds the rule and a one-line
15
+ why; what happened, when, and what the person said go on the miss entry (`--quote`).
14
16
  - **Before calling a skill done**, run
15
17
  `npx @supersuit/superskill doctor <skill-folder>`
16
18
  and fix everything it lists for the level the skill is meant to reach.