@supersuit/superskill 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,73 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.6.0 (2026-10-06)
4
+
5
+ **A skill's embodiment is what it has absorbed, not one run of it.** 0.5.0 made a golden from a
6
+ real run the top-level requirement. Its first real test reversed it: nine harvested candidates,
7
+ each a real accepted run, and none was the standard the skill should be held to. The tests and
8
+ fixes a skill has absorbed are. **Breaking for anyone at superskill today:** the top level now
9
+ needs a real-run record and a 0.6.0 `--run` (one that records its `skill_sha` and the triggers).
10
+
11
+ - **Superskill is now** every miss closed with a regression eval (`misses-closed`, which replaces
12
+ `no-stale-open-miss`: an open miss fails however recent it is), and `evals-pass`: the last
13
+ `--run` was against the current `SKILL.md`, the suite passed (default 90% of runs), every fixed
14
+ miss's regression eval passed every run, and the trigger evals were right (default 90%).
15
+ `run-evidence` and `run-fresh` stand.
16
+ - **`real-runs`:** a real-run record where one model+harness pair, alone, has at least 5 real runs
17
+ at 80% one-shot (no correction). Every pair is reported; pairs are never pooled, and a pair that
18
+ says `unknown` never counts. The record is a harness-neutral file, `superskill-real-runs/1`, read
19
+ from `--real-runs <dir>` / `SUPERSKILL_REAL_RUNS`, then `evals/real-runs.json`, then Freedom's
20
+ skill ledger live. It replaces `real-use`.
21
+ - **Configurable bar:** `--min-real-runs`, `--min-one-shot`, `--min-pass-rate`,
22
+ `--min-trigger-rate` (and `SUPERSKILL_MIN_*`). Never frontmatter: a skill cannot lower its own bar.
23
+ - **Goldens are optional evidence.** `golden-approved` never fails. A golden is an eval graded on
24
+ `goldens/<id>/expectations.json`, a checklist of expected behavior (grade outcomes, not paths);
25
+ `output.md` is the reference that proves the task solvable. Provenance rules still say which
26
+ goldens are real runs.
27
+ - **`doctor --run` runs the trigger evals** (did the harness load the skill exactly when it
28
+ should) and writes `skill_sha` and `triggers` into `latest.json`.
29
+ - **`history-in-skill` (warn, level skill):** a dated incident story in `SKILL.md` ("Earned
30
+ 2026-09-08", "(Gary, 2026-09-16: ...)", "on 2026-09-13 a session ...") belongs in `MISSES.md`,
31
+ which now holds a story's full entry: id, date, what happened, fix, eval, optional `Quote:`, with
32
+ indented continuation lines. Tuned on Freedom's 103 shipped skills. Still no line-count rule.
33
+ - `superskill miss` takes `--quote` and `--date`.
34
+ - Tests: 17 new; each guard broken on purpose and seen red (a stale-sha pass counting, pooled pairs
35
+ counting, an unknown pair counting, an open miss passing, a half-passing regression eval passing,
36
+ examples and provenance stamps flagged as history).
37
+
38
+ ## 0.5.0 (2026-10-06)
39
+
40
+ **Superskill needs a golden from a real run.** Before, any golden a person approved counted, so a
41
+ skill could reach the top level on examples its author invented. (Gary Sheng, on five invented
42
+ goldens written to lift Freedom's flagship skills: *"I'm just concerned about hallucination that
43
+ we're accepting just to get higher doctor ratings."*) **Breaking for anyone at superskill today:**
44
+ an approved golden without the new `PROVENANCE.json` no longer counts, and the doctor says which
45
+ golden and why.
46
+
47
+ - `goldens/<id>/PROVENANCE.json`: `source` (`real-run`, `synthetic`, `synthetic-reconstruction`),
48
+ the `run` it came from (session, commit or ledger id), and who `accepted` the output when it ran,
49
+ and when. Only `real-run` with all of that counts for `golden-approved`. An anonymized twin of a
50
+ private golden counts with `derived_from` and the anonymizer's `ANONYMIZED.json` receipt.
51
+ - **Private goldens.** `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`) on `doctor`,
52
+ `approve` and `doctor --run` also reads `<dir>/<skill-name>/<id>/`, for real runs too personal
53
+ to ship with the skill. `approve` writes the approval beside the golden it found.
54
+ - `init --from-session` writes a `real-run` provenance with `accepted` left empty, so the golden
55
+ counts only once someone records who accepted it.
56
+ - **`evals-real` (tested):** a `superskill init` `REPLACE:` placeholder, or the same prompt or
57
+ trigger query twice, is not a case. Measured in Freedom on 2026-10-05: the init scaffold copied up
58
+ to 3 evals and 10 triggers scored `tested` while testing nothing.
59
+ - **`doctor --run` is never recorded as a real use.** Every harness spawns its child with
60
+ `FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`. Measured 2026-10-05: three sandbox runs
61
+ landed in Freedom's skill ledger as perfect one-shot runs of a skill nobody had used.
62
+ - **`real-use` (info):** where the run ledger records what the person's next message made of each
63
+ run, the doctor prints how many of the last 30 days' judged runs they accepted, beside the level.
64
+ - `miss import --freedom-ledger` skips sandbox records, matches `freedom:<skill>` records, and
65
+ points a correction's "Should have" at the person's next message (session and time, never text).
66
+ It reads `FREEDOM_SKILL_LEDGER_HOME` when Freedom's ledger was re-pointed.
67
+ - Tests: 11 new; each guard broken on purpose and seen red (invented golden reaching superskill,
68
+ placeholders counting, duplicate prompts counting, a sandbox run becoming a miss, a twin with no
69
+ receipt counting, the ledger off switch missing from the run env, codex spawned without it).
70
+
3
71
  ## 0.4.0 (2026-09-29)
4
72
 
5
73
  **Approve from your phone.** `superskill approve` worked only at an interactive terminal, so an
package/README.md CHANGED
@@ -4,8 +4,9 @@ Score any agent skill folder as **skill**, **tested**, or **superskill**, and ge
4
4
  to-do list for the next level. Works on skills for Claude Code, Codex, or any harness that
5
5
  reads the [Agent Skills](https://agentskills.io) format. Zero dependencies, Node 20 or later.
6
6
 
7
- A superskill runs on frontier intelligence, is checked against examples you approved, and is
8
- fixed every time it gets something wrong. [What that means](https://supersuit.wiki/concepts/superskill);
7
+ A superskill is fixed every time it gets something wrong, with a test that keeps each fix fixed;
8
+ passes its whole eval suite on a current model; and does the job for real people without a
9
+ correction, often enough to count. [What that means](https://supersuit.wiki/concepts/superskill);
9
10
  [the standard](SPEC.md).
10
11
 
11
12
  ## 30 seconds
@@ -34,12 +35,17 @@ file error. `--json` prints one JSON document and nothing else.
34
35
 
35
36
  1. **skill**: a valid `SKILL.md` (name matches the folder, description of 1024 characters or
36
37
  fewer that says when to use it, hard rules above the compaction fold, headings in long bodies, nothing said twice), references one level deep, no
37
- hard-coded machine paths, nothing that reads like a prompt injection.
38
+ hard-coded machine paths, nothing that reads like a prompt injection. It warns on dated
39
+ incident stories in `SKILL.md`, which belong in `MISSES.md`.
38
40
  2. **tested**: at least three task evals with checks a machine can verify, and a trigger set of
39
41
  at least ten requests, some that should load the skill and some near-misses that should not.
40
- 3. **superskill**: a golden a person approved; no miss open longer than 14 days and every fixed
41
- miss guarded by an eval; a recent `--run` on file where the skill beats the same task done
42
- without it. "Recent" follows `metadata.cadence` (a weekly skill's proof lasts 30 days).
42
+ 3. **superskill**: no open miss, and every fixed miss guarded by a regression eval; a recent
43
+ `--run` against the `SKILL.md` that is there now, where the suite passes (90%), every
44
+ regression eval passes every run, the trigger set is 90% right, and the skill beats the same
45
+ task done without it; and a real-run record where one model+harness has at least 5 real runs
46
+ at 80% one-shot (no correction). Pairs are never pooled. "Recent" follows `metadata.cadence`
47
+ (a weekly skill's proof lasts 30 days). Goldens are optional evidence. Every threshold is a
48
+ flag (`--min-real-runs`, `--min-one-shot`, `--min-pass-rate`, `--min-trigger-rate`).
43
49
 
44
50
  Every rule and threshold is in [SPEC.md](SPEC.md).
45
51
 
@@ -52,7 +58,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
52
58
  | `superskill doctor <skill> --run [--harness claude\|codex] [--repeat 3] [--yes]` | Run the evals for real, with and without the skill, and write `evals/results/latest.json`. **The only command that spends model calls**; it prints an estimate and asks first. |
53
59
  | `superskill init <skill>` | Add missing `evals/`, `goldens/`, `MISSES.md`. Never overwrites. |
54
60
  | `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
55
- | `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
61
+ | `superskill miss <skill> "<what happened>" [--expected "..."] [--quote "..."] [--date YYYY-MM-DD]` | Log a time the skill got it wrong, or the story behind a rule it learned. |
56
62
  | `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
57
63
  | `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). At a terminal it asks for your name; from your phone, an agent records your tap with `--approved-by` and `--via`. Approvals accumulate. |
58
64
  | `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
@@ -70,9 +76,10 @@ my-skill/
70
76
  SKILL.md
71
77
  evals/evals.json task evals (Anthropic skill-creator format)
72
78
  evals/triggers.json should / should-not load (skill-creator format)
73
- goldens/<id>/ input.md, output.md, APPROVAL.json
74
- MISSES.md every miss, open or fixed with its eval
75
- evals/results/latest.json the last --run
79
+ goldens/<id>/ optional: input.md, expectations.json, output.md, PROVENANCE.json, APPROVAL.json
80
+ MISSES.md every miss and its story, open or fixed with its eval
81
+ evals/results/latest.json the last --run (which SKILL.md it proved, suite and triggers)
82
+ evals/real-runs.json real uses per model+harness, one-shot or not (or --real-runs <dir>)
76
83
  ```
77
84
 
78
85
  Harnesses ignore folders they do not know, so none of this changes how the skill loads.
@@ -84,13 +91,16 @@ Harnesses ignore folders they do not know, so none of this changes how the skill
84
91
  - Codex: the skill is linked at `.agents/skills/<name>`. Codex cannot switch skills off, so a
85
92
  copy installed in `~/.agents/skills` can leak into the baseline; move it aside while proving.
86
93
  - Machine checks (`contains:`, `regex:`, `file_exists:`) are free. Each plain-language expectation
87
- and each golden costs one grader call per run.
94
+ costs one grader call per run.
95
+ - Trigger evals run each query in `evals/triggers.json` with the skill installed and record
96
+ whether the harness loaded it (Claude Code: a `Skill` call or a read of its `SKILL.md`; Codex:
97
+ a read of its `SKILL.md`).
88
98
 
89
99
  ## Freedom
90
100
 
91
101
  Nothing here needs [Freedom](https://getfreedom.wiki). If a skill has Freedom's `HDSOP.md`, the
92
102
  doctor shows it as a bonus; `miss import --freedom-ledger` reads Freedom's run ledger as plain
93
- files. `superskill snippet` gives any agent the same habits with no Freedom installed.
103
+ files, and with no `evals/real-runs.json` the doctor reads the real-run record from it too. `superskill snippet` gives any agent the same habits with no Freedom installed.
94
104
 
95
105
  ## Releasing
96
106
 
package/SPEC.md CHANGED
@@ -1,13 +1,21 @@
1
1
  # The superskill standard
2
2
 
3
- **Version 0.4.0** (2026-09-29). The reference checker is `@supersuit/superskill`; where this
3
+ **Version 0.6.0** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
4
4
  document and the checker disagree, the checker has a bug.
5
5
 
6
- A **superskill** runs on frontier intelligence, is checked against examples a person approved,
7
- and is fixed every time it gets something wrong
6
+ A **superskill** runs on frontier intelligence, is fixed every time it gets something wrong, and
7
+ does the job for real people without a correction
8
8
  ([definition](https://supersuit.wiki/concepts/superskill)). This standard turns each clause into
9
9
  a file in the skill's own folder, so the evidence travels with the skill wherever it is copied.
10
10
 
11
+ **What a skill has absorbed is its embodiment, not one run of it (0.6.0).** Until 0.5.0 the top
12
+ level required a golden: one real run a person approved as the standard. A random accepted run
13
+ is not the ultimate embodiment of a skill; the misses it was fixed for, and the evals that keep
14
+ each fix fixed, are. This follows Anthropic's guidance on agent evals: build the suite from real
15
+ failures, hold regression evals near 100%, grade outcomes rather than paths, and use a reference
16
+ solution to prove a task is solvable, not as the answer to match. Goldens remain, as optional
17
+ evidence.
18
+
11
19
  ## Contents
12
20
 
13
21
  - [The clauses and their evidence](#the-clauses-and-their-evidence)
@@ -23,9 +31,10 @@ a file in the skill's own folder, so the evidence travels with the skill whereve
23
31
  | The clause | What proves it | Where it lives |
24
32
  |---|---|---|
25
33
  | A skill at all | A valid `SKILL.md` under the Agent Skills spec, plus the hygiene rules below | `SKILL.md` |
26
- | Checked against examples you approved | At least one golden: a real input, the output a person said was right, and a record of who approved it and when | `goldens/<id>/` |
27
- | Fixed every time it gets something wrong | A miss log where every miss is open (recently) or fixed with a regression eval that would catch it again | `MISSES.md` + `evals/evals.json` |
28
- | Runs on frontier intelligence | Its evals last passed on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
34
+ | Fixed every time it gets something wrong | A miss log where every miss is fixed, each with a regression eval that would catch it again, and the story of each fix | `MISSES.md` + `evals/evals.json` |
35
+ | Runs on frontier intelligence | Its whole suite and trigger set last passed against the current `SKILL.md`, on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
36
+ | Does the job for real | Enough real runs on one model+harness, enough of them one-shot (no correction) | `evals/real-runs.json` (or an operator's export, see below) |
37
+ | Optional: an example a person approved | A golden: a real input, the checklist a right answer meets, a reference output, where it came from | `goldens/<id>/` (or a private folder) |
29
38
 
30
39
  ```
31
40
  my-skill/
@@ -33,11 +42,14 @@ my-skill/
33
42
  references/ scripts/ ... (as the Agent Skills spec allows)
34
43
  evals/evals.json task evals (level: tested)
35
44
  evals/triggers.json trigger evals (level: tested)
36
- goldens/<id>/input.md approved examples (level: superskill)
45
+ MISSES.md miss log + history (level: superskill)
46
+ evals/results/latest.json last --run (level: superskill)
47
+ evals/real-runs.json real-run record (level: superskill)
48
+ goldens/<id>/input.md optional evidence
49
+ goldens/<id>/expectations.json
37
50
  goldens/<id>/output.md
51
+ goldens/<id>/PROVENANCE.json
38
52
  goldens/<id>/APPROVAL.json
39
- MISSES.md miss log (level: superskill)
40
- evals/results/latest.json last --run (level: superskill)
41
53
  ```
42
54
 
43
55
  ## Compatibility
@@ -77,6 +89,7 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
77
89
  | `body-size` | info | reports lines and estimated tokens (characters / 4) once the body passes about 5000 tokens. Length alone is never a defect. |
78
90
  | `rules-above-the-fold` | warn | in a body past about 5000 tokens, every hard rule (a shouted NEVER, ALWAYS, MUST, DO NOT, REFUSE, or a bolded **Never ...** command, outside code fences) appears in the first 5000 tokens, either there or restated there. After compaction Claude Code keeps only that much of each invoked skill. |
79
91
  | `reference-says-when` | warn | every link from SKILL.md to a markdown file sits on a line that says when to read it (before, when, if, for, read ...). Step files (`steps/<step>.md`) are the recommended way to keep a long skill's detail out of the always-loaded body: they are read fresh when the step comes up, so compaction does not lose them. |
92
+ | `history-in-skill` | warn | no dated incident story in the body outside code: "Earned 2026-09-08", "(Gary, 2026-09-16: ...)", "Wilson, 2026-09-05: \*"..."\*", "on 2026-09-13 a session ...", "until 2026-09-21 the flag ...", "measured 2026-09-20", "(2026-09-30, #324)", "- 2026-09-15 (#147): ...". Matched per paragraph, since a line wrap can split a name from its date. Not flagged: a line with "e.g." or "example", a date in a code span or fence, a heading, an HTML comment (a generator's provenance stamp), "as of <date>". The story goes in `MISSES.md` (below); `SKILL.md` keeps the rule and a one-line why. There is deliberately no line-count rule (retired 2026-09-28): length alone is never a defect. |
80
93
  | `navigable` | warn | a body over 300 lines has no run of more than 150 lines without a heading |
81
94
  | `no-repeated-paragraphs` | warn | no paragraph of 100+ characters appears twice |
82
95
  | `references-one-deep` | fail | a markdown file linked from `SKILL.md` links on to no further local file |
@@ -91,18 +104,27 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
91
104
  |---|---|---|
92
105
  | `evals-present` | fail | `evals/evals.json` parses and there are at least 3 cases (each golden with an input and an output counts as one) |
93
106
  | `evals-verifiable` | fail | every case in `evals.json` has a prompt and at least one expectation |
107
+ | `evals-real` | fail | no eval case and no trigger still holds a `superskill init` `REPLACE:` placeholder, no two cases share a prompt, and no two triggers share a query |
94
108
  | `triggers-present` | fail | `evals/triggers.json` has at least 10 queries, at least 3 that should load the skill and at least 3 near-misses that should not |
95
109
 
96
110
  ### Level 3: superskill
97
111
 
98
112
  | Rule | Severity | Threshold |
99
113
  |---|---|---|
100
- | `golden-approved` | fail | at least one golden has an approval with non-empty `approved_by` and a valid `approved_at`; info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
101
114
  | `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
102
- | `no-stale-open-miss` | fail | no miss has been open more than 14 days |
115
+ | `misses-closed` | fail | no miss is open. A miss is a known way the skill fails; a skill that still fails a known way is not at the top level, however recent the miss |
103
116
  | `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
117
+ | `evals-pass` | fail | the last `--run` was against the current `SKILL.md` (`skill_sha` equals the sha256 of `SKILL.md`); its suite passed at least the minimum share of runs with the skill (default 90%); every fixed miss's regression eval passed every run; and it ran the trigger evals and got at least the minimum share right (default 90%) |
104
118
  | `run-evidence` | fail | `evals/results/latest.json` exists and `with_skill.pass_rate` > `without_skill.pass_rate` |
105
119
  | `run-fresh` | fail / warn | the last run is younger than the cadence window: `daily` or `weekly` 30 days, `monthly` 60, `quarterly` 120, `yearly` 365, none declared 30 (info). A `yearly` skill always warns to `--run` before its next real use. An unknown cadence warns |
120
+ | `real-runs` | fail | a real-run record (below) in which at least one model+harness pair, on its own, has at least the minimum runs (default 5) and one-shot share (default 80%). Each pair is reported as info; **pairs are never pooled**, since a clean record on one model in one harness proves nothing about another. A pair whose model or harness is `unknown` is reported and never counts |
121
+ | `golden-approved` | info / warn | never fails. Reports each golden that is approved and from a real run (with its approval weight), names approved goldens that are not from a real run and why, names a golden with no `expectations.json`, and warns on an `APPROVAL.json` or `PROVENANCE.json` that does not parse |
122
+
123
+ **The thresholds are configurable, never by the skill itself.** A skill cannot lower its own bar,
124
+ so they are not frontmatter. `doctor` takes `--min-real-runs <n>`, `--min-one-shot <0..1>`,
125
+ `--min-pass-rate <0..1>` and `--min-trigger-rate <0..1>`, or the environment variables
126
+ `SUPERSKILL_MIN_REAL_RUNS`, `SUPERSKILL_MIN_ONE_SHOT`, `SUPERSKILL_MIN_PASS_RATE` and
127
+ `SUPERSKILL_MIN_TRIGGER_RATE`. A report scored with non-default thresholds should say so.
106
128
 
107
129
  `metadata.cadence` in `SKILL.md` frontmatter declares how often the skill really runs:
108
130
 
@@ -152,11 +174,18 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
152
174
  ]
153
175
  ```
154
176
 
155
- ### `goldens/<id>/`
177
+ ### `goldens/<id>/` (optional evidence)
178
+
179
+ **A golden is an eval whose grading is a checklist of expected behavior (0.6.0).** It is never
180
+ required for any level. Grade the outcome, not the path: the checklist says what a right answer
181
+ does, and the reference output proves the task is solvable and calibrates the grader. A golden
182
+ that exists is still held to the provenance rules below, because "a person accepted this when
183
+ it ran" is a claim that has to be true.
156
184
 
157
185
  - `input.md`: the real request.
158
- - `output.md` (or any other file that is not `input.*` or `APPROVAL.json`): the output a person
159
- said was right.
186
+ - `expectations.json`: an array of expectations (the same forms as `evals.json`), graded by
187
+ `--run`. Without it, `--run` grades by likeness to the reference output.
188
+ - `output.md` (or any other file that is not `input.*` or a metadata file): the reference output.
160
189
  - `APPROVAL.json`, written only by `superskill approve`: at an interactive terminal, or relayed by an agent with `--approved-by` and `--via` after the person approved with a tap (the `via` field records where):
161
190
 
162
191
  ```json
@@ -181,6 +210,35 @@ approval whose `note` is its rationale.
181
210
  A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
182
211
  output.
183
212
 
213
+ **`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run is
214
+ reported as real evidence. An invented input with an invented output puts "a person said this
215
+ was right" on something no person's work produced. Invented cases belong in `evals/evals.json`.
216
+
217
+ ```json
218
+ {
219
+ "source": "real-run",
220
+ "run": { "session": "88696e1b-...", "ledger_id": "inv_2026-09-28T02-28-29Z_ab12", "commit": null, "at": "2026-09-28T02:28:29Z" },
221
+ "accepted": { "by": "Ann Example", "at": "2026-09-28T03:22:51Z", "signal": "close", "evidence": "her next message after the run" }
222
+ }
223
+ ```
224
+
225
+ - `source`: `real-run`, `synthetic`, or `synthetic-reconstruction`. Only `real-run` counts.
226
+ - `run`: at least one of `session`, `commit`, `ledger_id` names the run it came from.
227
+ - `accepted`: who accepted the output **when it ran** (`by`), and when (`at`). This is not the
228
+ approval: accepting is what the person did at the time (moved on, closed, said go); approving
229
+ is saying afterwards that this example is the standard. Both are required.
230
+ - An **anonymized twin** of a private golden says `anonymized: true` and `derived_from` (a hash of
231
+ the original, never its path or content) in place of the run, and carries `ANONYMIZED.json`, the
232
+ anonymizer's receipt (`checker`, `checked_at`, `counts` by kind, `fingerprint`, never the
233
+ mapping). Without the receipt the twin does not count.
234
+ - `superskill init --from-session` writes `source: real-run` with `accepted` empty, so the golden
235
+ is not reported as real until someone records who accepted it.
236
+
237
+ **Private goldens.** A real run's input and output are usually about real people and should not
238
+ travel with a skill that is shared. `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`)
239
+ makes `doctor`, `approve` and `doctor --run` also read `<dir>/<skill-name>/<id>/`, laid out exactly
240
+ like `goldens/<id>/`. A miss is closed only by an eval or golden that ships with the skill.
241
+
184
242
  ### `MISSES.md`
185
243
 
186
244
  ```markdown
@@ -191,11 +249,29 @@ output.
191
249
  - Should have: Merged duplicates into one line.
192
250
  - Fix: a1b2c3d
193
251
  - Eval: m1
252
+ - Quote: "why is Atlas in here twice" (optional)
194
253
  - Source: freedom-ledger inv_2026-09-20T10-00-00Z_ab12 (optional)
195
254
  ```
196
255
 
197
256
  Headings are `## <id> · <YYYY-MM-DD> · <open|fixed>`; `|` or `-` also separate. Ids are `m1`,
198
- `m2`, ... Other headings are ignored.
257
+ `m2`, ... Other headings are ignored. A field runs on over indented lines that follow it.
258
+
259
+ **`MISSES.md` is where a skill's history lives (0.6.0).** `SKILL.md` is instructions, loaded on
260
+ every run; the story of the incident that earned a rule is history, and Anthropic's authoring
261
+ guidance is to keep time-sensitive content out of instructions. So a story becomes a miss entry:
262
+
263
+ | Field | Holds |
264
+ |---|---|
265
+ | id, date | `m<N>`, and the date the incident happened (not the date it was written down) |
266
+ | What happened | the incident, in full: what the skill did, what it cost, what was noticed |
267
+ | Should have | what the skill should have done |
268
+ | Fix | the commit, or the rule now in `SKILL.md` that the incident earned |
269
+ | Eval | the regression eval that would catch it again. Empty until one exists, and the doctor says so (`fixed-miss-has-eval`): a story with no eval is a fix nothing guards yet |
270
+ | Quote | optional: what the person said, verbatim |
271
+ | Source | optional: where it was recorded (a ledger id, an issue) |
272
+
273
+ `SKILL.md` keeps the rule and a one-line why. `superskill miss <skill> "<what>" --date <d> --quote
274
+ "<q>"` writes one.
199
275
 
200
276
  ### `evals/results/latest.json`
201
277
 
@@ -206,15 +282,53 @@ Written only by `superskill doctor --run`:
206
282
  "run_at": "2026-09-20T10:00:00.000Z",
207
283
  "harness": "claude",
208
284
  "model": "<model id the harness reported>",
285
+ "skill_sha": "<sha256 of the SKILL.md the run proved>",
209
286
  "cases": 4,
210
287
  "repeat": 3,
211
288
  "with_skill": { "pass_rate": 1.0, "mean_ms": 21000, "mean_tokens": 4100 },
212
289
  "without_skill": { "pass_rate": 0.33, "mean_ms": 18000, "mean_tokens": 3900 },
213
- "per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }]
290
+ "per_case": [{ "id": "1", "with_skill": { "runs": 3, "passes": 3 }, "without_skill": { "runs": 3, "passes": 1 }, "failures": [] }],
291
+ "triggers": { "cases": 12, "runs": 36, "passes": 35, "pass_rate": 0.972,
292
+ "per_query": [{ "query": "write my weekly status", "should_trigger": true, "runs": 3, "passes": 3 }] }
214
293
  }
215
294
  ```
216
295
 
217
296
  A run passes when every expectation of its case passes; `pass_rate` is passing runs over runs.
297
+ A trigger run passes when the harness loaded the skill exactly when `should_trigger` says
298
+ (Claude Code: a `Skill` tool call naming it or a read of its `SKILL.md`; Codex: a read of its
299
+ `SKILL.md`). A golden with `expectations.json` is graded on that checklist; one without is graded
300
+ by likeness to its reference output.
301
+
302
+ ### `evals/real-runs.json` (the real-run record)
303
+
304
+ Harness-neutral: any harness, ledger or script may write it, and the doctor only reads it.
305
+
306
+ ```json
307
+ {
308
+ "format": "superskill-real-runs/1",
309
+ "skill": "weekly-status",
310
+ "generated_at": "2026-10-06T12:00:00Z",
311
+ "source": "freedom-skill-ledger",
312
+ "pairs": [
313
+ { "model": "claude-opus-5-5", "harness": "claude-code", "runs": 12, "one_shot": 11,
314
+ "first_at": "2026-09-01T09:00:00Z", "last_at": "2026-10-05T18:00:00Z" }
315
+ ]
316
+ }
317
+ ```
318
+
319
+ - A **run** is one real use of the skill by a person. Never a sandbox run (`doctor --run`), never
320
+ a test.
321
+ - It is **one-shot** when the person needed no correction, rescue or redirect, it did not fail or
322
+ get abandoned, and it was not corrected after it handed back. A taste note is the person's
323
+ preference, not the skill's defect, and does not break one-shot.
324
+ - `model` is the model id the harness ran (as the transcript or environment reports it);
325
+ `harness` is `claude-code`, `codex`, or another harness's name. A writer that cannot tell
326
+ writes `unknown`, never a guess. Counts only: the file carries no input, output or names, so it
327
+ can ship with the skill.
328
+
329
+ **Where the doctor reads it**, first found wins: `<dir>/<skill-name>.json` where `<dir>` is
330
+ `--real-runs` or `SUPERSKILL_REAL_RUNS` (an operator's own export, kept out of the skill); then
331
+ `evals/real-runs.json` in the skill; then Freedom's skill ledger, read live (below).
218
332
 
219
333
  ## Collections and plugins
220
334
 
@@ -246,7 +360,17 @@ exist they are read as plain files:
246
360
  `{id, skill, started, outcome, interventions: [{kind, note|what}], errors}`. A record with an
247
361
  intervention of kind `redirect`, `correction` or `rescue`, or with `outcome: "failed"`, becomes
248
362
  an open miss dated from `started`; `taste` interventions are skipped; the ledger id is kept on
249
- a `Source:` line so a record is never imported twice.
363
+ a `Source:` line so a record is never imported twice. A record with `synthetic: true` (a
364
+ sandbox run) is skipped. A record corrected after it handed back (`corrected_after`, with
365
+ `next_turn_ref`) gets a "Should have" that points at the person's correction: the session and
366
+ the time, never their words.
367
+ - **The real-run record.** Freedom's skill ledger records, on every run, the `model` id and the
368
+ `harness` (`claude-code`, `codex`, ...), read from the session transcript or the environment and
369
+ `unknown` when neither says. Freedom exports it in the `superskill-real-runs/1` format above,
370
+ one file per skill, for `--real-runs`. With no export and no `evals/real-runs.json`, the doctor
371
+ reads the same ledger files live and counts them the same way (sandbox runs skipped).
372
+ - **`doctor --run` is never recorded as a use.** Every harness spawns its child with
373
+ `FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`, in a folder named `superskill-run-*`.
250
374
 
251
375
  ## Versioning
252
376
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@supersuit/superskill",
3
- "version": "0.4.0",
3
+ "version": "0.6.0",
4
4
  "description": "Score any agent skill folder as skill, tested, or superskill. An open standard and a zero-dependency CLI.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -5,8 +5,9 @@ import { createInterface } from "node:readline/promises";
5
5
  import { parseArgs, clock, UsageError } from "../args.mjs";
6
6
  import { skillDir, refuse } from "./common.mjs";
7
7
  import { readGoldens, approvalEntry, withApproval, BASES } from "../goldens.mjs";
8
+ import { parseSkillFile } from "../frontmatter.mjs";
8
9
 
9
- export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"]
10
+ export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"] [--private-goldens <dir>]
10
11
 
11
12
  Record that a person checked goldens/<id>/ and signs off on its output. At a terminal it asks
12
13
  for your name. From an agent (a phone tap on a board or review page), pass --approved-by
@@ -18,6 +19,11 @@ Every approval says WHY (--rationale, or asked) and WHAT IT RESTS ON (--basis, o
18
19
  outcome it produced a result someone can check; --evidence is required
19
20
  Approvals accumulate: two people approving, or a judgment approval later backed by an outcome,
20
21
  all stay on the record. Writes goldens/<id>/APPROVAL.json with the SKILL.md hash.
22
+
23
+ --private-goldens <dir> (or SUPERSKILL_PRIVATE_GOLDENS) also looks for the golden in
24
+ <dir>/<skill-name>/<id>/, where an operator keeps goldens made from real runs that are too
25
+ personal to travel with the skill. Approving records the approval; it never changes where a
26
+ golden came from (PROVENANCE.json), and only a golden from a real run counts for superskill.
21
27
  `;
22
28
 
23
29
  export async function run(argv) {
@@ -26,8 +32,12 @@ export async function run(argv) {
26
32
  const dir = skillDir(a._[0], "approve");
27
33
  const id = a._[1];
28
34
  if (!id) throw new UsageError("approve needs a golden id");
29
- const g = readGoldens(dir).find((x) => x.id === id);
30
- if (!g) return refuse(`no goldens/${id}/`);
35
+ const name = parseSkillFile(readFileSync(join(dir, "SKILL.md"), "utf8")).data.name || "";
36
+ const privateGoldens = a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null;
37
+ const found = readGoldens(dir, { privateGoldens, name }).filter((x) => x.id === id);
38
+ // The private copy wins when both exist: it is the one the operator was shown.
39
+ const g = found.find((x) => x.private) || found[0];
40
+ if (!g) return refuse(`no goldens/${id}/${privateGoldens ? ` (nor ${join(privateGoldens, name, id)})` : ""}`);
31
41
  if (g.input === null || g.output === null || !g.output.trim()) return refuse(`goldens/${id}/ needs an input file and a non-empty output file before it can be approved`);
32
42
  if (!(process.stdin.isTTY && process.stdout.isTTY)) {
33
43
  // Mobile first: a person approves with a tap (a board, a review page) and an agent records it.
@@ -39,7 +49,7 @@ export async function run(argv) {
39
49
  const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
40
50
  const made = approvalEntry({ name, rationale: a.flags.rationale || a.flags.note, basis: a.flags.basis || "judgment", evidence: a.flags.evidence, at: clock(a.flags).toISOString(), sha });
41
51
  if (made.error) return refuse(made.error);
42
- const p = join(dir, "goldens", id, "APPROVAL.json");
52
+ const p = join(g.dir, "APPROVAL.json");
43
53
  const existed = existsSync(p);
44
54
  writeFileSync(p, JSON.stringify(withApproval(existed ? g.approval : null, { ...made.entry, via }), null, 2) + "\n");
45
55
  process.stdout.write(`${existed ? "added an approval to" : "approved"} goldens/${id}/ by ${name} (${made.entry.basis}, via ${via})\n`);
@@ -57,7 +67,7 @@ export async function run(argv) {
57
67
  const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
58
68
  const made = approvalEntry({ name, rationale, basis, evidence, at: clock(a.flags).toISOString(), sha });
59
69
  if (made.error) return refuse(made.error);
60
- const p = join(dir, "goldens", id, "APPROVAL.json");
70
+ const p = join(g.dir, "APPROVAL.json");
61
71
  const existed = existsSync(p);
62
72
  const prior = existed ? g.approval : null;
63
73
  writeFileSync(p, JSON.stringify(withApproval(prior, made.entry), null, 2) + "\n");
@@ -8,8 +8,10 @@ import { changedFiles, touchedSkills, levelDrops } from "../changed.mjs";
8
8
  export const help = `superskill doctor <path...> [options]
9
9
 
10
10
  Score one skill, a folder of skills, or a plugin (skills/*/SKILL.md).
11
- Levels: skill (spec-valid, hygienic) < tested (evals + triggers) < superskill
12
- (approved golden, misses fixed with evals, a fresh --run that beats no-skill).
11
+ Levels: skill (spec-valid, hygienic) < tested (real evals + triggers) < superskill
12
+ (every miss fixed with a regression eval; the suite and triggers pass against this SKILL.md
13
+ in a fresh --run that beats no-skill; and a real-run record: at least 5 real runs, 80%
14
+ one-shot, on one model+harness). Goldens are optional evidence.
13
15
 
14
16
  Options:
15
17
  --level <skill|tested|superskill> target level for the exit code (default skill)
@@ -20,6 +22,15 @@ Options:
20
22
  --baseline-json <file> a previous --json; exit 1 if any skill's level dropped
21
23
  --run run the evals with and without the skill (costs
22
24
  model calls; see "superskill doctor --run --help")
25
+ --private-goldens <dir> also read goldens from <dir>/<skill-name>/<id>/ (or set
26
+ SUPERSKILL_PRIVATE_GOLDENS): real runs kept out of the skill
27
+ --real-runs <dir> read the real-run record from <dir>/<skill-name>.json (or
28
+ set SUPERSKILL_REAL_RUNS); else evals/real-runs.json, else
29
+ Freedom's skill ledger
30
+ --min-real-runs <n> real runs one model+harness needs (default 5)
31
+ --min-one-shot <0..1> one-shot share that pair needs (default 0.8)
32
+ --min-pass-rate <0..1> suite pass rate the last --run needs (default 0.9)
33
+ --min-trigger-rate <0..1> trigger evals right in the last --run (default 0.9)
23
34
  --now <iso date> evaluate dates as of this moment
24
35
  --help this text
25
36
 
@@ -34,7 +45,7 @@ export async function run(argv) {
34
45
  const level = a.flags.level || "skill";
35
46
  if (!LEVELS.includes(level)) throw new UsageError(`--level must be one of ${LEVELS.join(", ")}`);
36
47
  if (!a._.length) throw new UsageError("doctor needs a path");
37
- const opts = { level, now: clock(a.flags) };
48
+ const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null, ...thresholds(a.flags) };
38
49
  if (a.flags.changed) {
39
50
  const all = a._.flatMap((p) => findSkills(p));
40
51
  const files = a._.flatMap((p) => changedFiles(p, a.flags.base));
@@ -61,3 +72,21 @@ export async function run(argv) {
61
72
  }
62
73
  return result.ok ? 0 : 1;
63
74
  }
75
+
76
+ /** The superskill bar's thresholds, from flags then environment. Unset means the documented default. */
77
+ export function thresholds(flags, env = process.env) {
78
+ const pick = (flag, envName, lo, hi) => {
79
+ const raw = flags[flag] ?? env[envName];
80
+ if (raw === undefined || raw === "") return undefined;
81
+ const n = Number(raw);
82
+ if (!Number.isFinite(n) || n < lo || (hi !== null && n > hi)) throw new UsageError(`--${flag} must be a number${hi === null ? ` of at least ${lo}` : ` from ${lo} to ${hi}`}`);
83
+ return n;
84
+ };
85
+ return {
86
+ realRuns: flags["real-runs"] || env.SUPERSKILL_REAL_RUNS || null,
87
+ minRealRuns: pick("min-real-runs", "SUPERSKILL_MIN_REAL_RUNS", 1, null),
88
+ minOneShot: pick("min-one-shot", "SUPERSKILL_MIN_ONE_SHOT", 0, 1),
89
+ minPassRate: pick("min-pass-rate", "SUPERSKILL_MIN_PASS_RATE", 0, 1),
90
+ minTriggerRate: pick("min-trigger-rate", "SUPERSKILL_MIN_TRIGGER_RATE", 0, 1),
91
+ };
92
+ }
@@ -1,5 +1,5 @@
1
1
  import { existsSync, mkdirSync, writeFileSync, readFileSync } from "node:fs";
2
- import { join } from "node:path";
2
+ import { join, basename } from "node:path";
3
3
  import { parseArgs, UsageError } from "../args.mjs";
4
4
  import { skillDir } from "./common.mjs";
5
5
  import { MISSES_HEADER } from "../misses.mjs";
@@ -15,7 +15,9 @@ Never overwrites a file that exists.
15
15
  --from-session <file> build the first eval and golden candidate from the session where
16
16
  the job was done by hand. A Claude Code .jsonl gives the first
17
17
  request and the final answer; any other file is the request.
18
- The golden waits for a person: run \`superskill approve\`.
18
+ The golden waits for a person: fill in accepted.by and
19
+ accepted.at in its PROVENANCE.json (who accepted the
20
+ output when it ran), then \`superskill approve\`.
19
21
  `;
20
22
 
21
23
  const TRIGGER_EXAMPLE = [
@@ -65,6 +67,10 @@ export async function run(argv) {
65
67
  writeFileSync(evalsPath, JSON.stringify(doc, null, 2) + "\n");
66
68
  put(`goldens/${id}/input.md`, sessionCase.prompt + "\n");
67
69
  put(`goldens/${id}/output.md`, sessionCase.output ? sessionCase.output + "\n" : "");
70
+ // A real run, but nobody has said yet that its output was accepted. Until accepted.by and
71
+ // accepted.at are filled in (by the person, or by a tool that read their next message), this
72
+ // golden cannot count toward superskill however it is approved.
73
+ put(`goldens/${id}/PROVENANCE.json`, JSON.stringify({ source: "real-run", run: { session: basename(session), ledger_id: null, commit: null, at: null }, accepted: { by: null, at: null, signal: null, evidence: "" } }, null, 2) + "\n");
68
74
  process.stdout.write(`eval ${id} and golden candidate goldens/${id}/ written from ${session}\n`);
69
75
  process.stdout.write(`a person approves it with: superskill approve ${a._[0]} ${id}\n`);
70
76
  }
@@ -2,11 +2,12 @@ import { parseArgs, clock, UsageError } from "../args.mjs";
2
2
  import { skillDir, today } from "./common.mjs";
3
3
  import { readMisses, appendMisses, nextMissId } from "../misses.mjs";
4
4
 
5
- export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"]
5
+ export const help = `superskill miss <skill> "<what happened>" [--expected "<what should have>"] [--quote "<what they said>"] [--date YYYY-MM-DD]
6
6
  superskill miss import <skill> --freedom-ledger [--ledger <file>]
7
7
 
8
- Log a time the skill got something wrong. The entry opens today; an open miss older than
9
- 14 days blocks the superskill level. Close it with \`superskill fix\`.
8
+ Log a time the skill got something wrong. The entry opens today (or on --date, for a story
9
+ from the past); an open miss blocks the superskill level. Close it with \`superskill fix\`.
10
+ This is also where the story behind a rule goes: SKILL.md keeps the rule and a one-line why.
10
11
  `;
11
12
 
12
13
  export async function run(argv) {
@@ -18,7 +19,9 @@ export async function run(argv) {
18
19
  if (!what) throw new UsageError('miss needs a description: superskill miss <skill> "<what happened>"');
19
20
  const misses = readMisses(dir) || [];
20
21
  const id = nextMissId(misses);
21
- appendMisses(dir, [{ id, date: today(clock(a.flags)), status: "open", what, expected: a.flags.expected || "" }]);
22
+ const date = a.flags.date || today(clock(a.flags));
23
+ if (!/^\d{4}-\d{2}-\d{2}$/.test(date)) throw new UsageError("--date must be YYYY-MM-DD");
24
+ appendMisses(dir, [{ id, date, status: "open", what, expected: a.flags.expected || "", quote: a.flags.quote || "" }]);
22
25
  process.stdout.write(`${id} logged (open). Close it with: superskill fix ${a._[0]} ${id} --eval <id>\n`);
23
26
  return 0;
24
27
  }
package/src/doctor.mjs CHANGED
@@ -37,6 +37,7 @@ export function findSkills(path) {
37
37
  /** Score one loaded skill. */
38
38
  export function scoreSkill(dir, opts = {}) {
39
39
  const ctx = loadSkill(dir);
40
+ if (opts.privateGoldens) ctx.privateGoldens = opts.privateGoldens;
40
41
  const findings = runRules(ctx, opts.rules || allRules, opts);
41
42
  const level = computeLevel(findings);
42
43
  const up = nextLevel(level);