@supersuit/superskill 0.3.1 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,52 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.5.0 (2026-10-06)
4
+
5
+ **Superskill needs a golden from a real run.** Before, any golden a person approved counted, so a
6
+ skill could reach the top level on examples its author invented. (Gary Sheng, on five invented
7
+ goldens written to lift Freedom's flagship skills: *"I'm just concerned about hallucination that
8
+ we're accepting just to get higher doctor ratings."*) **Breaking for anyone at superskill today:**
9
+ an approved golden without the new `PROVENANCE.json` no longer counts, and the doctor says which
10
+ golden and why.
11
+
12
+ - `goldens/<id>/PROVENANCE.json`: `source` (`real-run`, `synthetic`, `synthetic-reconstruction`),
13
+ the `run` it came from (session, commit or ledger id), and who `accepted` the output when it ran,
14
+ and when. Only `real-run` with all of that counts for `golden-approved`. An anonymized twin of a
15
+ private golden counts with `derived_from` and the anonymizer's `ANONYMIZED.json` receipt.
16
+ - **Private goldens.** `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`) on `doctor`,
17
+ `approve` and `doctor --run` also reads `<dir>/<skill-name>/<id>/`, for real runs too personal
18
+ to ship with the skill. `approve` writes the approval beside the golden it found.
19
+ - `init --from-session` writes a `real-run` provenance with `accepted` left empty, so the golden
20
+ counts only once someone records who accepted it.
21
+ - **`evals-real` (tested):** a `superskill init` `REPLACE:` placeholder, or the same prompt or
22
+ trigger query twice, is not a case. Measured in Freedom on 2026-10-05: the init scaffold copied up
23
+ to 3 evals and 10 triggers scored `tested` while testing nothing.
24
+ - **`doctor --run` is never recorded as a real use.** Every harness spawns its child with
25
+ `FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`. Measured 2026-10-05: three sandbox runs
26
+ landed in Freedom's skill ledger as perfect one-shot runs of a skill nobody had used.
27
+ - **`real-use` (info):** where the run ledger records what the person's next message made of each
28
+ run, the doctor prints how many of the last 30 days' judged runs they accepted, beside the level.
29
+ - `miss import --freedom-ledger` skips sandbox records, matches `freedom:<skill>` records, and
30
+ points a correction's "Should have" at the person's next message (session and time, never text).
31
+ It reads `FREEDOM_SKILL_LEDGER_HOME` when Freedom's ledger was re-pointed.
32
+ - Tests: 11 new; each guard broken on purpose and seen red (invented golden reaching superskill,
33
+ placeholders counting, duplicate prompts counting, a sandbox run becoming a miss, a twin with no
34
+ receipt counting, the ledger off switch missing from the run env, codex spawned without it).
35
+
36
+ ## 0.4.0 (2026-09-29)
37
+
38
+ **Approve from your phone.** `superskill approve` worked only at an interactive terminal, so an
39
+ operator who reviews on a phone had to find a laptop to say yes. (Gary Sheng: *"The best way for
40
+ me to approve it is to click yes, approve, not you telling me to go to my terminal. I'm normally
41
+ mobile first."*) Away from a terminal it now accepts `--approved-by "<the person>"` and
42
+ `--via "<where they said yes>"`, after that person approved with a tap on a board or a review page.
43
+ Both are required and both are recorded (`via` on the approval), so an approval an agent relayed is
44
+ always distinguishable from one typed at a terminal, and an agent still cannot approve with no
45
+ named person and no channel. A rationale is required either way.
46
+
47
+ - Tests: refused with no name, refused with a name and no channel, refused with no rationale;
48
+ recorded with person, channel and basis. The channel requirement was mutated out and went red.
49
+
3
50
  ## 0.3.1 (2026-09-29)
4
51
 
5
52
  The 0.3.0 changes below, released. The v0.3.0 tag's publish refused on a red test (SPEC.md still
package/README.md CHANGED
@@ -37,7 +37,7 @@ file error. `--json` prints one JSON document and nothing else.
37
37
  hard-coded machine paths, nothing that reads like a prompt injection.
38
38
  2. **tested**: at least three task evals with checks a machine can verify, and a trigger set of
39
39
  at least ten requests, some that should load the skill and some near-misses that should not.
40
- 3. **superskill**: a golden a person approved; no miss open longer than 14 days and every fixed
40
+ 3. **superskill**: a golden from a real run, accepted when it ran and approved by a person; no miss open longer than 14 days and every fixed
41
41
  miss guarded by an eval; a recent `--run` on file where the skill beats the same task done
42
42
  without it. "Recent" follows `metadata.cadence` (a weekly skill's proof lasts 30 days).
43
43
 
@@ -54,7 +54,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
54
54
  | `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
55
55
  | `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
56
56
  | `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
57
- | `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). Terminal only, asks for your name, so an agent cannot approve its own output. Approvals accumulate. |
57
+ | `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). At a terminal it asks for your name; from your phone, an agent records your tap with `--approved-by` and `--via`. Approvals accumulate. |
58
58
  | `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
59
59
  | `superskill miss import <skill> --freedom-ledger [--ledger <file>]` | Import runs that needed correcting from Freedom's run ledger. |
60
60
  | `superskill snippet` | Print a block for `AGENTS.md` / `CLAUDE.md` that teaches any agent these habits. |
@@ -70,7 +70,7 @@ my-skill/
70
70
  SKILL.md
71
71
  evals/evals.json task evals (Anthropic skill-creator format)
72
72
  evals/triggers.json should / should-not load (skill-creator format)
73
- goldens/<id>/ input.md, output.md, APPROVAL.json
73
+ goldens/<id>/ input.md, output.md, PROVENANCE.json, APPROVAL.json
74
74
  MISSES.md every miss, open or fixed with its eval
75
75
  evals/results/latest.json the last --run
76
76
  ```
package/SPEC.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # The superskill standard
2
2
 
3
- **Version 0.3.1** (2026-09-29). The reference checker is `@supersuit/superskill`; where this
3
+ **Version 0.5.0** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
4
4
  document and the checker disagree, the checker has a bug.
5
5
 
6
6
  A **superskill** runs on frontier intelligence, is checked against examples a person approved,
@@ -23,7 +23,7 @@ a file in the skill's own folder, so the evidence travels with the skill whereve
23
23
  | The clause | What proves it | Where it lives |
24
24
  |---|---|---|
25
25
  | A skill at all | A valid `SKILL.md` under the Agent Skills spec, plus the hygiene rules below | `SKILL.md` |
26
- | Checked against examples you approved | At least one golden: a real input, the output a person said was right, and a record of who approved it and when | `goldens/<id>/` |
26
+ | Checked against examples you approved | At least one golden from a real run: the real input, the output the person accepted when it ran, where it came from, and a record of who approved it as a golden and when | `goldens/<id>/` (or a private folder, see below) |
27
27
  | Fixed every time it gets something wrong | A miss log where every miss is open (recently) or fixed with a regression eval that would catch it again | `MISSES.md` + `evals/evals.json` |
28
28
  | Runs on frontier intelligence | Its evals last passed on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
29
29
 
@@ -35,6 +35,7 @@ my-skill/
35
35
  evals/triggers.json trigger evals (level: tested)
36
36
  goldens/<id>/input.md approved examples (level: superskill)
37
37
  goldens/<id>/output.md
38
+ goldens/<id>/PROVENANCE.json
38
39
  goldens/<id>/APPROVAL.json
39
40
  MISSES.md miss log (level: superskill)
40
41
  evals/results/latest.json last --run (level: superskill)
@@ -91,13 +92,15 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
91
92
  |---|---|---|
92
93
  | `evals-present` | fail | `evals/evals.json` parses and there are at least 3 cases (each golden with an input and an output counts as one) |
93
94
  | `evals-verifiable` | fail | every case in `evals.json` has a prompt and at least one expectation |
95
+ | `evals-real` | fail | no eval case and no trigger still holds a `superskill init` `REPLACE:` placeholder, no two cases share a prompt, and no two triggers share a query |
94
96
  | `triggers-present` | fail | `evals/triggers.json` has at least 10 queries, at least 3 that should load the skill and at least 3 near-misses that should not |
95
97
 
96
98
  ### Level 3: superskill
97
99
 
98
100
  | Rule | Severity | Threshold |
99
101
  |---|---|---|
100
- | `golden-approved` | fail | at least one golden has an approval with non-empty `approved_by` and a valid `approved_at`; info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
102
+ | `golden-approved` | fail | at least one golden **from a real run** (its `PROVENANCE.json` says `source: real-run`, names the run, and names who accepted it and when; an anonymized twin also carries `ANONYMIZED.json`) has an approval with non-empty `approved_by` and a valid `approved_at`. An approved golden without that provenance is named and does not count. Info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
103
+ | `real-use` | info | when the run ledger records what the person's next message made of each run (Freedom's `next_turn`), the share of those runs in the last 30 days they accepted; sandbox runs are never counted |
101
104
  | `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
102
105
  | `no-stale-open-miss` | fail | no miss has been open more than 14 days |
103
106
  | `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
@@ -157,7 +160,7 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
157
160
  - `input.md`: the real request.
158
161
  - `output.md` (or any other file that is not `input.*` or `APPROVAL.json`): the output a person
159
162
  said was right.
160
- - `APPROVAL.json`, written only by `superskill approve` at an interactive terminal:
163
+ - `APPROVAL.json`, written only by `superskill approve`: at an interactive terminal, or relayed by an agent with `--approved-by` and `--via` after the person approved with a tap (the `via` field records where):
161
164
 
162
165
  ```json
163
166
  { "approvals": [
@@ -181,6 +184,38 @@ approval whose `note` is its rationale.
181
184
  A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
182
185
  output.
183
186
 
187
+ **`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run
188
+ counts toward superskill. An invented input with an invented output puts "a person said this was
189
+ right" on something no person's work produced, and a skill could then reach the top level on its
190
+ author's fiction. Invented cases still belong in `evals/evals.json`, where they hold a skill at
191
+ `tested`.
192
+
193
+ ```json
194
+ {
195
+ "source": "real-run",
196
+ "run": { "session": "88696e1b-...", "ledger_id": "inv_2026-09-28T02-28-29Z_ab12", "commit": null, "at": "2026-09-28T02:28:29Z" },
197
+ "accepted": { "by": "Ann Example", "at": "2026-09-28T03:22:51Z", "signal": "close", "evidence": "her next message after the run" }
198
+ }
199
+ ```
200
+
201
+ - `source`: `real-run`, `synthetic`, or `synthetic-reconstruction`. Only `real-run` counts.
202
+ - `run`: at least one of `session`, `commit`, `ledger_id` names the run it came from.
203
+ - `accepted`: who accepted the output **when it ran** (`by`), and when (`at`). This is not the
204
+ approval: accepting is what the person did at the time (moved on, closed, said go); approving
205
+ is saying afterwards that this example is the standard. Both are required.
206
+ - An **anonymized twin** of a private golden says `anonymized: true` and `derived_from` (a hash of
207
+ the original, never its path or content) in place of the run, and carries `ANONYMIZED.json`, the
208
+ anonymizer's receipt (`checker`, `checked_at`, `counts` by kind, `fingerprint`, never the
209
+ mapping). Without the receipt the twin does not count.
210
+ - `superskill init --from-session` writes `source: real-run` with `accepted` empty, so the golden
211
+ cannot count until someone records who accepted it.
212
+
213
+ **Private goldens.** A real run's input and output are usually about real people and should not
214
+ travel with a skill that is shared. `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`)
215
+ makes `doctor`, `approve` and `doctor --run` also read `<dir>/<skill-name>/<id>/`, laid out exactly
216
+ like `goldens/<id>/`. A private golden counts for the operator who holds it. A miss is closed only
217
+ by an eval or golden that ships with the skill.
218
+
184
219
  ### `MISSES.md`
185
220
 
186
221
  ```markdown
@@ -246,7 +281,12 @@ exist they are read as plain files:
246
281
  `{id, skill, started, outcome, interventions: [{kind, note|what}], errors}`. A record with an
247
282
  intervention of kind `redirect`, `correction` or `rescue`, or with `outcome: "failed"`, becomes
248
283
  an open miss dated from `started`; `taste` interventions are skipped; the ledger id is kept on
249
- a `Source:` line so a record is never imported twice.
284
+ a `Source:` line so a record is never imported twice. A record with `synthetic: true` (a
285
+ sandbox run) is skipped. A record corrected after it handed back (`corrected_after`, with
286
+ `next_turn_ref`) gets a "Should have" that points at the person's correction: the session and
287
+ the time, never their words.
288
+ - **`doctor --run` is never recorded as a use.** Every harness spawns its child with
289
+ `FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`, in a folder named `superskill-run-*`.
250
290
 
251
291
  ## Versioning
252
292
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@supersuit/superskill",
3
- "version": "0.3.1",
3
+ "version": "0.5.0",
4
4
  "description": "Score any agent skill folder as skill, tested, or superskill. An open standard and a zero-dependency CLI.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -5,17 +5,25 @@ import { createInterface } from "node:readline/promises";
5
5
  import { parseArgs, clock, UsageError } from "../args.mjs";
6
6
  import { skillDir, refuse } from "./common.mjs";
7
7
  import { readGoldens, approvalEntry, withApproval, BASES } from "../goldens.mjs";
8
+ import { parseSkillFile } from "../frontmatter.mjs";
8
9
 
9
- export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"]
10
+ export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"] [--private-goldens <dir>]
10
11
 
11
- Record that a person checked goldens/<id>/ and signs off on its output. Works only at an
12
- interactive terminal and asks for your name, so an agent cannot approve its own output.
12
+ Record that a person checked goldens/<id>/ and signs off on its output. At a terminal it asks
13
+ for your name. From an agent (a phone tap on a board or review page), pass --approved-by
14
+ "<the person>" and --via "<where they said yes>": both are required and both are recorded, so an
15
+ approval an agent relayed is always distinguishable from one typed at a terminal.
13
16
 
14
17
  Every approval says WHY (--rationale, or asked) and WHAT IT RESTS ON (--basis, or asked):
15
18
  judgment you read it and it is right (the default)
16
19
  outcome it produced a result someone can check; --evidence is required
17
20
  Approvals accumulate: two people approving, or a judgment approval later backed by an outcome,
18
21
  all stay on the record. Writes goldens/<id>/APPROVAL.json with the SKILL.md hash.
22
+
23
+ --private-goldens <dir> (or SUPERSKILL_PRIVATE_GOLDENS) also looks for the golden in
24
+ <dir>/<skill-name>/<id>/, where an operator keeps goldens made from real runs that are too
25
+ personal to travel with the skill. Approving records the approval; it never changes where a
26
+ golden came from (PROVENANCE.json), and only a golden from a real run counts for superskill.
19
27
  `;
20
28
 
21
29
  export async function run(argv) {
@@ -24,10 +32,29 @@ export async function run(argv) {
24
32
  const dir = skillDir(a._[0], "approve");
25
33
  const id = a._[1];
26
34
  if (!id) throw new UsageError("approve needs a golden id");
27
- if (!(process.stdin.isTTY && process.stdout.isTTY)) return refuse("approval needs a person at a terminal. Run this yourself, not through an agent.");
28
- const g = readGoldens(dir).find((x) => x.id === id);
29
- if (!g) return refuse(`no goldens/${id}/`);
35
+ const name = parseSkillFile(readFileSync(join(dir, "SKILL.md"), "utf8")).data.name || "";
36
+ const privateGoldens = a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null;
37
+ const found = readGoldens(dir, { privateGoldens, name }).filter((x) => x.id === id);
38
+ // The private copy wins when both exist: it is the one the operator was shown.
39
+ const g = found.find((x) => x.private) || found[0];
40
+ if (!g) return refuse(`no goldens/${id}/${privateGoldens ? ` (nor ${join(privateGoldens, name, id)})` : ""}`);
30
41
  if (g.input === null || g.output === null || !g.output.trim()) return refuse(`goldens/${id}/ needs an input file and a non-empty output file before it can be approved`);
42
+ if (!(process.stdin.isTTY && process.stdout.isTTY)) {
43
+ // Mobile first: a person approves with a tap (a board, a review page) and an agent records it.
44
+ // The record names the person AND the channel, so an approval relayed by an agent is never
45
+ // mistaken for one typed at a terminal. Without both, an agent could approve its own output.
46
+ const name = String(a.flags["approved-by"] || "").trim();
47
+ const via = String(a.flags.via || "").trim();
48
+ if (!name || !via) return refuse("approval needs a person. At a terminal it asks for your name; from an agent, pass --approved-by \"<the person>\" and --via \"<where they approved: the tap, the page, the message>\", after that person said yes");
49
+ const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
50
+ const made = approvalEntry({ name, rationale: a.flags.rationale || a.flags.note, basis: a.flags.basis || "judgment", evidence: a.flags.evidence, at: clock(a.flags).toISOString(), sha });
51
+ if (made.error) return refuse(made.error);
52
+ const p = join(g.dir, "APPROVAL.json");
53
+ const existed = existsSync(p);
54
+ writeFileSync(p, JSON.stringify(withApproval(existed ? g.approval : null, { ...made.entry, via }), null, 2) + "\n");
55
+ process.stdout.write(`${existed ? "added an approval to" : "approved"} goldens/${id}/ by ${name} (${made.entry.basis}, via ${via})\n`);
56
+ return 0;
57
+ }
31
58
  const rl = createInterface({ input: process.stdin, output: process.stdout });
32
59
  try {
33
60
  process.stdout.write(`\n--- goldens/${id}/${g.outputFile} ---\n${g.output.slice(0, 2000)}${g.output.length > 2000 ? "\n[...]" : ""}\n---\n`);
@@ -40,7 +67,7 @@ export async function run(argv) {
40
67
  const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
41
68
  const made = approvalEntry({ name, rationale, basis, evidence, at: clock(a.flags).toISOString(), sha });
42
69
  if (made.error) return refuse(made.error);
43
- const p = join(dir, "goldens", id, "APPROVAL.json");
70
+ const p = join(g.dir, "APPROVAL.json");
44
71
  const existed = existsSync(p);
45
72
  const prior = existed ? g.approval : null;
46
73
  writeFileSync(p, JSON.stringify(withApproval(prior, made.entry), null, 2) + "\n");
@@ -8,8 +8,9 @@ import { changedFiles, touchedSkills, levelDrops } from "../changed.mjs";
8
8
  export const help = `superskill doctor <path...> [options]
9
9
 
10
10
  Score one skill, a folder of skills, or a plugin (skills/*/SKILL.md).
11
- Levels: skill (spec-valid, hygienic) < tested (evals + triggers) < superskill
12
- (approved golden, misses fixed with evals, a fresh --run that beats no-skill).
11
+ Levels: skill (spec-valid, hygienic) < tested (real evals + triggers) < superskill
12
+ (an approved golden from a real run a person accepted, misses fixed with evals, a fresh
13
+ --run that beats no-skill).
13
14
 
14
15
  Options:
15
16
  --level <skill|tested|superskill> target level for the exit code (default skill)
@@ -20,6 +21,8 @@ Options:
20
21
  --baseline-json <file> a previous --json; exit 1 if any skill's level dropped
21
22
  --run run the evals with and without the skill (costs
22
23
  model calls; see "superskill doctor --run --help")
24
+ --private-goldens <dir> also read goldens from <dir>/<skill-name>/<id>/ (or set
25
+ SUPERSKILL_PRIVATE_GOLDENS): real runs kept out of the skill
23
26
  --now <iso date> evaluate dates as of this moment
24
27
  --help this text
25
28
 
@@ -34,7 +37,7 @@ export async function run(argv) {
34
37
  const level = a.flags.level || "skill";
35
38
  if (!LEVELS.includes(level)) throw new UsageError(`--level must be one of ${LEVELS.join(", ")}`);
36
39
  if (!a._.length) throw new UsageError("doctor needs a path");
37
- const opts = { level, now: clock(a.flags) };
40
+ const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null };
38
41
  if (a.flags.changed) {
39
42
  const all = a._.flatMap((p) => findSkills(p));
40
43
  const files = a._.flatMap((p) => changedFiles(p, a.flags.base));
@@ -1,5 +1,5 @@
1
1
  import { existsSync, mkdirSync, writeFileSync, readFileSync } from "node:fs";
2
- import { join } from "node:path";
2
+ import { join, basename } from "node:path";
3
3
  import { parseArgs, UsageError } from "../args.mjs";
4
4
  import { skillDir } from "./common.mjs";
5
5
  import { MISSES_HEADER } from "../misses.mjs";
@@ -15,7 +15,9 @@ Never overwrites a file that exists.
15
15
  --from-session <file> build the first eval and golden candidate from the session where
16
16
  the job was done by hand. A Claude Code .jsonl gives the first
17
17
  request and the final answer; any other file is the request.
18
- The golden waits for a person: run \`superskill approve\`.
18
+ The golden waits for a person: fill in accepted.by and
19
+ accepted.at in its PROVENANCE.json (who accepted the
20
+ output when it ran), then \`superskill approve\`.
19
21
  `;
20
22
 
21
23
  const TRIGGER_EXAMPLE = [
@@ -65,6 +67,10 @@ export async function run(argv) {
65
67
  writeFileSync(evalsPath, JSON.stringify(doc, null, 2) + "\n");
66
68
  put(`goldens/${id}/input.md`, sessionCase.prompt + "\n");
67
69
  put(`goldens/${id}/output.md`, sessionCase.output ? sessionCase.output + "\n" : "");
70
+ // A real run, but nobody has said yet that its output was accepted. Until accepted.by and
71
+ // accepted.at are filled in (by the person, or by a tool that read their next message), this
72
+ // golden cannot count toward superskill however it is approved.
73
+ put(`goldens/${id}/PROVENANCE.json`, JSON.stringify({ source: "real-run", run: { session: basename(session), ledger_id: null, commit: null, at: null }, accepted: { by: null, at: null, signal: null, evidence: "" } }, null, 2) + "\n");
68
74
  process.stdout.write(`eval ${id} and golden candidate goldens/${id}/ written from ${session}\n`);
69
75
  process.stdout.write(`a person approves it with: superskill approve ${a._[0]} ${id}\n`);
70
76
  }
package/src/doctor.mjs CHANGED
@@ -37,6 +37,7 @@ export function findSkills(path) {
37
37
  /** Score one loaded skill. */
38
38
  export function scoreSkill(dir, opts = {}) {
39
39
  const ctx = loadSkill(dir);
40
+ if (opts.privateGoldens) ctx.privateGoldens = opts.privateGoldens;
40
41
  const findings = runRules(ctx, opts.rules || allRules, opts);
41
42
  const level = computeLevel(findings);
42
43
  const up = nextLevel(level);
package/src/goldens.mjs CHANGED
@@ -1,38 +1,90 @@
1
1
  // goldens/<id>/: input.md, the approved output (output.md, or any other non-input file),
2
- // and APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
3
- // rationale, basis, evidence}] }, or the single-approval shape from before 0.3.0.
2
+ // APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
3
+ // rationale, basis, evidence}] } (or the single-approval shape from before 0.3.0),
4
+ // PROVENANCE.json saying where the example came from (0.5.0), and, for an anonymized twin of a
5
+ // private golden, ANONYMIZED.json, the anonymizer's receipt.
6
+ //
7
+ // A golden may also live OUTSIDE the skill, in a private folder the operator keeps
8
+ // (<private>/<skill-name>/<id>/), because a real run's input and output are usually about real
9
+ // people and should not travel with a skill that is shared. The doctor reads both.
4
10
  import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
5
11
  import { join } from "node:path";
6
12
 
7
- export function readGoldens(dir) {
8
- const root = join(dir, "goldens");
9
- if (!existsSync(root)) return [];
13
+ /** Files in a golden folder that describe it rather than being its input or output. */
14
+ export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json"]);
15
+
16
+ /** Where goldens are read from: the skill's own goldens/, then the private folder for its name. */
17
+ export function goldenRoots(dir, { privateGoldens = null, name = null } = {}) {
18
+ const roots = [{ root: join(dir, "goldens"), private: false }];
19
+ if (privateGoldens && name) roots.push({ root: join(privateGoldens, name), private: true });
20
+ return roots;
21
+ }
22
+
23
+ const readJsonFile = (p) => {
24
+ if (!existsSync(p)) return { value: null, error: null };
25
+ try { return { value: JSON.parse(readFileSync(p, "utf8")), error: null }; }
26
+ catch (e) { return { value: null, error: e.message }; }
27
+ };
28
+
29
+ export function readGoldens(dir, opts = {}) {
10
30
  const out = [];
11
- for (const id of readdirSync(root).sort()) {
12
- const gdir = join(root, id);
13
- try { if (!statSync(gdir).isDirectory()) continue; } catch { continue; }
14
- const files = readdirSync(gdir);
15
- const inputName = files.find((f) => /^input\./i.test(f));
16
- const outputName = files.find((f) => /^output\./i.test(f)) || files.find((f) => f !== inputName && f !== "APPROVAL.json" && !f.startsWith("."));
17
- let approval = null, approvalError = null;
18
- if (files.includes("APPROVAL.json")) {
19
- try { approval = JSON.parse(readFileSync(join(gdir, "APPROVAL.json"), "utf8")); }
20
- catch (e) { approvalError = e.message; }
31
+ for (const { root, private: priv } of goldenRoots(dir, opts)) {
32
+ if (!existsSync(root)) continue;
33
+ for (const id of readdirSync(root).sort()) {
34
+ const gdir = join(root, id);
35
+ try { if (!statSync(gdir).isDirectory()) continue; } catch { continue; }
36
+ const files = readdirSync(gdir).filter((f) => { try { return statSync(join(gdir, f)).isFile(); } catch { return false; } });
37
+ const inputName = files.find((f) => /^input\./i.test(f));
38
+ const outputName = files.find((f) => /^output\./i.test(f)) || files.find((f) => f !== inputName && !META_FILES.has(f) && !f.startsWith("."));
39
+ const appr = readJsonFile(join(gdir, "APPROVAL.json"));
40
+ const prov = readJsonFile(join(gdir, "PROVENANCE.json"));
41
+ const anon = readJsonFile(join(gdir, "ANONYMIZED.json"));
42
+ const provenance = prov.value;
43
+ out.push({
44
+ id,
45
+ dir: gdir,
46
+ private: priv,
47
+ input: inputName ? readFileSync(join(gdir, inputName), "utf8") : null,
48
+ outputFile: outputName || null,
49
+ output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
50
+ approval: appr.value,
51
+ approvals: approvalsOf(appr.value),
52
+ approvalError: appr.error,
53
+ provenance,
54
+ provenanceError: prov.error,
55
+ anonymized: anon.value,
56
+ origin: originOf(provenance, anon.value),
57
+ });
21
58
  }
22
- out.push({
23
- id,
24
- dir: gdir,
25
- input: inputName ? readFileSync(join(gdir, inputName), "utf8") : null,
26
- outputFile: outputName || null,
27
- output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
28
- approval,
29
- approvals: approvalsOf(approval),
30
- approvalError,
31
- });
32
59
  }
33
60
  return out;
34
61
  }
35
62
 
63
+ /** Where a golden's example came from. `real-run` is the only source that can reach superskill. */
64
+ export const SOURCES = ["real-run", "synthetic", "synthetic-reconstruction"];
65
+
66
+ /**
67
+ * Is this golden a real run a person accepted, and if not, why not?
68
+ *
69
+ * A real run names the run it came from (a session, a commit, a ledger id, or, for an
70
+ * anonymized twin, `derived_from`: a hash of the private original, never its content) and the
71
+ * person who accepted the output when it happened, and when. An anonymized twin must also carry
72
+ * the anonymizer's receipt, because "the original was accepted" and "this twin still says the
73
+ * same thing" are different claims.
74
+ */
75
+ export function originOf(prov, receipt = null) {
76
+ if (!prov || typeof prov !== "object") return { real: false, why: "no PROVENANCE.json" };
77
+ if (prov.source !== "real-run") return { real: false, why: `source is ${JSON.stringify(prov.source ?? null)}, not "real-run"` };
78
+ const run = prov.run && typeof prov.run === "object" ? prov.run : {};
79
+ const ref = ["session", "commit", "ledger_id"].map((k) => run[k]).concat(prov.derived_from).find((x) => typeof x === "string" && x.trim());
80
+ if (!ref) return { real: false, why: "names no run (run.session, run.commit, run.ledger_id or derived_from)" };
81
+ const acc = prov.accepted && typeof prov.accepted === "object" ? prov.accepted : {};
82
+ if (!(typeof acc.by === "string" && acc.by.trim())) return { real: false, why: "names no person who accepted the run (accepted.by)" };
83
+ if (!validDate(acc.at)) return { real: false, why: "has no valid accepted.at" };
84
+ if (prov.anonymized === true && !(receipt && typeof receipt === "object" && receipt.fingerprint)) return { real: false, why: "is an anonymized twin with no ANONYMIZED.json receipt" };
85
+ return { real: true, why: "" };
86
+ }
87
+
36
88
  /** What an approval rests on. `judgment`: the people who approved it read it and said it is right.
37
89
  * `outcome`: it produced a result in the world someone can check (a client landed, a call booked).
38
90
  * Being liked and being proven are different weights, and a golden says which it carries. */
@@ -54,6 +106,7 @@ export function approvalsOf(approval) {
54
106
  rationale: String(a.rationale ?? a.note ?? "").trim(),
55
107
  basis: a.basis === "outcome" ? "outcome" : "judgment",
56
108
  evidence: String(a.evidence ?? "").trim(),
109
+ via: String(a.via ?? "").trim(),
57
110
  }));
58
111
  }
59
112
 
@@ -81,3 +134,12 @@ export function withApproval(existing, entry) {
81
134
  }
82
135
 
83
136
  export const isApproved = (g) => approvalsOf(g.approval).length > 0;
137
+
138
+ /** Approved AND from a real run a person accepted: the only golden that counts for superskill. */
139
+ export const isRealApproved = (g) => isApproved(g) && (g.origin || originOf(g.provenance, g.anonymized)).real;
140
+
141
+ /** The options readGoldens needs for a loaded skill: its name and the private folder, if any. */
142
+ export const goldenOpts = (ctx, opts = {}) => ({
143
+ name: (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName,
144
+ privateGoldens: opts.privateGoldens ?? ctx.privateGoldens ?? process.env.SUPERSKILL_PRIVATE_GOLDENS ?? null,
145
+ });
package/src/ledger.mjs CHANGED
@@ -12,7 +12,10 @@ export function ledgerPaths(skillDir, skillName) {
12
12
  const out = [];
13
13
  const local = join(skillDir, "invocations.jsonl");
14
14
  if (existsSync(local)) out.push(local);
15
- const root = join(homedir(), ".freedom", "ledger", "skills");
15
+ // FREEDOM_SKILL_LEDGER_HOME is where Freedom itself writes when it is re-pointed (tests, a
16
+ // second profile); read the same place it writes.
17
+ const home = process.env.FREEDOM_SKILL_LEDGER_HOME || join(homedir(), ".freedom", "ledger");
18
+ const root = join(home, "skills");
16
19
  if (existsSync(root) && skillName) {
17
20
  for (const plugin of readdirSync(root).sort()) {
18
21
  const p = join(root, plugin, `${skillName}.jsonl`);
@@ -30,15 +33,48 @@ export function ledgerMisses(path, known, skillName) {
30
33
  let rec;
31
34
  try { rec = JSON.parse(line); } catch { continue; }
32
35
  if (!rec || !rec.id || known.has(rec.id)) continue;
33
- if (skillName && rec.skill && rec.skill !== skillName) continue;
36
+ // A sandbox run (superskill doctor --run) is a test of the skill, not a use of it.
37
+ if (rec.synthetic) continue;
38
+ if (skillName && rec.skill && bare(rec.skill) !== bare(skillName)) continue;
34
39
  const kinds = (Array.isArray(rec.interventions) ? rec.interventions : []).filter((i) => i && MISS_KINDS.has(i.kind));
35
40
  const failed = rec.outcome === "failed";
36
41
  if (!kinds.length && !failed) continue;
37
42
  const parts = kinds.map((i) => `${i.kind}${i.note || i.what ? `: ${i.note || i.what}` : ""}`);
38
- if (failed) parts.unshift(`run failed${Array.isArray(rec.errors) && rec.errors.length ? ` (${rec.errors.slice(0, 2).join("; ")})` : ""}`);
43
+ if (failed) parts.unshift(`run failed${Array.isArray(rec.errors) && rec.errors.length ? ` (${rec.errors.slice(0, 2).map((e) => (typeof e === "string" ? e : e?.kind || "error")).join("; ")})` : ""}`);
39
44
  const date = typeof rec.started === "string" && /^\d{4}-\d{2}-\d{2}/.test(rec.started) ? rec.started.slice(0, 10) : null;
40
- out.push({ date, status: "open", what: parts.join("; ").replace(/\s+/g, " ").slice(0, 300), expected: "", source: `freedom-ledger ${rec.id}` });
45
+ // A correction made in the operator's NEXT message is the best "should have" there is, and the
46
+ // ledger keeps only where it is, never its words: point at it, so the fixer reads it there.
47
+ const ref = rec.next_turn_ref && rec.session_id ? ` (the correction is the operator's message in session ${String(rec.session_id).slice(0, 8)} at ${rec.next_turn_ref.at || "?"}${Number.isFinite(rec.next_turn_ref.offset) ? `, transcript byte ${rec.next_turn_ref.offset}` : ""})` : "";
48
+ const expected = rec.corrected_after ? `what the operator asked for instead${ref}` : "";
49
+ out.push({ date, status: "open", what: parts.join("; ").replace(/\s+/g, " ").slice(0, 300), expected, source: `freedom-ledger ${rec.id}` });
41
50
  known.add(rec.id);
42
51
  }
43
52
  return out;
44
53
  }
54
+
55
+ const bare = (name) => String(name || "").split(":").pop();
56
+
57
+ /**
58
+ * How many runs the person accepted, from Freedom's `next_turn` verdict (the class of the first
59
+ * message after a run handed back). Runs with no verdict are left out of both sides, so a missing
60
+ * record can never read as an acceptance, and sandbox runs never count.
61
+ */
62
+ export function acceptedRate(paths, { now = new Date(), days = 30 } = {}) {
63
+ const since = now.getTime() - days * 86400000;
64
+ let judged = 0, accepted = 0, synthetic = 0;
65
+ for (const p of paths) {
66
+ let text = "";
67
+ try { text = readFileSync(p, "utf8"); } catch { continue; }
68
+ for (const line of text.split("\n")) {
69
+ if (!line.trim()) continue;
70
+ let rec; try { rec = JSON.parse(line); } catch { continue; }
71
+ const t = Date.parse(rec?.started || "");
72
+ if (!Number.isFinite(t) || t < since || t > now.getTime()) continue;
73
+ if (rec.synthetic) { synthetic++; continue; }
74
+ if (!rec.next_turn || rec.next_turn === "none") continue;
75
+ judged++;
76
+ if (["close", "go", "new_topic"].includes(rec.next_turn) && !rec.corrected_after && !["failed", "abandoned"].includes(rec.outcome)) accepted++;
77
+ }
78
+ }
79
+ return { judged, accepted, synthetic };
80
+ }
@@ -2,13 +2,15 @@
2
2
  // got something wrong, and proven recently on a current model against the no-skill baseline.
3
3
  import { createHash } from "node:crypto";
4
4
  import { defineRules } from "./define.mjs";
5
- import { readGoldens, isApproved, weightOf } from "../goldens.mjs";
5
+ import { readGoldens, isApproved, isRealApproved, weightOf, goldenOpts } from "../goldens.mjs";
6
+ import { ledgerPaths, acceptedRate } from "../ledger.mjs";
6
7
  import { readMisses } from "../misses.mjs";
7
8
  import { readEvals, readLatestRun } from "../evals.mjs";
8
9
 
9
10
  const f = (severity, message, fix) => ({ severity, message, fix });
10
11
  const DAY = 86400000;
11
12
  export const OPEN_MISS_DAYS = 14;
13
+ export const REAL_USE_DAYS = 30;
12
14
  /** How long a --run stays fresh, by metadata.cadence. */
13
15
  export const FRESH_DAYS = { daily: 30, weekly: 30, monthly: 60, quarterly: 120, yearly: 365 };
14
16
  const DEFAULT_FRESH = 30;
@@ -19,15 +21,27 @@ export const superskillRules = defineRules([
19
21
  {
20
22
  id: "golden-approved",
21
23
  level: "superskill",
22
- check(ctx) {
24
+ check(ctx, opts = {}) {
23
25
  if (ctx.error) return [];
24
- const goldens = readGoldens(ctx.dir);
25
- const approved = goldens.filter(isApproved);
26
+ const goldens = readGoldens(ctx.dir, goldenOpts(ctx, opts));
27
+ const approvedAny = goldens.filter(isApproved);
28
+ // 0.5.0: only a golden from a real run a person accepted counts. An invented example, however
29
+ // careful, puts "a person said this was right" on something no person's work produced, and a
30
+ // skill could then reach the top level on its author's fiction.
31
+ const approved = goldens.filter(isRealApproved);
26
32
  const out = [];
33
+ const where = (g) => (g.private ? `private golden ${g.id}` : `goldens/${g.id}`);
27
34
  for (const g of goldens.filter((x) => x.approvalError))
28
- out.push(f("fail", `goldens/${g.id}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
35
+ out.push(f("fail", `${where(g)}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
36
+ for (const g of goldens.filter((x) => x.provenanceError))
37
+ out.push(f("fail", `${where(g)}/PROVENANCE.json is not valid JSON`, "Rewrite it: source, run, accepted (see SPEC.md, goldens)."));
29
38
  if (!approved.length) {
30
- out.push(f("fail", goldens.length ? `${goldens.length} golden${goldens.length === 1 ? "" : "s"}, none approved by a person` : "no goldens", goldens.length ? "A person runs `superskill approve <skill> <golden>` at a terminal after checking the output." : "Save a real input and the output you would sign off on under goldens/<id>/, then `superskill approve`."));
39
+ if (approvedAny.length) {
40
+ const why = approvedAny.map((g) => `${g.id} ${g.origin.why}`).join("; ");
41
+ out.push(f("fail", `${approvedAny.length} approved golden${approvedAny.length === 1 ? "" : "s"}, none from a real run a person accepted (${why})`, "A golden counts toward superskill only when goldens/<id>/PROVENANCE.json says source: real-run, names the run (session, commit, ledger id, or derived_from for an anonymized twin) and who accepted it and when. Invented examples belong in evals/evals.json, where they hold the skill at tested."));
42
+ } else {
43
+ out.push(f("fail", goldens.length ? `${goldens.length} golden${goldens.length === 1 ? "" : "s"}, none approved by a person` : "no goldens", goldens.length ? "A person runs `superskill approve <skill> <golden>` at a terminal (or relays a tap with --approved-by and --via) after checking the output." : "Save a real run's input and the output a person accepted under goldens/<id>/ with PROVENANCE.json, then `superskill approve`."));
44
+ }
31
45
  return out;
32
46
  }
33
47
  const sha = createHash("sha256").update(ctx.raw).digest("hex");
@@ -43,6 +57,19 @@ export const superskillRules = defineRules([
43
57
  return out;
44
58
  },
45
59
  },
60
+ {
61
+ // The real number beside the level: of the runs whose next message is recorded, how many did
62
+ // the person accept. Info only. A level is evidence about examples; this is evidence about use.
63
+ id: "real-use",
64
+ level: "superskill",
65
+ check(ctx, { now = new Date() } = {}) {
66
+ if (ctx.error) return [];
67
+ const name = (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName;
68
+ const r = acceptedRate(ledgerPaths(ctx.dir, name), { now, days: REAL_USE_DAYS });
69
+ if (!r.judged) return [];
70
+ return [f("info", `real use, last ${REAL_USE_DAYS} days: ${r.accepted} of ${r.judged} judged runs accepted (${pct(r.accepted / r.judged)})${r.synthetic ? `; ${r.synthetic} sandbox run${r.synthetic === 1 ? "" : "s"} not counted` : ""}`, "")];
71
+ },
72
+ },
46
73
  {
47
74
  id: "misses-log-present",
48
75
  level: "superskill",
@@ -74,6 +101,7 @@ export const superskillRules = defineRules([
74
101
  const misses = (readMisses(ctx.dir) || []).filter((m) => m.status === "fixed");
75
102
  if (!misses.length) return [];
76
103
  const ids = new Set(readEvals(ctx.dir).cases.map((c) => String(c.id)));
104
+ // Shipped goldens only: a miss is closed by a check that travels with the skill.
77
105
  for (const g of readGoldens(ctx.dir)) ids.add(g.id);
78
106
  return misses
79
107
  .filter((m) => !m.eval || !ids.has(String(m.eval)))
@@ -2,7 +2,7 @@
2
2
  // with both should-load requests and near-misses that should not load the skill.
3
3
  import { defineRules } from "./define.mjs";
4
4
  import { readEvals, readTriggers } from "../evals.mjs";
5
- import { readGoldens } from "../goldens.mjs";
5
+ import { readGoldens, goldenOpts } from "../goldens.mjs";
6
6
 
7
7
  const f = (severity, message, fix) => ({ severity, message, fix });
8
8
  export const MIN_EVALS = 3, MIN_TRIGGERS = 10, MIN_EACH_SIDE = 3;
@@ -15,7 +15,7 @@ export const testedRules = defineRules([
15
15
  if (ctx.error) return [];
16
16
  const e = readEvals(ctx.dir);
17
17
  if (e.error) return [f("fail", e.error, "Fix evals/evals.json so it parses as skill-creator's {skill_name, evals: [...]}.")];
18
- const goldenCases = readGoldens(ctx.dir).filter((g) => g.input !== null && g.output !== null).length;
18
+ const goldenCases = readGoldens(ctx.dir, goldenOpts(ctx)).filter((g) => g.input !== null && g.output !== null).length;
19
19
  const n = e.cases.length + goldenCases;
20
20
  if (n < MIN_EVALS)
21
21
  return [f("fail", `${n} eval case${n === 1 ? "" : "s"} (need ${MIN_EVALS})`, "Add real requests to evals/evals.json (`superskill init` writes an example).")];
@@ -35,6 +35,29 @@ export const testedRules = defineRules([
35
35
  : [];
36
36
  },
37
37
  },
38
+ {
39
+ // MEASURED 2026-10-05 (freedom-dev): `superskill init` writes one example eval and two example
40
+ // triggers, each starting "REPLACE:", and a skill whose author copied those up to 3 and 10
41
+ // scored `tested` while testing nothing. A placeholder, or the same request twice, is not a case.
42
+ id: "evals-real",
43
+ level: "tested",
44
+ check(ctx) {
45
+ if (ctx.error) return [];
46
+ const e = readEvals(ctx.dir), t = readTriggers(ctx.dir);
47
+ const out = [];
48
+ const placeholder = (x) => /\bREPLACE:/.test(JSON.stringify(x));
49
+ const evalHits = e.cases.filter((c) => placeholder([c.prompt, c.expected_output, c.assertions])).map((c) => c.id);
50
+ if (evalHits.length) out.push(f("fail", `eval case${evalHits.length === 1 ? "" : "s"} ${evalHits.join(", ")} still hold${evalHits.length === 1 ? "s" : ""} a \`superskill init\` REPLACE: placeholder`, "Replace every example with a real request this skill handles and what a good answer must contain."));
51
+ const trigHits = t.triggers.filter((x) => placeholder(x.query)).length;
52
+ if (trigHits) out.push(f("fail", `${trigHits} trigger${trigHits === 1 ? "" : "s"} still hold a REPLACE: placeholder`, "Replace them with realistic requests, including near-misses that should not load the skill."));
53
+ const dup = (list) => list.filter((x, i) => x && list.indexOf(x) !== i);
54
+ const dp = [...new Set(dup(e.cases.map((c) => c.prompt.trim())))];
55
+ if (dp.length) out.push(f("fail", `${dp.length} eval prompt${dp.length === 1 ? " is" : "s are"} repeated: the same request twice is one case`, "Make each case a different request."));
56
+ const dq = [...new Set(dup(t.triggers.map((x) => x.query.trim().toLowerCase())))];
57
+ if (dq.length) out.push(f("fail", `${dq.length} trigger quer${dq.length === 1 ? "y is" : "ies are"} repeated`, "Make each trigger a different request."));
58
+ return out;
59
+ },
60
+ },
38
61
  {
39
62
  id: "triggers-present",
40
63
  level: "tested",
@@ -3,12 +3,13 @@
3
3
  // a copy installed at user level cannot leak into the baseline.
4
4
  import { spawnSync } from "node:child_process";
5
5
  import { prepareWorkspace } from "./workspace.mjs";
6
+ import { sandboxEnv } from "./env.mjs";
6
7
 
7
8
  export const name = "claude";
8
9
 
9
10
  function call(args, cwd) {
10
11
  const t0 = Date.now();
11
- const r = spawnSync("claude", args, { cwd, encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
12
+ const r = spawnSync("claude", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
12
13
  const ms = Date.now() - t0;
13
14
  if (r.error) throw Object.assign(new Error(`claude failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
14
15
  let doc = null;
package/src/run/codex.mjs CHANGED
@@ -5,6 +5,7 @@ import { spawnSync } from "node:child_process";
5
5
  import { readFileSync, existsSync } from "node:fs";
6
6
  import { join } from "node:path";
7
7
  import { prepareWorkspace } from "./workspace.mjs";
8
+ import { sandboxEnv } from "./env.mjs";
8
9
 
9
10
  export const name = "codex";
10
11
 
@@ -14,7 +15,7 @@ function call(prompt, cwd, model) {
14
15
  if (model) args.push("-m", model);
15
16
  args.push(prompt);
16
17
  const t0 = Date.now();
17
- const r = spawnSync("codex", args, { cwd, encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
18
+ const r = spawnSync("codex", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
18
19
  if (r.error) throw Object.assign(new Error(`codex failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
19
20
  const output = existsSync(out) ? readFileSync(out, "utf8") : r.stdout;
20
21
  const tok = (r.stderr + r.stdout).match(/tokens used[:\s]+([\d,]+)/i);
@@ -0,0 +1,17 @@
1
+ // The environment every sandboxed run gets, so a `doctor --run` session is never counted as a
2
+ // real use of the skill by whatever is recording real uses.
3
+ //
4
+ // MEASURED, NOT FEARED (2026-10-05): Freedom's skill ledger records every skill run from a
5
+ // harness hook. The doctor's own sandbox runs fired that hook, so three runs of a skill nobody had
6
+ // ever used for real landed in the operator's ledger as perfect one-shot runs. Left alone, every
7
+ // `doctor --run` raises the number that is meant to keep the doctor honest.
8
+ //
9
+ // Two signals, because a caller's recorder may honour either: FREEDOM_SKILL_LEDGER=off is the
10
+ // Freedom ledger's documented off switch, and SUPERSKILL_SANDBOX=1 names the run for any other
11
+ // recorder. The run folder's own name (SANDBOX_PREFIX) is the third, for a recorder that reads
12
+ // only the hook payload's cwd.
13
+ export const SANDBOX_PREFIX = "superskill-run-";
14
+
15
+ export function sandboxEnv(base = process.env) {
16
+ return { ...base, FREEDOM_SKILL_LEDGER: "off", SUPERSKILL_SANDBOX: "1" };
17
+ }
package/src/run/fake.mjs CHANGED
@@ -3,12 +3,13 @@
3
3
  // {output, tokens?, model?}. Never calls a model.
4
4
  import { spawnSync } from "node:child_process";
5
5
  import { prepareWorkspace } from "./workspace.mjs";
6
+ import { sandboxEnv } from "./env.mjs";
6
7
 
7
8
  export const name = "fake";
8
9
 
9
10
  function call(script, req) {
10
11
  const t0 = Date.now();
11
- const r = spawnSync(process.execPath, [script], { input: JSON.stringify(req), encoding: "utf8" });
12
+ const r = spawnSync(process.execPath, [script], { input: JSON.stringify(req), env: sandboxEnv(), encoding: "utf8" });
12
13
  if (r.status !== 0) throw Object.assign(new Error(`fake harness failed: ${r.stderr}`), { code: "SUPERSKILL" });
13
14
  const doc = JSON.parse(r.stdout);
14
15
  return { output: doc.output ?? "", tokens: doc.tokens ?? null, model: doc.model ?? "fake", ms: Date.now() - t0, failed: false };
package/src/run/index.mjs CHANGED
@@ -39,11 +39,12 @@ export function pickHarness(name) {
39
39
  }
40
40
 
41
41
  /** Every case the run will execute: evals.json cases plus goldens judged against their approved output. */
42
- export function collectCases(skillDir) {
42
+ export function collectCases(skillDir, { privateGoldens = process.env.SUPERSKILL_PRIVATE_GOLDENS || null } = {}) {
43
43
  const e = readEvals(skillDir);
44
+ const name = parseSkillFile(readFileSync(join(skillDir, "SKILL.md"), "utf8")).data.name || "";
44
45
  if (e.error) throw new UsageError(e.error);
45
46
  const cases = e.cases.filter((c) => c.prompt.trim()).map((c) => ({ id: String(c.id), prompt: c.prompt, files: c.files, assertions: c.assertions.length ? c.assertions : [c.expected_output].filter(Boolean) }));
46
- for (const g of readGoldens(skillDir)) {
47
+ for (const g of readGoldens(skillDir, { privateGoldens, name })) {
47
48
  if (!g.input || !g.output || !g.output.trim()) continue;
48
49
  cases.push({ id: `golden:${g.id}`, prompt: g.input, files: [], assertions: [`The output matches this approved output in substance (same facts, same shape; wording may differ):\n${g.output}`] });
49
50
  }
@@ -1,6 +1,7 @@
1
1
  import { mkdtempSync, mkdirSync, symlinkSync, cpSync, existsSync } from "node:fs";
2
2
  import { tmpdir } from "node:os";
3
3
  import { join, dirname } from "node:path";
4
+ import { SANDBOX_PREFIX } from "./env.mjs";
4
5
 
5
6
  /**
6
7
  * A fresh working folder for one run. With `linkAt` (e.g. ".claude/skills"), the skill is
@@ -8,7 +9,7 @@ import { join, dirname } from "node:path";
8
9
  * files (paths relative to the skill) are copied in at the same relative paths.
9
10
  */
10
11
  export function prepareWorkspace({ skillDir, skillName, files = [], linkAt = null }) {
11
- const cwd = mkdtempSync(join(tmpdir(), "superskill-run-"));
12
+ const cwd = mkdtempSync(join(tmpdir(), SANDBOX_PREFIX));
12
13
  if (linkAt) {
13
14
  const target = join(cwd, linkAt, skillName);
14
15
  mkdirSync(dirname(target), { recursive: true });