@supersuit/superskill 0.3.1 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +47 -0
- package/README.md +3 -3
- package/SPEC.md +45 -5
- package/package.json +1 -1
- package/src/commands/approve.mjs +34 -7
- package/src/commands/doctor.mjs +6 -3
- package/src/commands/init.mjs +8 -2
- package/src/doctor.mjs +1 -0
- package/src/goldens.mjs +87 -25
- package/src/ledger.mjs +40 -4
- package/src/rules/superskill.mjs +34 -6
- package/src/rules/tested.mjs +25 -2
- package/src/run/claude.mjs +2 -1
- package/src/run/codex.mjs +2 -1
- package/src/run/env.mjs +17 -0
- package/src/run/fake.mjs +2 -1
- package/src/run/index.mjs +3 -2
- package/src/run/workspace.mjs +2 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,52 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.5.0 (2026-10-06)
|
|
4
|
+
|
|
5
|
+
**Superskill needs a golden from a real run.** Before, any golden a person approved counted, so a
|
|
6
|
+
skill could reach the top level on examples its author invented. (Gary Sheng, on five invented
|
|
7
|
+
goldens written to lift Freedom's flagship skills: *"I'm just concerned about hallucination that
|
|
8
|
+
we're accepting just to get higher doctor ratings."*) **Breaking for anyone at superskill today:**
|
|
9
|
+
an approved golden without the new `PROVENANCE.json` no longer counts, and the doctor says which
|
|
10
|
+
golden and why.
|
|
11
|
+
|
|
12
|
+
- `goldens/<id>/PROVENANCE.json`: `source` (`real-run`, `synthetic`, `synthetic-reconstruction`),
|
|
13
|
+
the `run` it came from (session, commit or ledger id), and who `accepted` the output when it ran,
|
|
14
|
+
and when. Only `real-run` with all of that counts for `golden-approved`. An anonymized twin of a
|
|
15
|
+
private golden counts with `derived_from` and the anonymizer's `ANONYMIZED.json` receipt.
|
|
16
|
+
- **Private goldens.** `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`) on `doctor`,
|
|
17
|
+
`approve` and `doctor --run` also reads `<dir>/<skill-name>/<id>/`, for real runs too personal
|
|
18
|
+
to ship with the skill. `approve` writes the approval beside the golden it found.
|
|
19
|
+
- `init --from-session` writes a `real-run` provenance with `accepted` left empty, so the golden
|
|
20
|
+
counts only once someone records who accepted it.
|
|
21
|
+
- **`evals-real` (tested):** a `superskill init` `REPLACE:` placeholder, or the same prompt or
|
|
22
|
+
trigger query twice, is not a case. Measured in Freedom on 2026-10-05: the init scaffold copied up
|
|
23
|
+
to 3 evals and 10 triggers scored `tested` while testing nothing.
|
|
24
|
+
- **`doctor --run` is never recorded as a real use.** Every harness spawns its child with
|
|
25
|
+
`FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`. Measured 2026-10-05: three sandbox runs
|
|
26
|
+
landed in Freedom's skill ledger as perfect one-shot runs of a skill nobody had used.
|
|
27
|
+
- **`real-use` (info):** where the run ledger records what the person's next message made of each
|
|
28
|
+
run, the doctor prints how many of the last 30 days' judged runs they accepted, beside the level.
|
|
29
|
+
- `miss import --freedom-ledger` skips sandbox records, matches `freedom:<skill>` records, and
|
|
30
|
+
points a correction's "Should have" at the person's next message (session and time, never text).
|
|
31
|
+
It reads `FREEDOM_SKILL_LEDGER_HOME` when Freedom's ledger was re-pointed.
|
|
32
|
+
- Tests: 11 new; each guard broken on purpose and seen red (invented golden reaching superskill,
|
|
33
|
+
placeholders counting, duplicate prompts counting, a sandbox run becoming a miss, a twin with no
|
|
34
|
+
receipt counting, the ledger off switch missing from the run env, codex spawned without it).
|
|
35
|
+
|
|
36
|
+
## 0.4.0 (2026-09-29)
|
|
37
|
+
|
|
38
|
+
**Approve from your phone.** `superskill approve` worked only at an interactive terminal, so an
|
|
39
|
+
operator who reviews on a phone had to find a laptop to say yes. (Gary Sheng: *"The best way for
|
|
40
|
+
me to approve it is to click yes, approve, not you telling me to go to my terminal. I'm normally
|
|
41
|
+
mobile first."*) Away from a terminal it now accepts `--approved-by "<the person>"` and
|
|
42
|
+
`--via "<where they said yes>"`, after that person approved with a tap on a board or a review page.
|
|
43
|
+
Both are required and both are recorded (`via` on the approval), so an approval an agent relayed is
|
|
44
|
+
always distinguishable from one typed at a terminal, and an agent still cannot approve with no
|
|
45
|
+
named person and no channel. A rationale is required either way.
|
|
46
|
+
|
|
47
|
+
- Tests: refused with no name, refused with a name and no channel, refused with no rationale;
|
|
48
|
+
recorded with person, channel and basis. The channel requirement was mutated out and went red.
|
|
49
|
+
|
|
3
50
|
## 0.3.1 (2026-09-29)
|
|
4
51
|
|
|
5
52
|
The 0.3.0 changes below, released. The v0.3.0 tag's publish refused on a red test (SPEC.md still
|
package/README.md
CHANGED
|
@@ -37,7 +37,7 @@ file error. `--json` prints one JSON document and nothing else.
|
|
|
37
37
|
hard-coded machine paths, nothing that reads like a prompt injection.
|
|
38
38
|
2. **tested**: at least three task evals with checks a machine can verify, and a trigger set of
|
|
39
39
|
at least ten requests, some that should load the skill and some near-misses that should not.
|
|
40
|
-
3. **superskill**: a golden a
|
|
40
|
+
3. **superskill**: a golden from a real run, accepted when it ran and approved by a person; no miss open longer than 14 days and every fixed
|
|
41
41
|
miss guarded by an eval; a recent `--run` on file where the skill beats the same task done
|
|
42
42
|
without it. "Recent" follows `metadata.cadence` (a weekly skill's proof lasts 30 days).
|
|
43
43
|
|
|
@@ -54,7 +54,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
|
|
|
54
54
|
| `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
|
|
55
55
|
| `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
|
|
56
56
|
| `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
|
|
57
|
-
| `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence).
|
|
57
|
+
| `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). At a terminal it asks for your name; from your phone, an agent records your tap with `--approved-by` and `--via`. Approvals accumulate. |
|
|
58
58
|
| `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
|
|
59
59
|
| `superskill miss import <skill> --freedom-ledger [--ledger <file>]` | Import runs that needed correcting from Freedom's run ledger. |
|
|
60
60
|
| `superskill snippet` | Print a block for `AGENTS.md` / `CLAUDE.md` that teaches any agent these habits. |
|
|
@@ -70,7 +70,7 @@ my-skill/
|
|
|
70
70
|
SKILL.md
|
|
71
71
|
evals/evals.json task evals (Anthropic skill-creator format)
|
|
72
72
|
evals/triggers.json should / should-not load (skill-creator format)
|
|
73
|
-
goldens/<id>/ input.md, output.md, APPROVAL.json
|
|
73
|
+
goldens/<id>/ input.md, output.md, PROVENANCE.json, APPROVAL.json
|
|
74
74
|
MISSES.md every miss, open or fixed with its eval
|
|
75
75
|
evals/results/latest.json the last --run
|
|
76
76
|
```
|
package/SPEC.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# The superskill standard
|
|
2
2
|
|
|
3
|
-
**Version 0.
|
|
3
|
+
**Version 0.5.0** (2026-10-06). The reference checker is `@supersuit/superskill`; where this
|
|
4
4
|
document and the checker disagree, the checker has a bug.
|
|
5
5
|
|
|
6
6
|
A **superskill** runs on frontier intelligence, is checked against examples a person approved,
|
|
@@ -23,7 +23,7 @@ a file in the skill's own folder, so the evidence travels with the skill whereve
|
|
|
23
23
|
| The clause | What proves it | Where it lives |
|
|
24
24
|
|---|---|---|
|
|
25
25
|
| A skill at all | A valid `SKILL.md` under the Agent Skills spec, plus the hygiene rules below | `SKILL.md` |
|
|
26
|
-
| Checked against examples you approved | At least one golden
|
|
26
|
+
| Checked against examples you approved | At least one golden from a real run: the real input, the output the person accepted when it ran, where it came from, and a record of who approved it as a golden and when | `goldens/<id>/` (or a private folder, see below) |
|
|
27
27
|
| Fixed every time it gets something wrong | A miss log where every miss is open (recently) or fixed with a regression eval that would catch it again | `MISSES.md` + `evals/evals.json` |
|
|
28
28
|
| Runs on frontier intelligence | Its evals last passed on a current model, recently, and beat the same task run without the skill | `evals/results/latest.json` |
|
|
29
29
|
|
|
@@ -35,6 +35,7 @@ my-skill/
|
|
|
35
35
|
evals/triggers.json trigger evals (level: tested)
|
|
36
36
|
goldens/<id>/input.md approved examples (level: superskill)
|
|
37
37
|
goldens/<id>/output.md
|
|
38
|
+
goldens/<id>/PROVENANCE.json
|
|
38
39
|
goldens/<id>/APPROVAL.json
|
|
39
40
|
MISSES.md miss log (level: superskill)
|
|
40
41
|
evals/results/latest.json last --run (level: superskill)
|
|
@@ -91,13 +92,15 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
|
|
|
91
92
|
|---|---|---|
|
|
92
93
|
| `evals-present` | fail | `evals/evals.json` parses and there are at least 3 cases (each golden with an input and an output counts as one) |
|
|
93
94
|
| `evals-verifiable` | fail | every case in `evals.json` has a prompt and at least one expectation |
|
|
95
|
+
| `evals-real` | fail | no eval case and no trigger still holds a `superskill init` `REPLACE:` placeholder, no two cases share a prompt, and no two triggers share a query |
|
|
94
96
|
| `triggers-present` | fail | `evals/triggers.json` has at least 10 queries, at least 3 that should load the skill and at least 3 near-misses that should not |
|
|
95
97
|
|
|
96
98
|
### Level 3: superskill
|
|
97
99
|
|
|
98
100
|
| Rule | Severity | Threshold |
|
|
99
101
|
|---|---|---|
|
|
100
|
-
| `golden-approved` | fail | at least one golden has an approval with non-empty `approved_by` and a valid `approved_at
|
|
102
|
+
| `golden-approved` | fail | at least one golden **from a real run** (its `PROVENANCE.json` says `source: real-run`, names the run, and names who accepted it and when; an anonymized twin also carries `ANONYMIZED.json`) has an approval with non-empty `approved_by` and a valid `approved_at`. An approved golden without that provenance is named and does not count. Info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
|
|
103
|
+
| `real-use` | info | when the run ledger records what the person's next message made of each run (Freedom's `next_turn`), the share of those runs in the last 30 days they accepted; sandbox runs are never counted |
|
|
101
104
|
| `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
|
|
102
105
|
| `no-stale-open-miss` | fail | no miss has been open more than 14 days |
|
|
103
106
|
| `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
|
|
@@ -157,7 +160,7 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
|
|
|
157
160
|
- `input.md`: the real request.
|
|
158
161
|
- `output.md` (or any other file that is not `input.*` or `APPROVAL.json`): the output a person
|
|
159
162
|
said was right.
|
|
160
|
-
- `APPROVAL.json`, written only by `superskill approve
|
|
163
|
+
- `APPROVAL.json`, written only by `superskill approve`: at an interactive terminal, or relayed by an agent with `--approved-by` and `--via` after the person approved with a tap (the `via` field records where):
|
|
161
164
|
|
|
162
165
|
```json
|
|
163
166
|
{ "approvals": [
|
|
@@ -181,6 +184,38 @@ approval whose `note` is its rationale.
|
|
|
181
184
|
A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
|
|
182
185
|
output.
|
|
183
186
|
|
|
187
|
+
**`PROVENANCE.json` says where the example came from (0.5.0).** Only a golden from a real run
|
|
188
|
+
counts toward superskill. An invented input with an invented output puts "a person said this was
|
|
189
|
+
right" on something no person's work produced, and a skill could then reach the top level on its
|
|
190
|
+
author's fiction. Invented cases still belong in `evals/evals.json`, where they hold a skill at
|
|
191
|
+
`tested`.
|
|
192
|
+
|
|
193
|
+
```json
|
|
194
|
+
{
|
|
195
|
+
"source": "real-run",
|
|
196
|
+
"run": { "session": "88696e1b-...", "ledger_id": "inv_2026-09-28T02-28-29Z_ab12", "commit": null, "at": "2026-09-28T02:28:29Z" },
|
|
197
|
+
"accepted": { "by": "Ann Example", "at": "2026-09-28T03:22:51Z", "signal": "close", "evidence": "her next message after the run" }
|
|
198
|
+
}
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
- `source`: `real-run`, `synthetic`, or `synthetic-reconstruction`. Only `real-run` counts.
|
|
202
|
+
- `run`: at least one of `session`, `commit`, `ledger_id` names the run it came from.
|
|
203
|
+
- `accepted`: who accepted the output **when it ran** (`by`), and when (`at`). This is not the
|
|
204
|
+
approval: accepting is what the person did at the time (moved on, closed, said go); approving
|
|
205
|
+
is saying afterwards that this example is the standard. Both are required.
|
|
206
|
+
- An **anonymized twin** of a private golden says `anonymized: true` and `derived_from` (a hash of
|
|
207
|
+
the original, never its path or content) in place of the run, and carries `ANONYMIZED.json`, the
|
|
208
|
+
anonymizer's receipt (`checker`, `checked_at`, `counts` by kind, `fingerprint`, never the
|
|
209
|
+
mapping). Without the receipt the twin does not count.
|
|
210
|
+
- `superskill init --from-session` writes `source: real-run` with `accepted` empty, so the golden
|
|
211
|
+
cannot count until someone records who accepted it.
|
|
212
|
+
|
|
213
|
+
**Private goldens.** A real run's input and output are usually about real people and should not
|
|
214
|
+
travel with a skill that is shared. `--private-goldens <dir>` (or `SUPERSKILL_PRIVATE_GOLDENS`)
|
|
215
|
+
makes `doctor`, `approve` and `doctor --run` also read `<dir>/<skill-name>/<id>/`, laid out exactly
|
|
216
|
+
like `goldens/<id>/`. A private golden counts for the operator who holds it. A miss is closed only
|
|
217
|
+
by an eval or golden that ships with the skill.
|
|
218
|
+
|
|
184
219
|
### `MISSES.md`
|
|
185
220
|
|
|
186
221
|
```markdown
|
|
@@ -246,7 +281,12 @@ exist they are read as plain files:
|
|
|
246
281
|
`{id, skill, started, outcome, interventions: [{kind, note|what}], errors}`. A record with an
|
|
247
282
|
intervention of kind `redirect`, `correction` or `rescue`, or with `outcome: "failed"`, becomes
|
|
248
283
|
an open miss dated from `started`; `taste` interventions are skipped; the ledger id is kept on
|
|
249
|
-
a `Source:` line so a record is never imported twice.
|
|
284
|
+
a `Source:` line so a record is never imported twice. A record with `synthetic: true` (a
|
|
285
|
+
sandbox run) is skipped. A record corrected after it handed back (`corrected_after`, with
|
|
286
|
+
`next_turn_ref`) gets a "Should have" that points at the person's correction: the session and
|
|
287
|
+
the time, never their words.
|
|
288
|
+
- **`doctor --run` is never recorded as a use.** Every harness spawns its child with
|
|
289
|
+
`FREEDOM_SKILL_LEDGER=off` and `SUPERSKILL_SANDBOX=1`, in a folder named `superskill-run-*`.
|
|
250
290
|
|
|
251
291
|
## Versioning
|
|
252
292
|
|
package/package.json
CHANGED
package/src/commands/approve.mjs
CHANGED
|
@@ -5,17 +5,25 @@ import { createInterface } from "node:readline/promises";
|
|
|
5
5
|
import { parseArgs, clock, UsageError } from "../args.mjs";
|
|
6
6
|
import { skillDir, refuse } from "./common.mjs";
|
|
7
7
|
import { readGoldens, approvalEntry, withApproval, BASES } from "../goldens.mjs";
|
|
8
|
+
import { parseSkillFile } from "../frontmatter.mjs";
|
|
8
9
|
|
|
9
|
-
export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"]
|
|
10
|
+
export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"] [--private-goldens <dir>]
|
|
10
11
|
|
|
11
|
-
Record that a person checked goldens/<id>/ and signs off on its output.
|
|
12
|
-
|
|
12
|
+
Record that a person checked goldens/<id>/ and signs off on its output. At a terminal it asks
|
|
13
|
+
for your name. From an agent (a phone tap on a board or review page), pass --approved-by
|
|
14
|
+
"<the person>" and --via "<where they said yes>": both are required and both are recorded, so an
|
|
15
|
+
approval an agent relayed is always distinguishable from one typed at a terminal.
|
|
13
16
|
|
|
14
17
|
Every approval says WHY (--rationale, or asked) and WHAT IT RESTS ON (--basis, or asked):
|
|
15
18
|
judgment you read it and it is right (the default)
|
|
16
19
|
outcome it produced a result someone can check; --evidence is required
|
|
17
20
|
Approvals accumulate: two people approving, or a judgment approval later backed by an outcome,
|
|
18
21
|
all stay on the record. Writes goldens/<id>/APPROVAL.json with the SKILL.md hash.
|
|
22
|
+
|
|
23
|
+
--private-goldens <dir> (or SUPERSKILL_PRIVATE_GOLDENS) also looks for the golden in
|
|
24
|
+
<dir>/<skill-name>/<id>/, where an operator keeps goldens made from real runs that are too
|
|
25
|
+
personal to travel with the skill. Approving records the approval; it never changes where a
|
|
26
|
+
golden came from (PROVENANCE.json), and only a golden from a real run counts for superskill.
|
|
19
27
|
`;
|
|
20
28
|
|
|
21
29
|
export async function run(argv) {
|
|
@@ -24,10 +32,29 @@ export async function run(argv) {
|
|
|
24
32
|
const dir = skillDir(a._[0], "approve");
|
|
25
33
|
const id = a._[1];
|
|
26
34
|
if (!id) throw new UsageError("approve needs a golden id");
|
|
27
|
-
|
|
28
|
-
const
|
|
29
|
-
|
|
35
|
+
const name = parseSkillFile(readFileSync(join(dir, "SKILL.md"), "utf8")).data.name || "";
|
|
36
|
+
const privateGoldens = a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null;
|
|
37
|
+
const found = readGoldens(dir, { privateGoldens, name }).filter((x) => x.id === id);
|
|
38
|
+
// The private copy wins when both exist: it is the one the operator was shown.
|
|
39
|
+
const g = found.find((x) => x.private) || found[0];
|
|
40
|
+
if (!g) return refuse(`no goldens/${id}/${privateGoldens ? ` (nor ${join(privateGoldens, name, id)})` : ""}`);
|
|
30
41
|
if (g.input === null || g.output === null || !g.output.trim()) return refuse(`goldens/${id}/ needs an input file and a non-empty output file before it can be approved`);
|
|
42
|
+
if (!(process.stdin.isTTY && process.stdout.isTTY)) {
|
|
43
|
+
// Mobile first: a person approves with a tap (a board, a review page) and an agent records it.
|
|
44
|
+
// The record names the person AND the channel, so an approval relayed by an agent is never
|
|
45
|
+
// mistaken for one typed at a terminal. Without both, an agent could approve its own output.
|
|
46
|
+
const name = String(a.flags["approved-by"] || "").trim();
|
|
47
|
+
const via = String(a.flags.via || "").trim();
|
|
48
|
+
if (!name || !via) return refuse("approval needs a person. At a terminal it asks for your name; from an agent, pass --approved-by \"<the person>\" and --via \"<where they approved: the tap, the page, the message>\", after that person said yes");
|
|
49
|
+
const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
|
|
50
|
+
const made = approvalEntry({ name, rationale: a.flags.rationale || a.flags.note, basis: a.flags.basis || "judgment", evidence: a.flags.evidence, at: clock(a.flags).toISOString(), sha });
|
|
51
|
+
if (made.error) return refuse(made.error);
|
|
52
|
+
const p = join(g.dir, "APPROVAL.json");
|
|
53
|
+
const existed = existsSync(p);
|
|
54
|
+
writeFileSync(p, JSON.stringify(withApproval(existed ? g.approval : null, { ...made.entry, via }), null, 2) + "\n");
|
|
55
|
+
process.stdout.write(`${existed ? "added an approval to" : "approved"} goldens/${id}/ by ${name} (${made.entry.basis}, via ${via})\n`);
|
|
56
|
+
return 0;
|
|
57
|
+
}
|
|
31
58
|
const rl = createInterface({ input: process.stdin, output: process.stdout });
|
|
32
59
|
try {
|
|
33
60
|
process.stdout.write(`\n--- goldens/${id}/${g.outputFile} ---\n${g.output.slice(0, 2000)}${g.output.length > 2000 ? "\n[...]" : ""}\n---\n`);
|
|
@@ -40,7 +67,7 @@ export async function run(argv) {
|
|
|
40
67
|
const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
|
|
41
68
|
const made = approvalEntry({ name, rationale, basis, evidence, at: clock(a.flags).toISOString(), sha });
|
|
42
69
|
if (made.error) return refuse(made.error);
|
|
43
|
-
const p = join(dir, "
|
|
70
|
+
const p = join(g.dir, "APPROVAL.json");
|
|
44
71
|
const existed = existsSync(p);
|
|
45
72
|
const prior = existed ? g.approval : null;
|
|
46
73
|
writeFileSync(p, JSON.stringify(withApproval(prior, made.entry), null, 2) + "\n");
|
package/src/commands/doctor.mjs
CHANGED
|
@@ -8,8 +8,9 @@ import { changedFiles, touchedSkills, levelDrops } from "../changed.mjs";
|
|
|
8
8
|
export const help = `superskill doctor <path...> [options]
|
|
9
9
|
|
|
10
10
|
Score one skill, a folder of skills, or a plugin (skills/*/SKILL.md).
|
|
11
|
-
Levels: skill (spec-valid, hygienic) < tested (evals + triggers) < superskill
|
|
12
|
-
(approved golden, misses fixed with evals, a fresh
|
|
11
|
+
Levels: skill (spec-valid, hygienic) < tested (real evals + triggers) < superskill
|
|
12
|
+
(an approved golden from a real run a person accepted, misses fixed with evals, a fresh
|
|
13
|
+
--run that beats no-skill).
|
|
13
14
|
|
|
14
15
|
Options:
|
|
15
16
|
--level <skill|tested|superskill> target level for the exit code (default skill)
|
|
@@ -20,6 +21,8 @@ Options:
|
|
|
20
21
|
--baseline-json <file> a previous --json; exit 1 if any skill's level dropped
|
|
21
22
|
--run run the evals with and without the skill (costs
|
|
22
23
|
model calls; see "superskill doctor --run --help")
|
|
24
|
+
--private-goldens <dir> also read goldens from <dir>/<skill-name>/<id>/ (or set
|
|
25
|
+
SUPERSKILL_PRIVATE_GOLDENS): real runs kept out of the skill
|
|
23
26
|
--now <iso date> evaluate dates as of this moment
|
|
24
27
|
--help this text
|
|
25
28
|
|
|
@@ -34,7 +37,7 @@ export async function run(argv) {
|
|
|
34
37
|
const level = a.flags.level || "skill";
|
|
35
38
|
if (!LEVELS.includes(level)) throw new UsageError(`--level must be one of ${LEVELS.join(", ")}`);
|
|
36
39
|
if (!a._.length) throw new UsageError("doctor needs a path");
|
|
37
|
-
const opts = { level, now: clock(a.flags) };
|
|
40
|
+
const opts = { level, now: clock(a.flags), privateGoldens: a.flags["private-goldens"] || process.env.SUPERSKILL_PRIVATE_GOLDENS || null };
|
|
38
41
|
if (a.flags.changed) {
|
|
39
42
|
const all = a._.flatMap((p) => findSkills(p));
|
|
40
43
|
const files = a._.flatMap((p) => changedFiles(p, a.flags.base));
|
package/src/commands/init.mjs
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { existsSync, mkdirSync, writeFileSync, readFileSync } from "node:fs";
|
|
2
|
-
import { join } from "node:path";
|
|
2
|
+
import { join, basename } from "node:path";
|
|
3
3
|
import { parseArgs, UsageError } from "../args.mjs";
|
|
4
4
|
import { skillDir } from "./common.mjs";
|
|
5
5
|
import { MISSES_HEADER } from "../misses.mjs";
|
|
@@ -15,7 +15,9 @@ Never overwrites a file that exists.
|
|
|
15
15
|
--from-session <file> build the first eval and golden candidate from the session where
|
|
16
16
|
the job was done by hand. A Claude Code .jsonl gives the first
|
|
17
17
|
request and the final answer; any other file is the request.
|
|
18
|
-
The golden waits for a person:
|
|
18
|
+
The golden waits for a person: fill in accepted.by and
|
|
19
|
+
accepted.at in its PROVENANCE.json (who accepted the
|
|
20
|
+
output when it ran), then \`superskill approve\`.
|
|
19
21
|
`;
|
|
20
22
|
|
|
21
23
|
const TRIGGER_EXAMPLE = [
|
|
@@ -65,6 +67,10 @@ export async function run(argv) {
|
|
|
65
67
|
writeFileSync(evalsPath, JSON.stringify(doc, null, 2) + "\n");
|
|
66
68
|
put(`goldens/${id}/input.md`, sessionCase.prompt + "\n");
|
|
67
69
|
put(`goldens/${id}/output.md`, sessionCase.output ? sessionCase.output + "\n" : "");
|
|
70
|
+
// A real run, but nobody has said yet that its output was accepted. Until accepted.by and
|
|
71
|
+
// accepted.at are filled in (by the person, or by a tool that read their next message), this
|
|
72
|
+
// golden cannot count toward superskill however it is approved.
|
|
73
|
+
put(`goldens/${id}/PROVENANCE.json`, JSON.stringify({ source: "real-run", run: { session: basename(session), ledger_id: null, commit: null, at: null }, accepted: { by: null, at: null, signal: null, evidence: "" } }, null, 2) + "\n");
|
|
68
74
|
process.stdout.write(`eval ${id} and golden candidate goldens/${id}/ written from ${session}\n`);
|
|
69
75
|
process.stdout.write(`a person approves it with: superskill approve ${a._[0]} ${id}\n`);
|
|
70
76
|
}
|
package/src/doctor.mjs
CHANGED
|
@@ -37,6 +37,7 @@ export function findSkills(path) {
|
|
|
37
37
|
/** Score one loaded skill. */
|
|
38
38
|
export function scoreSkill(dir, opts = {}) {
|
|
39
39
|
const ctx = loadSkill(dir);
|
|
40
|
+
if (opts.privateGoldens) ctx.privateGoldens = opts.privateGoldens;
|
|
40
41
|
const findings = runRules(ctx, opts.rules || allRules, opts);
|
|
41
42
|
const level = computeLevel(findings);
|
|
42
43
|
const up = nextLevel(level);
|
package/src/goldens.mjs
CHANGED
|
@@ -1,38 +1,90 @@
|
|
|
1
1
|
// goldens/<id>/: input.md, the approved output (output.md, or any other non-input file),
|
|
2
|
-
//
|
|
3
|
-
// rationale, basis, evidence}] }
|
|
2
|
+
// APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
|
|
3
|
+
// rationale, basis, evidence}] } (or the single-approval shape from before 0.3.0),
|
|
4
|
+
// PROVENANCE.json saying where the example came from (0.5.0), and, for an anonymized twin of a
|
|
5
|
+
// private golden, ANONYMIZED.json, the anonymizer's receipt.
|
|
6
|
+
//
|
|
7
|
+
// A golden may also live OUTSIDE the skill, in a private folder the operator keeps
|
|
8
|
+
// (<private>/<skill-name>/<id>/), because a real run's input and output are usually about real
|
|
9
|
+
// people and should not travel with a skill that is shared. The doctor reads both.
|
|
4
10
|
import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
|
|
5
11
|
import { join } from "node:path";
|
|
6
12
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
13
|
+
/** Files in a golden folder that describe it rather than being its input or output. */
|
|
14
|
+
export const META_FILES = new Set(["APPROVAL.json", "PROVENANCE.json", "ANONYMIZED.json"]);
|
|
15
|
+
|
|
16
|
+
/** Where goldens are read from: the skill's own goldens/, then the private folder for its name. */
|
|
17
|
+
export function goldenRoots(dir, { privateGoldens = null, name = null } = {}) {
|
|
18
|
+
const roots = [{ root: join(dir, "goldens"), private: false }];
|
|
19
|
+
if (privateGoldens && name) roots.push({ root: join(privateGoldens, name), private: true });
|
|
20
|
+
return roots;
|
|
21
|
+
}
|
|
22
|
+
|
|
23
|
+
const readJsonFile = (p) => {
|
|
24
|
+
if (!existsSync(p)) return { value: null, error: null };
|
|
25
|
+
try { return { value: JSON.parse(readFileSync(p, "utf8")), error: null }; }
|
|
26
|
+
catch (e) { return { value: null, error: e.message }; }
|
|
27
|
+
};
|
|
28
|
+
|
|
29
|
+
export function readGoldens(dir, opts = {}) {
|
|
10
30
|
const out = [];
|
|
11
|
-
for (const
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
31
|
+
for (const { root, private: priv } of goldenRoots(dir, opts)) {
|
|
32
|
+
if (!existsSync(root)) continue;
|
|
33
|
+
for (const id of readdirSync(root).sort()) {
|
|
34
|
+
const gdir = join(root, id);
|
|
35
|
+
try { if (!statSync(gdir).isDirectory()) continue; } catch { continue; }
|
|
36
|
+
const files = readdirSync(gdir).filter((f) => { try { return statSync(join(gdir, f)).isFile(); } catch { return false; } });
|
|
37
|
+
const inputName = files.find((f) => /^input\./i.test(f));
|
|
38
|
+
const outputName = files.find((f) => /^output\./i.test(f)) || files.find((f) => f !== inputName && !META_FILES.has(f) && !f.startsWith("."));
|
|
39
|
+
const appr = readJsonFile(join(gdir, "APPROVAL.json"));
|
|
40
|
+
const prov = readJsonFile(join(gdir, "PROVENANCE.json"));
|
|
41
|
+
const anon = readJsonFile(join(gdir, "ANONYMIZED.json"));
|
|
42
|
+
const provenance = prov.value;
|
|
43
|
+
out.push({
|
|
44
|
+
id,
|
|
45
|
+
dir: gdir,
|
|
46
|
+
private: priv,
|
|
47
|
+
input: inputName ? readFileSync(join(gdir, inputName), "utf8") : null,
|
|
48
|
+
outputFile: outputName || null,
|
|
49
|
+
output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
|
|
50
|
+
approval: appr.value,
|
|
51
|
+
approvals: approvalsOf(appr.value),
|
|
52
|
+
approvalError: appr.error,
|
|
53
|
+
provenance,
|
|
54
|
+
provenanceError: prov.error,
|
|
55
|
+
anonymized: anon.value,
|
|
56
|
+
origin: originOf(provenance, anon.value),
|
|
57
|
+
});
|
|
21
58
|
}
|
|
22
|
-
out.push({
|
|
23
|
-
id,
|
|
24
|
-
dir: gdir,
|
|
25
|
-
input: inputName ? readFileSync(join(gdir, inputName), "utf8") : null,
|
|
26
|
-
outputFile: outputName || null,
|
|
27
|
-
output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
|
|
28
|
-
approval,
|
|
29
|
-
approvals: approvalsOf(approval),
|
|
30
|
-
approvalError,
|
|
31
|
-
});
|
|
32
59
|
}
|
|
33
60
|
return out;
|
|
34
61
|
}
|
|
35
62
|
|
|
63
|
+
/** Where a golden's example came from. `real-run` is the only source that can reach superskill. */
|
|
64
|
+
export const SOURCES = ["real-run", "synthetic", "synthetic-reconstruction"];
|
|
65
|
+
|
|
66
|
+
/**
|
|
67
|
+
* Is this golden a real run a person accepted, and if not, why not?
|
|
68
|
+
*
|
|
69
|
+
* A real run names the run it came from (a session, a commit, a ledger id, or, for an
|
|
70
|
+
* anonymized twin, `derived_from`: a hash of the private original, never its content) and the
|
|
71
|
+
* person who accepted the output when it happened, and when. An anonymized twin must also carry
|
|
72
|
+
* the anonymizer's receipt, because "the original was accepted" and "this twin still says the
|
|
73
|
+
* same thing" are different claims.
|
|
74
|
+
*/
|
|
75
|
+
export function originOf(prov, receipt = null) {
|
|
76
|
+
if (!prov || typeof prov !== "object") return { real: false, why: "no PROVENANCE.json" };
|
|
77
|
+
if (prov.source !== "real-run") return { real: false, why: `source is ${JSON.stringify(prov.source ?? null)}, not "real-run"` };
|
|
78
|
+
const run = prov.run && typeof prov.run === "object" ? prov.run : {};
|
|
79
|
+
const ref = ["session", "commit", "ledger_id"].map((k) => run[k]).concat(prov.derived_from).find((x) => typeof x === "string" && x.trim());
|
|
80
|
+
if (!ref) return { real: false, why: "names no run (run.session, run.commit, run.ledger_id or derived_from)" };
|
|
81
|
+
const acc = prov.accepted && typeof prov.accepted === "object" ? prov.accepted : {};
|
|
82
|
+
if (!(typeof acc.by === "string" && acc.by.trim())) return { real: false, why: "names no person who accepted the run (accepted.by)" };
|
|
83
|
+
if (!validDate(acc.at)) return { real: false, why: "has no valid accepted.at" };
|
|
84
|
+
if (prov.anonymized === true && !(receipt && typeof receipt === "object" && receipt.fingerprint)) return { real: false, why: "is an anonymized twin with no ANONYMIZED.json receipt" };
|
|
85
|
+
return { real: true, why: "" };
|
|
86
|
+
}
|
|
87
|
+
|
|
36
88
|
/** What an approval rests on. `judgment`: the people who approved it read it and said it is right.
|
|
37
89
|
* `outcome`: it produced a result in the world someone can check (a client landed, a call booked).
|
|
38
90
|
* Being liked and being proven are different weights, and a golden says which it carries. */
|
|
@@ -54,6 +106,7 @@ export function approvalsOf(approval) {
|
|
|
54
106
|
rationale: String(a.rationale ?? a.note ?? "").trim(),
|
|
55
107
|
basis: a.basis === "outcome" ? "outcome" : "judgment",
|
|
56
108
|
evidence: String(a.evidence ?? "").trim(),
|
|
109
|
+
via: String(a.via ?? "").trim(),
|
|
57
110
|
}));
|
|
58
111
|
}
|
|
59
112
|
|
|
@@ -81,3 +134,12 @@ export function withApproval(existing, entry) {
|
|
|
81
134
|
}
|
|
82
135
|
|
|
83
136
|
export const isApproved = (g) => approvalsOf(g.approval).length > 0;
|
|
137
|
+
|
|
138
|
+
/** Approved AND from a real run a person accepted: the only golden that counts for superskill. */
|
|
139
|
+
export const isRealApproved = (g) => isApproved(g) && (g.origin || originOf(g.provenance, g.anonymized)).real;
|
|
140
|
+
|
|
141
|
+
/** The options readGoldens needs for a loaded skill: its name and the private folder, if any. */
|
|
142
|
+
export const goldenOpts = (ctx, opts = {}) => ({
|
|
143
|
+
name: (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName,
|
|
144
|
+
privateGoldens: opts.privateGoldens ?? ctx.privateGoldens ?? process.env.SUPERSKILL_PRIVATE_GOLDENS ?? null,
|
|
145
|
+
});
|
package/src/ledger.mjs
CHANGED
|
@@ -12,7 +12,10 @@ export function ledgerPaths(skillDir, skillName) {
|
|
|
12
12
|
const out = [];
|
|
13
13
|
const local = join(skillDir, "invocations.jsonl");
|
|
14
14
|
if (existsSync(local)) out.push(local);
|
|
15
|
-
|
|
15
|
+
// FREEDOM_SKILL_LEDGER_HOME is where Freedom itself writes when it is re-pointed (tests, a
|
|
16
|
+
// second profile); read the same place it writes.
|
|
17
|
+
const home = process.env.FREEDOM_SKILL_LEDGER_HOME || join(homedir(), ".freedom", "ledger");
|
|
18
|
+
const root = join(home, "skills");
|
|
16
19
|
if (existsSync(root) && skillName) {
|
|
17
20
|
for (const plugin of readdirSync(root).sort()) {
|
|
18
21
|
const p = join(root, plugin, `${skillName}.jsonl`);
|
|
@@ -30,15 +33,48 @@ export function ledgerMisses(path, known, skillName) {
|
|
|
30
33
|
let rec;
|
|
31
34
|
try { rec = JSON.parse(line); } catch { continue; }
|
|
32
35
|
if (!rec || !rec.id || known.has(rec.id)) continue;
|
|
33
|
-
|
|
36
|
+
// A sandbox run (superskill doctor --run) is a test of the skill, not a use of it.
|
|
37
|
+
if (rec.synthetic) continue;
|
|
38
|
+
if (skillName && rec.skill && bare(rec.skill) !== bare(skillName)) continue;
|
|
34
39
|
const kinds = (Array.isArray(rec.interventions) ? rec.interventions : []).filter((i) => i && MISS_KINDS.has(i.kind));
|
|
35
40
|
const failed = rec.outcome === "failed";
|
|
36
41
|
if (!kinds.length && !failed) continue;
|
|
37
42
|
const parts = kinds.map((i) => `${i.kind}${i.note || i.what ? `: ${i.note || i.what}` : ""}`);
|
|
38
|
-
if (failed) parts.unshift(`run failed${Array.isArray(rec.errors) && rec.errors.length ? ` (${rec.errors.slice(0, 2).join("; ")})` : ""}`);
|
|
43
|
+
if (failed) parts.unshift(`run failed${Array.isArray(rec.errors) && rec.errors.length ? ` (${rec.errors.slice(0, 2).map((e) => (typeof e === "string" ? e : e?.kind || "error")).join("; ")})` : ""}`);
|
|
39
44
|
const date = typeof rec.started === "string" && /^\d{4}-\d{2}-\d{2}/.test(rec.started) ? rec.started.slice(0, 10) : null;
|
|
40
|
-
|
|
45
|
+
// A correction made in the operator's NEXT message is the best "should have" there is, and the
|
|
46
|
+
// ledger keeps only where it is, never its words: point at it, so the fixer reads it there.
|
|
47
|
+
const ref = rec.next_turn_ref && rec.session_id ? ` (the correction is the operator's message in session ${String(rec.session_id).slice(0, 8)} at ${rec.next_turn_ref.at || "?"}${Number.isFinite(rec.next_turn_ref.offset) ? `, transcript byte ${rec.next_turn_ref.offset}` : ""})` : "";
|
|
48
|
+
const expected = rec.corrected_after ? `what the operator asked for instead${ref}` : "";
|
|
49
|
+
out.push({ date, status: "open", what: parts.join("; ").replace(/\s+/g, " ").slice(0, 300), expected, source: `freedom-ledger ${rec.id}` });
|
|
41
50
|
known.add(rec.id);
|
|
42
51
|
}
|
|
43
52
|
return out;
|
|
44
53
|
}
|
|
54
|
+
|
|
55
|
+
const bare = (name) => String(name || "").split(":").pop();
|
|
56
|
+
|
|
57
|
+
/**
|
|
58
|
+
* How many runs the person accepted, from Freedom's `next_turn` verdict (the class of the first
|
|
59
|
+
* message after a run handed back). Runs with no verdict are left out of both sides, so a missing
|
|
60
|
+
* record can never read as an acceptance, and sandbox runs never count.
|
|
61
|
+
*/
|
|
62
|
+
export function acceptedRate(paths, { now = new Date(), days = 30 } = {}) {
|
|
63
|
+
const since = now.getTime() - days * 86400000;
|
|
64
|
+
let judged = 0, accepted = 0, synthetic = 0;
|
|
65
|
+
for (const p of paths) {
|
|
66
|
+
let text = "";
|
|
67
|
+
try { text = readFileSync(p, "utf8"); } catch { continue; }
|
|
68
|
+
for (const line of text.split("\n")) {
|
|
69
|
+
if (!line.trim()) continue;
|
|
70
|
+
let rec; try { rec = JSON.parse(line); } catch { continue; }
|
|
71
|
+
const t = Date.parse(rec?.started || "");
|
|
72
|
+
if (!Number.isFinite(t) || t < since || t > now.getTime()) continue;
|
|
73
|
+
if (rec.synthetic) { synthetic++; continue; }
|
|
74
|
+
if (!rec.next_turn || rec.next_turn === "none") continue;
|
|
75
|
+
judged++;
|
|
76
|
+
if (["close", "go", "new_topic"].includes(rec.next_turn) && !rec.corrected_after && !["failed", "abandoned"].includes(rec.outcome)) accepted++;
|
|
77
|
+
}
|
|
78
|
+
}
|
|
79
|
+
return { judged, accepted, synthetic };
|
|
80
|
+
}
|
package/src/rules/superskill.mjs
CHANGED
|
@@ -2,13 +2,15 @@
|
|
|
2
2
|
// got something wrong, and proven recently on a current model against the no-skill baseline.
|
|
3
3
|
import { createHash } from "node:crypto";
|
|
4
4
|
import { defineRules } from "./define.mjs";
|
|
5
|
-
import { readGoldens, isApproved, weightOf } from "../goldens.mjs";
|
|
5
|
+
import { readGoldens, isApproved, isRealApproved, weightOf, goldenOpts } from "../goldens.mjs";
|
|
6
|
+
import { ledgerPaths, acceptedRate } from "../ledger.mjs";
|
|
6
7
|
import { readMisses } from "../misses.mjs";
|
|
7
8
|
import { readEvals, readLatestRun } from "../evals.mjs";
|
|
8
9
|
|
|
9
10
|
const f = (severity, message, fix) => ({ severity, message, fix });
|
|
10
11
|
const DAY = 86400000;
|
|
11
12
|
export const OPEN_MISS_DAYS = 14;
|
|
13
|
+
export const REAL_USE_DAYS = 30;
|
|
12
14
|
/** How long a --run stays fresh, by metadata.cadence. */
|
|
13
15
|
export const FRESH_DAYS = { daily: 30, weekly: 30, monthly: 60, quarterly: 120, yearly: 365 };
|
|
14
16
|
const DEFAULT_FRESH = 30;
|
|
@@ -19,15 +21,27 @@ export const superskillRules = defineRules([
|
|
|
19
21
|
{
|
|
20
22
|
id: "golden-approved",
|
|
21
23
|
level: "superskill",
|
|
22
|
-
check(ctx) {
|
|
24
|
+
check(ctx, opts = {}) {
|
|
23
25
|
if (ctx.error) return [];
|
|
24
|
-
const goldens = readGoldens(ctx.dir);
|
|
25
|
-
const
|
|
26
|
+
const goldens = readGoldens(ctx.dir, goldenOpts(ctx, opts));
|
|
27
|
+
const approvedAny = goldens.filter(isApproved);
|
|
28
|
+
// 0.5.0: only a golden from a real run a person accepted counts. An invented example, however
|
|
29
|
+
// careful, puts "a person said this was right" on something no person's work produced, and a
|
|
30
|
+
// skill could then reach the top level on its author's fiction.
|
|
31
|
+
const approved = goldens.filter(isRealApproved);
|
|
26
32
|
const out = [];
|
|
33
|
+
const where = (g) => (g.private ? `private golden ${g.id}` : `goldens/${g.id}`);
|
|
27
34
|
for (const g of goldens.filter((x) => x.approvalError))
|
|
28
|
-
out.push(f("fail",
|
|
35
|
+
out.push(f("fail", `${where(g)}/APPROVAL.json is not valid JSON`, `Re-record it with \`superskill approve . ${g.id}\`.`));
|
|
36
|
+
for (const g of goldens.filter((x) => x.provenanceError))
|
|
37
|
+
out.push(f("fail", `${where(g)}/PROVENANCE.json is not valid JSON`, "Rewrite it: source, run, accepted (see SPEC.md, goldens)."));
|
|
29
38
|
if (!approved.length) {
|
|
30
|
-
|
|
39
|
+
if (approvedAny.length) {
|
|
40
|
+
const why = approvedAny.map((g) => `${g.id} ${g.origin.why}`).join("; ");
|
|
41
|
+
out.push(f("fail", `${approvedAny.length} approved golden${approvedAny.length === 1 ? "" : "s"}, none from a real run a person accepted (${why})`, "A golden counts toward superskill only when goldens/<id>/PROVENANCE.json says source: real-run, names the run (session, commit, ledger id, or derived_from for an anonymized twin) and who accepted it and when. Invented examples belong in evals/evals.json, where they hold the skill at tested."));
|
|
42
|
+
} else {
|
|
43
|
+
out.push(f("fail", goldens.length ? `${goldens.length} golden${goldens.length === 1 ? "" : "s"}, none approved by a person` : "no goldens", goldens.length ? "A person runs `superskill approve <skill> <golden>` at a terminal (or relays a tap with --approved-by and --via) after checking the output." : "Save a real run's input and the output a person accepted under goldens/<id>/ with PROVENANCE.json, then `superskill approve`."));
|
|
44
|
+
}
|
|
31
45
|
return out;
|
|
32
46
|
}
|
|
33
47
|
const sha = createHash("sha256").update(ctx.raw).digest("hex");
|
|
@@ -43,6 +57,19 @@ export const superskillRules = defineRules([
|
|
|
43
57
|
return out;
|
|
44
58
|
},
|
|
45
59
|
},
|
|
60
|
+
{
|
|
61
|
+
// The real number beside the level: of the runs whose next message is recorded, how many did
|
|
62
|
+
// the person accept. Info only. A level is evidence about examples; this is evidence about use.
|
|
63
|
+
id: "real-use",
|
|
64
|
+
level: "superskill",
|
|
65
|
+
check(ctx, { now = new Date() } = {}) {
|
|
66
|
+
if (ctx.error) return [];
|
|
67
|
+
const name = (typeof ctx.data?.name === "string" && ctx.data.name) || ctx.folderName;
|
|
68
|
+
const r = acceptedRate(ledgerPaths(ctx.dir, name), { now, days: REAL_USE_DAYS });
|
|
69
|
+
if (!r.judged) return [];
|
|
70
|
+
return [f("info", `real use, last ${REAL_USE_DAYS} days: ${r.accepted} of ${r.judged} judged runs accepted (${pct(r.accepted / r.judged)})${r.synthetic ? `; ${r.synthetic} sandbox run${r.synthetic === 1 ? "" : "s"} not counted` : ""}`, "")];
|
|
71
|
+
},
|
|
72
|
+
},
|
|
46
73
|
{
|
|
47
74
|
id: "misses-log-present",
|
|
48
75
|
level: "superskill",
|
|
@@ -74,6 +101,7 @@ export const superskillRules = defineRules([
|
|
|
74
101
|
const misses = (readMisses(ctx.dir) || []).filter((m) => m.status === "fixed");
|
|
75
102
|
if (!misses.length) return [];
|
|
76
103
|
const ids = new Set(readEvals(ctx.dir).cases.map((c) => String(c.id)));
|
|
104
|
+
// Shipped goldens only: a miss is closed by a check that travels with the skill.
|
|
77
105
|
for (const g of readGoldens(ctx.dir)) ids.add(g.id);
|
|
78
106
|
return misses
|
|
79
107
|
.filter((m) => !m.eval || !ids.has(String(m.eval)))
|
package/src/rules/tested.mjs
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
// with both should-load requests and near-misses that should not load the skill.
|
|
3
3
|
import { defineRules } from "./define.mjs";
|
|
4
4
|
import { readEvals, readTriggers } from "../evals.mjs";
|
|
5
|
-
import { readGoldens } from "../goldens.mjs";
|
|
5
|
+
import { readGoldens, goldenOpts } from "../goldens.mjs";
|
|
6
6
|
|
|
7
7
|
const f = (severity, message, fix) => ({ severity, message, fix });
|
|
8
8
|
export const MIN_EVALS = 3, MIN_TRIGGERS = 10, MIN_EACH_SIDE = 3;
|
|
@@ -15,7 +15,7 @@ export const testedRules = defineRules([
|
|
|
15
15
|
if (ctx.error) return [];
|
|
16
16
|
const e = readEvals(ctx.dir);
|
|
17
17
|
if (e.error) return [f("fail", e.error, "Fix evals/evals.json so it parses as skill-creator's {skill_name, evals: [...]}.")];
|
|
18
|
-
const goldenCases = readGoldens(ctx.dir).filter((g) => g.input !== null && g.output !== null).length;
|
|
18
|
+
const goldenCases = readGoldens(ctx.dir, goldenOpts(ctx)).filter((g) => g.input !== null && g.output !== null).length;
|
|
19
19
|
const n = e.cases.length + goldenCases;
|
|
20
20
|
if (n < MIN_EVALS)
|
|
21
21
|
return [f("fail", `${n} eval case${n === 1 ? "" : "s"} (need ${MIN_EVALS})`, "Add real requests to evals/evals.json (`superskill init` writes an example).")];
|
|
@@ -35,6 +35,29 @@ export const testedRules = defineRules([
|
|
|
35
35
|
: [];
|
|
36
36
|
},
|
|
37
37
|
},
|
|
38
|
+
{
|
|
39
|
+
// MEASURED 2026-10-05 (freedom-dev): `superskill init` writes one example eval and two example
|
|
40
|
+
// triggers, each starting "REPLACE:", and a skill whose author copied those up to 3 and 10
|
|
41
|
+
// scored `tested` while testing nothing. A placeholder, or the same request twice, is not a case.
|
|
42
|
+
id: "evals-real",
|
|
43
|
+
level: "tested",
|
|
44
|
+
check(ctx) {
|
|
45
|
+
if (ctx.error) return [];
|
|
46
|
+
const e = readEvals(ctx.dir), t = readTriggers(ctx.dir);
|
|
47
|
+
const out = [];
|
|
48
|
+
const placeholder = (x) => /\bREPLACE:/.test(JSON.stringify(x));
|
|
49
|
+
const evalHits = e.cases.filter((c) => placeholder([c.prompt, c.expected_output, c.assertions])).map((c) => c.id);
|
|
50
|
+
if (evalHits.length) out.push(f("fail", `eval case${evalHits.length === 1 ? "" : "s"} ${evalHits.join(", ")} still hold${evalHits.length === 1 ? "s" : ""} a \`superskill init\` REPLACE: placeholder`, "Replace every example with a real request this skill handles and what a good answer must contain."));
|
|
51
|
+
const trigHits = t.triggers.filter((x) => placeholder(x.query)).length;
|
|
52
|
+
if (trigHits) out.push(f("fail", `${trigHits} trigger${trigHits === 1 ? "" : "s"} still hold a REPLACE: placeholder`, "Replace them with realistic requests, including near-misses that should not load the skill."));
|
|
53
|
+
const dup = (list) => list.filter((x, i) => x && list.indexOf(x) !== i);
|
|
54
|
+
const dp = [...new Set(dup(e.cases.map((c) => c.prompt.trim())))];
|
|
55
|
+
if (dp.length) out.push(f("fail", `${dp.length} eval prompt${dp.length === 1 ? " is" : "s are"} repeated: the same request twice is one case`, "Make each case a different request."));
|
|
56
|
+
const dq = [...new Set(dup(t.triggers.map((x) => x.query.trim().toLowerCase())))];
|
|
57
|
+
if (dq.length) out.push(f("fail", `${dq.length} trigger quer${dq.length === 1 ? "y is" : "ies are"} repeated`, "Make each trigger a different request."));
|
|
58
|
+
return out;
|
|
59
|
+
},
|
|
60
|
+
},
|
|
38
61
|
{
|
|
39
62
|
id: "triggers-present",
|
|
40
63
|
level: "tested",
|
package/src/run/claude.mjs
CHANGED
|
@@ -3,12 +3,13 @@
|
|
|
3
3
|
// a copy installed at user level cannot leak into the baseline.
|
|
4
4
|
import { spawnSync } from "node:child_process";
|
|
5
5
|
import { prepareWorkspace } from "./workspace.mjs";
|
|
6
|
+
import { sandboxEnv } from "./env.mjs";
|
|
6
7
|
|
|
7
8
|
export const name = "claude";
|
|
8
9
|
|
|
9
10
|
function call(args, cwd) {
|
|
10
11
|
const t0 = Date.now();
|
|
11
|
-
const r = spawnSync("claude", args, { cwd, encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
12
|
+
const r = spawnSync("claude", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
12
13
|
const ms = Date.now() - t0;
|
|
13
14
|
if (r.error) throw Object.assign(new Error(`claude failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
|
|
14
15
|
let doc = null;
|
package/src/run/codex.mjs
CHANGED
|
@@ -5,6 +5,7 @@ import { spawnSync } from "node:child_process";
|
|
|
5
5
|
import { readFileSync, existsSync } from "node:fs";
|
|
6
6
|
import { join } from "node:path";
|
|
7
7
|
import { prepareWorkspace } from "./workspace.mjs";
|
|
8
|
+
import { sandboxEnv } from "./env.mjs";
|
|
8
9
|
|
|
9
10
|
export const name = "codex";
|
|
10
11
|
|
|
@@ -14,7 +15,7 @@ function call(prompt, cwd, model) {
|
|
|
14
15
|
if (model) args.push("-m", model);
|
|
15
16
|
args.push(prompt);
|
|
16
17
|
const t0 = Date.now();
|
|
17
|
-
const r = spawnSync("codex", args, { cwd, encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
18
|
+
const r = spawnSync("codex", args, { cwd, env: sandboxEnv(), encoding: "utf8", maxBuffer: 64 * 1024 * 1024, timeout: 15 * 60 * 1000 });
|
|
18
19
|
if (r.error) throw Object.assign(new Error(`codex failed to start: ${r.error.message}`), { code: "SUPERSKILL" });
|
|
19
20
|
const output = existsSync(out) ? readFileSync(out, "utf8") : r.stdout;
|
|
20
21
|
const tok = (r.stderr + r.stdout).match(/tokens used[:\s]+([\d,]+)/i);
|
package/src/run/env.mjs
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
// The environment every sandboxed run gets, so a `doctor --run` session is never counted as a
|
|
2
|
+
// real use of the skill by whatever is recording real uses.
|
|
3
|
+
//
|
|
4
|
+
// MEASURED, NOT FEARED (2026-10-05): Freedom's skill ledger records every skill run from a
|
|
5
|
+
// harness hook. The doctor's own sandbox runs fired that hook, so three runs of a skill nobody had
|
|
6
|
+
// ever used for real landed in the operator's ledger as perfect one-shot runs. Left alone, every
|
|
7
|
+
// `doctor --run` raises the number that is meant to keep the doctor honest.
|
|
8
|
+
//
|
|
9
|
+
// Two signals, because a caller's recorder may honour either: FREEDOM_SKILL_LEDGER=off is the
|
|
10
|
+
// Freedom ledger's documented off switch, and SUPERSKILL_SANDBOX=1 names the run for any other
|
|
11
|
+
// recorder. The run folder's own name (SANDBOX_PREFIX) is the third, for a recorder that reads
|
|
12
|
+
// only the hook payload's cwd.
|
|
13
|
+
export const SANDBOX_PREFIX = "superskill-run-";
|
|
14
|
+
|
|
15
|
+
export function sandboxEnv(base = process.env) {
|
|
16
|
+
return { ...base, FREEDOM_SKILL_LEDGER: "off", SUPERSKILL_SANDBOX: "1" };
|
|
17
|
+
}
|
package/src/run/fake.mjs
CHANGED
|
@@ -3,12 +3,13 @@
|
|
|
3
3
|
// {output, tokens?, model?}. Never calls a model.
|
|
4
4
|
import { spawnSync } from "node:child_process";
|
|
5
5
|
import { prepareWorkspace } from "./workspace.mjs";
|
|
6
|
+
import { sandboxEnv } from "./env.mjs";
|
|
6
7
|
|
|
7
8
|
export const name = "fake";
|
|
8
9
|
|
|
9
10
|
function call(script, req) {
|
|
10
11
|
const t0 = Date.now();
|
|
11
|
-
const r = spawnSync(process.execPath, [script], { input: JSON.stringify(req), encoding: "utf8" });
|
|
12
|
+
const r = spawnSync(process.execPath, [script], { input: JSON.stringify(req), env: sandboxEnv(), encoding: "utf8" });
|
|
12
13
|
if (r.status !== 0) throw Object.assign(new Error(`fake harness failed: ${r.stderr}`), { code: "SUPERSKILL" });
|
|
13
14
|
const doc = JSON.parse(r.stdout);
|
|
14
15
|
return { output: doc.output ?? "", tokens: doc.tokens ?? null, model: doc.model ?? "fake", ms: Date.now() - t0, failed: false };
|
package/src/run/index.mjs
CHANGED
|
@@ -39,11 +39,12 @@ export function pickHarness(name) {
|
|
|
39
39
|
}
|
|
40
40
|
|
|
41
41
|
/** Every case the run will execute: evals.json cases plus goldens judged against their approved output. */
|
|
42
|
-
export function collectCases(skillDir) {
|
|
42
|
+
export function collectCases(skillDir, { privateGoldens = process.env.SUPERSKILL_PRIVATE_GOLDENS || null } = {}) {
|
|
43
43
|
const e = readEvals(skillDir);
|
|
44
|
+
const name = parseSkillFile(readFileSync(join(skillDir, "SKILL.md"), "utf8")).data.name || "";
|
|
44
45
|
if (e.error) throw new UsageError(e.error);
|
|
45
46
|
const cases = e.cases.filter((c) => c.prompt.trim()).map((c) => ({ id: String(c.id), prompt: c.prompt, files: c.files, assertions: c.assertions.length ? c.assertions : [c.expected_output].filter(Boolean) }));
|
|
46
|
-
for (const g of readGoldens(skillDir)) {
|
|
47
|
+
for (const g of readGoldens(skillDir, { privateGoldens, name })) {
|
|
47
48
|
if (!g.input || !g.output || !g.output.trim()) continue;
|
|
48
49
|
cases.push({ id: `golden:${g.id}`, prompt: g.input, files: [], assertions: [`The output matches this approved output in substance (same facts, same shape; wording may differ):\n${g.output}`] });
|
|
49
50
|
}
|
package/src/run/workspace.mjs
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { mkdtempSync, mkdirSync, symlinkSync, cpSync, existsSync } from "node:fs";
|
|
2
2
|
import { tmpdir } from "node:os";
|
|
3
3
|
import { join, dirname } from "node:path";
|
|
4
|
+
import { SANDBOX_PREFIX } from "./env.mjs";
|
|
4
5
|
|
|
5
6
|
/**
|
|
6
7
|
* A fresh working folder for one run. With `linkAt` (e.g. ".claude/skills"), the skill is
|
|
@@ -8,7 +9,7 @@ import { join, dirname } from "node:path";
|
|
|
8
9
|
* files (paths relative to the skill) are copied in at the same relative paths.
|
|
9
10
|
*/
|
|
10
11
|
export function prepareWorkspace({ skillDir, skillName, files = [], linkAt = null }) {
|
|
11
|
-
const cwd = mkdtempSync(join(tmpdir(),
|
|
12
|
+
const cwd = mkdtempSync(join(tmpdir(), SANDBOX_PREFIX));
|
|
12
13
|
if (linkAt) {
|
|
13
14
|
const target = join(cwd, linkAt, skillName);
|
|
14
15
|
mkdirSync(dirname(target), { recursive: true });
|