@supersuit/superskill 0.2.2 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +25 -0
- package/README.md +1 -1
- package/SPEC.md +18 -3
- package/package.json +30 -7
- package/src/commands/approve.mjs +17 -6
- package/src/goldens.mjs +51 -2
- package/src/rules/superskill.mjs +8 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,30 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.3.1 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
The 0.3.0 changes below, released. The v0.3.0 tag's publish refused on a red test (SPEC.md still
|
|
6
|
+
named 0.2.2), so nothing was published under 0.3.0.
|
|
7
|
+
|
|
8
|
+
## 0.3.0 (2026-09-29, not published)
|
|
9
|
+
|
|
10
|
+
**A golden's approval says why, and what it rests on.** Before, an approval recorded who and when,
|
|
11
|
+
with an optional note, so "a person approved it" read the same whether they liked it or it had
|
|
12
|
+
landed a client. (Gary Sheng, on essays he and his co-founder had approved: *"golden approval also
|
|
13
|
+
needs to carry rationale and weight. Right now the approval is based on Wilson and Gary being happy,
|
|
14
|
+
not proven results of landing clients."*)
|
|
15
|
+
|
|
16
|
+
- `superskill approve` now requires a rationale and asks for a basis: `judgment` (you read it and it
|
|
17
|
+
is right, the default) or `outcome` (it produced a result someone can check; `--evidence` is
|
|
18
|
+
required). New flags `--rationale` (`--note` still works), `--basis`, `--evidence`.
|
|
19
|
+
- Approvals accumulate in `APPROVAL.json` under `approvals`, so two people approving, and a later
|
|
20
|
+
outcome, are all kept. The newest is mirrored at the top level for older readers; a pre-0.3.0 file
|
|
21
|
+
reads as one judgment approval.
|
|
22
|
+
- `doctor` reports each golden's weight (`g1: 2 judgment, 1 outcome`) and says plainly when every
|
|
23
|
+
approval is judgment only. Informational: the level still needs one approved golden, since many
|
|
24
|
+
skills have no measurable outcome.
|
|
25
|
+
- Tests: rationale and evidence refusals, legacy normalization, accumulation, and the doctor's
|
|
26
|
+
judgment-only finding appearing and clearing. Both new guards were mutated and went red.
|
|
27
|
+
|
|
3
28
|
## 0.2.2 (2026-09-29)
|
|
4
29
|
|
|
5
30
|
- An inline flow map (`scope: { form: essay, audience: builders, purpose: persuade }`, `check: {
|
package/README.md
CHANGED
|
@@ -54,7 +54,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
|
|
|
54
54
|
| `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
|
|
55
55
|
| `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
|
|
56
56
|
| `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
|
|
57
|
-
| `superskill approve <skill> <golden
|
|
57
|
+
| `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). Terminal only, asks for your name, so an agent cannot approve its own output. Approvals accumulate. |
|
|
58
58
|
| `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
|
|
59
59
|
| `superskill miss import <skill> --freedom-ledger [--ledger <file>]` | Import runs that needed correcting from Freedom's run ledger. |
|
|
60
60
|
| `superskill snippet` | Print a block for `AGENTS.md` / `CLAUDE.md` that teaches any agent these habits. |
|
package/SPEC.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# The superskill standard
|
|
2
2
|
|
|
3
|
-
**Version 0.
|
|
3
|
+
**Version 0.3.1** (2026-09-29). The reference checker is `@supersuit/superskill`; where this
|
|
4
4
|
document and the checker disagree, the checker has a bug.
|
|
5
5
|
|
|
6
6
|
A **superskill** runs on frontier intelligence, is checked against examples a person approved,
|
|
@@ -97,7 +97,7 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
|
|
|
97
97
|
|
|
98
98
|
| Rule | Severity | Threshold |
|
|
99
99
|
|---|---|---|
|
|
100
|
-
| `golden-approved` | fail | at least one golden has an
|
|
100
|
+
| `golden-approved` | fail | at least one golden has an approval with non-empty `approved_by` and a valid `approved_at`; info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
|
|
101
101
|
| `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
|
|
102
102
|
| `no-stale-open-miss` | fail | no miss has been open more than 14 days |
|
|
103
103
|
| `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
|
|
@@ -160,9 +160,24 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
|
|
|
160
160
|
- `APPROVAL.json`, written only by `superskill approve` at an interactive terminal:
|
|
161
161
|
|
|
162
162
|
```json
|
|
163
|
-
{ "
|
|
163
|
+
{ "approvals": [
|
|
164
|
+
{ "approved_by": "Ann Example", "approved_at": "2026-09-10T15:00:00.000Z", "skill_sha": "<sha256 of SKILL.md>",
|
|
165
|
+
"rationale": "Exactly the shape I send my manager.", "basis": "judgment" },
|
|
166
|
+
{ "approved_by": "Ann Example", "approved_at": "2026-09-20T15:00:00.000Z", "skill_sha": "<sha256 of SKILL.md>",
|
|
167
|
+
"rationale": "My manager adopted it as the team template.", "basis": "outcome",
|
|
168
|
+
"evidence": "Sent 2026-09-19; adopted as the template in the team wiki" }
|
|
169
|
+
] }
|
|
164
170
|
```
|
|
165
171
|
|
|
172
|
+
**Every approval carries a rationale and a basis, because being liked and being proven are
|
|
173
|
+
different weights.** `judgment`: a person read the output and says it is right. `outcome`: the
|
|
174
|
+
output produced a result in the world that someone can check (a client landed, a call booked, a
|
|
175
|
+
template adopted), and `evidence` says what happened and where to check it. Approvals accumulate:
|
|
176
|
+
two people approving, and later an outcome, all stay on the record. The newest approval is also
|
|
177
|
+
mirrored at the top level (`approved_by`, `approved_at`, `skill_sha`, `note`), so a reader written
|
|
178
|
+
before 0.3.0 still sees it. A pre-0.3.0 file with a single approval reads as one `judgment`
|
|
179
|
+
approval whose `note` is its rationale.
|
|
180
|
+
|
|
166
181
|
A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
|
|
167
182
|
output.
|
|
168
183
|
|
package/package.json
CHANGED
|
@@ -1,19 +1,42 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@supersuit/superskill",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.1",
|
|
4
4
|
"description": "Score any agent skill folder as skill, tested, or superskill. An open standard and a zero-dependency CLI.",
|
|
5
5
|
"type": "module",
|
|
6
|
-
"bin": {
|
|
6
|
+
"bin": {
|
|
7
|
+
"superskill": "bin/superskill.mjs"
|
|
8
|
+
},
|
|
7
9
|
"exports": {
|
|
8
10
|
"./yaml": "./src/frontmatter.mjs",
|
|
9
11
|
"./package.json": "./package.json"
|
|
10
12
|
},
|
|
11
|
-
"files": [
|
|
12
|
-
|
|
13
|
-
|
|
13
|
+
"files": [
|
|
14
|
+
"bin/",
|
|
15
|
+
"src/",
|
|
16
|
+
"SPEC.md",
|
|
17
|
+
"README.md",
|
|
18
|
+
"CHANGELOG.md",
|
|
19
|
+
"LICENSE"
|
|
20
|
+
],
|
|
21
|
+
"scripts": {
|
|
22
|
+
"test": "node --test test/*.test.mjs"
|
|
23
|
+
},
|
|
24
|
+
"engines": {
|
|
25
|
+
"node": ">=20"
|
|
26
|
+
},
|
|
14
27
|
"license": "MIT",
|
|
15
|
-
"repository": {
|
|
28
|
+
"repository": {
|
|
29
|
+
"type": "git",
|
|
30
|
+
"url": "git+https://github.com/SupersuitUp/superskill.git"
|
|
31
|
+
},
|
|
16
32
|
"homepage": "https://supersuit.wiki/concepts/superskill",
|
|
17
|
-
"keywords": [
|
|
33
|
+
"keywords": [
|
|
34
|
+
"agent-skills",
|
|
35
|
+
"skills",
|
|
36
|
+
"claude-code",
|
|
37
|
+
"codex",
|
|
38
|
+
"evals",
|
|
39
|
+
"superskill"
|
|
40
|
+
],
|
|
18
41
|
"dependencies": {}
|
|
19
42
|
}
|
package/src/commands/approve.mjs
CHANGED
|
@@ -4,13 +4,18 @@ import { join } from "node:path";
|
|
|
4
4
|
import { createInterface } from "node:readline/promises";
|
|
5
5
|
import { parseArgs, clock, UsageError } from "../args.mjs";
|
|
6
6
|
import { skillDir, refuse } from "./common.mjs";
|
|
7
|
-
import { readGoldens } from "../goldens.mjs";
|
|
7
|
+
import { readGoldens, approvalEntry, withApproval, BASES } from "../goldens.mjs";
|
|
8
8
|
|
|
9
|
-
export const help = `superskill approve <skill> <golden-id> [--
|
|
9
|
+
export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"]
|
|
10
10
|
|
|
11
11
|
Record that a person checked goldens/<id>/ and signs off on its output. Works only at an
|
|
12
12
|
interactive terminal and asks for your name, so an agent cannot approve its own output.
|
|
13
|
-
|
|
13
|
+
|
|
14
|
+
Every approval says WHY (--rationale, or asked) and WHAT IT RESTS ON (--basis, or asked):
|
|
15
|
+
judgment you read it and it is right (the default)
|
|
16
|
+
outcome it produced a result someone can check; --evidence is required
|
|
17
|
+
Approvals accumulate: two people approving, or a judgment approval later backed by an outcome,
|
|
18
|
+
all stay on the record. Writes goldens/<id>/APPROVAL.json with the SKILL.md hash.
|
|
14
19
|
`;
|
|
15
20
|
|
|
16
21
|
export async function run(argv) {
|
|
@@ -28,12 +33,18 @@ export async function run(argv) {
|
|
|
28
33
|
process.stdout.write(`\n--- goldens/${id}/${g.outputFile} ---\n${g.output.slice(0, 2000)}${g.output.length > 2000 ? "\n[...]" : ""}\n---\n`);
|
|
29
34
|
const name = (await rl.question("Your name (blank to cancel): ")).trim();
|
|
30
35
|
if (!name) return refuse("not approved");
|
|
31
|
-
const
|
|
36
|
+
const rationale = String(a.flags.rationale || a.flags.note || (await rl.question("Why is this right? ")).trim());
|
|
37
|
+
let basis = a.flags.basis;
|
|
38
|
+
if (!basis) basis = /^o/i.test((await rl.question("What does this rest on: (j)udgment, you read it and it is right, or (o)utcome, it produced a result someone can check? [j] ")).trim()) ? "outcome" : "judgment";
|
|
39
|
+
const evidence = basis === "outcome" ? String(a.flags.evidence || (await rl.question("What happened, and where can someone check it? ")).trim()) : "";
|
|
32
40
|
const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
|
|
41
|
+
const made = approvalEntry({ name, rationale, basis, evidence, at: clock(a.flags).toISOString(), sha });
|
|
42
|
+
if (made.error) return refuse(made.error);
|
|
33
43
|
const p = join(dir, "goldens", id, "APPROVAL.json");
|
|
34
44
|
const existed = existsSync(p);
|
|
35
|
-
|
|
36
|
-
|
|
45
|
+
const prior = existed ? g.approval : null;
|
|
46
|
+
writeFileSync(p, JSON.stringify(withApproval(prior, made.entry), null, 2) + "\n");
|
|
47
|
+
process.stdout.write(`${existed ? "added an approval to" : "approved"} goldens/${id}/ by ${name} (${basis})\n`);
|
|
37
48
|
return 0;
|
|
38
49
|
} finally {
|
|
39
50
|
rl.close();
|
package/src/goldens.mjs
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
// goldens/<id>/: input.md, the approved output (output.md, or any other non-input file),
|
|
2
|
-
// and APPROVAL.json {approved_by, approved_at, skill_sha,
|
|
2
|
+
// and APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
|
|
3
|
+
// rationale, basis, evidence}] }, or the single-approval shape from before 0.3.0.
|
|
3
4
|
import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
|
|
4
5
|
import { join } from "node:path";
|
|
5
6
|
|
|
@@ -25,10 +26,58 @@ export function readGoldens(dir) {
|
|
|
25
26
|
outputFile: outputName || null,
|
|
26
27
|
output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
|
|
27
28
|
approval,
|
|
29
|
+
approvals: approvalsOf(approval),
|
|
28
30
|
approvalError,
|
|
29
31
|
});
|
|
30
32
|
}
|
|
31
33
|
return out;
|
|
32
34
|
}
|
|
33
35
|
|
|
34
|
-
|
|
36
|
+
/** What an approval rests on. `judgment`: the people who approved it read it and said it is right.
|
|
37
|
+
* `outcome`: it produced a result in the world someone can check (a client landed, a call booked).
|
|
38
|
+
* Being liked and being proven are different weights, and a golden says which it carries. */
|
|
39
|
+
export const BASES = ["judgment", "outcome"];
|
|
40
|
+
|
|
41
|
+
const validDate = (d) => Boolean(d) && !Number.isNaN(new Date(d).getTime());
|
|
42
|
+
|
|
43
|
+
/** Every approval on a golden, normalized. Reads the 0.3.0 `{ approvals: [...] }` shape and the
|
|
44
|
+
* single-approval shape before it (its `note` becomes the rationale, its basis judgment). */
|
|
45
|
+
export function approvalsOf(approval) {
|
|
46
|
+
if (!approval || typeof approval !== "object") return [];
|
|
47
|
+
const list = Array.isArray(approval.approvals) ? approval.approvals : [approval];
|
|
48
|
+
return list
|
|
49
|
+
.filter((a) => a && typeof a.approved_by === "string" && a.approved_by.trim() && validDate(a.approved_at))
|
|
50
|
+
.map((a) => ({
|
|
51
|
+
approved_by: a.approved_by.trim(),
|
|
52
|
+
approved_at: a.approved_at,
|
|
53
|
+
skill_sha: a.skill_sha || null,
|
|
54
|
+
rationale: String(a.rationale ?? a.note ?? "").trim(),
|
|
55
|
+
basis: a.basis === "outcome" ? "outcome" : "judgment",
|
|
56
|
+
evidence: String(a.evidence ?? "").trim(),
|
|
57
|
+
}));
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
/** How much a golden's approval weighs: how many people vouched for it, and how many results it produced. */
|
|
61
|
+
export function weightOf(g) {
|
|
62
|
+
const list = g.approvals || approvalsOf(g.approval);
|
|
63
|
+
return { judgment: list.filter((a) => a.basis === "judgment").length, outcome: list.filter((a) => a.basis === "outcome").length };
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
/** A new approval, or the reason it is refused. An approval says why, and an outcome says what happened. */
|
|
67
|
+
export function approvalEntry({ name, rationale, basis = "judgment", evidence = "", at, sha }) {
|
|
68
|
+
if (!String(name || "").trim()) return { error: "not approved: no name" };
|
|
69
|
+
if (!String(rationale || "").trim()) return { error: "an approval says why the output is right; give a rationale" };
|
|
70
|
+
if (!BASES.includes(basis)) return { error: `basis must be one of ${BASES.join(", ")}` };
|
|
71
|
+
if (basis === "outcome" && !String(evidence || "").trim()) return { error: "an outcome approval needs evidence: what happened, and where someone can check it" };
|
|
72
|
+
return { entry: { approved_by: name.trim(), approved_at: at, skill_sha: sha, rationale: rationale.trim(), basis, ...(basis === "outcome" ? { evidence: evidence.trim() } : {}) } };
|
|
73
|
+
}
|
|
74
|
+
|
|
75
|
+
/** The APPROVAL.json to write: every earlier approval kept, the new one appended, and the newest
|
|
76
|
+
* mirrored at the top level so a reader written before 0.3.0 still sees an approval. */
|
|
77
|
+
export function withApproval(existing, entry) {
|
|
78
|
+
const approvals = [...approvalsOf(existing), entry];
|
|
79
|
+
const { approved_by, approved_at, skill_sha, rationale } = entry;
|
|
80
|
+
return { approved_by, approved_at, skill_sha, note: rationale, approvals };
|
|
81
|
+
}
|
|
82
|
+
|
|
83
|
+
export const isApproved = (g) => approvalsOf(g.approval).length > 0;
|
package/src/rules/superskill.mjs
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
// got something wrong, and proven recently on a current model against the no-skill baseline.
|
|
3
3
|
import { createHash } from "node:crypto";
|
|
4
4
|
import { defineRules } from "./define.mjs";
|
|
5
|
-
import { readGoldens, isApproved } from "../goldens.mjs";
|
|
5
|
+
import { readGoldens, isApproved, weightOf } from "../goldens.mjs";
|
|
6
6
|
import { readMisses } from "../misses.mjs";
|
|
7
7
|
import { readEvals, readLatestRun } from "../evals.mjs";
|
|
8
8
|
|
|
@@ -33,6 +33,13 @@ export const superskillRules = defineRules([
|
|
|
33
33
|
const sha = createHash("sha256").update(ctx.raw).digest("hex");
|
|
34
34
|
const stale = approved.filter((g) => g.approval.skill_sha && g.approval.skill_sha !== sha).map((g) => g.id);
|
|
35
35
|
if (stale.length) out.push(f("info", `golden${stale.length === 1 ? "" : "s"} ${stale.join(", ")} approved against an earlier SKILL.md`, "Re-run the golden and re-approve if the output still holds."));
|
|
36
|
+
// Liked is not proven. Say which weight the approvals carry, so "a person approved it" is
|
|
37
|
+
// never read as "it worked in the world".
|
|
38
|
+
const w = approved.map((g) => ({ id: g.id, ...weightOf(g) }));
|
|
39
|
+
const proven = w.filter((x) => x.outcome > 0);
|
|
40
|
+
const summary = w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ");
|
|
41
|
+
if (!proven.length) out.push(f("info", `approved on judgment only, no outcome recorded yet (${summary})`, "When a golden produces a real result, record it: `superskill approve <skill> <golden> --basis outcome --evidence \"<what happened, where to check>\"`."));
|
|
42
|
+
else out.push(f("info", `approval weight: ${summary}`, ""));
|
|
36
43
|
return out;
|
|
37
44
|
},
|
|
38
45
|
},
|