@supersuit/superskill 0.2.1 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -0
- package/README.md +1 -1
- package/SPEC.md +18 -3
- package/package.json +30 -7
- package/src/commands/approve.mjs +17 -6
- package/src/frontmatter.mjs +96 -14
- package/src/goldens.mjs +51 -2
- package/src/rules/superskill.mjs +8 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,49 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.3.1 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
The 0.3.0 changes below, released. The v0.3.0 tag's publish refused on a red test (SPEC.md still
|
|
6
|
+
named 0.2.2), so nothing was published under 0.3.0.
|
|
7
|
+
|
|
8
|
+
## 0.3.0 (2026-09-29, not published)
|
|
9
|
+
|
|
10
|
+
**A golden's approval says why, and what it rests on.** Before, an approval recorded who and when,
|
|
11
|
+
with an optional note, so "a person approved it" read the same whether they liked it or it had
|
|
12
|
+
landed a client. (Gary Sheng, on essays he and his co-founder had approved: *"golden approval also
|
|
13
|
+
needs to carry rationale and weight. Right now the approval is based on Wilson and Gary being happy,
|
|
14
|
+
not proven results of landing clients."*)
|
|
15
|
+
|
|
16
|
+
- `superskill approve` now requires a rationale and asks for a basis: `judgment` (you read it and it
|
|
17
|
+
is right, the default) or `outcome` (it produced a result someone can check; `--evidence` is
|
|
18
|
+
required). New flags `--rationale` (`--note` still works), `--basis`, `--evidence`.
|
|
19
|
+
- Approvals accumulate in `APPROVAL.json` under `approvals`, so two people approving, and a later
|
|
20
|
+
outcome, are all kept. The newest is mirrored at the top level for older readers; a pre-0.3.0 file
|
|
21
|
+
reads as one judgment approval.
|
|
22
|
+
- `doctor` reports each golden's weight (`g1: 2 judgment, 1 outcome`) and says plainly when every
|
|
23
|
+
approval is judgment only. Informational: the level still needs one approved golden, since many
|
|
24
|
+
skills have no measurable outcome.
|
|
25
|
+
- Tests: rationale and evidence refusals, legacy normalization, accumulation, and the doctor's
|
|
26
|
+
judgment-only finding appearing and clearing. Both new guards were mutated and went red.
|
|
27
|
+
|
|
28
|
+
## 0.2.2 (2026-09-29)
|
|
29
|
+
|
|
30
|
+
- An inline flow map (`scope: { form: essay, audience: builders, purpose: persuade }`, `check: {
|
|
31
|
+
station: every segment carries a label }`) now reads as an object instead of the whole `{ ...
|
|
32
|
+
}` coming back as a string. Nested inline maps and inline lists inside inline maps work too
|
|
33
|
+
(`speech: { uses: [a, b], never: [c] }`).
|
|
34
|
+
- An inline flow list followed by a same-line comment (`conditions: [r1, r2, r3] # 5 to 10
|
|
35
|
+
ids`) now reads as a list. Before this, the inline-list check required the raw value to END in
|
|
36
|
+
`]`, and a trailing comment broke that, so the whole line came back as a string.
|
|
37
|
+
- Both are read by a quote-aware character scanner rather than a naive split, so a comma,
|
|
38
|
+
bracket, brace or hash inside a quoted value inside a flow collection stays text
|
|
39
|
+
(`{ note: "a, b] } # c" }`). Malformed flow syntax (unbalanced brackets) still never throws: it
|
|
40
|
+
falls back to the raw string, same as an unrecognized value always has.
|
|
41
|
+
- A block list item that is itself a bare inline flow map or list (`- { station: fine, severity:
|
|
42
|
+
fail }`, `- [a, b]`) now reads as an object or a list. Before this, the KEY regex that decides
|
|
43
|
+
whether `- key: value` opens a block submap matched on the first colon inside the braces
|
|
44
|
+
(reading `"{ station"` as the key), so the item came back as `{ "{ station": "fine, severity:
|
|
45
|
+
fail }" }`. `- key: { ... }` still opens a block submap as before.
|
|
46
|
+
|
|
3
47
|
## 0.2.1 (2026-09-28)
|
|
4
48
|
|
|
5
49
|
- An unquoted value that is only a comment now reads as empty, which is what YAML means. Before
|
package/README.md
CHANGED
|
@@ -54,7 +54,7 @@ Every rule and threshold is in [SPEC.md](SPEC.md).
|
|
|
54
54
|
| `superskill init <skill> --from-session <transcript>` | Turn the session where you did the job by hand into the first eval and a golden candidate (Claude Code `.jsonl`, or any text file as the request). |
|
|
55
55
|
| `superskill miss <skill> "<what happened>" [--expected "..."]` | Log a time the skill got it wrong. |
|
|
56
56
|
| `superskill fix <skill> <miss-id> --eval <id> [--commit <sha>]` | Close a miss. Refuses without an eval that exists. |
|
|
57
|
-
| `superskill approve <skill> <golden
|
|
57
|
+
| `superskill approve <skill> <golden> [--basis judgment\|outcome] [--rationale ...] [--evidence ...]` | A person signs off on a golden, saying why and what it rests on: `judgment` (it reads right) or `outcome` (it produced a checkable result, with evidence). Terminal only, asks for your name, so an agent cannot approve its own output. Approvals accumulate. |
|
|
58
58
|
| `superskill collection <folder...> [--budget <chars>] [--overlap 0.5]` | Listing budget used, descriptions that get cut off, pairs of skills an agent could confuse (with near-miss triggers to add). |
|
|
59
59
|
| `superskill miss import <skill> --freedom-ledger [--ledger <file>]` | Import runs that needed correcting from Freedom's run ledger. |
|
|
60
60
|
| `superskill snippet` | Print a block for `AGENTS.md` / `CLAUDE.md` that teaches any agent these habits. |
|
package/SPEC.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# The superskill standard
|
|
2
2
|
|
|
3
|
-
**Version 0.
|
|
3
|
+
**Version 0.3.1** (2026-09-29). The reference checker is `@supersuit/superskill`; where this
|
|
4
4
|
document and the checker disagree, the checker has a bug.
|
|
5
5
|
|
|
6
6
|
A **superskill** runs on frontier intelligence, is checked against examples a person approved,
|
|
@@ -97,7 +97,7 @@ A line in a bundled file containing `superskill-ignore` is skipped by `no-absolu
|
|
|
97
97
|
|
|
98
98
|
| Rule | Severity | Threshold |
|
|
99
99
|
|---|---|---|
|
|
100
|
-
| `golden-approved` | fail | at least one golden has an
|
|
100
|
+
| `golden-approved` | fail | at least one golden has an approval with non-empty `approved_by` and a valid `approved_at`; info when it was approved against an earlier `SKILL.md`; info naming the weight (judgment and outcome approvals per golden), and saying so plainly when no golden has an outcome yet |
|
|
101
101
|
| `misses-log-present` | fail | `MISSES.md` exists (it may have no entries) |
|
|
102
102
|
| `no-stale-open-miss` | fail | no miss has been open more than 14 days |
|
|
103
103
|
| `fixed-miss-has-eval` | fail | every fixed miss names an eval id present in `evals.json` or `goldens/` |
|
|
@@ -160,9 +160,24 @@ A bare array of cases is accepted on read, as is `assertions` for `expectations`
|
|
|
160
160
|
- `APPROVAL.json`, written only by `superskill approve` at an interactive terminal:
|
|
161
161
|
|
|
162
162
|
```json
|
|
163
|
-
{ "
|
|
163
|
+
{ "approvals": [
|
|
164
|
+
{ "approved_by": "Ann Example", "approved_at": "2026-09-10T15:00:00.000Z", "skill_sha": "<sha256 of SKILL.md>",
|
|
165
|
+
"rationale": "Exactly the shape I send my manager.", "basis": "judgment" },
|
|
166
|
+
{ "approved_by": "Ann Example", "approved_at": "2026-09-20T15:00:00.000Z", "skill_sha": "<sha256 of SKILL.md>",
|
|
167
|
+
"rationale": "My manager adopted it as the team template.", "basis": "outcome",
|
|
168
|
+
"evidence": "Sent 2026-09-19; adopted as the template in the team wiki" }
|
|
169
|
+
] }
|
|
164
170
|
```
|
|
165
171
|
|
|
172
|
+
**Every approval carries a rationale and a basis, because being liked and being proven are
|
|
173
|
+
different weights.** `judgment`: a person read the output and says it is right. `outcome`: the
|
|
174
|
+
output produced a result in the world that someone can check (a client landed, a call booked, a
|
|
175
|
+
template adopted), and `evidence` says what happened and where to check it. Approvals accumulate:
|
|
176
|
+
two people approving, and later an outcome, all stay on the record. The newest approval is also
|
|
177
|
+
mirrored at the top level (`approved_by`, `approved_at`, `skill_sha`, `note`), so a reader written
|
|
178
|
+
before 0.3.0 still sees it. A pre-0.3.0 file with a single approval reads as one `judgment`
|
|
179
|
+
approval whose `note` is its rationale.
|
|
180
|
+
|
|
166
181
|
A golden is also an eval: `--run` judges the skill's output for `input.md` against the approved
|
|
167
182
|
output.
|
|
168
183
|
|
package/package.json
CHANGED
|
@@ -1,19 +1,42 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@supersuit/superskill",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.3.1",
|
|
4
4
|
"description": "Score any agent skill folder as skill, tested, or superskill. An open standard and a zero-dependency CLI.",
|
|
5
5
|
"type": "module",
|
|
6
|
-
"bin": {
|
|
6
|
+
"bin": {
|
|
7
|
+
"superskill": "bin/superskill.mjs"
|
|
8
|
+
},
|
|
7
9
|
"exports": {
|
|
8
10
|
"./yaml": "./src/frontmatter.mjs",
|
|
9
11
|
"./package.json": "./package.json"
|
|
10
12
|
},
|
|
11
|
-
"files": [
|
|
12
|
-
|
|
13
|
-
|
|
13
|
+
"files": [
|
|
14
|
+
"bin/",
|
|
15
|
+
"src/",
|
|
16
|
+
"SPEC.md",
|
|
17
|
+
"README.md",
|
|
18
|
+
"CHANGELOG.md",
|
|
19
|
+
"LICENSE"
|
|
20
|
+
],
|
|
21
|
+
"scripts": {
|
|
22
|
+
"test": "node --test test/*.test.mjs"
|
|
23
|
+
},
|
|
24
|
+
"engines": {
|
|
25
|
+
"node": ">=20"
|
|
26
|
+
},
|
|
14
27
|
"license": "MIT",
|
|
15
|
-
"repository": {
|
|
28
|
+
"repository": {
|
|
29
|
+
"type": "git",
|
|
30
|
+
"url": "git+https://github.com/SupersuitUp/superskill.git"
|
|
31
|
+
},
|
|
16
32
|
"homepage": "https://supersuit.wiki/concepts/superskill",
|
|
17
|
-
"keywords": [
|
|
33
|
+
"keywords": [
|
|
34
|
+
"agent-skills",
|
|
35
|
+
"skills",
|
|
36
|
+
"claude-code",
|
|
37
|
+
"codex",
|
|
38
|
+
"evals",
|
|
39
|
+
"superskill"
|
|
40
|
+
],
|
|
18
41
|
"dependencies": {}
|
|
19
42
|
}
|
package/src/commands/approve.mjs
CHANGED
|
@@ -4,13 +4,18 @@ import { join } from "node:path";
|
|
|
4
4
|
import { createInterface } from "node:readline/promises";
|
|
5
5
|
import { parseArgs, clock, UsageError } from "../args.mjs";
|
|
6
6
|
import { skillDir, refuse } from "./common.mjs";
|
|
7
|
-
import { readGoldens } from "../goldens.mjs";
|
|
7
|
+
import { readGoldens, approvalEntry, withApproval, BASES } from "../goldens.mjs";
|
|
8
8
|
|
|
9
|
-
export const help = `superskill approve <skill> <golden-id> [--
|
|
9
|
+
export const help = `superskill approve <skill> <golden-id> [--rationale "<why it is right>"] [--basis judgment|outcome] [--evidence "<what happened, where to check>"]
|
|
10
10
|
|
|
11
11
|
Record that a person checked goldens/<id>/ and signs off on its output. Works only at an
|
|
12
12
|
interactive terminal and asks for your name, so an agent cannot approve its own output.
|
|
13
|
-
|
|
13
|
+
|
|
14
|
+
Every approval says WHY (--rationale, or asked) and WHAT IT RESTS ON (--basis, or asked):
|
|
15
|
+
judgment you read it and it is right (the default)
|
|
16
|
+
outcome it produced a result someone can check; --evidence is required
|
|
17
|
+
Approvals accumulate: two people approving, or a judgment approval later backed by an outcome,
|
|
18
|
+
all stay on the record. Writes goldens/<id>/APPROVAL.json with the SKILL.md hash.
|
|
14
19
|
`;
|
|
15
20
|
|
|
16
21
|
export async function run(argv) {
|
|
@@ -28,12 +33,18 @@ export async function run(argv) {
|
|
|
28
33
|
process.stdout.write(`\n--- goldens/${id}/${g.outputFile} ---\n${g.output.slice(0, 2000)}${g.output.length > 2000 ? "\n[...]" : ""}\n---\n`);
|
|
29
34
|
const name = (await rl.question("Your name (blank to cancel): ")).trim();
|
|
30
35
|
if (!name) return refuse("not approved");
|
|
31
|
-
const
|
|
36
|
+
const rationale = String(a.flags.rationale || a.flags.note || (await rl.question("Why is this right? ")).trim());
|
|
37
|
+
let basis = a.flags.basis;
|
|
38
|
+
if (!basis) basis = /^o/i.test((await rl.question("What does this rest on: (j)udgment, you read it and it is right, or (o)utcome, it produced a result someone can check? [j] ")).trim()) ? "outcome" : "judgment";
|
|
39
|
+
const evidence = basis === "outcome" ? String(a.flags.evidence || (await rl.question("What happened, and where can someone check it? ")).trim()) : "";
|
|
32
40
|
const sha = createHash("sha256").update(readFileSync(join(dir, "SKILL.md"))).digest("hex");
|
|
41
|
+
const made = approvalEntry({ name, rationale, basis, evidence, at: clock(a.flags).toISOString(), sha });
|
|
42
|
+
if (made.error) return refuse(made.error);
|
|
33
43
|
const p = join(dir, "goldens", id, "APPROVAL.json");
|
|
34
44
|
const existed = existsSync(p);
|
|
35
|
-
|
|
36
|
-
|
|
45
|
+
const prior = existed ? g.approval : null;
|
|
46
|
+
writeFileSync(p, JSON.stringify(withApproval(prior, made.entry), null, 2) + "\n");
|
|
47
|
+
process.stdout.write(`${existed ? "added an approval to" : "approved"} goldens/${id}/ by ${name} (${basis})\n`);
|
|
37
48
|
return 0;
|
|
38
49
|
} finally {
|
|
39
50
|
rl.close();
|
package/src/frontmatter.mjs
CHANGED
|
@@ -31,6 +31,13 @@ const indentOf = (l) => l.length - l.trimStart().length;
|
|
|
31
31
|
// reads as empty, the same as no value at all, which is what YAML means. A comment-only value
|
|
32
32
|
// that is followed by a more-indented block still opens that nested map or list, exactly as
|
|
33
33
|
// `key:` with nothing after it does.
|
|
34
|
+
// 0.2.2: an inline flow map (`scope: { form: essay, audience: builders }`) reads as an object,
|
|
35
|
+
// nested to any depth (a flow map inside a flow map, a flow list inside a flow map). And an
|
|
36
|
+
// inline flow list followed by a same-line comment (`conditions: [r1, r2, r3] # 5 to 10 ids`)
|
|
37
|
+
// reads as a list instead of the whole `[...] # ...` text, because the flow value is now parsed
|
|
38
|
+
// character-by-character (quote-aware) instead of by checking whether the raw value ends in `]`.
|
|
39
|
+
// Malformed flow syntax (unbalanced brackets) never throws: it falls back to the raw string, the
|
|
40
|
+
// same as an unrecognized value always has.
|
|
34
41
|
export function parseYamlSubset(lines) {
|
|
35
42
|
return parseMap(lines, 0, 0, true)[0];
|
|
36
43
|
}
|
|
@@ -103,6 +110,16 @@ function parseList(lines, i, indent) {
|
|
|
103
110
|
} else out.push("");
|
|
104
111
|
continue;
|
|
105
112
|
}
|
|
113
|
+
if (content.startsWith("{") || content.startsWith("[")) {
|
|
114
|
+
// A bare flow map/list item (`- { station: fine, severity: fail }`, `- [a, b]`) is not a
|
|
115
|
+
// "- key: value" line, even though the KEY regex below would happily match on the first
|
|
116
|
+
// colon inside the braces (reading "{ station" as the key, which produced garbage). Try
|
|
117
|
+
// the flow parse first; when it is malformed, `inlineOrScalar` falls back to the raw
|
|
118
|
+
// string via `scalar()`, same as an unrecognized top-level value always has. Either way
|
|
119
|
+
// this never falls through to the "- key: value" submap heuristic below, because a flow
|
|
120
|
+
// collection's opening bracket can never legitimately be a map key.
|
|
121
|
+
out.push(inlineOrScalar(content)); continue;
|
|
122
|
+
}
|
|
106
123
|
if (KEY.test(content) && !/^\[.*\]$/.test(content)) {
|
|
107
124
|
// "- key: value" opens a map whose keys sit where this content starts.
|
|
108
125
|
const sub = [" ".repeat(at) + content, ...lines.slice(i)];
|
|
@@ -124,10 +141,88 @@ function readBlock(lines, i, keyIndent, style) {
|
|
|
124
141
|
}
|
|
125
142
|
|
|
126
143
|
function inlineOrScalar(rest) {
|
|
127
|
-
if (rest.startsWith("[")
|
|
144
|
+
if (rest.startsWith("[") || rest.startsWith("{")) {
|
|
145
|
+
const flow = tryParseFlow(rest);
|
|
146
|
+
if (flow !== undefined) return flow;
|
|
147
|
+
}
|
|
128
148
|
return scalar(rest);
|
|
129
149
|
}
|
|
130
150
|
|
|
151
|
+
// A flow collection (`[...]` or `{...}`) parsed character-by-character so quotes can protect a
|
|
152
|
+
// `,` `]` `}` or `#` from being read as structure, and so trailing whitespace plus a same-line
|
|
153
|
+
// comment after the closing bracket does not fall the whole value back to a raw string. Returns
|
|
154
|
+
// `undefined` (never throws) when `rest` is not a clean flow value: unbalanced brackets, an
|
|
155
|
+
// unterminated quote, or trailing content that is neither blank nor a comment. The caller falls
|
|
156
|
+
// back to `scalar(rest)` in every one of those cases, same as an unrecognized value always has.
|
|
157
|
+
function tryParseFlow(rest) {
|
|
158
|
+
let parsed;
|
|
159
|
+
try {
|
|
160
|
+
parsed = readFlowCollection(rest, 0);
|
|
161
|
+
} catch {
|
|
162
|
+
return undefined;
|
|
163
|
+
}
|
|
164
|
+
const trailing = rest.slice(parsed.end);
|
|
165
|
+
if (trailing.trim() === "" || /^\s+#/.test(trailing)) return parsed.value;
|
|
166
|
+
return undefined;
|
|
167
|
+
}
|
|
168
|
+
|
|
169
|
+
function readFlowCollection(s, i) {
|
|
170
|
+
const open = s[i];
|
|
171
|
+
const close = open === "{" ? "}" : "]";
|
|
172
|
+
const isMap = open === "{";
|
|
173
|
+
i = skipFlowWs(s, i + 1);
|
|
174
|
+
if (s[i] === close) return { value: isMap ? {} : [], end: i + 1 };
|
|
175
|
+
const items = [];
|
|
176
|
+
for (;;) {
|
|
177
|
+
i = skipFlowWs(s, i);
|
|
178
|
+
if (i >= s.length) throw new Error("unterminated flow collection");
|
|
179
|
+
let key;
|
|
180
|
+
if (isMap) {
|
|
181
|
+
const k = readFlowToken(s, i, [":"]);
|
|
182
|
+
key = k.text.trim().replace(/^["']|["']$/g, "");
|
|
183
|
+
i = skipFlowWs(s, k.end + 1);
|
|
184
|
+
}
|
|
185
|
+
let value;
|
|
186
|
+
if (s[i] === "{" || s[i] === "[") {
|
|
187
|
+
const nested = readFlowCollection(s, i);
|
|
188
|
+
value = nested.value; i = nested.end;
|
|
189
|
+
} else {
|
|
190
|
+
const v = readFlowToken(s, i, [",", close]);
|
|
191
|
+
value = scalar(v.text);
|
|
192
|
+
i = v.end;
|
|
193
|
+
}
|
|
194
|
+
items.push(isMap ? [key, value] : value);
|
|
195
|
+
i = skipFlowWs(s, i);
|
|
196
|
+
if (s[i] === ",") { i = skipFlowWs(s, i + 1); if (s[i] === close) { i++; break; } continue; }
|
|
197
|
+
if (s[i] === close) { i++; break; }
|
|
198
|
+
throw new Error(`expected ',' or '${close}'`);
|
|
199
|
+
}
|
|
200
|
+
const value = isMap ? Object.fromEntries(items) : items.filter((v) => v !== "");
|
|
201
|
+
return { value, end: i };
|
|
202
|
+
}
|
|
203
|
+
|
|
204
|
+
function skipFlowWs(s, i) {
|
|
205
|
+
while (i < s.length && /\s/.test(s[i])) i++;
|
|
206
|
+
return i;
|
|
207
|
+
}
|
|
208
|
+
|
|
209
|
+
// Reads raw text from `i` up to (but not including) the first unquoted occurrence of a char in
|
|
210
|
+
// `stopChars`, honoring both quote styles so a stop char inside quotes stays text. Throws (never
|
|
211
|
+
// returns a partial token) when the string runs out before a stop char is found outside quotes,
|
|
212
|
+
// which is what an unbalanced bracket or an unterminated quote looks like from here.
|
|
213
|
+
function readFlowToken(s, i, stopChars) {
|
|
214
|
+
let text = "";
|
|
215
|
+
let q = null;
|
|
216
|
+
while (i < s.length) {
|
|
217
|
+
const ch = s[i];
|
|
218
|
+
if (q) { text += ch; if (ch === q) q = null; i++; continue; }
|
|
219
|
+
if (ch === '"' || ch === "'") { q = ch; text += ch; i++; continue; }
|
|
220
|
+
if (stopChars.includes(ch)) return { text, end: i };
|
|
221
|
+
text += ch; i++;
|
|
222
|
+
}
|
|
223
|
+
throw new Error("unterminated flow token");
|
|
224
|
+
}
|
|
225
|
+
|
|
131
226
|
function foldLines(lines) {
|
|
132
227
|
let s = "";
|
|
133
228
|
for (const l of lines) {
|
|
@@ -137,19 +232,6 @@ function foldLines(lines) {
|
|
|
137
232
|
return s;
|
|
138
233
|
}
|
|
139
234
|
|
|
140
|
-
function splitInline(s) {
|
|
141
|
-
const parts = [];
|
|
142
|
-
let cur = "", q = null;
|
|
143
|
-
for (const ch of s) {
|
|
144
|
-
if (q) { cur += ch; if (ch === q) q = null; }
|
|
145
|
-
else if (ch === '"' || ch === "'") { q = ch; cur += ch; }
|
|
146
|
-
else if (ch === ",") { parts.push(cur); cur = ""; }
|
|
147
|
-
else cur += ch;
|
|
148
|
-
}
|
|
149
|
-
parts.push(cur);
|
|
150
|
-
return parts.map((p) => p.trim());
|
|
151
|
-
}
|
|
152
|
-
|
|
153
235
|
function scalar(raw) {
|
|
154
236
|
const s = String(raw).trim();
|
|
155
237
|
if (s.startsWith('"')) {
|
package/src/goldens.mjs
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
// goldens/<id>/: input.md, the approved output (output.md, or any other non-input file),
|
|
2
|
-
// and APPROVAL.json {approved_by, approved_at, skill_sha,
|
|
2
|
+
// and APPROVAL.json written by a person: { approvals: [{approved_by, approved_at, skill_sha,
|
|
3
|
+
// rationale, basis, evidence}] }, or the single-approval shape from before 0.3.0.
|
|
3
4
|
import { readdirSync, readFileSync, existsSync, statSync } from "node:fs";
|
|
4
5
|
import { join } from "node:path";
|
|
5
6
|
|
|
@@ -25,10 +26,58 @@ export function readGoldens(dir) {
|
|
|
25
26
|
outputFile: outputName || null,
|
|
26
27
|
output: outputName ? readFileSync(join(gdir, outputName), "utf8") : null,
|
|
27
28
|
approval,
|
|
29
|
+
approvals: approvalsOf(approval),
|
|
28
30
|
approvalError,
|
|
29
31
|
});
|
|
30
32
|
}
|
|
31
33
|
return out;
|
|
32
34
|
}
|
|
33
35
|
|
|
34
|
-
|
|
36
|
+
/** What an approval rests on. `judgment`: the people who approved it read it and said it is right.
|
|
37
|
+
* `outcome`: it produced a result in the world someone can check (a client landed, a call booked).
|
|
38
|
+
* Being liked and being proven are different weights, and a golden says which it carries. */
|
|
39
|
+
export const BASES = ["judgment", "outcome"];
|
|
40
|
+
|
|
41
|
+
const validDate = (d) => Boolean(d) && !Number.isNaN(new Date(d).getTime());
|
|
42
|
+
|
|
43
|
+
/** Every approval on a golden, normalized. Reads the 0.3.0 `{ approvals: [...] }` shape and the
|
|
44
|
+
* single-approval shape before it (its `note` becomes the rationale, its basis judgment). */
|
|
45
|
+
export function approvalsOf(approval) {
|
|
46
|
+
if (!approval || typeof approval !== "object") return [];
|
|
47
|
+
const list = Array.isArray(approval.approvals) ? approval.approvals : [approval];
|
|
48
|
+
return list
|
|
49
|
+
.filter((a) => a && typeof a.approved_by === "string" && a.approved_by.trim() && validDate(a.approved_at))
|
|
50
|
+
.map((a) => ({
|
|
51
|
+
approved_by: a.approved_by.trim(),
|
|
52
|
+
approved_at: a.approved_at,
|
|
53
|
+
skill_sha: a.skill_sha || null,
|
|
54
|
+
rationale: String(a.rationale ?? a.note ?? "").trim(),
|
|
55
|
+
basis: a.basis === "outcome" ? "outcome" : "judgment",
|
|
56
|
+
evidence: String(a.evidence ?? "").trim(),
|
|
57
|
+
}));
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
/** How much a golden's approval weighs: how many people vouched for it, and how many results it produced. */
|
|
61
|
+
export function weightOf(g) {
|
|
62
|
+
const list = g.approvals || approvalsOf(g.approval);
|
|
63
|
+
return { judgment: list.filter((a) => a.basis === "judgment").length, outcome: list.filter((a) => a.basis === "outcome").length };
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
/** A new approval, or the reason it is refused. An approval says why, and an outcome says what happened. */
|
|
67
|
+
export function approvalEntry({ name, rationale, basis = "judgment", evidence = "", at, sha }) {
|
|
68
|
+
if (!String(name || "").trim()) return { error: "not approved: no name" };
|
|
69
|
+
if (!String(rationale || "").trim()) return { error: "an approval says why the output is right; give a rationale" };
|
|
70
|
+
if (!BASES.includes(basis)) return { error: `basis must be one of ${BASES.join(", ")}` };
|
|
71
|
+
if (basis === "outcome" && !String(evidence || "").trim()) return { error: "an outcome approval needs evidence: what happened, and where someone can check it" };
|
|
72
|
+
return { entry: { approved_by: name.trim(), approved_at: at, skill_sha: sha, rationale: rationale.trim(), basis, ...(basis === "outcome" ? { evidence: evidence.trim() } : {}) } };
|
|
73
|
+
}
|
|
74
|
+
|
|
75
|
+
/** The APPROVAL.json to write: every earlier approval kept, the new one appended, and the newest
|
|
76
|
+
* mirrored at the top level so a reader written before 0.3.0 still sees an approval. */
|
|
77
|
+
export function withApproval(existing, entry) {
|
|
78
|
+
const approvals = [...approvalsOf(existing), entry];
|
|
79
|
+
const { approved_by, approved_at, skill_sha, rationale } = entry;
|
|
80
|
+
return { approved_by, approved_at, skill_sha, note: rationale, approvals };
|
|
81
|
+
}
|
|
82
|
+
|
|
83
|
+
export const isApproved = (g) => approvalsOf(g.approval).length > 0;
|
package/src/rules/superskill.mjs
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
// got something wrong, and proven recently on a current model against the no-skill baseline.
|
|
3
3
|
import { createHash } from "node:crypto";
|
|
4
4
|
import { defineRules } from "./define.mjs";
|
|
5
|
-
import { readGoldens, isApproved } from "../goldens.mjs";
|
|
5
|
+
import { readGoldens, isApproved, weightOf } from "../goldens.mjs";
|
|
6
6
|
import { readMisses } from "../misses.mjs";
|
|
7
7
|
import { readEvals, readLatestRun } from "../evals.mjs";
|
|
8
8
|
|
|
@@ -33,6 +33,13 @@ export const superskillRules = defineRules([
|
|
|
33
33
|
const sha = createHash("sha256").update(ctx.raw).digest("hex");
|
|
34
34
|
const stale = approved.filter((g) => g.approval.skill_sha && g.approval.skill_sha !== sha).map((g) => g.id);
|
|
35
35
|
if (stale.length) out.push(f("info", `golden${stale.length === 1 ? "" : "s"} ${stale.join(", ")} approved against an earlier SKILL.md`, "Re-run the golden and re-approve if the output still holds."));
|
|
36
|
+
// Liked is not proven. Say which weight the approvals carry, so "a person approved it" is
|
|
37
|
+
// never read as "it worked in the world".
|
|
38
|
+
const w = approved.map((g) => ({ id: g.id, ...weightOf(g) }));
|
|
39
|
+
const proven = w.filter((x) => x.outcome > 0);
|
|
40
|
+
const summary = w.map((x) => `${x.id}: ${x.judgment} judgment, ${x.outcome} outcome`).join("; ");
|
|
41
|
+
if (!proven.length) out.push(f("info", `approved on judgment only, no outcome recorded yet (${summary})`, "When a golden produces a real result, record it: `superskill approve <skill> <golden> --basis outcome --evidence \"<what happened, where to check>\"`."));
|
|
42
|
+
else out.push(f("info", `approval weight: ${summary}`, ""));
|
|
36
43
|
return out;
|
|
37
44
|
},
|
|
38
45
|
},
|