@supersuit/hyperspec 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/CHANGELOG.md +91 -0
  2. package/README.md +37 -0
  3. package/SPEC.md +2 -2
  4. package/WRITING.md +542 -10
  5. package/bin/hyperspec.mjs +183 -0
  6. package/examples/writing/essay/judge/doctor.packet.json +108 -0
  7. package/examples/writing/essay/judge/lineup.packet.json +64 -0
  8. package/examples/writing/essay/judge/persona.packet.json +73 -0
  9. package/examples/writing/essay/judge/reader.packet.json +93 -0
  10. package/examples/writing/essay/learn/first-draft.md +84 -0
  11. package/examples/writing/essay/learn/learn.packet.json +106 -0
  12. package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
  13. package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
  14. package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
  15. package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
  16. package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
  17. package/examples/writing/essay.hyperspec.md +6 -1
  18. package/examples/writing/story/judge/attribution.packet.json +194 -0
  19. package/examples/writing/story/judge/doctor.packet.json +108 -0
  20. package/examples/writing/story/judge/knowledge.packet.json +77 -0
  21. package/examples/writing/story/judge/persona.packet.json +73 -0
  22. package/examples/writing/story/judge/reader.packet.json +94 -0
  23. package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
  24. package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
  25. package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
  26. package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
  27. package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
  28. package/examples/writing/story.hyperspec.md +7 -3
  29. package/package.json +1 -1
  30. package/src/check.mjs +75 -129
  31. package/src/draft.mjs +26 -0
  32. package/src/judge.mjs +386 -0
  33. package/src/judges/attribution.mjs +360 -0
  34. package/src/judges/doctor.mjs +126 -0
  35. package/src/judges/index.mjs +31 -0
  36. package/src/judges/knowledge.mjs +111 -0
  37. package/src/judges/lineup.mjs +272 -0
  38. package/src/judges/persona.mjs +137 -0
  39. package/src/judges/reader.mjs +111 -0
  40. package/src/learn.mjs +422 -0
  41. package/src/ledger.mjs +108 -0
  42. package/src/sentences.mjs +81 -0
  43. package/src/stations/claims.mjs +44 -39
  44. package/src/stations/quotes.mjs +6 -4
  45. package/src/writing.mjs +1 -1
package/WRITING.md CHANGED
@@ -219,7 +219,7 @@ writing:
219
219
  goal:
220
220
  from: plans to run the meeting from their own list
221
221
  to: hands the meeting to the report
222
- next_if_worked: copies the three questions into the invite
222
+ next_if_worked: writes the three questions on a card
223
223
  change:
224
224
  kind: action # belief | action | feeling
225
225
  text: the reader asks the three questions and waits
@@ -913,8 +913,9 @@ letters and is not a common function word such as "the", so a speaker recorded a
913
913
  interviewed" is named only by all three words. Attribution needs a declared speaker: a name that
914
914
  is no segment's `speaker` attributes nothing, so start a `speaker` with the person's name, as
915
915
  the essay example does with `dana, an engineering manager`. A spec with `fiction: true` skips
916
- the station: a character's dialogue is invented rather than quoted from a material, and a later
917
- release checks it against each character's own lines.
916
+ the station: a character's dialogue is invented rather than quoted from a material. The
917
+ `attribution` judge tests it against each character's own lines instead (see
918
+ [Judging a draft](#judging-a-draft)).
918
919
 
919
920
  | Id | Severity | Meaning |
920
921
  |---|---|---|
@@ -1009,6 +1010,531 @@ the same draft:
1009
1010
 
1010
1011
  A ledger path that leads outside the spec's folder is not written, and `check` prints a warning.
1011
1012
 
1013
+ ## Judging a draft
1014
+
1015
+ `check` runs the stations that are plain functions of the spec and the draft. The rest of a
1016
+ spec's checks are rubrics: whether the doctor would pass each goal condition, whether a reader
1017
+ would get lost, whether the voice can be told from the writer's own. Those are judgments, and
1018
+ hyperspec calls no model, so it does not make them. It makes them checkable instead. For each
1019
+ judgment station, `judge prepare` writes a packet: the rubric from the spec, fixed
1020
+ instructions, the inputs the judge reads, and the exact shape of the answer. An outside judge
1021
+ (your agent, any model, or a person) fills in a verdict. `judge record` validates it, derives
1022
+ the station's status from it by a fixed rule, and records it in the runs ledger. The judge
1023
+ decides; hyperspec checks that every passage the judge quotes is in the draft, scores
1024
+ every blind test against an answer key the judge never saw, and keeps the record.
1025
+
1026
+ ```bash
1027
+ npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
1028
+ ```
1029
+
1030
+ ```
1031
+ essay/judge/doctor.packet.json
1032
+ essay/judge/lineup.packet.json
1033
+ essay/judge/lineup.key.json
1034
+ essay/judge/reader.packet.json
1035
+ essay/judge/persona.packet.json
1036
+ attribution: skip (the spec is not fiction; attribution applies only with fiction: true)
1037
+ knowledge: skip (the spec is not fiction; knowledge applies only with fiction: true)
1038
+ ```
1039
+
1040
+ Like `check`, it lints the spec first: a spec that fails lint, or is blocked on an open
1041
+ decision, gets no packet, and `prepare` exits with lint's own code. Then it writes one
1042
+ `<station>.packet.json` for each station that applies, in a fixed order (`doctor, lineup,
1043
+ reader, persona, attribution, knowledge`), and prints a `skip` line with the reason for each
1044
+ that does not. `--only doctor,reader` prepares just those. The `--out` folder must already
1045
+ exist. `prepare` refuses to overwrite any file it would write, naming every one, and then writes
1046
+ nothing; `--force` replaces them. The worked examples ship the packets this writes, so add
1047
+ `--force` to write them again there. The same spec and draft, with the same goldens and claims
1048
+ ledger, always give byte-identical packets. A station that throws while building its packet
1049
+ prints `no packet` with a `judge-<name>-crashed` finding; the other packets are still written,
1050
+ and `prepare` exits 1.
1051
+
1052
+ **Hand the judge only the `*.packet.json` files, never the `--out` folder.** Two stations are
1053
+ blind tests with a right answer. `lineup` writes which candidate is the draft's to
1054
+ `lineup.key.json`, and `attribution` writes each line's true speaker to
1055
+ `attribution.key.json`, both beside the packets, for a person to read. A judge who can see them
1056
+ is not judging. `record` never reads either file as truth: it builds the key again from the spec
1057
+ and the draft, so an edited key changes nothing.
1058
+
1059
+ **Give the two blind packets to a judge in a fresh context with no access to the draft.** Every
1060
+ packet names its spec and its draft by path, and a judge that can open files can open those: the
1061
+ draft's own text shows which lineup passage is the draft's, and its speech tags give every
1062
+ attribution answer. Both stations' instructions tell the judge to decide from the packet's inputs
1063
+ alone and open no file the packet names, and a judge with file access can still ignore that, so
1064
+ the instruction is not a guarantee. Paste the packet into a new conversation, or hand it to a
1065
+ person who has not read the draft, rather than to an agent working in the folder the draft is in.
1066
+ The other four packets carry the draft in their inputs and hide nothing, which is why each blind
1067
+ packet goes to its own context, apart from the other packets as well as from the folder: a judge
1068
+ that has read the doctor, reader or persona packet has read the draft.
1069
+
1070
+ Exit codes for `prepare`: **0** written; **1** a station crashed; **2** usage: no spec path, no
1071
+ `--draft`, no `--out`, an `--out` that is missing or not a folder, a draft that cannot be read, a
1072
+ spec without `profile: writing`, an `--only` that names no known judge, or a file that exists
1073
+ without `--force`; and lint's own **1** or **3** when the spec is not ready. `--json` prints the
1074
+ result, and a usage error prints `{ "spec", "draft", "out", "error" }`.
1075
+
1076
+ ### The packet
1077
+
1078
+ Every packet is JSON with these fields, in this order:
1079
+
1080
+ | Field | Holds |
1081
+ |---|---|
1082
+ | `hyperspec_judge` | the packet format's version, `"0.1"` |
1083
+ | `station` | the station's name |
1084
+ | `spec`, `draft` | the paths exactly as `prepare` was given them |
1085
+ | `spec_sha256`, `draft_sha256` | the SHA-256 of each file's bytes |
1086
+ | `rubric` | the `check.rubric` of the block the station reads, verbatim |
1087
+ | `instructions` | fixed text for the station, the same for every spec |
1088
+ | `inputs` | what the judge reads; each station below lists its own |
1089
+ | `verdict_schema` | the verdict's exact shape, as a small JSON Schema |
1090
+
1091
+ A station that needs a rubric skips when its block has none, and a station whose block is
1092
+ deferred to a decision skips too. The paths are kept as given so the packet is the same on every
1093
+ machine, which means **`record` must run in the folder `prepare` ran in**. Run from anywhere else,
1094
+ it cannot find the spec and exits 2.
1095
+
1096
+ ### The evidence rule
1097
+
1098
+ Every verdict field that cites the draft (the doctor's `evidence`, the reader's `lost_at` and
1099
+ `stopped_at`, the persona's `breaks`, the knowledge `leaks`) is a span copied from the draft. A
1100
+ span counts only when:
1101
+
1102
+ - it has at least three words, where a word is a run of letters and digits (so "It's" is two);
1103
+ - it is in the draft after both are normalized: every run of whitespace, line breaks included,
1104
+ becomes one space, and curly, low and angle quotation marks and apostrophes
1105
+ (`‘ ’ ‚ ‛ “ ” „ ‟ ‹ › « »`) become straight ones. Primes (`′ ″`) are not quotation marks and
1106
+ are left alone, case is kept, and nothing else changes, so a changed word is not a quotation;
1107
+ - it matches whole words: a span that starts or ends with a letter or digit may not start or end
1108
+ inside a word of the draft.
1109
+
1110
+ A span that fails makes the verdict invalid (`judge-evidence-missing`, `-too-short` or
1111
+ `-not-found`), and nothing is recorded. The lineup and attribution verdicts carry no evidence:
1112
+ they are blind tests, scored against their keys.
1113
+
1114
+ ### Recording a verdict
1115
+
1116
+ ```bash
1117
+ npx @supersuit/hyperspec judge record essay/judge/lineup.packet.json --verdict essay/sample-verdicts/lineup.verdict.json
1118
+ ```
1119
+
1120
+ ```
1121
+ lineup: fail
1122
+ fail [judge-lineup-picked] the judge picked the draft's passage (D) out of 4 candidates with confidence 0.6: D is the only passage that rests on a survey figure, and it opens by pointing at something outside itself (the survey); A, B and C each turn one claim into an instruction in the second person, with no numbers. (line 32)
1123
+ fix: Revise this passage toward the goldens' voice, where the reason points; if the goldens do not cover this kind of passage, add one that does. Then prepare and judge again.
1124
+ verdict: not-improved (failing stations: lineup)
1125
+ ```
1126
+
1127
+ `record` trusts nothing in the packet file. First it hashes the spec and the draft again; if
1128
+ either no longer matches the hash the packet recorded, the verdict is stale (`judge-stale`,
1129
+ "the draft does not match the hash the packet recorded: it changed since prepare, or the packet
1130
+ was edited"). Then it rebuilds the packet from the spec, the draft and the station's other
1131
+ inputs on disk (the DNA scope, its `scope.md` and goldens, for `lineup`; the claims ledger for
1132
+ `persona`) and
1133
+ requires the file to be exactly those bytes:
1134
+
1135
+ - a packet whose inputs no longer match what those other files produce is stale, and the message
1136
+ names them. Nothing on disk can tell a changed golden from a hand-edited packet, so it says
1137
+ "changed since the packet was prepared, or the packet was edited" and claims neither;
1138
+ - a station that no longer applies because of those files (the claims ledger deleted, say) is
1139
+ stale in the same way, and says why it no longer applies;
1140
+ - anything else (an edited condition, a reformatted file, a hash made to match a changed draft, a
1141
+ station that no longer applies for another reason) is `judge-packet-altered`.
1142
+
1143
+ Either way nothing is recorded. A stale packet prints `<station>: stale packet, nothing
1144
+ recorded`, as `learn record` does, and with `--json` both commands mark it `"invalid": true,
1145
+ "stale": true`, so a script can test `invalid` alone. The fix is to run `judge prepare` again
1146
+ with `--force` and judge the new packet, or, for a station that no longer applies, to restore
1147
+ the file it reads. Only then is the verdict read: it must be JSON (one leading
1148
+ byte order mark is ignored), in the shape the packet gives, with every evidence span found.
1149
+ Every problem is named, and an invalid verdict records nothing. The validator ignores fields it
1150
+ does not know, so a verdict can carry a note of its own; the worked examples' sample verdicts
1151
+ each carry one, `sample`, saying what they are.
1152
+
1153
+ A valid verdict gives the station's status, its findings (a warning prints under the status and
1154
+ never fails it), and for `attribution` one summary line. Then one line goes to the runs ledger.
1155
+
1156
+ Exit codes for `record`: **0** the station passed; **1** it failed, or the verdict is invalid,
1157
+ or the packet is stale or altered; **2** usage: no packet path, no `--verdict`, a packet or verdict
1158
+ file that cannot be read, a file that is not a judge packet or names an unknown judge, or a
1159
+ packet whose spec (which must carry `profile: writing`) or draft cannot be read. `--json` prints
1160
+ the result, and a usage error prints `{ "packet", "verdict", "error" }`.
1161
+
1162
+ In the tables below, **fail** fails the station, **warn** is printed and never fails it,
1163
+ **invalid** refuses the verdict, and **stale** refuses the packet; an invalid or stale verdict
1164
+ records nothing.
1165
+
1166
+ ### doctor
1167
+
1168
+ Grades the draft against every goal condition, and asks whether the reader would take the next
1169
+ step now. Applies when `writing.goal` is written and its check has a rubric. Inputs: `goal`
1170
+ (`from`, `to`, `next_if_worked`, and `change` with its `kind` and `text`), `conditions` (each
1171
+ condition id with its requirement's `text` and `fails_when`) and the `draft`. The verdict is
1172
+ `{ conditions: [{ id, pass, evidence, note }], would_take_next_step, evidence }`: every condition
1173
+ id exactly once, `pass` and `would_take_next_step` true or false, a `note` on every condition
1174
+ (one that fails must say why), and the top-level `evidence` for the passage that decided the next
1175
+ step. It passes when every condition passes and the reader would take the next step.
1176
+
1177
+ | Id | Kind | Meaning |
1178
+ |---|---|---|
1179
+ | `judge-doctor-condition` | fail | a condition fails; gives the judge's note, at its evidence's line |
1180
+ | `judge-doctor-next-step` | fail | the reader would not take `goal.next_if_worked` now |
1181
+ | `judge-doctor-condition-missing` | invalid | a condition in the packet has no entry |
1182
+ | `judge-doctor-condition-unknown` | invalid | the verdict grades an id that is not a condition in the packet |
1183
+ | `judge-doctor-condition-duplicate` | invalid | a condition is graded more than once |
1184
+
1185
+ ### lineup
1186
+
1187
+ A blind test of voice. Applies when `writing.dna` names a `scope_dir` whose goldens can be read
1188
+ and its check has a rubric, at least one golden has a prose paragraph, and the draft has a prose
1189
+ paragraph that is not already a golden's. A draft that is one paragraph and nothing else is
1190
+ skipped: its candidate would be the whole file, and the packet's `draft_sha256` would identify it. A prose paragraph is a run of non-blank lines that are
1191
+ all plain text: headings and code fences end one, and a block holding a list item, a quotation, a
1192
+ table row, a thematic break or HTML is left out whole, as is indented code.
1193
+
1194
+ Every candidate is one prose paragraph reflowed onto a single line, so none can be told by its
1195
+ formatting. The target length is the median, in characters, of every prose paragraph of every
1196
+ golden in the scope. The first three goldens by file name that have one each give their
1197
+ paragraph closest to that length, and the draft gives its paragraph closest to it, skipping any
1198
+ that is word for word a golden's; ties go to the earliest. The candidates are shuffled with a
1199
+ seed derived from the draft's full text (the SHA-256 of `hyperspec lineup seed`, a line break, and
1200
+ the text), which the packet does not carry, so the packet cannot reveal the order: not even its
1201
+ `draft_sha256`, which is a different hash. The same draft always gets the same labels, labeled A
1202
+ to D. Inputs: `scope` (the scope's `writer`, `form`, `audience` and `purpose`) and `candidates`,
1203
+ each `{ label, text }`; no path and no source. `lineup.key.json` records the draft's label, the
1204
+ draft line its paragraph starts on, and where every candidate came from.
1205
+
1206
+ The verdict is `{ pick, confidence, reason }`: a label, a number from 0 to 1, and what decided it.
1207
+ It passes when the pick is not the draft's passage: the judge could not tell. A judge who picks at
1208
+ random also passes, three times in four with four candidates, so one lineup is weak evidence of a
1209
+ voice. A failing lineup is the stronger signal, and its reason says where to look.
1210
+
1211
+ | Id | Kind | Meaning |
1212
+ |---|---|---|
1213
+ | `judge-lineup-picked` | fail | the judge picked the draft's passage; gives its confidence and reason, at the paragraph's line |
1214
+ | `judge-lineup-pick-unknown` | invalid | the pick is not one of the lineup's labels |
1215
+
1216
+ ### reader
1217
+
1218
+ Reads the draft as the audience block's reader. Applies when `writing.audience` is written and its
1219
+ check has a rubric. Inputs: `audience` (`who`, `funnel_now`, `knows`, `terms`, `believes_now`,
1220
+ `wants`, `reads_on`, `reader`) and the `draft`. The verdict is `{ lost_at: [{ evidence, why }],
1221
+ stopped_at, would_take_next_step, next_step }`, where `stopped_at` is `{ evidence, why }` or
1222
+ `null` when the reader read to the end, and must be present either way. The reader names its own
1223
+ next step: it is not shown `goal.next_if_worked`, so a reader that would act and a doctor that
1224
+ says the reader would not take the spec's next step can both be right. In the essay example they
1225
+ agree: both name the card. It passes when the reader read to the end and would take its next step
1226
+ now.
1227
+
1228
+ | Id | Kind | Meaning |
1229
+ |---|---|---|
1230
+ | `judge-reader-lost` | warn | the reader got lost here, and why |
1231
+ | `judge-reader-stopped` | fail | the reader stopped reading here, and why |
1232
+ | `judge-reader-next-step` | fail | the reader would not take its next step now |
1233
+
1234
+ ### persona
1235
+
1236
+ Reads the draft as the persona block's speaker, against the claims ledger. Applies when
1237
+ `writing.persona` is written and its check has a rubric, and the claims ledger, if
1238
+ `sources.ledger` names one, can be read. Inputs: `persona` (`identity`, `stance`, `may_assert`,
1239
+ `will_not_say`), `claims` (the text of every well-formed line of the claims ledger, in order) and
1240
+ the `draft`. With no ledger declared, `claims` is `null`: facts cannot be checked against
1241
+ sources, so the instructions say not to report one, the schema leaves that kind out, and a break
1242
+ of that kind is invalid. The verdict is `{ breaks: [{ evidence, kind, why }] }`, where `kind` is
1243
+ `stance` (the voice leaves its stance), `assertion` (it asserts something outside `may_assert`),
1244
+ `will_not_say` or `unsourced_fact` (a fact no claim holds). It passes when there is no break.
1245
+
1246
+ | Id | Kind | Meaning |
1247
+ |---|---|---|
1248
+ | `judge-persona-break` | fail | the voice breaks here; the message names the kind and gives the judge's why |
1249
+ | `judge-persona-kind-unknown` | invalid | a break's kind is not one of the four |
1250
+ | `judge-persona-no-ledger` | invalid | a break is `unsourced_fact`, and the spec declares no claims ledger |
1251
+
1252
+ ### attribution
1253
+
1254
+ Fiction only: a blind test of whether the characters' voices can be told apart. The packet holds
1255
+ the draft's dialogue lines with the narration and the speaker removed, and each speaking
1256
+ character's `speech` block, `golden_lines` and `rejected_lines`; it carries no draft. The judge
1257
+ names a speaker for every line, and `record` scores the answers against the true speakers, which
1258
+ only `attribution.key.json` holds. Applies when `fiction: true`, at least two characters have a
1259
+ speech block and one of them has a check rubric, and the draft has attributable lines from at
1260
+ least two speakers.
1261
+
1262
+ A dialogue line is a double-quoted span, straight or curly, inside one paragraph, outside code. A
1263
+ quote split by a speech tag (`"Twenty minutes," Ines said, "then we fold it."`) is one line when
1264
+ its first part ends in a comma and the narration between the parts is exactly one tag and its
1265
+ comma, nothing else. Narration that holds anything more joins nothing: in `"Leave it there," Ines
1266
+ said, and Theo muttered, "No chance at all."` the first part is Ines's, and the second is left
1267
+ out, since a second speaker brought in by a beat, a pronoun or a verb off the list cannot be read
1268
+ mechanically. The second part is left out whatever ends that narration, even another tag (`Ines
1269
+ said, and Theo said,`), and so is the second part of `"Twenty minutes," Ines said, wiping her
1270
+ hands, "then we fold it."`. A line's speaker comes only from a speech tag:
1271
+ narration in the same paragraph directly after the closing mark (`"...," Ines said`) or directly
1272
+ before the opening mark, ending in a comma or colon (`Ines said, "..."`). A tag is a subject and
1273
+ one of the verbs said, asked, told, replied, called, whispered, shouted, answered, added and went
1274
+ on, or their present tense (says, asks, goes on). The subject is:
1275
+
1276
+ - a character's id, or its `name` if it has one: whole words, any case, and an id's words may be
1277
+ joined by a space, a hyphen or an underscore, so `old-man` is named by "old man";
1278
+ - "I", when `persona.identity` is `character:<id>`: the narrator speaks;
1279
+ - "she" or "he", when exactly two characters have speech blocks and one of them narrates: the
1280
+ other one speaks. After a quote the pronoun must be lower case.
1281
+
1282
+ The verb may come first ("said Ines") only for said, replied, whispered, shouted and went on (and
1283
+ their present tense), since "Ines told Theo" names the person spoken to. Anything else leaves the
1284
+ line out, and every doubt does: no tag, a possessive ("Ines's") or an action beat ("Theo nodded"),
1285
+ a subject that is not a character with a speech block, tags that name two different speakers, or a
1286
+ line of three words or more that repeats, or is part of, a golden or rejected line, which would
1287
+ give its speaker away. So does the second quote in
1288
+ `"Not a bakery," she said. "They want the room."`, since no tag sits against it. The key is never wrong, and recall pays for it: the worked
1289
+ story has 36 dialogue lines, and 19 of them are in the test, 10 by Ines and 9 by Theo. The packet
1290
+ counts the lines left out, and the key lists each with its reason.
1291
+
1292
+ Inputs: `characters`, `lines` (each `{ id, text }`, numbered L1 up in draft order, whitespace
1293
+ reflowed) and `excluded`, the count left out. The verdict is `{ lines: [{ id, speaker }] }`,
1294
+ every line id exactly once, a speaker by id or name, in any case. Accuracy is taken per speaker
1295
+ and averaged, so naming one character for every line cannot pass: the station passes when that
1296
+ mean is 80 percent or more, compared exactly, with percentages rounded down. This is the worked
1297
+ story's sample:
1298
+
1299
+ ```
1300
+ attribution: pass
1301
+ accuracy per speaker: ines 10/10 (100%), theo 9/9 (100%); mean 100%, passing at 80%; 17 dialogue lines left out (8 repeating a golden or rejected line, 9 with no speech tag)
1302
+ verdict: one-shot
1303
+ ```
1304
+
1305
+ | Id | Kind | Meaning |
1306
+ |---|---|---|
1307
+ | `judge-attribution-accuracy` | fail | the mean accuracy over speakers is under 80%; gives each speaker's |
1308
+ | `judge-attribution-miss` | warn | a line was given to the wrong character, at its draft line |
1309
+ | `judge-attribution-speaker-unknown` | invalid | a speaker is not a character in the packet |
1310
+ | `judge-attribution-line-missing` | invalid | a line in the packet has no answer |
1311
+ | `judge-attribution-line-unknown` | invalid | the verdict answers an id that is not a line in the packet |
1312
+ | `judge-attribution-line-duplicate` | invalid | a line is answered more than once |
1313
+
1314
+ ### knowledge
1315
+
1316
+ Fiction only: whether anyone knows something before their timeline gives it to them. Applies when
1317
+ `fiction: true` and at least one character has a knowledge timeline (an entry with both `by` and
1318
+ `knows`) and one of those characters has a check rubric. Inputs: `characters` (each with its
1319
+ `knowledge` entries) and the `draft`. The verdict is `{ leaks: [{ character, evidence,
1320
+ knows_too_early }] }`, and it passes when there is no leak.
1321
+
1322
+ hyperspec does not read the draft's structure: `by` goes into the packet as the spec wrote it,
1323
+ and the judge maps it onto the draft. The worked story's timelines say `scene-1` to `scene-4`,
1324
+ and its draft's four headings are times, 3:40 to 6:55, in the scene list's order, so a judge maps
1325
+ them by order. Write `by` as something a reader of the draft can find.
1326
+
1327
+ | Id | Kind | Meaning |
1328
+ |---|---|---|
1329
+ | `judge-knowledge-leak` | fail | a character knows something too early, at its evidence's line |
1330
+ | `judge-knowledge-character-unknown` | invalid | a leak names a character with no timeline in the packet |
1331
+
1332
+ ### Any judgment station
1333
+
1334
+ | Id | Kind | Meaning |
1335
+ |---|---|---|
1336
+ | `judge-verdict-not-json` | invalid | the verdict file is not JSON |
1337
+ | `judge-verdict-shape` | invalid | a field is missing, of the wrong type or out of range, or a required note, reason or why is empty |
1338
+ | `judge-evidence-missing` | invalid | an evidence field is empty or not a string |
1339
+ | `judge-evidence-too-short` | invalid | an evidence span has fewer than three words |
1340
+ | `judge-evidence-not-found` | invalid | an evidence span is not in the draft, word for word |
1341
+ | `judge-stale` | stale | the spec, the draft, or a file the station reads changed since the packet was prepared, or the packet was edited |
1342
+ | `judge-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
1343
+ | `judge-<name>-crashed` | invalid | the station threw; at `prepare` its packet is not written, at `record` nothing is recorded |
1344
+
1345
+ ### Judge lines in the runs ledger
1346
+
1347
+ Each recorded verdict appends one line to the spec's `improvement.ledger`:
1348
+
1349
+ ```json
1350
+ {"at":"2026-09-29T16:27:36.883Z","kind":"judge","station":"lineup","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","packet_sha256":"<sha256 of the packet>","inputs_sha256":"<sha256 of its inputs>","status":"fail","verdict":"not-improved","reason":"failing stations: lineup"}
1351
+ ```
1352
+
1353
+ `draft` is relative to the spec's folder, as in a check line. `packet_sha256` is the SHA-256 of
1354
+ the packet the judge was shown, taken with its two paths written as the ledger writes them
1355
+ (relative to the spec's folder), so the same packet prepared from another folder, or as
1356
+ `./draft.md`, hashes the same. `inputs_sha256` is the SHA-256 of the packet's `inputs` alone. The
1357
+ verdict follows the same rules as a check line, compared with
1358
+ the most recent earlier judge line for the same station and the same draft, with three more:
1359
+
1360
+ - **What changed** is judged by the packet. The draft or the spec is named when its bytes
1361
+ changed. When neither did and the packet's inputs did, the files the station reads besides them
1362
+ are named: `the DNA scope (<scope_dir>: scope.md and goldens)` for `lineup`, `the claims ledger
1363
+ (<path>)` for `persona`. When only the rest of the packet changed, which a later hyperspec
1364
+ release can do by rewording a station's instructions, the reason says `the packet's fixed text
1365
+ (hyperspec's instructions or format) changed`. So adding the golden a failing lineup asked for, or the claims a
1366
+ failing persona asked for, and passing on the new packet is `improved`, "the DNA scope
1367
+ (dna/essay-new-managers-teach: scope.md and goldens) changed; stations now pass: lineup". An
1368
+ improved judge line always names what changed, "draft changed; stations now pass: doctor".
1369
+ - **one-shot** also needs these draft bytes never to have been judged by this station before,
1370
+ under any name. A copy or a rename of a judged draft gets `not-improved`, with the reason
1371
+ "these draft bytes were judged before as <path>: <status>".
1372
+ - **improved** also needs the packet to have changed since the failing line. A judge can answer
1373
+ differently about an identical packet, and that is not the work improving: the verdict is
1374
+ `not-improved`, "the verdict changed; nothing the judge was shown changed".
1375
+
1376
+ Otherwise the reasons are check's, with `judgment` where check says `check`: `failing stations:
1377
+ doctor`, `no change since the last passing judgment`, `no change since the last judgment; still
1378
+ failing: doctor`, or `draft changed; every station still passes`. Check and judge each read only
1379
+ their own lines, and learn reads none, so judging a draft never changes what `check` says about
1380
+ it, or the reverse. Judge lines keep lint's test 9 passing, as check lines do.
1381
+
1382
+ ### The worked examples
1383
+
1384
+ Both examples ship the packets `prepare` writes for their drafts, in `essay/judge/` and
1385
+ `story/judge/`, with no key beside them, and one sample verdict per packet in
1386
+ `essay/sample-verdicts/` and `story/sample-verdicts/`. The samples were filled in by hand, as
1387
+ one careful judge would, to show the shape and what `record` does with it; each says so in its
1388
+ `sample` field. Another judge may answer differently. A test records every sample on a fresh
1389
+ copy on every release.
1390
+
1391
+ | Example | Station | Sample | What the judge found |
1392
+ |---|---|---|---|
1393
+ | essay | doctor | pass | every condition holds, and the reader would write the three questions on a card, the goal's next step and the draft's close |
1394
+ | essay | lineup | fail | the draft's passage is the only one resting on a survey figure; the goldens hold no numbers, so they do not cover this kind of passage |
1395
+ | essay | reader | pass | read to the end, lost nowhere; the reader's own next step is the card |
1396
+ | essay | persona | pass | the mentor stance holds, and every figure is in the claims ledger |
1397
+ | story | doctor | pass | every condition holds, and a reader would look for the author's other stories |
1398
+ | story | reader | pass | read to the end; lost for a moment at "proving cabinet" and "peel", two warnings |
1399
+ | story | persona | fail | three process details (the deck oven's heat-up time, the rolls' bake time, how the starter is fed) are in no claim, and the rubric allows none outside the ledger |
1400
+ | story | attribution | pass | all 19 lines named right by voice alone |
1401
+ | story | knowledge | pass | neither character knows anything early |
1402
+
1403
+ Both examples pass every station of `check`. Each failure here is something no deterministic
1404
+ station can see.
1405
+
1406
+ ## Learning from edits
1407
+
1408
+ A factory writes a first draft, and a person edits it into the draft they approve. Every edit is
1409
+ something the spec did not say, or did not say well enough. `learn prepare` diffs the two drafts
1410
+ sentence by sentence and writes a packet of the edits; an outside judge names, for each edit, the
1411
+ one block of the spec that would have prevented it; `learn record` validates that verdict, counts
1412
+ the edits by block, and names one next move for the block with the most. It never edits the spec:
1413
+ the move is a suggestion for whoever keeps it.
1414
+
1415
+ ```bash
1416
+ npx @supersuit/hyperspec learn prepare essay.hyperspec.md --first essay/learn/first-draft.md --approved essay/draft.md --out essay/learn
1417
+ npx @supersuit/hyperspec learn record essay/learn/learn.packet.json --verdict essay/sample-verdicts/learn.verdict.json
1418
+ ```
1419
+
1420
+ ```
1421
+ essay/learn/learn.packet.json
1422
+ 5 edits over 8 sentences: 3 replaced, 2 deleted
1423
+ ```
1424
+
1425
+ ```
1426
+ learn: 5 edits classified
1427
+ dna 2 edits, 4 sentences
1428
+ materials 1 edit, 2 sentences
1429
+ persona 1 edit, 1 sentence
1430
+ none 1 edit, 1 sentence
1431
+ next move: dna, 4 of 8 sentences (2 of 5 edits): add a golden or a style rule
1432
+ verdict: not-improved (edits by block: dna 2 edits (4 sentences), materials 1 edit (2 sentences), persona 1 edit (1 sentence), none 1 edit (1 sentence); not yet applied to the spec)
1433
+ ```
1434
+
1435
+ `prepare` lints the spec first, exactly as `judge prepare` does, then writes
1436
+ `<out>/learn.packet.json` and prints how many edits it found, or `no edits: the first draft and the
1437
+ approved draft match; nothing to learn`. The rules on `--out`, `--force` and the paths are the
1438
+ judge's: the folder must exist, an existing packet needs `--force`, and `record` runs in the
1439
+ folder `prepare` ran in. The essay example ships this pair and its packet in `essay/learn/`, so
1440
+ add `--force` to write the packet again there.
1441
+
1442
+ Exit codes: `prepare` **0** written, **2** usage (no spec path, `--first`, `--approved` or
1443
+ `--out`, a folder that is missing, a draft that cannot be read, a spec without `profile: writing`,
1444
+ a packet that exists without `--force`, or drafts too large to diff), and lint's own **1** or
1445
+ **3** when the spec is not ready; `record` **0** recorded, **1** the verdict is invalid or the
1446
+ packet stale or altered (nothing recorded), **2** usage.
1447
+
1448
+ ### Sentence units
1449
+
1450
+ The drafts are compared one unit at a time, and a unit is one sentence, one heading, or one list
1451
+ item:
1452
+
1453
+ - A blank line always ends a unit. When the first line of a paragraph is a Markdown heading, that
1454
+ line is a unit of its own; a later line that starts with `#` is a hard wrap in the text, and is
1455
+ read as text.
1456
+ - Other lines are split into sentences at `.`, `!` or `?` followed by whitespace, never inside a
1457
+ quotation that opens and closes on one line (straight or curly), so a quoted passage of several
1458
+ sentences on one line is one unit. Each list item starts a unit.
1459
+ - A unit that ends in a common abbreviation (Mr., Mrs., Ms., Dr., Prof., Sr., Jr., St., vs., cf.,
1460
+ e.g., i.e.) or a single capital initial ("J.") runs on into the next, and so does a unit
1461
+ followed by one that starts with a lower-case letter ("the U.S. economy").
1462
+
1463
+ ### The edits
1464
+
1465
+ Units are compared with their whitespace collapsed and matched by a longest common subsequence.
1466
+ Every run of unmatched units is one edit, a hunk: `deleted` (only in the first draft), `inserted`
1467
+ (only in the approved one) or `replaced`. A hunk never crosses a paragraph break or a heading, so
1468
+ a paragraph rewritten from end to end is one hunk and a change in two paragraphs is two. A hunk's
1469
+ texts are the drafts' own words from its first unit to its last, and its `sentences` is the
1470
+ larger of its two unit counts. A run that reads the same once whitespace is collapsed is a
1471
+ reflow, not an edit, so rewrapping lines, or joining or splitting paragraphs without changing a
1472
+ word, makes no hunk. The drafts' common start and end are set aside first; if what is left would
1473
+ need more than 10,000,000 comparisons, `prepare` stops with a usage error naming both sentence
1474
+ counts. Learn from a chapter or a scene at a time.
1475
+
1476
+ ### The learn packet
1477
+
1478
+ `{ "hyperspec_learn": "0.1", "spec", "spec_sha256", "first", "first_sha256", "approved",
1479
+ "approved_sha256", "blocks", "instructions", "hunks", "verdict_schema" }`. `blocks` lists the
1480
+ writing blocks the spec has written, in schema order, then `none`; a deferred block is not
1481
+ listed. Each hunk is `{ id, kind, first, approved, sentences }`, with ids E1 up, and `null` on
1482
+ the side that has no text. The packet carries no draft beyond its hunks: the instructions tell
1483
+ the judge to read the spec at its path. A learn packet has no answer key.
1484
+
1485
+ The verdict is `{ edits: [{ id, block, why }] }`: every hunk id exactly once, a `block` from the
1486
+ packet's `blocks`, and a `why` saying what that block should have said. `none` means no block of
1487
+ the spec could have prevented the edit, a typo say. As with a judge packet, `record` hashes the
1488
+ spec and both drafts again, rebuilds the packet, and requires the file to be exactly those bytes
1489
+ before it reads the verdict.
1490
+
1491
+ ### The tally and the next move
1492
+
1493
+ The tally counts, for each block the verdict names, its edits and the sentences they touched.
1494
+ Blocks are ordered by sentences, then edits, then the order of the table below with `none` last,
1495
+ and the next move goes to the first block that is not `none`, so a paragraph rewritten from end to
1496
+ end weighs as the sentences it rewrote. There is one move per block:
1497
+
1498
+ | Block | Next move |
1499
+ |---|---|
1500
+ | `materials` | mark or add the material the edit drew on |
1501
+ | `dna` | add a golden or a style rule |
1502
+ | `persona` | tighten the persona's stance, may_assert or will_not_say |
1503
+ | `audience` | extend the audience's knows or terms |
1504
+ | `goal` | tighten a goal condition's fails_when |
1505
+ | `form` | adjust the form block's length or shape |
1506
+ | `spine` | restate the spine's claim or its order |
1507
+ | `sources` | add or cite a source in the claims ledger |
1508
+ | `characters` | extend a character's speech or knowledge |
1509
+
1510
+ When every edit is `none`, the move is `none`: no block could have prevented any edit, so the spec
1511
+ has nothing to learn from the pair.
1512
+
1513
+ `record` appends one line to the runs ledger, `{ at, kind: "learn", first, first_sha256,
1514
+ approved, approved_sha256, spec_sha256, verdict, reason, tally }`, with both drafts' paths
1515
+ relative to the spec's folder. Its verdict is always `not-improved`: the spec has not changed yet,
1516
+ and the reason gives the counts by block. Learn reads no earlier line, and check and judge ignore
1517
+ learn lines.
1518
+
1519
+ | Id | Kind | Meaning |
1520
+ |---|---|---|
1521
+ | `learn-verdict-not-json` | invalid | the verdict file is not JSON |
1522
+ | `learn-verdict-shape` | invalid | the verdict is not `{ edits: [...] }`, or an entry is not an object with an id |
1523
+ | `learn-edit-unknown` | invalid | an id is not a hunk in the packet |
1524
+ | `learn-edit-duplicate` | invalid | a hunk is classified more than once |
1525
+ | `learn-edit-missing` | invalid | a hunk is not classified |
1526
+ | `learn-block-unknown` | invalid | a block is not one of the nine blocks or `none` |
1527
+ | `learn-block-absent` | invalid | a block is one the spec has not written |
1528
+ | `learn-why-missing` | invalid | a why is empty or not a string |
1529
+ | `learn-stale` | stale | the spec or a draft no longer matches the hash the packet recorded: it changed since prepare, or the packet was edited |
1530
+ | `learn-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
1531
+
1532
+ The essay example's pair is a first draft that differs from `essay/draft.md` by five edits: a
1533
+ paragraph of hedged advice and a hedged sentence, where the approved draft gives instructions; a
1534
+ claim about what most managers do; the walking aside from the voice memo; and a typo. The sample
1535
+ verdict puts the two hedges on `dna`, since no golden or style rule shows advice given flat, and
1536
+ the tally sends the next move there.
1537
+
1012
1538
  ## Deferring a block
1013
1539
 
1014
1540
  A block can be deferred, never silently missing. A required block that is absent fails test 1
@@ -1079,12 +1605,18 @@ both drafts pass every station of `check`: the essay with one dna warning, descr
1079
1605
  since it is fiction. A test lints both
1080
1606
  specs and checks both drafts on every release, so they cannot drift from the tool.
1081
1607
 
1608
+ Each also ships the packets `judge prepare` writes for its draft, in `essay/judge/` and
1609
+ `story/judge/`, and one sample verdict per packet in `essay/sample-verdicts/` and
1610
+ `story/sample-verdicts/`, filled in by hand and marked as samples; what each found is under
1611
+ [The worked examples](#the-worked-examples). The essay adds a learn pair in `essay/learn/`: a
1612
+ first draft, the packet `learn prepare` writes comparing it with `essay/draft.md`, and a sample
1613
+ learn verdict (see [Learning from edits](#learning-from-edits)). A test checks that every packet
1614
+ is what `prepare` writes now and records every sample.
1615
+
1082
1616
  ## What later versions add
1083
1617
 
1084
- This release is the schema, its lint, marked materials, scoped DNA, and `check` with seven
1085
- deterministic stations. Next come the judgment stations: the simulated reader, the blind lineup,
1086
- the persona judge and the doctor. hyperspec calls no model, so `check` will write each one as a
1087
- packet, the draft and the rubric and the materials the judge needs, for an outside judge to fill
1088
- in, and read the filled packet back as a station result. After that, a learn step that reads the
1089
- runs ledger for the stations that keep failing and the changes that made them pass, so a fix
1090
- lands in the spec or the skill that wrote the draft rather than in one draft.
1618
+ This release is the schema, its lint, marked materials, scoped DNA, `check` with seven
1619
+ deterministic stations, six judgment stations written as packets for an outside judge, and learn.
1620
+ Next: lineups over several passages of one draft, so that one lucky pick carries less weight, and
1621
+ a learn step that reads the runs ledger across drafts for the stations that keep failing and the
1622
+ changes that made them pass, beside what one pair of drafts shows.