@supersuit/hyperspec 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/CHANGELOG.md +166 -0
  2. package/README.md +62 -1
  3. package/SPEC.md +4 -4
  4. package/WRITING.md +770 -9
  5. package/bin/hyperspec.mjs +252 -0
  6. package/examples/writing/essay/claims.jsonl +9 -0
  7. package/examples/writing/essay/draft.md +82 -0
  8. package/examples/writing/essay/judge/doctor.packet.json +108 -0
  9. package/examples/writing/essay/judge/lineup.packet.json +64 -0
  10. package/examples/writing/essay/judge/persona.packet.json +73 -0
  11. package/examples/writing/essay/judge/reader.packet.json +93 -0
  12. package/examples/writing/essay/learn/first-draft.md +84 -0
  13. package/examples/writing/essay/learn/learn.packet.json +106 -0
  14. package/examples/writing/essay/materials/interview-notes.md.segments.jsonl +2 -2
  15. package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
  16. package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
  17. package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
  18. package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
  19. package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
  20. package/examples/writing/essay.hyperspec.md +23 -7
  21. package/examples/writing/story/claims.jsonl +9 -0
  22. package/examples/writing/story/draft.md +267 -0
  23. package/examples/writing/story/judge/attribution.packet.json +194 -0
  24. package/examples/writing/story/judge/doctor.packet.json +108 -0
  25. package/examples/writing/story/judge/knowledge.packet.json +77 -0
  26. package/examples/writing/story/judge/persona.packet.json +73 -0
  27. package/examples/writing/story/judge/reader.packet.json +94 -0
  28. package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
  29. package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
  30. package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
  31. package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
  32. package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
  33. package/examples/writing/story.hyperspec.md +24 -5
  34. package/package.json +1 -1
  35. package/src/check.mjs +191 -0
  36. package/src/dna.mjs +4 -1
  37. package/src/draft.mjs +26 -0
  38. package/src/judge.mjs +386 -0
  39. package/src/judges/attribution.mjs +360 -0
  40. package/src/judges/doctor.mjs +126 -0
  41. package/src/judges/index.mjs +31 -0
  42. package/src/judges/knowledge.mjs +111 -0
  43. package/src/judges/lineup.mjs +272 -0
  44. package/src/judges/persona.mjs +137 -0
  45. package/src/judges/reader.mjs +111 -0
  46. package/src/learn.mjs +422 -0
  47. package/src/ledger.mjs +108 -0
  48. package/src/sentences.mjs +81 -0
  49. package/src/stations/claims.mjs +155 -0
  50. package/src/stations/dna.mjs +126 -0
  51. package/src/stations/form.mjs +115 -0
  52. package/src/stations/index.mjs +27 -0
  53. package/src/stations/links.mjs +275 -0
  54. package/src/stations/private.mjs +117 -0
  55. package/src/stations/quotes.mjs +170 -0
  56. package/src/stations/terms.mjs +131 -0
  57. package/src/stations/util.mjs +99 -0
  58. package/src/writing-fields.mjs +26 -2
  59. package/src/writing.mjs +1 -1
package/WRITING.md CHANGED
@@ -209,6 +209,8 @@ writing:
209
209
  wants: a plan for the first meeting
210
210
  reads_on: a phone, in the ten minutes before the meeting
211
211
  reader: person # person | agent
212
+ terms: # optional: terms the piece uses that the reader may not know
213
+ - skip-level
212
214
  check:
213
215
  station: term check against knows
214
216
  rubric: simulated reader reports where it got lost
@@ -217,7 +219,7 @@ writing:
217
219
  goal:
218
220
  from: plans to run the meeting from their own list
219
221
  to: hands the meeting to the report
220
- next_if_worked: copies the three questions into the invite
222
+ next_if_worked: writes the three questions on a card
221
223
  change:
222
224
  kind: action # belief | action | feeling
223
225
  text: the reader asks the three questions and waits
@@ -317,6 +319,17 @@ cannot pass a presence check. A placeholder is a whole value, trimmed and in any
317
319
  marks, or an ellipsis, optionally followed by a trailing `.`, `:` or `!`. Real text that starts
318
320
  with one of those, such as `TODO: write the opening`, counts as present, and so does `none`.
319
321
 
322
+ `audience.terms` is checked by the `terms` station in `hyperspec check`, and that check is a
323
+ **mechanical proxy, not an understanding of meaning**: it looks for a definition-SHAPED phrase
324
+ near the term's first appearance (the word `is`, `means`, `refers to`, a colon within a few words,
325
+ or an immediate parenthetical), not for whether that phrase defines the term. A sentence
326
+ like "A hyperspec is mentioned here" reads as a definition of "hyperspec" by this rule, because
327
+ `is` immediately follows the word, even though nothing about the term is explained. This is
328
+ deliberate and known, not a bug to fix later in this station: reading for meaning is a judgment
329
+ call, and hyperspec's deterministic stations do not make judgment calls. A later release adds a
330
+ simulated-reader station that reads for meaning instead of shape; `terms` stays the fast,
331
+ mechanical first pass.
332
+
320
333
  ## The test mapping
321
334
 
322
335
  Each row lists what the writing profile adds to that test. The core conditions in
@@ -324,7 +337,7 @@ Each row lists what the writing profile adds to that test. The core conditions i
324
337
 
325
338
  | Test | A writing spec fails it when |
326
339
  |---|---|
327
- | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered; `dna.scope_dir` is present and is a placeholder; with `dna.scope_dir`, its `scope.md` is missing, unreadable or lacks a field, its writer, form, audience or purpose differs from the spec's, or its `goldens/` folder is missing or empty, or holds a golden that cannot be read, whose frontmatter never closes, or that has no passage. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
340
+ | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered; `dna.scope_dir` is present and is a placeholder; with `dna.scope_dir`, its `scope.md` is missing, unreadable or lacks a field, its writer, form, audience or purpose differs from the spec's, or its `goldens/` folder is missing or empty, or holds a golden that cannot be read, whose frontmatter never closes, or that has no passage; `audience.terms`, when present, holds a non-string entry or has no real entries at all. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
328
341
  | 2 every requirement can fail | `goal.conditions` lists fewer than five or more than ten distinct ids, lists an id twice, or names an id that is not a top-level requirement |
329
342
  | 3 every requirement names its check | a block or a character has no `check` with a `station` or a `rubric` |
330
343
  | 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items`, or a segment that is not in that material's segments file; a segment's text does not match its material word for word; a material changed after it was marked; a claim segment has no `source` and no `own`, a story no `teller`, a quote no `speaker`; with `dna.scope_dir`, a golden in the scope has no `approved_by`, an approver that starts `agent:`, or no `source` |
@@ -786,6 +799,742 @@ source, and a `features.json` that `dna measure` wrote. A test measures the fold
786
799
  release and requires the same bytes, so the example cannot drift from the tool. The short story
787
800
  beside it lists its goldens in the spec with no scope folder, the 0.4 shape, which still passes.
788
801
 
802
+ ## Checking a draft
803
+
804
+ Once a spec lints clean and a draft exists, `check` runs the spec's deterministic stations
805
+ against the draft:
806
+
807
+ ```bash
808
+ npx @supersuit/hyperspec check essay.hyperspec.md --draft essay/draft.md
809
+ ```
810
+
811
+ It lints the spec first. A spec that fails lint, or is blocked on an open decision, runs no
812
+ station and exits with lint's own code, because a draft cannot be checked against a spec that is
813
+ not ready. Then it runs seven stations in a fixed order and prints one line for each: `pass`,
814
+ `fail` with its findings, or `skip` with the reason. A warning prints under its station and never
815
+ fails it. This is the essay example's draft:
816
+
817
+ ```
818
+ form: pass
819
+ terms: pass
820
+ claims: pass
821
+ quotes: pass
822
+ private: pass
823
+ dna: pass
824
+ warn [station-dna-drift] first_person_singular_rate is 22.892 in the draft; the scope's goldens measure 0, band 0 to 5
825
+ fix: Bring first_person_singular_rate back inside the band, or, if the scope no longer describes this writer, re-measure it with better goldens.
826
+ links: pass
827
+ verdict: one-shot
828
+ ```
829
+
830
+ `--only form,terms` runs just those stations, still in the fixed order, and its ledger line is
831
+ marked partial (see [The runs ledger](#the-runs-ledger)). `--json` prints the whole result, every
832
+ finding included; a spec that is not ready prints lint's result with `lintBlocked: true` instead,
833
+ and a usage error prints `{ "spec", "draft", "error" }`. A finding names the draft line it points
834
+ at where there is one, quotes at most 80 characters of the draft, and never prints an absolute
835
+ path. A UTF-8 byte order mark at the start of the draft is ignored.
836
+
837
+ Exit codes: **0** every station that ran passed (a skip or a warning does not fail it); **1** a
838
+ station failed; **2** usage: no spec path, no `--draft`, a draft that cannot be read, a spec
839
+ without `profile: writing`, or an `--only` that names no known station; and lint's own **1** or
840
+ **3** when the spec is not ready.
841
+
842
+ Every station is a plain function of the spec and the draft. None of them calls a model, and none
843
+ of them touches the network. What each one checks, and what it cannot:
844
+
845
+ ### form
846
+
847
+ Counts the draft's words, by the same word definition `dna measure` uses, against
848
+ `form.length`. Only `unit: words` is measured; any other unit skips the whole station rather than
849
+ checking half of it. Every `required_parts` entry must appear as an ATX heading (`#` to
850
+ `######`, indented at most three spaces, closing `#`s allowed) whose text equals the part,
851
+ ignoring case, or as a line that starts with the part and a colon, for the fields a form fills in
852
+ place (`To:`, `Subject:`). An underlined (Setext) heading does not count, and neither does
853
+ anything inside a code block. It cannot tell whether the section under a heading does what the
854
+ part is for, so write `required_parts` as the headings the piece will carry,
855
+ as both examples do.
856
+
857
+ | Id | Severity | Meaning |
858
+ |---|---|---|
859
+ | `station-form-length` | fail | the word count is outside `form.length`; the message gives the count and the range |
860
+ | `station-form-required-part-<part>` | fail | a required part appears as neither a heading nor a `part:` line |
861
+
862
+ ### terms
863
+
864
+ Reads the optional `audience.terms`: the words the piece uses that its reader may not know. For
865
+ each term not also in `audience.knows`, it finds the term's first appearance (whole word, ignoring
866
+ case) and looks for a definition in that sentence or the next: the term followed within six words
867
+ by `is`, `means` or `refers to`, a colon among those words, or a parenthesis right after the term.
868
+ This is a mechanical proxy for a definition, not a reading of one: "A hyperspec is mentioned here"
869
+ passes. A term the draft never uses is not flagged, code blocks and inline code are ignored, and
870
+ with no `terms` list the station skips.
871
+
872
+ | Id | Severity | Meaning |
873
+ |---|---|---|
874
+ | `station-terms-undefined-<term>` | fail | the term's first appearance has no definition in that sentence or the next |
875
+
876
+ ### claims
877
+
878
+ Reads the claims ledger at `sources.ledger`: JSONL, one claim per line, each with the claim's
879
+ `text` exactly as the draft says it, a `source`, and optionally a `span`, the words in the source
880
+ that support it. The examples cite a segment as the source, the same `material#segment` form the
881
+ spine uses:
882
+
883
+ ```jsonl
884
+ {"text":"11 of 41 said at least one of their one-on-ones in the last quarter was mostly project status.","source":"survey#s4","span":"11 of 41 said at least one of their one-on-ones in the last quarter was mostly project status."}
885
+ ```
886
+
887
+ Every claim's text must still appear in the draft word for word, with whitespace and quote
888
+ characters normalized and case kept; otherwise the ledger is stale. Every claim needs a real
889
+ source; one without fails, or warns under `unsourced_claim: warn`. A missing ledger fails. The
890
+ station does not decide what counts as a factual claim: the ledger is the list of claims, so a
891
+ factual sentence left out of it passes unseen. Nor does it read the source to see whether it says
892
+ what the claim says.
893
+
894
+ | Id | Severity | Meaning |
895
+ |---|---|---|
896
+ | `station-claims-ledger-missing` | fail | the ledger file does not exist or cannot be read |
897
+ | `station-claims-json-line-<n>` | fail | ledger line n is not JSON, not an object, or has no `text` |
898
+ | `station-claims-stale` | fail | the claim on a ledger line no longer appears in the draft |
899
+ | `station-claims-unsourced` | fail, or warn under `unsourced_claim: warn` | the claim on a ledger line has no source, or only a placeholder; points at the draft line where the claim appears |
900
+
901
+ ### quotes
902
+
903
+ Every span in double quotation marks, straight or curly, of four words or more must appear word
904
+ for word in a `quote` or `story` segment of a marked material. Quote characters and whitespace
905
+ are normalized, case is kept, and a comma or period just inside the closing mark is dropped,
906
+ because that punctuation is the writer's; a `?` or `!` is kept, because adding one changes what
907
+ was said. Shorter spans are not checked, since two or three quoted words are as often a title as
908
+ a quotation. A private segment is never a source for a quote. When the sentence around a quote
909
+ names a speaker, the quote must come from a quote segment with that `speaker`. A speaker is named
910
+ by the full `speaker` value, hyphens read as spaces, or by its first word, so `maria-lopez` is
911
+ named by "Maria Lopez" and by "Maria". The first word alone counts only when it has two or more
912
+ letters and is not a common function word such as "the", so a speaker recorded as "the manager
913
+ interviewed" is named only by all three words. Attribution needs a declared speaker: a name that
914
+ is no segment's `speaker` attributes nothing, so start a `speaker` with the person's name, as
915
+ the essay example does with `dana, an engineering manager`. A spec with `fiction: true` skips
916
+ the station: a character's dialogue is invented rather than quoted from a material. The
917
+ `attribution` judge tests it against each character's own lines instead (see
918
+ [Judging a draft](#judging-a-draft)).
919
+
920
+ | Id | Severity | Meaning |
921
+ |---|---|---|
922
+ | `station-quotes-unmatched` | fail | a quoted span is in no quote or story segment |
923
+ | `station-quotes-misattributed` | fail | the sentence names a speaker, and the span is in no quote segment by that speaker |
924
+
925
+ ### private
926
+
927
+ No run of eight or more consecutive words from any `private` segment may appear in the draft,
928
+ compared by words with case and punctuation ignored. A private segment of four to seven words is
929
+ checked whole. One under four words is not checked, because two or three words match ordinary
930
+ prose; the station reports how many it skipped, as one warning that never quotes them. It cannot
931
+ catch a paraphrase, or a leak shorter than the run.
932
+
933
+ | Id | Severity | Meaning |
934
+ |---|---|---|
935
+ | `station-private-leak` | fail | the draft repeats a run from a private segment; names the material, the segment and the run |
936
+ | `station-private-short-skipped` | warn | private segments under four words were not checked; gives the count |
937
+
938
+ ### dna
939
+
940
+ Runs when `dna.scope_dir` is set and its `features.json` is current. It measures the draft the
941
+ way `dna measure` measures goldens and compares the sentence length mean, both paragraph length
942
+ means, every per-1000-word punctuation rate, and the contraction and person rates with the
943
+ scope's. For a scope value v, a draft value outside v ÷ 1.5 to the larger of v × 1.5 and v + 5 is
944
+ reported with both values. An em dash in a draft whose scope has none is its own finding,
945
+ pointing at the first one. Both
946
+ are warnings and the station never fails: it measures, and whether a draft sounds like its writer
947
+ is a judgment. The essay's warning is an example of what to read: its goldens are instructions in
948
+ the second person, and the essay tells the author's own story in the first. With no `scope_dir`,
949
+ or a `features.json` that is missing or stale, the station skips and says which.
950
+
951
+ | Id | Severity | Meaning |
952
+ |---|---|---|
953
+ | `station-dna-drift` | warn | a feature is outside its band; gives the draft's value, the scope's and the band |
954
+ | `station-dna-em-dash` | warn | the draft uses em dashes and the scope's goldens use none |
955
+
956
+ ### links
957
+
958
+ Every Markdown link (inline, reference, collapsed and shortcut) and every bare URL. An `http` or
959
+ `https` URL must parse and name a host, a `mailto:` link must carry an address, and any other
960
+ scheme fails. A relative link must resolve to a file, relative to the draft's own folder; the
961
+ part after `#` is not checked. A link that starts with `/` is relative to a site root the station
962
+ cannot see, so it warns. A full or collapsed reference, `[text][label]` or `[label][]`, needs a
963
+ definition for its label. A bare `[label]` is a link only when that label has a definition;
964
+ otherwise it is ordinary text, as Markdown renders it, so an editorial `[sic]`, a task list's
965
+ `[x]` and a numbered note `[1]` pass. Code blocks and inline code are ignored. It never touches
966
+ the network, so it cannot tell you a URL is live.
967
+
968
+ | Id | Severity | Meaning |
969
+ |---|---|---|
970
+ | `station-links-malformed` | fail | an http or https URL with no host (a bare `https://` included), or a `mailto:` with no address |
971
+ | `station-links-bad-scheme` | fail | a scheme other than http, https or mailto |
972
+ | `station-links-broken-relative` | fail | a relative link names no file beside the draft |
973
+ | `station-links-root-relative` | warn | a link starting with `/`, which cannot be resolved without the site |
974
+ | `station-links-undefined-reference` | fail | a full or collapsed reference link whose label has no definition |
975
+
976
+ ### Any station
977
+
978
+ | Id | Severity | Meaning |
979
+ |---|---|---|
980
+ | `station-<name>-crashed` | fail | the station threw; the message is the error's, with any absolute path shortened; the other stations and the ledger line still run |
981
+
982
+ ### The runs ledger
983
+
984
+ Each `check` appends one line to the spec's `improvement.ledger`, the same file lint's test 9
985
+ reads:
986
+
987
+ ```json
988
+ {"at":"2026-09-29T13:21:37.330Z","kind":"check","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","stations":{"form":"pass","terms":"pass","claims":"pass","quotes":"pass","private":"pass","dna":"pass","links":"pass"},"verdict":"one-shot"}
989
+ ```
990
+
991
+ `draft` is the draft's path relative to the spec's folder, however you spelled it, so one draft
992
+ has one history. `draft_sha256` and `spec_sha256` hash the two files' bytes; the files the spec
993
+ names (materials, the claims ledger, a scope folder) are not hashed, so "changed" below means the
994
+ draft or the spec. `stations` holds each station's status.
995
+
996
+ A run with `--only` is partial: its line carries `partial: true`, its verdict is `not-improved`
997
+ with the reason `partial run: <stations>`, and later verdicts ignore it, so a subset never claims
998
+ the verdict for the whole draft. A full run is compared with the most recent earlier full line for
999
+ the same draft:
1000
+
1001
+ - **one-shot**: there is none, and every station passes.
1002
+ - **improved**: that line failed and every station passes now; `change` names exactly the
1003
+ stations that failed then and pass now.
1004
+ - **not-improved** otherwise, with a `reason` that says which case it is: `failing stations: ...`
1005
+ on a first check that fails; `no change since the last passing check`; `draft changed; every
1006
+ station still passes` (or `spec changed`, or `spec and draft changed`); `still failing: ...`,
1007
+ after `no change since the last check;` or after what changed, when every failing station failed
1008
+ last time too; `failing stations: ...` after what changed when a station fails that passed last
1009
+ time; and `stations that failed last time now skip: ...` when a spec change stopped them running.
1010
+
1011
+ A ledger path that leads outside the spec's folder is not written, and `check` prints a warning.
1012
+
1013
+ ## Judging a draft
1014
+
1015
+ `check` runs the stations that are plain functions of the spec and the draft. The rest of a
1016
+ spec's checks are rubrics: whether the doctor would pass each goal condition, whether a reader
1017
+ would get lost, whether the voice can be told from the writer's own. Those are judgments, and
1018
+ hyperspec calls no model, so it does not make them. It makes them checkable instead. For each
1019
+ judgment station, `judge prepare` writes a packet: the rubric from the spec, fixed
1020
+ instructions, the inputs the judge reads, and the exact shape of the answer. An outside judge
1021
+ (your agent, any model, or a person) fills in a verdict. `judge record` validates it, derives
1022
+ the station's status from it by a fixed rule, and records it in the runs ledger. The judge
1023
+ decides; hyperspec checks that every passage the judge quotes is in the draft, scores
1024
+ every blind test against an answer key the judge never saw, and keeps the record.
1025
+
1026
+ ```bash
1027
+ npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
1028
+ ```
1029
+
1030
+ ```
1031
+ essay/judge/doctor.packet.json
1032
+ essay/judge/lineup.packet.json
1033
+ essay/judge/lineup.key.json
1034
+ essay/judge/reader.packet.json
1035
+ essay/judge/persona.packet.json
1036
+ attribution: skip (the spec is not fiction; attribution applies only with fiction: true)
1037
+ knowledge: skip (the spec is not fiction; knowledge applies only with fiction: true)
1038
+ ```
1039
+
1040
+ Like `check`, it lints the spec first: a spec that fails lint, or is blocked on an open
1041
+ decision, gets no packet, and `prepare` exits with lint's own code. Then it writes one
1042
+ `<station>.packet.json` for each station that applies, in a fixed order (`doctor, lineup,
1043
+ reader, persona, attribution, knowledge`), and prints a `skip` line with the reason for each
1044
+ that does not. `--only doctor,reader` prepares just those. The `--out` folder must already
1045
+ exist. `prepare` refuses to overwrite any file it would write, naming every one, and then writes
1046
+ nothing; `--force` replaces them. The worked examples ship the packets this writes, so add
1047
+ `--force` to write them again there. The same spec and draft, with the same goldens and claims
1048
+ ledger, always give byte-identical packets. A station that throws while building its packet
1049
+ prints `no packet` with a `judge-<name>-crashed` finding; the other packets are still written,
1050
+ and `prepare` exits 1.
1051
+
1052
+ **Hand the judge only the `*.packet.json` files, never the `--out` folder.** Two stations are
1053
+ blind tests with a right answer. `lineup` writes which candidate is the draft's to
1054
+ `lineup.key.json`, and `attribution` writes each line's true speaker to
1055
+ `attribution.key.json`, both beside the packets, for a person to read. A judge who can see them
1056
+ is not judging. `record` never reads either file as truth: it builds the key again from the spec
1057
+ and the draft, so an edited key changes nothing.
1058
+
1059
+ **Give the two blind packets to a judge in a fresh context with no access to the draft.** Every
1060
+ packet names its spec and its draft by path, and a judge that can open files can open those: the
1061
+ draft's own text shows which lineup passage is the draft's, and its speech tags give every
1062
+ attribution answer. Both stations' instructions tell the judge to decide from the packet's inputs
1063
+ alone and open no file the packet names, and a judge with file access can still ignore that, so
1064
+ the instruction is not a guarantee. Paste the packet into a new conversation, or hand it to a
1065
+ person who has not read the draft, rather than to an agent working in the folder the draft is in.
1066
+ The other four packets carry the draft in their inputs and hide nothing, which is why each blind
1067
+ packet goes to its own context, apart from the other packets as well as from the folder: a judge
1068
+ that has read the doctor, reader or persona packet has read the draft.
1069
+
1070
+ Exit codes for `prepare`: **0** written; **1** a station crashed; **2** usage: no spec path, no
1071
+ `--draft`, no `--out`, an `--out` that is missing or not a folder, a draft that cannot be read, a
1072
+ spec without `profile: writing`, an `--only` that names no known judge, or a file that exists
1073
+ without `--force`; and lint's own **1** or **3** when the spec is not ready. `--json` prints the
1074
+ result, and a usage error prints `{ "spec", "draft", "out", "error" }`.
1075
+
1076
+ ### The packet
1077
+
1078
+ Every packet is JSON with these fields, in this order:
1079
+
1080
+ | Field | Holds |
1081
+ |---|---|
1082
+ | `hyperspec_judge` | the packet format's version, `"0.1"` |
1083
+ | `station` | the station's name |
1084
+ | `spec`, `draft` | the paths exactly as `prepare` was given them |
1085
+ | `spec_sha256`, `draft_sha256` | the SHA-256 of each file's bytes |
1086
+ | `rubric` | the `check.rubric` of the block the station reads, verbatim |
1087
+ | `instructions` | fixed text for the station, the same for every spec |
1088
+ | `inputs` | what the judge reads; each station below lists its own |
1089
+ | `verdict_schema` | the verdict's exact shape, as a small JSON Schema |
1090
+
1091
+ A station that needs a rubric skips when its block has none, and a station whose block is
1092
+ deferred to a decision skips too. The paths are kept as given so the packet is the same on every
1093
+ machine, which means **`record` must run in the folder `prepare` ran in**. Run from anywhere else,
1094
+ it cannot find the spec and exits 2.
1095
+
1096
+ ### The evidence rule
1097
+
1098
+ Every verdict field that cites the draft (the doctor's `evidence`, the reader's `lost_at` and
1099
+ `stopped_at`, the persona's `breaks`, the knowledge `leaks`) is a span copied from the draft. A
1100
+ span counts only when:
1101
+
1102
+ - it has at least three words, where a word is a run of letters and digits (so "It's" is two);
1103
+ - it is in the draft after both are normalized: every run of whitespace, line breaks included,
1104
+ becomes one space, and curly, low and angle quotation marks and apostrophes
1105
+ (`‘ ’ ‚ ‛ “ ” „ ‟ ‹ › « »`) become straight ones. Primes (`′ ″`) are not quotation marks and
1106
+ are left alone, case is kept, and nothing else changes, so a changed word is not a quotation;
1107
+ - it matches whole words: a span that starts or ends with a letter or digit may not start or end
1108
+ inside a word of the draft.
1109
+
1110
+ A span that fails makes the verdict invalid (`judge-evidence-missing`, `-too-short` or
1111
+ `-not-found`), and nothing is recorded. The lineup and attribution verdicts carry no evidence:
1112
+ they are blind tests, scored against their keys.
1113
+
1114
+ ### Recording a verdict
1115
+
1116
+ ```bash
1117
+ npx @supersuit/hyperspec judge record essay/judge/lineup.packet.json --verdict essay/sample-verdicts/lineup.verdict.json
1118
+ ```
1119
+
1120
+ ```
1121
+ lineup: fail
1122
+ fail [judge-lineup-picked] the judge picked the draft's passage (D) out of 4 candidates with confidence 0.6: D is the only passage that rests on a survey figure, and it opens by pointing at something outside itself (the survey); A, B and C each turn one claim into an instruction in the second person, with no numbers. (line 32)
1123
+ fix: Revise this passage toward the goldens' voice, where the reason points; if the goldens do not cover this kind of passage, add one that does. Then prepare and judge again.
1124
+ verdict: not-improved (failing stations: lineup)
1125
+ ```
1126
+
1127
+ `record` trusts nothing in the packet file. First it hashes the spec and the draft again; if
1128
+ either no longer matches the hash the packet recorded, the verdict is stale (`judge-stale`,
1129
+ "the draft does not match the hash the packet recorded: it changed since prepare, or the packet
1130
+ was edited"). Then it rebuilds the packet from the spec, the draft and the station's other
1131
+ inputs on disk (the DNA scope, its `scope.md` and goldens, for `lineup`; the claims ledger for
1132
+ `persona`) and
1133
+ requires the file to be exactly those bytes:
1134
+
1135
+ - a packet whose inputs no longer match what those other files produce is stale, and the message
1136
+ names them. Nothing on disk can tell a changed golden from a hand-edited packet, so it says
1137
+ "changed since the packet was prepared, or the packet was edited" and claims neither;
1138
+ - a station that no longer applies because of those files (the claims ledger deleted, say) is
1139
+ stale in the same way, and says why it no longer applies;
1140
+ - anything else (an edited condition, a reformatted file, a hash made to match a changed draft, a
1141
+ station that no longer applies for another reason) is `judge-packet-altered`.
1142
+
1143
+ Either way nothing is recorded. A stale packet prints `<station>: stale packet, nothing
1144
+ recorded`, as `learn record` does, and with `--json` both commands mark it `"invalid": true,
1145
+ "stale": true`, so a script can test `invalid` alone. The fix is to run `judge prepare` again
1146
+ with `--force` and judge the new packet, or, for a station that no longer applies, to restore
1147
+ the file it reads. Only then is the verdict read: it must be JSON (one leading
1148
+ byte order mark is ignored), in the shape the packet gives, with every evidence span found.
1149
+ Every problem is named, and an invalid verdict records nothing. The validator ignores fields it
1150
+ does not know, so a verdict can carry a note of its own; the worked examples' sample verdicts
1151
+ each carry one, `sample`, saying what they are.
1152
+
1153
+ A valid verdict gives the station's status, its findings (a warning prints under the status and
1154
+ never fails it), and for `attribution` one summary line. Then one line goes to the runs ledger.
1155
+
1156
+ Exit codes for `record`: **0** the station passed; **1** it failed, or the verdict is invalid,
1157
+ or the packet is stale or altered; **2** usage: no packet path, no `--verdict`, a packet or verdict
1158
+ file that cannot be read, a file that is not a judge packet or names an unknown judge, or a
1159
+ packet whose spec (which must carry `profile: writing`) or draft cannot be read. `--json` prints
1160
+ the result, and a usage error prints `{ "packet", "verdict", "error" }`.
1161
+
1162
+ In the tables below, **fail** fails the station, **warn** is printed and never fails it,
1163
+ **invalid** refuses the verdict, and **stale** refuses the packet; an invalid or stale verdict
1164
+ records nothing.
1165
+
1166
+ ### doctor
1167
+
1168
+ Grades the draft against every goal condition, and asks whether the reader would take the next
1169
+ step now. Applies when `writing.goal` is written and its check has a rubric. Inputs: `goal`
1170
+ (`from`, `to`, `next_if_worked`, and `change` with its `kind` and `text`), `conditions` (each
1171
+ condition id with its requirement's `text` and `fails_when`) and the `draft`. The verdict is
1172
+ `{ conditions: [{ id, pass, evidence, note }], would_take_next_step, evidence }`: every condition
1173
+ id exactly once, `pass` and `would_take_next_step` true or false, a `note` on every condition
1174
+ (one that fails must say why), and the top-level `evidence` for the passage that decided the next
1175
+ step. It passes when every condition passes and the reader would take the next step.
1176
+
1177
+ | Id | Kind | Meaning |
1178
+ |---|---|---|
1179
+ | `judge-doctor-condition` | fail | a condition fails; gives the judge's note, at its evidence's line |
1180
+ | `judge-doctor-next-step` | fail | the reader would not take `goal.next_if_worked` now |
1181
+ | `judge-doctor-condition-missing` | invalid | a condition in the packet has no entry |
1182
+ | `judge-doctor-condition-unknown` | invalid | the verdict grades an id that is not a condition in the packet |
1183
+ | `judge-doctor-condition-duplicate` | invalid | a condition is graded more than once |
1184
+
1185
+ ### lineup
1186
+
1187
+ A blind test of voice. Applies when `writing.dna` names a `scope_dir` whose goldens can be read
1188
+ and its check has a rubric, at least one golden has a prose paragraph, and the draft has a prose
1189
+ paragraph that is not already a golden's. A draft that is one paragraph and nothing else is
1190
+ skipped: its candidate would be the whole file, and the packet's `draft_sha256` would identify it. A prose paragraph is a run of non-blank lines that are
1191
+ all plain text: headings and code fences end one, and a block holding a list item, a quotation, a
1192
+ table row, a thematic break or HTML is left out whole, as is indented code.
1193
+
1194
+ Every candidate is one prose paragraph reflowed onto a single line, so none can be told by its
1195
+ formatting. The target length is the median, in characters, of every prose paragraph of every
1196
+ golden in the scope. The first three goldens by file name that have one each give their
1197
+ paragraph closest to that length, and the draft gives its paragraph closest to it, skipping any
1198
+ that is word for word a golden's; ties go to the earliest. The candidates are shuffled with a
1199
+ seed derived from the draft's full text (the SHA-256 of `hyperspec lineup seed`, a line break, and
1200
+ the text), which the packet does not carry, so the packet cannot reveal the order: not even its
1201
+ `draft_sha256`, which is a different hash. The same draft always gets the same labels, labeled A
1202
+ to D. Inputs: `scope` (the scope's `writer`, `form`, `audience` and `purpose`) and `candidates`,
1203
+ each `{ label, text }`; no path and no source. `lineup.key.json` records the draft's label, the
1204
+ draft line its paragraph starts on, and where every candidate came from.
1205
+
1206
+ The verdict is `{ pick, confidence, reason }`: a label, a number from 0 to 1, and what decided it.
1207
+ It passes when the pick is not the draft's passage: the judge could not tell. A judge who picks at
1208
+ random also passes, three times in four with four candidates, so one lineup is weak evidence of a
1209
+ voice. A failing lineup is the stronger signal, and its reason says where to look.
1210
+
1211
+ | Id | Kind | Meaning |
1212
+ |---|---|---|
1213
+ | `judge-lineup-picked` | fail | the judge picked the draft's passage; gives its confidence and reason, at the paragraph's line |
1214
+ | `judge-lineup-pick-unknown` | invalid | the pick is not one of the lineup's labels |
1215
+
1216
+ ### reader
1217
+
1218
+ Reads the draft as the audience block's reader. Applies when `writing.audience` is written and its
1219
+ check has a rubric. Inputs: `audience` (`who`, `funnel_now`, `knows`, `terms`, `believes_now`,
1220
+ `wants`, `reads_on`, `reader`) and the `draft`. The verdict is `{ lost_at: [{ evidence, why }],
1221
+ stopped_at, would_take_next_step, next_step }`, where `stopped_at` is `{ evidence, why }` or
1222
+ `null` when the reader read to the end, and must be present either way. The reader names its own
1223
+ next step: it is not shown `goal.next_if_worked`, so a reader that would act and a doctor that
1224
+ says the reader would not take the spec's next step can both be right. In the essay example they
1225
+ agree: both name the card. It passes when the reader read to the end and would take its next step
1226
+ now.
1227
+
1228
+ | Id | Kind | Meaning |
1229
+ |---|---|---|
1230
+ | `judge-reader-lost` | warn | the reader got lost here, and why |
1231
+ | `judge-reader-stopped` | fail | the reader stopped reading here, and why |
1232
+ | `judge-reader-next-step` | fail | the reader would not take its next step now |
1233
+
1234
+ ### persona
1235
+
1236
+ Reads the draft as the persona block's speaker, against the claims ledger. Applies when
1237
+ `writing.persona` is written and its check has a rubric, and the claims ledger, if
1238
+ `sources.ledger` names one, can be read. Inputs: `persona` (`identity`, `stance`, `may_assert`,
1239
+ `will_not_say`), `claims` (the text of every well-formed line of the claims ledger, in order) and
1240
+ the `draft`. With no ledger declared, `claims` is `null`: facts cannot be checked against
1241
+ sources, so the instructions say not to report one, the schema leaves that kind out, and a break
1242
+ of that kind is invalid. The verdict is `{ breaks: [{ evidence, kind, why }] }`, where `kind` is
1243
+ `stance` (the voice leaves its stance), `assertion` (it asserts something outside `may_assert`),
1244
+ `will_not_say` or `unsourced_fact` (a fact no claim holds). It passes when there is no break.
1245
+
1246
+ | Id | Kind | Meaning |
1247
+ |---|---|---|
1248
+ | `judge-persona-break` | fail | the voice breaks here; the message names the kind and gives the judge's why |
1249
+ | `judge-persona-kind-unknown` | invalid | a break's kind is not one of the four |
1250
+ | `judge-persona-no-ledger` | invalid | a break is `unsourced_fact`, and the spec declares no claims ledger |
1251
+
1252
+ ### attribution
1253
+
1254
+ Fiction only: a blind test of whether the characters' voices can be told apart. The packet holds
1255
+ the draft's dialogue lines with the narration and the speaker removed, and each speaking
1256
+ character's `speech` block, `golden_lines` and `rejected_lines`; it carries no draft. The judge
1257
+ names a speaker for every line, and `record` scores the answers against the true speakers, which
1258
+ only `attribution.key.json` holds. Applies when `fiction: true`, at least two characters have a
1259
+ speech block and one of them has a check rubric, and the draft has attributable lines from at
1260
+ least two speakers.
1261
+
1262
+ A dialogue line is a double-quoted span, straight or curly, inside one paragraph, outside code. A
1263
+ quote split by a speech tag (`"Twenty minutes," Ines said, "then we fold it."`) is one line when
1264
+ its first part ends in a comma and the narration between the parts is exactly one tag and its
1265
+ comma, nothing else. Narration that holds anything more joins nothing: in `"Leave it there," Ines
1266
+ said, and Theo muttered, "No chance at all."` the first part is Ines's, and the second is left
1267
+ out, since a second speaker brought in by a beat, a pronoun or a verb off the list cannot be read
1268
+ mechanically. The second part is left out whatever ends that narration, even another tag (`Ines
1269
+ said, and Theo said,`), and so is the second part of `"Twenty minutes," Ines said, wiping her
1270
+ hands, "then we fold it."`. A line's speaker comes only from a speech tag:
1271
+ narration in the same paragraph directly after the closing mark (`"...," Ines said`) or directly
1272
+ before the opening mark, ending in a comma or colon (`Ines said, "..."`). A tag is a subject and
1273
+ one of the verbs said, asked, told, replied, called, whispered, shouted, answered, added and went
1274
+ on, or their present tense (says, asks, goes on). The subject is:
1275
+
1276
+ - a character's id, or its `name` if it has one: whole words, any case, and an id's words may be
1277
+ joined by a space, a hyphen or an underscore, so `old-man` is named by "old man";
1278
+ - "I", when `persona.identity` is `character:<id>`: the narrator speaks;
1279
+ - "she" or "he", when exactly two characters have speech blocks and one of them narrates: the
1280
+ other one speaks. After a quote the pronoun must be lower case.
1281
+
1282
+ The verb may come first ("said Ines") only for said, replied, whispered, shouted and went on (and
1283
+ their present tense), since "Ines told Theo" names the person spoken to. Anything else leaves the
1284
+ line out, and every doubt does: no tag, a possessive ("Ines's") or an action beat ("Theo nodded"),
1285
+ a subject that is not a character with a speech block, tags that name two different speakers, or a
1286
+ line of three words or more that repeats, or is part of, a golden or rejected line, which would
1287
+ give its speaker away. So does the second quote in
1288
+ `"Not a bakery," she said. "They want the room."`, since no tag sits against it. The key is never wrong, and recall pays for it: the worked
1289
+ story has 36 dialogue lines, and 19 of them are in the test, 10 by Ines and 9 by Theo. The packet
1290
+ counts the lines left out, and the key lists each with its reason.
1291
+
1292
+ Inputs: `characters`, `lines` (each `{ id, text }`, numbered L1 up in draft order, whitespace
1293
+ reflowed) and `excluded`, the count left out. The verdict is `{ lines: [{ id, speaker }] }`,
1294
+ every line id exactly once, a speaker by id or name, in any case. Accuracy is taken per speaker
1295
+ and averaged, so naming one character for every line cannot pass: the station passes when that
1296
+ mean is 80 percent or more, compared exactly, with percentages rounded down. This is the worked
1297
+ story's sample:
1298
+
1299
+ ```
1300
+ attribution: pass
1301
+ accuracy per speaker: ines 10/10 (100%), theo 9/9 (100%); mean 100%, passing at 80%; 17 dialogue lines left out (8 repeating a golden or rejected line, 9 with no speech tag)
1302
+ verdict: one-shot
1303
+ ```
1304
+
1305
+ | Id | Kind | Meaning |
1306
+ |---|---|---|
1307
+ | `judge-attribution-accuracy` | fail | the mean accuracy over speakers is under 80%; gives each speaker's |
1308
+ | `judge-attribution-miss` | warn | a line was given to the wrong character, at its draft line |
1309
+ | `judge-attribution-speaker-unknown` | invalid | a speaker is not a character in the packet |
1310
+ | `judge-attribution-line-missing` | invalid | a line in the packet has no answer |
1311
+ | `judge-attribution-line-unknown` | invalid | the verdict answers an id that is not a line in the packet |
1312
+ | `judge-attribution-line-duplicate` | invalid | a line is answered more than once |
1313
+
1314
+ ### knowledge
1315
+
1316
+ Fiction only: whether anyone knows something before their timeline gives it to them. Applies when
1317
+ `fiction: true` and at least one character has a knowledge timeline (an entry with both `by` and
1318
+ `knows`) and one of those characters has a check rubric. Inputs: `characters` (each with its
1319
+ `knowledge` entries) and the `draft`. The verdict is `{ leaks: [{ character, evidence,
1320
+ knows_too_early }] }`, and it passes when there is no leak.
1321
+
1322
+ hyperspec does not read the draft's structure: `by` goes into the packet as the spec wrote it,
1323
+ and the judge maps it onto the draft. The worked story's timelines say `scene-1` to `scene-4`,
1324
+ and its draft's four headings are times, 3:40 to 6:55, in the scene list's order, so a judge maps
1325
+ them by order. Write `by` as something a reader of the draft can find.
1326
+
1327
+ | Id | Kind | Meaning |
1328
+ |---|---|---|
1329
+ | `judge-knowledge-leak` | fail | a character knows something too early, at its evidence's line |
1330
+ | `judge-knowledge-character-unknown` | invalid | a leak names a character with no timeline in the packet |
1331
+
1332
+ ### Any judgment station
1333
+
1334
+ | Id | Kind | Meaning |
1335
+ |---|---|---|
1336
+ | `judge-verdict-not-json` | invalid | the verdict file is not JSON |
1337
+ | `judge-verdict-shape` | invalid | a field is missing, of the wrong type or out of range, or a required note, reason or why is empty |
1338
+ | `judge-evidence-missing` | invalid | an evidence field is empty or not a string |
1339
+ | `judge-evidence-too-short` | invalid | an evidence span has fewer than three words |
1340
+ | `judge-evidence-not-found` | invalid | an evidence span is not in the draft, word for word |
1341
+ | `judge-stale` | stale | the spec, the draft, or a file the station reads changed since the packet was prepared, or the packet was edited |
1342
+ | `judge-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
1343
+ | `judge-<name>-crashed` | invalid | the station threw; at `prepare` its packet is not written, at `record` nothing is recorded |
1344
+
1345
+ ### Judge lines in the runs ledger
1346
+
1347
+ Each recorded verdict appends one line to the spec's `improvement.ledger`:
1348
+
1349
+ ```json
1350
+ {"at":"2026-09-29T16:27:36.883Z","kind":"judge","station":"lineup","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","packet_sha256":"<sha256 of the packet>","inputs_sha256":"<sha256 of its inputs>","status":"fail","verdict":"not-improved","reason":"failing stations: lineup"}
1351
+ ```
1352
+
1353
+ `draft` is relative to the spec's folder, as in a check line. `packet_sha256` is the SHA-256 of
1354
+ the packet the judge was shown, taken with its two paths written as the ledger writes them
1355
+ (relative to the spec's folder), so the same packet prepared from another folder, or as
1356
+ `./draft.md`, hashes the same. `inputs_sha256` is the SHA-256 of the packet's `inputs` alone. The
1357
+ verdict follows the same rules as a check line, compared with
1358
+ the most recent earlier judge line for the same station and the same draft, with three more:
1359
+
1360
+ - **What changed** is judged by the packet. The draft or the spec is named when its bytes
1361
+ changed. When neither did and the packet's inputs did, the files the station reads besides them
1362
+ are named: `the DNA scope (<scope_dir>: scope.md and goldens)` for `lineup`, `the claims ledger
1363
+ (<path>)` for `persona`. When only the rest of the packet changed, which a later hyperspec
1364
+ release can do by rewording a station's instructions, the reason says `the packet's fixed text
1365
+ (hyperspec's instructions or format) changed`. So adding the golden a failing lineup asked for, or the claims a
1366
+ failing persona asked for, and passing on the new packet is `improved`, "the DNA scope
1367
+ (dna/essay-new-managers-teach: scope.md and goldens) changed; stations now pass: lineup". An
1368
+ improved judge line always names what changed, "draft changed; stations now pass: doctor".
1369
+ - **one-shot** also needs these draft bytes never to have been judged by this station before,
1370
+ under any name. A copy or a rename of a judged draft gets `not-improved`, with the reason
1371
+ "these draft bytes were judged before as <path>: <status>".
1372
+ - **improved** also needs the packet to have changed since the failing line. A judge can answer
1373
+ differently about an identical packet, and that is not the work improving: the verdict is
1374
+ `not-improved`, "the verdict changed; nothing the judge was shown changed".
1375
+
1376
+ Otherwise the reasons are check's, with `judgment` where check says `check`: `failing stations:
1377
+ doctor`, `no change since the last passing judgment`, `no change since the last judgment; still
1378
+ failing: doctor`, or `draft changed; every station still passes`. Check and judge each read only
1379
+ their own lines, and learn reads none, so judging a draft never changes what `check` says about
1380
+ it, or the reverse. Judge lines keep lint's test 9 passing, as check lines do.
1381
+
1382
+ ### The worked examples
1383
+
1384
+ Both examples ship the packets `prepare` writes for their drafts, in `essay/judge/` and
1385
+ `story/judge/`, with no key beside them, and one sample verdict per packet in
1386
+ `essay/sample-verdicts/` and `story/sample-verdicts/`. The samples were filled in by hand, as
1387
+ one careful judge would, to show the shape and what `record` does with it; each says so in its
1388
+ `sample` field. Another judge may answer differently. A test records every sample on a fresh
1389
+ copy on every release.
1390
+
1391
+ | Example | Station | Sample | What the judge found |
1392
+ |---|---|---|---|
1393
+ | essay | doctor | pass | every condition holds, and the reader would write the three questions on a card, the goal's next step and the draft's close |
1394
+ | essay | lineup | fail | the draft's passage is the only one resting on a survey figure; the goldens hold no numbers, so they do not cover this kind of passage |
1395
+ | essay | reader | pass | read to the end, lost nowhere; the reader's own next step is the card |
1396
+ | essay | persona | pass | the mentor stance holds, and every figure is in the claims ledger |
1397
+ | story | doctor | pass | every condition holds, and a reader would look for the author's other stories |
1398
+ | story | reader | pass | read to the end; lost for a moment at "proving cabinet" and "peel", two warnings |
1399
+ | story | persona | fail | three process details (the deck oven's heat-up time, the rolls' bake time, how the starter is fed) are in no claim, and the rubric allows none outside the ledger |
1400
+ | story | attribution | pass | all 19 lines named right by voice alone |
1401
+ | story | knowledge | pass | neither character knows anything early |
1402
+
1403
+ Both examples pass every station of `check`. Each failure here is something no deterministic
1404
+ station can see.
1405
+
1406
+ ## Learning from edits
1407
+
1408
+ A factory writes a first draft, and a person edits it into the draft they approve. Every edit is
1409
+ something the spec did not say, or did not say well enough. `learn prepare` diffs the two drafts
1410
+ sentence by sentence and writes a packet of the edits; an outside judge names, for each edit, the
1411
+ one block of the spec that would have prevented it; `learn record` validates that verdict, counts
1412
+ the edits by block, and names one next move for the block with the most. It never edits the spec:
1413
+ the move is a suggestion for whoever keeps it.
1414
+
1415
+ ```bash
1416
+ npx @supersuit/hyperspec learn prepare essay.hyperspec.md --first essay/learn/first-draft.md --approved essay/draft.md --out essay/learn
1417
+ npx @supersuit/hyperspec learn record essay/learn/learn.packet.json --verdict essay/sample-verdicts/learn.verdict.json
1418
+ ```
1419
+
1420
+ ```
1421
+ essay/learn/learn.packet.json
1422
+ 5 edits over 8 sentences: 3 replaced, 2 deleted
1423
+ ```
1424
+
1425
+ ```
1426
+ learn: 5 edits classified
1427
+ dna 2 edits, 4 sentences
1428
+ materials 1 edit, 2 sentences
1429
+ persona 1 edit, 1 sentence
1430
+ none 1 edit, 1 sentence
1431
+ next move: dna, 4 of 8 sentences (2 of 5 edits): add a golden or a style rule
1432
+ verdict: not-improved (edits by block: dna 2 edits (4 sentences), materials 1 edit (2 sentences), persona 1 edit (1 sentence), none 1 edit (1 sentence); not yet applied to the spec)
1433
+ ```
1434
+
1435
+ `prepare` lints the spec first, exactly as `judge prepare` does, then writes
1436
+ `<out>/learn.packet.json` and prints how many edits it found, or `no edits: the first draft and the
1437
+ approved draft match; nothing to learn`. The rules on `--out`, `--force` and the paths are the
1438
+ judge's: the folder must exist, an existing packet needs `--force`, and `record` runs in the
1439
+ folder `prepare` ran in. The essay example ships this pair and its packet in `essay/learn/`, so
1440
+ add `--force` to write the packet again there.
1441
+
1442
+ Exit codes: `prepare` **0** written, **2** usage (no spec path, `--first`, `--approved` or
1443
+ `--out`, a folder that is missing, a draft that cannot be read, a spec without `profile: writing`,
1444
+ a packet that exists without `--force`, or drafts too large to diff), and lint's own **1** or
1445
+ **3** when the spec is not ready; `record` **0** recorded, **1** the verdict is invalid or the
1446
+ packet stale or altered (nothing recorded), **2** usage.
1447
+
1448
+ ### Sentence units
1449
+
1450
+ The drafts are compared one unit at a time, and a unit is one sentence, one heading, or one list
1451
+ item:
1452
+
1453
+ - A blank line always ends a unit. When the first line of a paragraph is a Markdown heading, that
1454
+ line is a unit of its own; a later line that starts with `#` is a hard wrap in the text, and is
1455
+ read as text.
1456
+ - Other lines are split into sentences at `.`, `!` or `?` followed by whitespace, never inside a
1457
+ quotation that opens and closes on one line (straight or curly), so a quoted passage of several
1458
+ sentences on one line is one unit. Each list item starts a unit.
1459
+ - A unit that ends in a common abbreviation (Mr., Mrs., Ms., Dr., Prof., Sr., Jr., St., vs., cf.,
1460
+ e.g., i.e.) or a single capital initial ("J.") runs on into the next, and so does a unit
1461
+ followed by one that starts with a lower-case letter ("the U.S. economy").
1462
+
1463
+ ### The edits
1464
+
1465
+ Units are compared with their whitespace collapsed and matched by a longest common subsequence.
1466
+ Every run of unmatched units is one edit, a hunk: `deleted` (only in the first draft), `inserted`
1467
+ (only in the approved one) or `replaced`. A hunk never crosses a paragraph break or a heading, so
1468
+ a paragraph rewritten from end to end is one hunk and a change in two paragraphs is two. A hunk's
1469
+ texts are the drafts' own words from its first unit to its last, and its `sentences` is the
1470
+ larger of its two unit counts. A run that reads the same once whitespace is collapsed is a
1471
+ reflow, not an edit, so rewrapping lines, or joining or splitting paragraphs without changing a
1472
+ word, makes no hunk. The drafts' common start and end are set aside first; if what is left would
1473
+ need more than 10,000,000 comparisons, `prepare` stops with a usage error naming both sentence
1474
+ counts. Learn from a chapter or a scene at a time.
1475
+
1476
+ ### The learn packet
1477
+
1478
+ `{ "hyperspec_learn": "0.1", "spec", "spec_sha256", "first", "first_sha256", "approved",
1479
+ "approved_sha256", "blocks", "instructions", "hunks", "verdict_schema" }`. `blocks` lists the
1480
+ writing blocks the spec has written, in schema order, then `none`; a deferred block is not
1481
+ listed. Each hunk is `{ id, kind, first, approved, sentences }`, with ids E1 up, and `null` on
1482
+ the side that has no text. The packet carries no draft beyond its hunks: the instructions tell
1483
+ the judge to read the spec at its path. A learn packet has no answer key.
1484
+
1485
+ The verdict is `{ edits: [{ id, block, why }] }`: every hunk id exactly once, a `block` from the
1486
+ packet's `blocks`, and a `why` saying what that block should have said. `none` means no block of
1487
+ the spec could have prevented the edit, a typo say. As with a judge packet, `record` hashes the
1488
+ spec and both drafts again, rebuilds the packet, and requires the file to be exactly those bytes
1489
+ before it reads the verdict.
1490
+
1491
+ ### The tally and the next move
1492
+
1493
+ The tally counts, for each block the verdict names, its edits and the sentences they touched.
1494
+ Blocks are ordered by sentences, then edits, then the order of the table below with `none` last,
1495
+ and the next move goes to the first block that is not `none`, so a paragraph rewritten from end to
1496
+ end weighs as the sentences it rewrote. There is one move per block:
1497
+
1498
+ | Block | Next move |
1499
+ |---|---|
1500
+ | `materials` | mark or add the material the edit drew on |
1501
+ | `dna` | add a golden or a style rule |
1502
+ | `persona` | tighten the persona's stance, may_assert or will_not_say |
1503
+ | `audience` | extend the audience's knows or terms |
1504
+ | `goal` | tighten a goal condition's fails_when |
1505
+ | `form` | adjust the form block's length or shape |
1506
+ | `spine` | restate the spine's claim or its order |
1507
+ | `sources` | add or cite a source in the claims ledger |
1508
+ | `characters` | extend a character's speech or knowledge |
1509
+
1510
+ When every edit is `none`, the move is `none`: no block could have prevented any edit, so the spec
1511
+ has nothing to learn from the pair.
1512
+
1513
+ `record` appends one line to the runs ledger, `{ at, kind: "learn", first, first_sha256,
1514
+ approved, approved_sha256, spec_sha256, verdict, reason, tally }`, with both drafts' paths
1515
+ relative to the spec's folder. Its verdict is always `not-improved`: the spec has not changed yet,
1516
+ and the reason gives the counts by block. Learn reads no earlier line, and check and judge ignore
1517
+ learn lines.
1518
+
1519
+ | Id | Kind | Meaning |
1520
+ |---|---|---|
1521
+ | `learn-verdict-not-json` | invalid | the verdict file is not JSON |
1522
+ | `learn-verdict-shape` | invalid | the verdict is not `{ edits: [...] }`, or an entry is not an object with an id |
1523
+ | `learn-edit-unknown` | invalid | an id is not a hunk in the packet |
1524
+ | `learn-edit-duplicate` | invalid | a hunk is classified more than once |
1525
+ | `learn-edit-missing` | invalid | a hunk is not classified |
1526
+ | `learn-block-unknown` | invalid | a block is not one of the nine blocks or `none` |
1527
+ | `learn-block-absent` | invalid | a block is one the spec has not written |
1528
+ | `learn-why-missing` | invalid | a why is empty or not a string |
1529
+ | `learn-stale` | stale | the spec or a draft no longer matches the hash the packet recorded: it changed since prepare, or the packet was edited |
1530
+ | `learn-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
1531
+
1532
+ The essay example's pair is a first draft that differs from `essay/draft.md` by five edits: a
1533
+ paragraph of hedged advice and a hedged sentence, where the approved draft gives instructions; a
1534
+ claim about what most managers do; the walking aside from the voice memo; and a typo. The sample
1535
+ verdict puts the two hedges on `dna`, since no golden or style rule shows advice given flat, and
1536
+ the tally sends the next move there.
1537
+
789
1538
  ## Deferring a block
790
1539
 
791
1540
  A block can be deferred, never silently missing. A required block that is absent fails test 1
@@ -849,13 +1598,25 @@ the field it needs, and every spine claim cites the segments that support it. Ea
849
1598
  keeps the boundaries `segments init` wrote, in paragraph mode for prose and sentence mode for
850
1599
  bulleted notes, so you can re-run it and compare.
851
1600
 
852
- Both lint `pass (9/9)` with `writing: 9/9 blocks complete` and no findings. A test runs them on
853
- every release, so they cannot drift from the linter.
1601
+ Both lint `pass (9/9)` with `writing: 9/9 blocks complete` and no findings. Each also ships a
1602
+ draft written to it, `essay/draft.md` and `story/draft.md`, with its claims ledger beside it, and
1603
+ both drafts pass every station of `check`: the essay with one dna warning, described under
1604
+ [dna](#dna), and the story with dna skipped, since it names no scope folder, and quotes skipped,
1605
+ since it is fiction. A test lints both
1606
+ specs and checks both drafts on every release, so they cannot drift from the tool.
1607
+
1608
+ Each also ships the packets `judge prepare` writes for its draft, in `essay/judge/` and
1609
+ `story/judge/`, and one sample verdict per packet in `essay/sample-verdicts/` and
1610
+ `story/sample-verdicts/`, filled in by hand and marked as samples; what each found is under
1611
+ [The worked examples](#the-worked-examples). The essay adds a learn pair in `essay/learn/`: a
1612
+ first draft, the packet `learn prepare` writes comparing it with `essay/draft.md`, and a sample
1613
+ learn verdict (see [Learning from edits](#learning-from-edits)). A test checks that every packet
1614
+ is what `prepare` writes now and records every sample.
854
1615
 
855
1616
  ## What later versions add
856
1617
 
857
- This release is the schema, its lint, marked materials, and scoped DNA with measured features.
858
- Later versions build on it in order. The first compares a draft against its scope: its features
859
- beside the scope's features, and a blind lineup in which a judge sees a generated passage among
860
- the scope's goldens and tries to pick it out. Then the stations themselves, running the checks
861
- each block names and grading drafts against the goal.
1618
+ This release is the schema, its lint, marked materials, scoped DNA, `check` with seven
1619
+ deterministic stations, six judgment stations written as packets for an outside judge, and learn.
1620
+ Next: lineups over several passages of one draft, so that one lucky pick carries less weight, and
1621
+ a learn step that reads the runs ledger across drafts for the stations that keep failing and the
1622
+ changes that made them pass, beside what one pair of drafts shows.