@supersuit/hyperspec 0.6.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/CHANGELOG.md +127 -0
  2. package/README.md +71 -8
  3. package/SPEC.md +2 -2
  4. package/WRITING.md +680 -21
  5. package/bin/hyperspec.mjs +188 -4
  6. package/examples/writing/course/claims.jsonl +0 -0
  7. package/examples/writing/course/goldens/lesson.md +1 -0
  8. package/examples/writing/course/materials/brief.md +9 -0
  9. package/examples/writing/course/materials/brief.md.segments.jsonl +6 -0
  10. package/examples/writing/course/outline.md +11 -0
  11. package/examples/writing/course/part-1.md +47 -0
  12. package/examples/writing/course/part-2.md +40 -0
  13. package/examples/writing/course/runs.jsonl +0 -0
  14. package/examples/writing/course.hyperspec.md +205 -0
  15. package/examples/writing/essay/judge/doctor.packet.json +108 -0
  16. package/examples/writing/essay/judge/lineup.packet.json +64 -0
  17. package/examples/writing/essay/judge/persona.packet.json +73 -0
  18. package/examples/writing/essay/judge/reader.packet.json +93 -0
  19. package/examples/writing/essay/learn/first-draft.md +84 -0
  20. package/examples/writing/essay/learn/learn.packet.json +106 -0
  21. package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
  22. package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
  23. package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
  24. package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
  25. package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
  26. package/examples/writing/essay.hyperspec.md +6 -1
  27. package/examples/writing/story/judge/attribution.packet.json +194 -0
  28. package/examples/writing/story/judge/doctor.packet.json +108 -0
  29. package/examples/writing/story/judge/knowledge.packet.json +77 -0
  30. package/examples/writing/story/judge/persona.packet.json +73 -0
  31. package/examples/writing/story/judge/reader.packet.json +94 -0
  32. package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
  33. package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
  34. package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
  35. package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
  36. package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
  37. package/examples/writing/story.hyperspec.md +7 -3
  38. package/package.json +1 -1
  39. package/src/check.mjs +96 -132
  40. package/src/draft.mjs +26 -0
  41. package/src/judge.mjs +386 -0
  42. package/src/judges/attribution.mjs +360 -0
  43. package/src/judges/doctor.mjs +126 -0
  44. package/src/judges/index.mjs +31 -0
  45. package/src/judges/knowledge.mjs +111 -0
  46. package/src/judges/lineup.mjs +272 -0
  47. package/src/judges/persona.mjs +137 -0
  48. package/src/judges/reader.mjs +111 -0
  49. package/src/learn.mjs +422 -0
  50. package/src/ledger.mjs +108 -0
  51. package/src/sentences.mjs +81 -0
  52. package/src/sequence-draft.mjs +75 -0
  53. package/src/stations/claims.mjs +44 -39
  54. package/src/stations/index.mjs +3 -1
  55. package/src/stations/links.mjs +11 -3
  56. package/src/stations/quotes.mjs +6 -4
  57. package/src/stations/sequence.mjs +275 -0
  58. package/src/writing-fields.mjs +34 -0
  59. package/src/writing.mjs +1 -1
package/CHANGELOG.md CHANGED
@@ -1,5 +1,132 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.8.0 (2026-09-29)
4
+
5
+ A work read in order can now be checked as one. Every station so far held ONE piece to its spec; a
6
+ course, a primer or a textbook makes promises ACROSS its pieces (Lesson 5 is written for someone who
7
+ has read Lessons 1 to 4 and nothing else), and nothing checked them. Declare `writing.form.sequence`
8
+ and the new `sequence` station does, deterministically, on every `check`. Promoted from a checker
9
+ written for one book, so every sequential work gets it by construction.
10
+
11
+ - `writing.form.sequence`, optional, every key optional: `unit` (the heading word, default
12
+ `Lesson`), `files` (the work's files in reading order; a `*` in a file name matches, sorted by
13
+ number), `sections` (default "After this lesson you can", "New terms", "Try this"),
14
+ `terms_section` (default "New terms"), `outline` (numbered items promising `*Terms: a, b.*`),
15
+ `knows` (words a lesson may use before one defines them, beside `audience.knows`) and `teaser`
16
+ (default "Next,": a closing line naming what is coming). Lint refuses a key that is present and
17
+ hollow under test 1, and a `files` entry matching nothing or an `outline` that is not a file under
18
+ test 6.
19
+ - The `sequence` station, run last: every unit carries its sections; every term is defined in
20
+ exactly one unit; no unit uses a term before the unit that defines it (code, the unit's own terms
21
+ section and its closing teaser are not uses, and a `(from Lesson N)` reminder defines nothing);
22
+ each unit defines what the outline promises, for the units the draft holds; unit numbers increase;
23
+ a pointer to a later unit is a warning. Findings: `station-sequence-no-units`, `-numbering`,
24
+ `-missing-section`, `-defined-twice`, `-used-before-defined`, `-outline-unreadable`,
25
+ `-outline-unkept` (fail) and `-forward-pointer` (warn). A spec with no `sequence` skips it.
26
+ - `hyperspec check <spec>` needs no `--draft` when the spec lists `sequence.files`: the files,
27
+ joined in order with each one's frontmatter blanked, are the draft, and every station reads them.
28
+ A finding names the file and its own line (`(course/part-1.md line 25)`; `file` and `line` in
29
+ `--json`), a relative link resolves beside the file that holds it, and the ledger keys the run by
30
+ the `files` entry as written, so the work keeps one history as parts are added. `--draft` still
31
+ names one file, which may hold every lesson.
32
+ - A third worked example, `examples/writing/course.hyperspec.md`: four lessons in two part files
33
+ with an outline, passing every station with two forward-pointer warnings.
34
+
35
+ **Behavior change:** `check` runs eight stations, so `--json` and the ledger's `stations` carry
36
+ `sequence` (as `skip` for every spec without a sequence). No other station, finding id or schema
37
+ field changed, and a 0.7 ledger's verdicts read as they did.
38
+
39
+ ## 0.7.0 (2026-09-29)
40
+
41
+ A spec can now have its judgments made and recorded. `check` covers what a function of the spec
42
+ and the draft can decide; the rest of a spec's checks are rubrics, and until this release nothing
43
+ ran them. hyperspec still calls no model. `hyperspec judge prepare` writes one packet per judgment
44
+ station, holding the rubric, fixed instructions, the inputs and the exact shape of the answer, for
45
+ an outside judge to fill: your agent, any model, or a person. `hyperspec judge record` checks the
46
+ verdict, derives the station's status from it by a fixed rule, and records it in the runs ledger.
47
+ Every passage a judge quotes must be in the draft, and the two blind tests are scored
48
+ against answer keys the judge never sees. `hyperspec learn` works from the other end: given the
49
+ first draft a factory wrote and the draft a person approved, it lists the edits, has a judge name
50
+ the spec block that should have prevented each one, and names one next move.
51
+
52
+ **No behavior change for lint or check.** Every 0.6.0 test passes unchanged, no finding id changed,
53
+ and no schema field was added. The runs ledger gains two kinds of line, `judge` and `learn`; check
54
+ reads only its own, so a 0.6 ledger's check verdicts are exactly what they were, and the new lines
55
+ keep lint's test 9 passing.
56
+
57
+ - `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force] [--json]`
58
+ lints the spec first, as `check` does, then writes `<station>.packet.json` for each station that
59
+ applies, in the order `doctor, lineup, reader, persona, attribution, knowledge`, and prints a
60
+ skip line with the reason for each that does not. Packets keep the paths as given and are
61
+ byte-identical for the same files. It refuses to overwrite any file without `--force`, naming
62
+ each. Exit 0 written, 1 a station crashed, 2 usage, or lint's own 1 or 3.
63
+ - The two blind tests write an answer key beside their packets, `lineup.key.json` and
64
+ `attribution.key.json`, for a person to read. Hand a judge only the `*.packet.json` files, and
65
+ each of the two blind packets to its own fresh context, with no access to the draft and apart
66
+ from the other packets, which carry it: their
67
+ instructions say to decide from the packet's inputs alone and open no file it names, but a judge
68
+ that can open the draft can always cheat. `record` never reads a key: it builds it again.
69
+ - `hyperspec judge record <packet> --verdict <file> [--json]` hashes the spec and the draft again
70
+ and rebuilds the packet from them and from the files the station reads (the DNA goldens, the
71
+ claims ledger). A file that changed makes the packet stale (`judge-stale`), and a packet that is
72
+ not those exact bytes is refused (`judge-packet-altered`); either way nothing is recorded. A
73
+ stale packet has `learn record`'s shape, `{ "invalid": true, "stale": true }` with `--json` and
74
+ "stale packet, nothing recorded" on the terminal. It must run in the folder `prepare` ran in. Then the verdict is validated, every problem named. Exit
75
+ 0 the station passed, 1 it failed or the verdict was refused, 2 usage.
76
+ - The evidence rule: every span a verdict cites is at least three words and appears in the draft,
77
+ on whole words, after whitespace runs become one space and curly, low and angle quotation marks
78
+ and apostrophes become straight ones (primes do not).
79
+ - Six stations. `doctor`: every goal condition, and whether the reader would take
80
+ `goal.next_if_worked` now. `lineup`: the draft's paragraph nearest the goldens' median length
81
+ beside up to three goldens' paragraphs, each reflowed to one line and shuffled with a seed
82
+ derived from the draft's full text, which the packet does not carry, so the packet cannot reveal
83
+ the order (a draft that is one paragraph and nothing else skips, since its hash would name its
84
+ candidate); it passes when the judge picks a golden, which a random pick does three times in four.
85
+ `reader`: where the audience's reader got lost (warnings) or stopped, and whether they would
86
+ take their own next step. `persona`: every break of stance, assertion, `will_not_say` or an
87
+ unsourced fact, against the claims ledger's texts (`null` with no ledger, and then no
88
+ unsourced fact may be reported). `attribution`, fiction only: the speaker of each dialogue line
89
+ whose speech tag names one, scored per speaker and averaged, passing at 80 percent; a line whose
90
+ speaker is in doubt (a split quote joins only across exactly one speech tag, and the part after
91
+ any longer narration is left out), or that repeats a golden or rejected line, is left out and
92
+ counted, so the key is never wrong (the worked story tests 19 of its 36 lines). `knowledge`, fiction only: every
93
+ place a character knows something before their timeline gives it to them.
94
+ - Judge ledger lines carry `packet_sha256`, the hash of what the judge was shown, and
95
+ `inputs_sha256`, the hash of its inputs, and follow check's rules against the last judge line
96
+ for the same station and draft, saying `judgment` where check says `check` ("no change since the
97
+ last passing judgment"), with three more: what changed is judged by the packet, naming the
98
+ draft, the spec, the file the station reads besides them when the inputs changed (the DNA scope,
99
+ the claims ledger), or the packet's fixed text when only that changed, so adding the golden or
100
+ the claims a failing station asked for and passing is improved; one-shot needs draft bytes never judged by that
101
+ station under any name; and improved needs the packet to have changed, since a judge answering
102
+ differently about the same packet is not the work improving ("the verdict changed; nothing the
103
+ judge was shown changed").
104
+ - `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]`
105
+ splits both drafts into sentence units (a paragraph's opening heading, each list item and each
106
+ sentence, with common abbreviations and lower-case continuations joined), diffs them, and
107
+ writes `learn.packet.json` with every edit as a hunk (`deleted`, `inserted` or `replaced`, with
108
+ both texts and a sentence count). A hunk never crosses a paragraph or a heading, and reflowed
109
+ text is no edit. Past 10,000,000 comparisons after the common start and end are set aside, it is
110
+ a usage error.
111
+ - `hyperspec learn record <packet> --verdict <file>` checks that every hunk is classified once by a
112
+ block the spec has written, or `none`, counts edits and sentences per block, and names one move
113
+ for the block with the most sentences, from a fixed table (`dna`: add a golden or a style rule).
114
+ It appends a `learn` line to the runs ledger, always `not-improved` since the spec has not
115
+ changed yet, and never edits the spec.
116
+ - The worked examples ship their packets (`essay/judge/`, `story/judge/`), with no key, and one
117
+ sample verdict per packet (`*/sample-verdicts/`), filled in by hand and marked `sample`. Two
118
+ of them fail, honestly: the essay's lineup (the draft's passage is the only one resting on a
119
+ figure) and the story's persona (three process details no claim holds). The essay's goal now
120
+ names the card its draft ends on as the reader's next step (`goal.next_if_worked`, matching the
121
+ spec's `agenda-card` decision), so its doctor sample passes. The essay adds a learn pair in `essay/learn/` and a sample verdict whose
122
+ tally sends the next move to `dna`. A test holds every packet to what `prepare` writes and records
123
+ every sample.
124
+ - WRITING.md gains "Judging a draft" and "Learning from edits": the commands and exit codes, the
125
+ packet, the evidence rule, what `record` refuses, each station's rules and findings table (held
126
+ by a test to the ids the code raises and their kind), the ledger rules, the samples and what they
127
+ found, sentence units, the diff, the tally and its move table. The quotes station's section
128
+ now points at `attribution` for fiction's dialogue.
129
+
3
130
  ## 0.6.0 (2026-09-29)
4
131
 
5
132
  A spec can now check a draft. Until this release hyperspec could tell you whether a writing spec
package/README.md CHANGED
@@ -28,7 +28,11 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
28
28
  | `hyperspec segments init <material> --id <mid> [--out F] [--by paragraph\|sentence]` | Split a material into segments to label. Refuses to overwrite an existing file. |
29
29
  | `hyperspec dna init <scope-dir> --writer W --form F --audience A --purpose P` | Start a writer-DNA scope folder. Refuses to overwrite an existing `scope.md`. |
30
30
  | `hyperspec dna measure <scope-dir>` | Check every golden in a scope and write its measured features. |
31
- | `hyperspec check <spec> --draft <file> [--only a,b]` | Run a writing spec's deterministic stations against a draft. |
31
+ | `hyperspec check <spec> [--draft <file>] [--only a,b]` | Run a writing spec's deterministic stations against a draft, or, for a sequential work, against its files in reading order. |
32
+ | `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force]` | Write one packet per judgment station, for an outside judge to fill. |
33
+ | `hyperspec judge record <packet> --verdict <file>` | Check a judge's verdict against its packet, derive the station's status, and record it. |
34
+ | `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]` | Write the edits between a first draft and the approved one, for a judge to classify by spec block. |
35
+ | `hyperspec learn record <packet> --verdict <file>` | Count the classified edits by block and name one next move. |
32
36
  | `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
33
37
  | `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
34
38
  | `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
@@ -47,6 +51,12 @@ error or a file that cannot be read, has broken frontmatter, or is not a hypersp
47
51
  error, a draft that cannot be read, or a spec without `profile: writing`. A spec that is not ready to check against exits with lint's
48
52
  own code, 1 or 3, and no station runs.
49
53
 
54
+ `hyperspec judge record` exits 0 when the station passed and 1 when it failed or the verdict was
55
+ refused (invalid, or its packet stale or edited), and `hyperspec learn record` 0 when it recorded
56
+ and 1 when it refused. `judge prepare` and `learn prepare` exit 0 when they wrote their packets and
57
+ with lint's own code when the spec is not ready, and `judge prepare` exits 1 when a station could
58
+ not build its packet. All four exit 2 on a usage error.
59
+
50
60
  The recipe commands use the same numbers: 0 ok, 1 a check failed or the child regressed, 2
51
61
  usage or unreadable input, 3 pending, when `regenerate` has stages waiting for a runner.
52
62
  `regenerate` also exits 1 when a stage it ran reported a failing verdict, or none. A stage it
@@ -135,9 +145,10 @@ npx @supersuit/hyperspec init story.hyperspec.md --profile writing --form "short
135
145
 
136
146
  The skeleton fails until every placeholder is real and its four open questions (whose voice,
137
147
  who speaks, who reads, what changes) are answered. The blocks, every field, and which test
138
- each rule reports under are in [WRITING.md](WRITING.md). Two complete specs that pass with
139
- nothing to warn ship in `examples/writing/`: an essay for new managers, and a short story with
140
- two characters whose voices a judge can tell apart. Each comes with a draft written to it.
148
+ each rule reports under are in [WRITING.md](WRITING.md). Three complete specs that pass with
149
+ nothing to warn ship in `examples/writing/`: an essay for new managers, a short story with
150
+ two characters whose voices a judge can tell apart, and a four-lesson course. Each comes with a
151
+ draft written to it.
141
152
 
142
153
  ### Marking materials
143
154
 
@@ -183,23 +194,75 @@ it did in 0.4. The folder shape, every feature, and every finding are in
183
194
 
184
195
  ### Checking a draft
185
196
 
186
- Once a draft exists, `check` holds it to its spec with seven stations, none of which calls a
197
+ Once a draft exists, `check` holds it to its spec with eight stations, none of which calls a
187
198
  model or touches the network: `form` (length and required parts), `terms` (every word in the new
188
199
  optional `writing.audience.terms` is defined where it first appears), `claims` (the claims
189
200
  ledger still matches the draft, and every claim has a source), `quotes` (in nonfiction, every quotation of four
190
201
  words or more is word for word in a marked quote), `private` (no run of eight words from a
191
- private segment), `dna` (the draft's measured style beside its scope's, as warnings) and `links`
192
- (well-formed, and relative links resolve).
202
+ private segment), `dna` (the draft's measured style beside its scope's, as warnings), `links`
203
+ (well-formed, and relative links resolve) and `sequence` (for a work read in order; see below).
193
204
 
194
205
  ```bash
195
206
  npx @supersuit/hyperspec check essay.hyperspec.md --draft essay/draft.md
196
207
  ```
197
208
 
198
209
  It lints the spec first, prints each station's pass, fail or skip, and appends one line to the
199
- spec's runs ledger with a verdict. Both examples ship a draft that passes. What each station
210
+ spec's runs ledger with a verdict. Every example ships a draft that passes. What each station
200
211
  checks and cannot check, and every finding, are in
201
212
  [WRITING.md](WRITING.md#checking-a-draft).
202
213
 
214
+ ### Sequential works
215
+
216
+ A course, a primer or a textbook promises something no single piece can check: Lesson 5 uses only
217
+ words Lessons 1 to 4 defined. Add `sequence:` to `writing.form` and the `sequence` station checks
218
+ it across the whole work: every lesson carries its sections ("After this lesson you can", "New
219
+ terms", "Try this" by default), every term is defined in exactly one lesson, no lesson uses a term
220
+ before the lesson that defines it (code, the part's closing teaser and words the reader already
221
+ knows are exempt), each lesson defines the terms the outline promises, and a pointer to a later
222
+ lesson is a warning. List the work's files and `check` needs no `--draft`:
223
+
224
+ ```yaml
225
+ sequence:
226
+ files:
227
+ - course/part-*.md
228
+ outline: course/outline.md
229
+ ```
230
+
231
+ ```bash
232
+ npx @supersuit/hyperspec check course.hyperspec.md
233
+ ```
234
+
235
+ The parts are read in order, a finding names the part and its line, and a new part matching the
236
+ pattern is picked up without touching the spec. Every key, how a lesson is read, and every finding
237
+ are in [WRITING.md](WRITING.md#sequential-works).
238
+
239
+ ### Judging a draft and learning from edits
240
+
241
+ The rest of a spec's checks are judgments: whether each goal condition holds, where a reader gets
242
+ lost, whether a passage can be told from the writer's goldens, whether the persona holds, and in
243
+ fiction whether the characters' voices can be told apart and whether anyone knows something too
244
+ early. hyperspec never calls a model. `judge prepare` writes a packet per station for an outside
245
+ judge, your agent, any model or a person, and `judge record` checks the verdict (every quoted
246
+ passage must be in the draft, and the blind tests are scored against keys the judge never
247
+ sees), derives the station's status, and records it in the runs ledger.
248
+
249
+ ```bash
250
+ npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
251
+ npx @supersuit/hyperspec judge record essay/judge/doctor.packet.json --verdict doctor.verdict.json
252
+ ```
253
+
254
+ Hand the judge the `*.packet.json` files only: the answer keys are written beside them. Give the
255
+ two blind packets (`lineup`, `attribution`) to a judge in a fresh context with no access to the
256
+ draft, such as a new conversation, never to an agent working in the draft's folder: a judge that
257
+ can open the draft can always find the answer, whatever the packet tells it. Give each blind
258
+ packet its own context, apart from the other packets too, since those carry the draft. `learn`
259
+ closes the loop from the other end. Given the first draft a factory wrote and the draft a person
260
+ approved, `learn prepare` lists the edits, a judge names the spec block that should have prevented
261
+ each, and `learn record` counts them by block and names one next move, such as "add a golden or a
262
+ style rule". It never edits the spec. Both examples ship their packets and hand-filled sample
263
+ verdicts. The packet shapes, every station's rules and findings, and the learn tally are in
264
+ [WRITING.md](WRITING.md#judging-a-draft).
265
+
203
266
  ## The format
204
267
 
205
268
  A hyperspec is a markdown file with a YAML frontmatter block: `decisions`, `requirements`,
package/SPEC.md CHANGED
@@ -97,7 +97,7 @@ examples:
97
97
  - path: examples/minimal.hyperspec.md
98
98
  why: the smallest spec that passes all nine tests
99
99
  resume:
100
- next_action: collect adopter issues on 0.6, the check command included, and cut 0.7 from them
100
+ next_action: collect adopter issues on 0.7, the judge and learn commands included, and cut 0.8 from them
101
101
  feedback:
102
102
  issues: https://github.com/SupersuitUp/hyperspec/issues
103
103
  fork: MIT; fork it for your own purposes and say so in your SPEC
@@ -109,7 +109,7 @@ improvement:
109
109
 
110
110
  A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
111
111
 
112
- **Version 0.6.0** (2026-09-29)
112
+ **Version 0.8.0** (2026-09-29)
113
113
 
114
114
  ## What makes a spec a hyperspec
115
115