@supersuit/hyperspec 0.6.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/CHANGELOG.md +91 -0
  2. package/README.md +37 -0
  3. package/SPEC.md +2 -2
  4. package/WRITING.md +542 -10
  5. package/bin/hyperspec.mjs +183 -0
  6. package/examples/writing/essay/judge/doctor.packet.json +108 -0
  7. package/examples/writing/essay/judge/lineup.packet.json +64 -0
  8. package/examples/writing/essay/judge/persona.packet.json +73 -0
  9. package/examples/writing/essay/judge/reader.packet.json +93 -0
  10. package/examples/writing/essay/learn/first-draft.md +84 -0
  11. package/examples/writing/essay/learn/learn.packet.json +106 -0
  12. package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
  13. package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
  14. package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
  15. package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
  16. package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
  17. package/examples/writing/essay.hyperspec.md +6 -1
  18. package/examples/writing/story/judge/attribution.packet.json +194 -0
  19. package/examples/writing/story/judge/doctor.packet.json +108 -0
  20. package/examples/writing/story/judge/knowledge.packet.json +77 -0
  21. package/examples/writing/story/judge/persona.packet.json +73 -0
  22. package/examples/writing/story/judge/reader.packet.json +94 -0
  23. package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
  24. package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
  25. package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
  26. package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
  27. package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
  28. package/examples/writing/story.hyperspec.md +7 -3
  29. package/package.json +1 -1
  30. package/src/check.mjs +75 -129
  31. package/src/draft.mjs +26 -0
  32. package/src/judge.mjs +386 -0
  33. package/src/judges/attribution.mjs +360 -0
  34. package/src/judges/doctor.mjs +126 -0
  35. package/src/judges/index.mjs +31 -0
  36. package/src/judges/knowledge.mjs +111 -0
  37. package/src/judges/lineup.mjs +272 -0
  38. package/src/judges/persona.mjs +137 -0
  39. package/src/judges/reader.mjs +111 -0
  40. package/src/learn.mjs +422 -0
  41. package/src/ledger.mjs +108 -0
  42. package/src/sentences.mjs +81 -0
  43. package/src/stations/claims.mjs +44 -39
  44. package/src/stations/quotes.mjs +6 -4
  45. package/src/writing.mjs +1 -1
package/CHANGELOG.md CHANGED
@@ -1,5 +1,96 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.7.0 (2026-09-29)
4
+
5
+ A spec can now have its judgments made and recorded. `check` covers what a function of the spec
6
+ and the draft can decide; the rest of a spec's checks are rubrics, and until this release nothing
7
+ ran them. hyperspec still calls no model. `hyperspec judge prepare` writes one packet per judgment
8
+ station, holding the rubric, fixed instructions, the inputs and the exact shape of the answer, for
9
+ an outside judge to fill: your agent, any model, or a person. `hyperspec judge record` checks the
10
+ verdict, derives the station's status from it by a fixed rule, and records it in the runs ledger.
11
+ Every passage a judge quotes must be in the draft, and the two blind tests are scored
12
+ against answer keys the judge never sees. `hyperspec learn` works from the other end: given the
13
+ first draft a factory wrote and the draft a person approved, it lists the edits, has a judge name
14
+ the spec block that should have prevented each one, and names one next move.
15
+
16
+ **No behavior change for lint or check.** Every 0.6.0 test passes unchanged, no finding id changed,
17
+ and no schema field was added. The runs ledger gains two kinds of line, `judge` and `learn`; check
18
+ reads only its own, so a 0.6 ledger's check verdicts are exactly what they were, and the new lines
19
+ keep lint's test 9 passing.
20
+
21
+ - `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force] [--json]`
22
+ lints the spec first, as `check` does, then writes `<station>.packet.json` for each station that
23
+ applies, in the order `doctor, lineup, reader, persona, attribution, knowledge`, and prints a
24
+ skip line with the reason for each that does not. Packets keep the paths as given and are
25
+ byte-identical for the same files. It refuses to overwrite any file without `--force`, naming
26
+ each. Exit 0 written, 1 a station crashed, 2 usage, or lint's own 1 or 3.
27
+ - The two blind tests write an answer key beside their packets, `lineup.key.json` and
28
+ `attribution.key.json`, for a person to read. Hand a judge only the `*.packet.json` files, and
29
+ each of the two blind packets to its own fresh context, with no access to the draft and apart
30
+ from the other packets, which carry it: their
31
+ instructions say to decide from the packet's inputs alone and open no file it names, but a judge
32
+ that can open the draft can always cheat. `record` never reads a key: it builds it again.
33
+ - `hyperspec judge record <packet> --verdict <file> [--json]` hashes the spec and the draft again
34
+ and rebuilds the packet from them and from the files the station reads (the DNA goldens, the
35
+ claims ledger). A file that changed makes the packet stale (`judge-stale`), and a packet that is
36
+ not those exact bytes is refused (`judge-packet-altered`); either way nothing is recorded. A
37
+ stale packet has `learn record`'s shape, `{ "invalid": true, "stale": true }` with `--json` and
38
+ "stale packet, nothing recorded" on the terminal. It must run in the folder `prepare` ran in. Then the verdict is validated, every problem named. Exit
39
+ 0 the station passed, 1 it failed or the verdict was refused, 2 usage.
40
+ - The evidence rule: every span a verdict cites is at least three words and appears in the draft,
41
+ on whole words, after whitespace runs become one space and curly, low and angle quotation marks
42
+ and apostrophes become straight ones (primes do not).
43
+ - Six stations. `doctor`: every goal condition, and whether the reader would take
44
+ `goal.next_if_worked` now. `lineup`: the draft's paragraph nearest the goldens' median length
45
+ beside up to three goldens' paragraphs, each reflowed to one line and shuffled with a seed
46
+ derived from the draft's full text, which the packet does not carry, so the packet cannot reveal
47
+ the order (a draft that is one paragraph and nothing else skips, since its hash would name its
48
+ candidate); it passes when the judge picks a golden, which a random pick does three times in four.
49
+ `reader`: where the audience's reader got lost (warnings) or stopped, and whether they would
50
+ take their own next step. `persona`: every break of stance, assertion, `will_not_say` or an
51
+ unsourced fact, against the claims ledger's texts (`null` with no ledger, and then no
52
+ unsourced fact may be reported). `attribution`, fiction only: the speaker of each dialogue line
53
+ whose speech tag names one, scored per speaker and averaged, passing at 80 percent; a line whose
54
+ speaker is in doubt (a split quote joins only across exactly one speech tag, and the part after
55
+ any longer narration is left out), or that repeats a golden or rejected line, is left out and
56
+ counted, so the key is never wrong (the worked story tests 19 of its 36 lines). `knowledge`, fiction only: every
57
+ place a character knows something before their timeline gives it to them.
58
+ - Judge ledger lines carry `packet_sha256`, the hash of what the judge was shown, and
59
+ `inputs_sha256`, the hash of its inputs, and follow check's rules against the last judge line
60
+ for the same station and draft, saying `judgment` where check says `check` ("no change since the
61
+ last passing judgment"), with three more: what changed is judged by the packet, naming the
62
+ draft, the spec, the file the station reads besides them when the inputs changed (the DNA scope,
63
+ the claims ledger), or the packet's fixed text when only that changed, so adding the golden or
64
+ the claims a failing station asked for and passing is improved; one-shot needs draft bytes never judged by that
65
+ station under any name; and improved needs the packet to have changed, since a judge answering
66
+ differently about the same packet is not the work improving ("the verdict changed; nothing the
67
+ judge was shown changed").
68
+ - `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]`
69
+ splits both drafts into sentence units (a paragraph's opening heading, each list item and each
70
+ sentence, with common abbreviations and lower-case continuations joined), diffs them, and
71
+ writes `learn.packet.json` with every edit as a hunk (`deleted`, `inserted` or `replaced`, with
72
+ both texts and a sentence count). A hunk never crosses a paragraph or a heading, and reflowed
73
+ text is no edit. Past 10,000,000 comparisons after the common start and end are set aside, it is
74
+ a usage error.
75
+ - `hyperspec learn record <packet> --verdict <file>` checks that every hunk is classified once by a
76
+ block the spec has written, or `none`, counts edits and sentences per block, and names one move
77
+ for the block with the most sentences, from a fixed table (`dna`: add a golden or a style rule).
78
+ It appends a `learn` line to the runs ledger, always `not-improved` since the spec has not
79
+ changed yet, and never edits the spec.
80
+ - The worked examples ship their packets (`essay/judge/`, `story/judge/`), with no key, and one
81
+ sample verdict per packet (`*/sample-verdicts/`), filled in by hand and marked `sample`. Two
82
+ of them fail, honestly: the essay's lineup (the draft's passage is the only one resting on a
83
+ figure) and the story's persona (three process details no claim holds). The essay's goal now
84
+ names the card its draft ends on as the reader's next step (`goal.next_if_worked`, matching the
85
+ spec's `agenda-card` decision), so its doctor sample passes. The essay adds a learn pair in `essay/learn/` and a sample verdict whose
86
+ tally sends the next move to `dna`. A test holds every packet to what `prepare` writes and records
87
+ every sample.
88
+ - WRITING.md gains "Judging a draft" and "Learning from edits": the commands and exit codes, the
89
+ packet, the evidence rule, what `record` refuses, each station's rules and findings table (held
90
+ by a test to the ids the code raises and their kind), the ledger rules, the samples and what they
91
+ found, sentence units, the diff, the tally and its move table. The quotes station's section
92
+ now points at `attribution` for fiction's dialogue.
93
+
3
94
  ## 0.6.0 (2026-09-29)
4
95
 
5
96
  A spec can now check a draft. Until this release hyperspec could tell you whether a writing spec
package/README.md CHANGED
@@ -29,6 +29,10 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
29
29
  | `hyperspec dna init <scope-dir> --writer W --form F --audience A --purpose P` | Start a writer-DNA scope folder. Refuses to overwrite an existing `scope.md`. |
30
30
  | `hyperspec dna measure <scope-dir>` | Check every golden in a scope and write its measured features. |
31
31
  | `hyperspec check <spec> --draft <file> [--only a,b]` | Run a writing spec's deterministic stations against a draft. |
32
+ | `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force]` | Write one packet per judgment station, for an outside judge to fill. |
33
+ | `hyperspec judge record <packet> --verdict <file>` | Check a judge's verdict against its packet, derive the station's status, and record it. |
34
+ | `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]` | Write the edits between a first draft and the approved one, for a judge to classify by spec block. |
35
+ | `hyperspec learn record <packet> --verdict <file>` | Count the classified edits by block and name one next move. |
32
36
  | `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
33
37
  | `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
34
38
  | `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
@@ -47,6 +51,12 @@ error or a file that cannot be read, has broken frontmatter, or is not a hypersp
47
51
  error, a draft that cannot be read, or a spec without `profile: writing`. A spec that is not ready to check against exits with lint's
48
52
  own code, 1 or 3, and no station runs.
49
53
 
54
+ `hyperspec judge record` exits 0 when the station passed and 1 when it failed or the verdict was
55
+ refused (invalid, or its packet stale or edited), and `hyperspec learn record` 0 when it recorded
56
+ and 1 when it refused. `judge prepare` and `learn prepare` exit 0 when they wrote their packets and
57
+ with lint's own code when the spec is not ready, and `judge prepare` exits 1 when a station could
58
+ not build its packet. All four exit 2 on a usage error.
59
+
50
60
  The recipe commands use the same numbers: 0 ok, 1 a check failed or the child regressed, 2
51
61
  usage or unreadable input, 3 pending, when `regenerate` has stages waiting for a runner.
52
62
  `regenerate` also exits 1 when a stage it ran reported a failing verdict, or none. A stage it
@@ -200,6 +210,33 @@ spec's runs ledger with a verdict. Both examples ship a draft that passes. What
200
210
  checks and cannot check, and every finding, are in
201
211
  [WRITING.md](WRITING.md#checking-a-draft).
202
212
 
213
+ ### Judging a draft and learning from edits
214
+
215
+ The rest of a spec's checks are judgments: whether each goal condition holds, where a reader gets
216
+ lost, whether a passage can be told from the writer's goldens, whether the persona holds, and in
217
+ fiction whether the characters' voices can be told apart and whether anyone knows something too
218
+ early. hyperspec never calls a model. `judge prepare` writes a packet per station for an outside
219
+ judge, your agent, any model or a person, and `judge record` checks the verdict (every quoted
220
+ passage must be in the draft, and the blind tests are scored against keys the judge never
221
+ sees), derives the station's status, and records it in the runs ledger.
222
+
223
+ ```bash
224
+ npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
225
+ npx @supersuit/hyperspec judge record essay/judge/doctor.packet.json --verdict doctor.verdict.json
226
+ ```
227
+
228
+ Hand the judge the `*.packet.json` files only: the answer keys are written beside them. Give the
229
+ two blind packets (`lineup`, `attribution`) to a judge in a fresh context with no access to the
230
+ draft, such as a new conversation, never to an agent working in the draft's folder: a judge that
231
+ can open the draft can always find the answer, whatever the packet tells it. Give each blind
232
+ packet its own context, apart from the other packets too, since those carry the draft. `learn`
233
+ closes the loop from the other end. Given the first draft a factory wrote and the draft a person
234
+ approved, `learn prepare` lists the edits, a judge names the spec block that should have prevented
235
+ each, and `learn record` counts them by block and names one next move, such as "add a golden or a
236
+ style rule". It never edits the spec. Both examples ship their packets and hand-filled sample
237
+ verdicts. The packet shapes, every station's rules and findings, and the learn tally are in
238
+ [WRITING.md](WRITING.md#judging-a-draft).
239
+
203
240
  ## The format
204
241
 
205
242
  A hyperspec is a markdown file with a YAML frontmatter block: `decisions`, `requirements`,
package/SPEC.md CHANGED
@@ -97,7 +97,7 @@ examples:
97
97
  - path: examples/minimal.hyperspec.md
98
98
  why: the smallest spec that passes all nine tests
99
99
  resume:
100
- next_action: collect adopter issues on 0.6, the check command included, and cut 0.7 from them
100
+ next_action: collect adopter issues on 0.7, the judge and learn commands included, and cut 0.8 from them
101
101
  feedback:
102
102
  issues: https://github.com/SupersuitUp/hyperspec/issues
103
103
  fork: MIT; fork it for your own purposes and say so in your SPEC
@@ -109,7 +109,7 @@ improvement:
109
109
 
110
110
  A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
111
111
 
112
- **Version 0.6.0** (2026-09-29)
112
+ **Version 0.7.0** (2026-09-29)
113
113
 
114
114
  ## What makes a spec a hyperspec
115
115