@supersuit/hyperspec 0.6.0 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +127 -0
- package/README.md +71 -8
- package/SPEC.md +2 -2
- package/WRITING.md +680 -21
- package/bin/hyperspec.mjs +188 -4
- package/examples/writing/course/claims.jsonl +0 -0
- package/examples/writing/course/goldens/lesson.md +1 -0
- package/examples/writing/course/materials/brief.md +9 -0
- package/examples/writing/course/materials/brief.md.segments.jsonl +6 -0
- package/examples/writing/course/outline.md +11 -0
- package/examples/writing/course/part-1.md +47 -0
- package/examples/writing/course/part-2.md +40 -0
- package/examples/writing/course/runs.jsonl +0 -0
- package/examples/writing/course.hyperspec.md +205 -0
- package/examples/writing/essay/judge/doctor.packet.json +108 -0
- package/examples/writing/essay/judge/lineup.packet.json +64 -0
- package/examples/writing/essay/judge/persona.packet.json +73 -0
- package/examples/writing/essay/judge/reader.packet.json +93 -0
- package/examples/writing/essay/learn/first-draft.md +84 -0
- package/examples/writing/essay/learn/learn.packet.json +106 -0
- package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
- package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
- package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
- package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
- package/examples/writing/essay.hyperspec.md +6 -1
- package/examples/writing/story/judge/attribution.packet.json +194 -0
- package/examples/writing/story/judge/doctor.packet.json +108 -0
- package/examples/writing/story/judge/knowledge.packet.json +77 -0
- package/examples/writing/story/judge/persona.packet.json +73 -0
- package/examples/writing/story/judge/reader.packet.json +94 -0
- package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
- package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
- package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
- package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
- package/examples/writing/story.hyperspec.md +7 -3
- package/package.json +1 -1
- package/src/check.mjs +96 -132
- package/src/draft.mjs +26 -0
- package/src/judge.mjs +386 -0
- package/src/judges/attribution.mjs +360 -0
- package/src/judges/doctor.mjs +126 -0
- package/src/judges/index.mjs +31 -0
- package/src/judges/knowledge.mjs +111 -0
- package/src/judges/lineup.mjs +272 -0
- package/src/judges/persona.mjs +137 -0
- package/src/judges/reader.mjs +111 -0
- package/src/learn.mjs +422 -0
- package/src/ledger.mjs +108 -0
- package/src/sentences.mjs +81 -0
- package/src/sequence-draft.mjs +75 -0
- package/src/stations/claims.mjs +44 -39
- package/src/stations/index.mjs +3 -1
- package/src/stations/links.mjs +11 -3
- package/src/stations/quotes.mjs +6 -4
- package/src/stations/sequence.mjs +275 -0
- package/src/writing-fields.mjs +34 -0
- package/src/writing.mjs +1 -1
package/WRITING.md
CHANGED
|
@@ -219,7 +219,7 @@ writing:
|
|
|
219
219
|
goal:
|
|
220
220
|
from: plans to run the meeting from their own list
|
|
221
221
|
to: hands the meeting to the report
|
|
222
|
-
next_if_worked:
|
|
222
|
+
next_if_worked: writes the three questions on a card
|
|
223
223
|
change:
|
|
224
224
|
kind: action # belief | action | feeling
|
|
225
225
|
text: the reader asks the three questions and waits
|
|
@@ -245,6 +245,10 @@ writing:
|
|
|
245
245
|
- a close
|
|
246
246
|
stations: # may be empty
|
|
247
247
|
- the three questions render as a numbered list
|
|
248
|
+
sequence: # optional: a work read in order; see Sequential works
|
|
249
|
+
files: # optional: its files in reading order; a * in a file name matches
|
|
250
|
+
- course/part-*.md
|
|
251
|
+
outline: course/outline.md # optional: the outline that promises each unit's terms
|
|
248
252
|
check:
|
|
249
253
|
station: structure and length check
|
|
250
254
|
source: form decision
|
|
@@ -337,12 +341,12 @@ Each row lists what the writing profile adds to that test. The core conditions i
|
|
|
337
341
|
|
|
338
342
|
| Test | A writing spec fails it when |
|
|
339
343
|
|---|---|
|
|
340
|
-
| 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered; `dna.scope_dir` is present and is a placeholder; with `dna.scope_dir`, its `scope.md` is missing, unreadable or lacks a field, its writer, form, audience or purpose differs from the spec's, or its `goldens/` folder is missing or empty, or holds a golden that cannot be read, whose frontmatter never closes, or that has no passage; `audience.terms`, when present, holds a non-string entry or has no real entries at all. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
|
|
344
|
+
| 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered; `dna.scope_dir` is present and is a placeholder; with `dna.scope_dir`, its `scope.md` is missing, unreadable or lacks a field, its writer, form, audience or purpose differs from the spec's, or its `goldens/` folder is missing or empty, or holds a golden that cannot be read, whose frontmatter never closes, or that has no passage; `audience.terms`, when present, holds a non-string entry or has no real entries at all; `form.sequence` is present and is not a map, or has `unit`, `terms_section`, `teaser` or `outline` empty or a placeholder, or has `files`, `sections` or `knows` that is not a list of real entries. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
|
|
341
345
|
| 2 every requirement can fail | `goal.conditions` lists fewer than five or more than ten distinct ids, lists an id twice, or names an id that is not a top-level requirement |
|
|
342
346
|
| 3 every requirement names its check | a block or a character has no `check` with a `station` or a `rubric` |
|
|
343
347
|
| 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items`, or a segment that is not in that material's segments file; a segment's text does not match its material word for word; a material changed after it was marked; a claim segment has no `source` and no `own`, a story no `teller`, a quote no `speaker`; with `dna.scope_dir`, a golden in the scope has no `approved_by`, an approver that starts `agent:`, or no `source` |
|
|
344
348
|
| 5 negative space is specified | `persona.will_not_say` is empty; `persona.facts_from` is anything other than `sources`; a spine claim cites a `private` or a `question` segment; with `dna.scope_dir`, a golden the spec lists is not one of the scope's goldens (it lives outside the scope's `goldens/` folder, is a symlink that resolves outside it, or sits in a subfolder, is `README.md` or is not a `.md` file), or the scope's `goldens/` folder resolves outside the scope |
|
|
345
|
-
| 6 examples outrank adjectives | a golden has no `why`; a material, `dna.rules`, golden
|
|
349
|
+
| 6 examples outrank adjectives | a golden has no `why`; a material, `dna.rules`, golden, character `entity` or `form.sequence.outline` path does not exist or is not a file; a `form.sequence.files` entry matches no file; a character has no golden lines or no rejected lines, or has the same line in both (compared trimmed and case-folded); with `dna.scope_dir`, a golden in the scope has no `why`, or the scope's `features.json` is missing or is not what `dna measure` would write now |
|
|
346
350
|
| 7 a stranger can resume it | `writing.progress` exists. An unknown `profile:` is a warning |
|
|
347
351
|
| 8 its adopters can push back on it | nothing further; the core rule applies |
|
|
348
352
|
| 9 it improves itself | nothing further; the core rule applies |
|
|
@@ -810,7 +814,7 @@ npx @supersuit/hyperspec check essay.hyperspec.md --draft essay/draft.md
|
|
|
810
814
|
|
|
811
815
|
It lints the spec first. A spec that fails lint, or is blocked on an open decision, runs no
|
|
812
816
|
station and exits with lint's own code, because a draft cannot be checked against a spec that is
|
|
813
|
-
not ready. Then it runs
|
|
817
|
+
not ready. Then it runs eight stations in a fixed order and prints one line for each: `pass`,
|
|
814
818
|
`fail` with its findings, or `skip` with the reason. A warning prints under its station and never
|
|
815
819
|
fails it. This is the essay example's draft:
|
|
816
820
|
|
|
@@ -824,6 +828,7 @@ dna: pass
|
|
|
824
828
|
warn [station-dna-drift] first_person_singular_rate is 22.892 in the draft; the scope's goldens measure 0, band 0 to 5
|
|
825
829
|
fix: Bring first_person_singular_rate back inside the band, or, if the scope no longer describes this writer, re-measure it with better goldens.
|
|
826
830
|
links: pass
|
|
831
|
+
sequence: skip (the spec declares no writing.form.sequence)
|
|
827
832
|
verdict: one-shot
|
|
828
833
|
```
|
|
829
834
|
|
|
@@ -832,10 +837,13 @@ marked partial (see [The runs ledger](#the-runs-ledger)). `--json` prints the wh
|
|
|
832
837
|
finding included; a spec that is not ready prints lint's result with `lintBlocked: true` instead,
|
|
833
838
|
and a usage error prints `{ "spec", "draft", "error" }`. A finding names the draft line it points
|
|
834
839
|
at where there is one, quotes at most 80 characters of the draft, and never prints an absolute
|
|
835
|
-
path. A UTF-8 byte order mark at the start of the draft is ignored.
|
|
840
|
+
path. A UTF-8 byte order mark at the start of the draft is ignored. A spec that lists
|
|
841
|
+
`form.sequence.files` needs no `--draft`: its files, joined in reading order, are the draft, and a
|
|
842
|
+
finding names the file and its own line in it (see [Sequential works](#sequential-works)).
|
|
836
843
|
|
|
837
844
|
Exit codes: **0** every station that ran passed (a skip or a warning does not fail it); **1** a
|
|
838
|
-
station failed; **2** usage: no spec path, no `--draft
|
|
845
|
+
station failed; **2** usage: no spec path, no `--draft` for a spec that lists no
|
|
846
|
+
`form.sequence.files`, a draft that cannot be read, a spec
|
|
839
847
|
without `profile: writing`, or an `--only` that names no known station; and lint's own **1** or
|
|
840
848
|
**3** when the spec is not ready.
|
|
841
849
|
|
|
@@ -913,8 +921,9 @@ letters and is not a common function word such as "the", so a speaker recorded a
|
|
|
913
921
|
interviewed" is named only by all three words. Attribution needs a declared speaker: a name that
|
|
914
922
|
is no segment's `speaker` attributes nothing, so start a `speaker` with the person's name, as
|
|
915
923
|
the essay example does with `dana, an engineering manager`. A spec with `fiction: true` skips
|
|
916
|
-
the station: a character's dialogue is invented rather than quoted from a material
|
|
917
|
-
|
|
924
|
+
the station: a character's dialogue is invented rather than quoted from a material. The
|
|
925
|
+
`attribution` judge tests it against each character's own lines instead (see
|
|
926
|
+
[Judging a draft](#judging-a-draft)).
|
|
918
927
|
|
|
919
928
|
| Id | Severity | Meaning |
|
|
920
929
|
|---|---|---|
|
|
@@ -972,6 +981,25 @@ the network, so it cannot tell you a URL is live.
|
|
|
972
981
|
| `station-links-root-relative` | warn | a link starting with `/`, which cannot be resolved without the site |
|
|
973
982
|
| `station-links-undefined-reference` | fail | a full or collapsed reference link whose label has no definition |
|
|
974
983
|
|
|
984
|
+
### sequence
|
|
985
|
+
|
|
986
|
+
Runs when the spec declares `form.sequence`, and holds a work read in order to what its reader
|
|
987
|
+
depends on: every lesson carries its sections, every term is defined once, no lesson uses a term
|
|
988
|
+
before the lesson that defines it, and each lesson defines what the outline promises. How a unit is
|
|
989
|
+
read, and each guard, are under [Sequential works](#sequential-works). With no `form.sequence` the
|
|
990
|
+
station skips. It reads words, never meaning: a term used in another sense still counts as a use.
|
|
991
|
+
|
|
992
|
+
| Id | Severity | Meaning |
|
|
993
|
+
|---|---|---|
|
|
994
|
+
| `station-sequence-no-units` | fail | the draft has no `<unit> <n>` heading |
|
|
995
|
+
| `station-sequence-numbering` | fail | a unit's number is not greater than the one before it |
|
|
996
|
+
| `station-sequence-missing-section` | fail | a unit lacks one of `sections`; names the unit and the section |
|
|
997
|
+
| `station-sequence-defined-twice` | fail | a term is defined in two units' terms sections; names both |
|
|
998
|
+
| `station-sequence-used-before-defined` | fail | a unit uses a term before the unit that defines it; points at the first use |
|
|
999
|
+
| `station-sequence-outline-unreadable` | fail | `outline` is set and the file cannot be read |
|
|
1000
|
+
| `station-sequence-outline-unkept` | fail | the outline promises a term in a unit that does not define it |
|
|
1001
|
+
| `station-sequence-forward-pointer` | warn | a unit mentions a later unit by number, once per pair |
|
|
1002
|
+
|
|
975
1003
|
### Any station
|
|
976
1004
|
|
|
977
1005
|
| Id | Severity | Meaning |
|
|
@@ -984,11 +1012,13 @@ Each `check` appends one line to the spec's `improvement.ledger`, the same file
|
|
|
984
1012
|
reads:
|
|
985
1013
|
|
|
986
1014
|
```json
|
|
987
|
-
{"at":"2026-09-29T13:21:37.330Z","kind":"check","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","stations":{"form":"pass","terms":"pass","claims":"pass","quotes":"pass","private":"pass","dna":"pass","links":"pass"},"verdict":"one-shot"}
|
|
1015
|
+
{"at":"2026-09-29T13:21:37.330Z","kind":"check","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","stations":{"form":"pass","terms":"pass","claims":"pass","quotes":"pass","private":"pass","dna":"pass","links":"pass","sequence":"skip"},"verdict":"one-shot"}
|
|
988
1016
|
```
|
|
989
1017
|
|
|
990
1018
|
`draft` is the draft's path relative to the spec's folder, however you spelled it, so one draft
|
|
991
|
-
has one history.
|
|
1019
|
+
has one history. A work checked from `form.sequence.files` is keyed by that list as the spec writes
|
|
1020
|
+
it, joined with ", ", so the work keeps one history as parts are added, and its `draft_sha256`
|
|
1021
|
+
covers every file's name and bytes. `draft_sha256` and `spec_sha256` hash the two files' bytes; the files the spec
|
|
992
1022
|
names (materials, the claims ledger, a scope folder) are not hashed, so "changed" below means the
|
|
993
1023
|
draft or the spec. `stations` holds each station's status.
|
|
994
1024
|
|
|
@@ -1009,6 +1039,625 @@ the same draft:
|
|
|
1009
1039
|
|
|
1010
1040
|
A ledger path that leads outside the spec's folder is not written, and `check` prints a warning.
|
|
1011
1041
|
|
|
1042
|
+
## Sequential works
|
|
1043
|
+
|
|
1044
|
+
A course, a primer, a textbook, a book of lessons: a work read in order makes a promise no single
|
|
1045
|
+
piece can check. Lesson 5 is written for someone who has read Lessons 1 to 4 and nothing else, so
|
|
1046
|
+
every word it uses was defined there. That promise lives across pieces, and it breaks one edit at a
|
|
1047
|
+
time: a term moves, a lesson is reordered, a sentence borrows a word from a lesson the reader has
|
|
1048
|
+
not reached. Declare the work a sequence and the `sequence` station checks the promise on every run.
|
|
1049
|
+
|
|
1050
|
+
### Declaring a sequence
|
|
1051
|
+
|
|
1052
|
+
Add `sequence:` to `writing.form`. Every key is optional; this is the course example's, with the
|
|
1053
|
+
four that have defaults written out:
|
|
1054
|
+
|
|
1055
|
+
```yaml
|
|
1056
|
+
sequence:
|
|
1057
|
+
unit: Lesson
|
|
1058
|
+
files:
|
|
1059
|
+
- course/part-*.md
|
|
1060
|
+
sections:
|
|
1061
|
+
- After this lesson you can
|
|
1062
|
+
- New terms
|
|
1063
|
+
- Try this
|
|
1064
|
+
terms_section: New terms
|
|
1065
|
+
outline: course/outline.md
|
|
1066
|
+
teaser: Next,
|
|
1067
|
+
```
|
|
1068
|
+
|
|
1069
|
+
| Key | Default | What it is |
|
|
1070
|
+
|---|---|---|
|
|
1071
|
+
| `unit` | `Lesson` | the word each unit's heading starts with: `## Lesson 3: Title` |
|
|
1072
|
+
| `files` | none | the work's files in reading order, relative to the spec. A `*` in a file name matches within its folder, sorted by number, so `part-2.md` comes before `part-10.md` and a new part is picked up without editing the spec |
|
|
1073
|
+
| `sections` | the three shown | what every unit carries, each as a `**Label:**` line or a heading |
|
|
1074
|
+
| `terms_section` | `New terms` | the section whose list defines the unit's terms |
|
|
1075
|
+
| `outline` | none | an outline whose numbered items promise each unit's terms as `*Terms: a, b.*` |
|
|
1076
|
+
| `knows` | none | words a unit may use before a unit defines them, beside `audience.knows` |
|
|
1077
|
+
| `teaser` | `Next,` | a closing line starting `**<teaser>` names what is coming; it and everything after it in the unit are exempt from the order guard and the pointer warning |
|
|
1078
|
+
|
|
1079
|
+
### How a unit is read
|
|
1080
|
+
|
|
1081
|
+
A unit starts at an ATX heading that begins with `unit` and a number, and runs to the next unit
|
|
1082
|
+
heading, the next heading of its own level or above, or the end of its file. So a part's own
|
|
1083
|
+
heading and introduction belong to no lesson, and neither does a file's YAML frontmatter, which is
|
|
1084
|
+
blanked before anything reads the file. A heading inside fenced code is not a heading.
|
|
1085
|
+
|
|
1086
|
+
The terms section is its label line and the list under it, to the first blank line after the list.
|
|
1087
|
+
Each item defines every bold term before its first `:**`, so `- **Claude Code**, **Codex** and
|
|
1088
|
+
**Claude Cowork:** ...` defines three. A term is compared in lower case, with code and emphasis
|
|
1089
|
+
marks and any parenthetical dropped. An item that says `(from Lesson 3)` reminds the reader of an
|
|
1090
|
+
earlier term and defines nothing.
|
|
1091
|
+
|
|
1092
|
+
A use is a whole-word match, ignoring case, where a hyphen joins a word: "context-aware" does not
|
|
1093
|
+
use "context", and "skills" does not use "skill". Code is not prose: fenced blocks and inline code
|
|
1094
|
+
never count as a use, so `@supersuit/superskill` does not use "supersuit". A unit's own terms
|
|
1095
|
+
section is not a use either. Listing a word in `knows` or `audience.knows` is a decision that the
|
|
1096
|
+
reader already has it, and it is the only way a unit may use a word before the unit that defines it.
|
|
1097
|
+
|
|
1098
|
+
### Checking a sequence
|
|
1099
|
+
|
|
1100
|
+
With `files` listed, `check` needs no `--draft`: the files, joined in reading order, are the draft,
|
|
1101
|
+
and every station reads them together. A finding names the file and its own line.
|
|
1102
|
+
|
|
1103
|
+
```bash
|
|
1104
|
+
npx @supersuit/hyperspec check course.hyperspec.md
|
|
1105
|
+
```
|
|
1106
|
+
|
|
1107
|
+
```
|
|
1108
|
+
form: pass
|
|
1109
|
+
terms: skip (writing.audience.terms is empty or not set)
|
|
1110
|
+
claims: pass
|
|
1111
|
+
quotes: pass
|
|
1112
|
+
private: pass
|
|
1113
|
+
dna: skip (writing.dna.scope_dir is not set)
|
|
1114
|
+
links: pass
|
|
1115
|
+
sequence: pass
|
|
1116
|
+
warn [station-sequence-forward-pointer] Lesson 1 points forward to Lesson 4 (course/part-1.md line 25)
|
|
1117
|
+
fix: Keep it a pointer ("more in Lesson 4"): Lesson 1 must make sense to a reader who has not read Lesson 4.
|
|
1118
|
+
warn [station-sequence-forward-pointer] Lesson 3 points forward to Lesson 4 (course/part-2.md line 24)
|
|
1119
|
+
fix: Keep it a pointer ("more in Lesson 4"): Lesson 3 must make sense to a reader who has not read Lesson 4.
|
|
1120
|
+
verdict: one-shot
|
|
1121
|
+
```
|
|
1122
|
+
|
|
1123
|
+
`--draft` still works on a sequence spec: it names one file, which may hold every lesson under
|
|
1124
|
+
repeated headings. Checked from `files`, `form.length` and `required_parts` apply to the whole work,
|
|
1125
|
+
so write `required_parts` as the part headings. A relative link resolves beside the file that holds
|
|
1126
|
+
it.
|
|
1127
|
+
|
|
1128
|
+
The outline guard checks only the units the draft holds, so a work is checked while it is being
|
|
1129
|
+
written: an outline that promises Lessons 1 to 28 checks a draft of Lessons 1 to 20 without
|
|
1130
|
+
complaint. A pointer to a later unit is a warning, never a failure: "more in Lesson 12" is fine as
|
|
1131
|
+
long as the lesson makes sense without it.
|
|
1132
|
+
|
|
1133
|
+
It cannot tell whether a definition is good, whether a term is used in the sense it was defined in,
|
|
1134
|
+
or whether a quiz tests what the lessons taught. The judges still take one `--draft` file.
|
|
1135
|
+
|
|
1136
|
+
## Judging a draft
|
|
1137
|
+
|
|
1138
|
+
`check` runs the stations that are plain functions of the spec and the draft. The rest of a
|
|
1139
|
+
spec's checks are rubrics: whether the doctor would pass each goal condition, whether a reader
|
|
1140
|
+
would get lost, whether the voice can be told from the writer's own. Those are judgments, and
|
|
1141
|
+
hyperspec calls no model, so it does not make them. It makes them checkable instead. For each
|
|
1142
|
+
judgment station, `judge prepare` writes a packet: the rubric from the spec, fixed
|
|
1143
|
+
instructions, the inputs the judge reads, and the exact shape of the answer. An outside judge
|
|
1144
|
+
(your agent, any model, or a person) fills in a verdict. `judge record` validates it, derives
|
|
1145
|
+
the station's status from it by a fixed rule, and records it in the runs ledger. The judge
|
|
1146
|
+
decides; hyperspec checks that every passage the judge quotes is in the draft, scores
|
|
1147
|
+
every blind test against an answer key the judge never saw, and keeps the record.
|
|
1148
|
+
|
|
1149
|
+
```bash
|
|
1150
|
+
npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
|
|
1151
|
+
```
|
|
1152
|
+
|
|
1153
|
+
```
|
|
1154
|
+
essay/judge/doctor.packet.json
|
|
1155
|
+
essay/judge/lineup.packet.json
|
|
1156
|
+
essay/judge/lineup.key.json
|
|
1157
|
+
essay/judge/reader.packet.json
|
|
1158
|
+
essay/judge/persona.packet.json
|
|
1159
|
+
attribution: skip (the spec is not fiction; attribution applies only with fiction: true)
|
|
1160
|
+
knowledge: skip (the spec is not fiction; knowledge applies only with fiction: true)
|
|
1161
|
+
```
|
|
1162
|
+
|
|
1163
|
+
Like `check`, it lints the spec first: a spec that fails lint, or is blocked on an open
|
|
1164
|
+
decision, gets no packet, and `prepare` exits with lint's own code. Then it writes one
|
|
1165
|
+
`<station>.packet.json` for each station that applies, in a fixed order (`doctor, lineup,
|
|
1166
|
+
reader, persona, attribution, knowledge`), and prints a `skip` line with the reason for each
|
|
1167
|
+
that does not. `--only doctor,reader` prepares just those. The `--out` folder must already
|
|
1168
|
+
exist. `prepare` refuses to overwrite any file it would write, naming every one, and then writes
|
|
1169
|
+
nothing; `--force` replaces them. The worked examples ship the packets this writes, so add
|
|
1170
|
+
`--force` to write them again there. The same spec and draft, with the same goldens and claims
|
|
1171
|
+
ledger, always give byte-identical packets. A station that throws while building its packet
|
|
1172
|
+
prints `no packet` with a `judge-<name>-crashed` finding; the other packets are still written,
|
|
1173
|
+
and `prepare` exits 1.
|
|
1174
|
+
|
|
1175
|
+
**Hand the judge only the `*.packet.json` files, never the `--out` folder.** Two stations are
|
|
1176
|
+
blind tests with a right answer. `lineup` writes which candidate is the draft's to
|
|
1177
|
+
`lineup.key.json`, and `attribution` writes each line's true speaker to
|
|
1178
|
+
`attribution.key.json`, both beside the packets, for a person to read. A judge who can see them
|
|
1179
|
+
is not judging. `record` never reads either file as truth: it builds the key again from the spec
|
|
1180
|
+
and the draft, so an edited key changes nothing.
|
|
1181
|
+
|
|
1182
|
+
**Give the two blind packets to a judge in a fresh context with no access to the draft.** Every
|
|
1183
|
+
packet names its spec and its draft by path, and a judge that can open files can open those: the
|
|
1184
|
+
draft's own text shows which lineup passage is the draft's, and its speech tags give every
|
|
1185
|
+
attribution answer. Both stations' instructions tell the judge to decide from the packet's inputs
|
|
1186
|
+
alone and open no file the packet names, and a judge with file access can still ignore that, so
|
|
1187
|
+
the instruction is not a guarantee. Paste the packet into a new conversation, or hand it to a
|
|
1188
|
+
person who has not read the draft, rather than to an agent working in the folder the draft is in.
|
|
1189
|
+
The other four packets carry the draft in their inputs and hide nothing, which is why each blind
|
|
1190
|
+
packet goes to its own context, apart from the other packets as well as from the folder: a judge
|
|
1191
|
+
that has read the doctor, reader or persona packet has read the draft.
|
|
1192
|
+
|
|
1193
|
+
Exit codes for `prepare`: **0** written; **1** a station crashed; **2** usage: no spec path, no
|
|
1194
|
+
`--draft`, no `--out`, an `--out` that is missing or not a folder, a draft that cannot be read, a
|
|
1195
|
+
spec without `profile: writing`, an `--only` that names no known judge, or a file that exists
|
|
1196
|
+
without `--force`; and lint's own **1** or **3** when the spec is not ready. `--json` prints the
|
|
1197
|
+
result, and a usage error prints `{ "spec", "draft", "out", "error" }`.
|
|
1198
|
+
|
|
1199
|
+
### The packet
|
|
1200
|
+
|
|
1201
|
+
Every packet is JSON with these fields, in this order:
|
|
1202
|
+
|
|
1203
|
+
| Field | Holds |
|
|
1204
|
+
|---|---|
|
|
1205
|
+
| `hyperspec_judge` | the packet format's version, `"0.1"` |
|
|
1206
|
+
| `station` | the station's name |
|
|
1207
|
+
| `spec`, `draft` | the paths exactly as `prepare` was given them |
|
|
1208
|
+
| `spec_sha256`, `draft_sha256` | the SHA-256 of each file's bytes |
|
|
1209
|
+
| `rubric` | the `check.rubric` of the block the station reads, verbatim |
|
|
1210
|
+
| `instructions` | fixed text for the station, the same for every spec |
|
|
1211
|
+
| `inputs` | what the judge reads; each station below lists its own |
|
|
1212
|
+
| `verdict_schema` | the verdict's exact shape, as a small JSON Schema |
|
|
1213
|
+
|
|
1214
|
+
A station that needs a rubric skips when its block has none, and a station whose block is
|
|
1215
|
+
deferred to a decision skips too. The paths are kept as given so the packet is the same on every
|
|
1216
|
+
machine, which means **`record` must run in the folder `prepare` ran in**. Run from anywhere else,
|
|
1217
|
+
it cannot find the spec and exits 2.
|
|
1218
|
+
|
|
1219
|
+
### The evidence rule
|
|
1220
|
+
|
|
1221
|
+
Every verdict field that cites the draft (the doctor's `evidence`, the reader's `lost_at` and
|
|
1222
|
+
`stopped_at`, the persona's `breaks`, the knowledge `leaks`) is a span copied from the draft. A
|
|
1223
|
+
span counts only when:
|
|
1224
|
+
|
|
1225
|
+
- it has at least three words, where a word is a run of letters and digits (so "It's" is two);
|
|
1226
|
+
- it is in the draft after both are normalized: every run of whitespace, line breaks included,
|
|
1227
|
+
becomes one space, and curly, low and angle quotation marks and apostrophes
|
|
1228
|
+
(`‘ ’ ‚ ‛ “ ” „ ‟ ‹ › « »`) become straight ones. Primes (`′ ″`) are not quotation marks and
|
|
1229
|
+
are left alone, case is kept, and nothing else changes, so a changed word is not a quotation;
|
|
1230
|
+
- it matches whole words: a span that starts or ends with a letter or digit may not start or end
|
|
1231
|
+
inside a word of the draft.
|
|
1232
|
+
|
|
1233
|
+
A span that fails makes the verdict invalid (`judge-evidence-missing`, `-too-short` or
|
|
1234
|
+
`-not-found`), and nothing is recorded. The lineup and attribution verdicts carry no evidence:
|
|
1235
|
+
they are blind tests, scored against their keys.
|
|
1236
|
+
|
|
1237
|
+
### Recording a verdict
|
|
1238
|
+
|
|
1239
|
+
```bash
|
|
1240
|
+
npx @supersuit/hyperspec judge record essay/judge/lineup.packet.json --verdict essay/sample-verdicts/lineup.verdict.json
|
|
1241
|
+
```
|
|
1242
|
+
|
|
1243
|
+
```
|
|
1244
|
+
lineup: fail
|
|
1245
|
+
fail [judge-lineup-picked] the judge picked the draft's passage (D) out of 4 candidates with confidence 0.6: D is the only passage that rests on a survey figure, and it opens by pointing at something outside itself (the survey); A, B and C each turn one claim into an instruction in the second person, with no numbers. (line 32)
|
|
1246
|
+
fix: Revise this passage toward the goldens' voice, where the reason points; if the goldens do not cover this kind of passage, add one that does. Then prepare and judge again.
|
|
1247
|
+
verdict: not-improved (failing stations: lineup)
|
|
1248
|
+
```
|
|
1249
|
+
|
|
1250
|
+
`record` trusts nothing in the packet file. First it hashes the spec and the draft again; if
|
|
1251
|
+
either no longer matches the hash the packet recorded, the verdict is stale (`judge-stale`,
|
|
1252
|
+
"the draft does not match the hash the packet recorded: it changed since prepare, or the packet
|
|
1253
|
+
was edited"). Then it rebuilds the packet from the spec, the draft and the station's other
|
|
1254
|
+
inputs on disk (the DNA scope, its `scope.md` and goldens, for `lineup`; the claims ledger for
|
|
1255
|
+
`persona`) and
|
|
1256
|
+
requires the file to be exactly those bytes:
|
|
1257
|
+
|
|
1258
|
+
- a packet whose inputs no longer match what those other files produce is stale, and the message
|
|
1259
|
+
names them. Nothing on disk can tell a changed golden from a hand-edited packet, so it says
|
|
1260
|
+
"changed since the packet was prepared, or the packet was edited" and claims neither;
|
|
1261
|
+
- a station that no longer applies because of those files (the claims ledger deleted, say) is
|
|
1262
|
+
stale in the same way, and says why it no longer applies;
|
|
1263
|
+
- anything else (an edited condition, a reformatted file, a hash made to match a changed draft, a
|
|
1264
|
+
station that no longer applies for another reason) is `judge-packet-altered`.
|
|
1265
|
+
|
|
1266
|
+
Either way nothing is recorded. A stale packet prints `<station>: stale packet, nothing
|
|
1267
|
+
recorded`, as `learn record` does, and with `--json` both commands mark it `"invalid": true,
|
|
1268
|
+
"stale": true`, so a script can test `invalid` alone. The fix is to run `judge prepare` again
|
|
1269
|
+
with `--force` and judge the new packet, or, for a station that no longer applies, to restore
|
|
1270
|
+
the file it reads. Only then is the verdict read: it must be JSON (one leading
|
|
1271
|
+
byte order mark is ignored), in the shape the packet gives, with every evidence span found.
|
|
1272
|
+
Every problem is named, and an invalid verdict records nothing. The validator ignores fields it
|
|
1273
|
+
does not know, so a verdict can carry a note of its own; the worked examples' sample verdicts
|
|
1274
|
+
each carry one, `sample`, saying what they are.
|
|
1275
|
+
|
|
1276
|
+
A valid verdict gives the station's status, its findings (a warning prints under the status and
|
|
1277
|
+
never fails it), and for `attribution` one summary line. Then one line goes to the runs ledger.
|
|
1278
|
+
|
|
1279
|
+
Exit codes for `record`: **0** the station passed; **1** it failed, or the verdict is invalid,
|
|
1280
|
+
or the packet is stale or altered; **2** usage: no packet path, no `--verdict`, a packet or verdict
|
|
1281
|
+
file that cannot be read, a file that is not a judge packet or names an unknown judge, or a
|
|
1282
|
+
packet whose spec (which must carry `profile: writing`) or draft cannot be read. `--json` prints
|
|
1283
|
+
the result, and a usage error prints `{ "packet", "verdict", "error" }`.
|
|
1284
|
+
|
|
1285
|
+
In the tables below, **fail** fails the station, **warn** is printed and never fails it,
|
|
1286
|
+
**invalid** refuses the verdict, and **stale** refuses the packet; an invalid or stale verdict
|
|
1287
|
+
records nothing.
|
|
1288
|
+
|
|
1289
|
+
### doctor
|
|
1290
|
+
|
|
1291
|
+
Grades the draft against every goal condition, and asks whether the reader would take the next
|
|
1292
|
+
step now. Applies when `writing.goal` is written and its check has a rubric. Inputs: `goal`
|
|
1293
|
+
(`from`, `to`, `next_if_worked`, and `change` with its `kind` and `text`), `conditions` (each
|
|
1294
|
+
condition id with its requirement's `text` and `fails_when`) and the `draft`. The verdict is
|
|
1295
|
+
`{ conditions: [{ id, pass, evidence, note }], would_take_next_step, evidence }`: every condition
|
|
1296
|
+
id exactly once, `pass` and `would_take_next_step` true or false, a `note` on every condition
|
|
1297
|
+
(one that fails must say why), and the top-level `evidence` for the passage that decided the next
|
|
1298
|
+
step. It passes when every condition passes and the reader would take the next step.
|
|
1299
|
+
|
|
1300
|
+
| Id | Kind | Meaning |
|
|
1301
|
+
|---|---|---|
|
|
1302
|
+
| `judge-doctor-condition` | fail | a condition fails; gives the judge's note, at its evidence's line |
|
|
1303
|
+
| `judge-doctor-next-step` | fail | the reader would not take `goal.next_if_worked` now |
|
|
1304
|
+
| `judge-doctor-condition-missing` | invalid | a condition in the packet has no entry |
|
|
1305
|
+
| `judge-doctor-condition-unknown` | invalid | the verdict grades an id that is not a condition in the packet |
|
|
1306
|
+
| `judge-doctor-condition-duplicate` | invalid | a condition is graded more than once |
|
|
1307
|
+
|
|
1308
|
+
### lineup
|
|
1309
|
+
|
|
1310
|
+
A blind test of voice. Applies when `writing.dna` names a `scope_dir` whose goldens can be read
|
|
1311
|
+
and its check has a rubric, at least one golden has a prose paragraph, and the draft has a prose
|
|
1312
|
+
paragraph that is not already a golden's. A draft that is one paragraph and nothing else is
|
|
1313
|
+
skipped: its candidate would be the whole file, and the packet's `draft_sha256` would identify it. A prose paragraph is a run of non-blank lines that are
|
|
1314
|
+
all plain text: headings and code fences end one, and a block holding a list item, a quotation, a
|
|
1315
|
+
table row, a thematic break or HTML is left out whole, as is indented code.
|
|
1316
|
+
|
|
1317
|
+
Every candidate is one prose paragraph reflowed onto a single line, so none can be told by its
|
|
1318
|
+
formatting. The target length is the median, in characters, of every prose paragraph of every
|
|
1319
|
+
golden in the scope. The first three goldens by file name that have one each give their
|
|
1320
|
+
paragraph closest to that length, and the draft gives its paragraph closest to it, skipping any
|
|
1321
|
+
that is word for word a golden's; ties go to the earliest. The candidates are shuffled with a
|
|
1322
|
+
seed derived from the draft's full text (the SHA-256 of `hyperspec lineup seed`, a line break, and
|
|
1323
|
+
the text), which the packet does not carry, so the packet cannot reveal the order: not even its
|
|
1324
|
+
`draft_sha256`, which is a different hash. The same draft always gets the same labels, labeled A
|
|
1325
|
+
to D. Inputs: `scope` (the scope's `writer`, `form`, `audience` and `purpose`) and `candidates`,
|
|
1326
|
+
each `{ label, text }`; no path and no source. `lineup.key.json` records the draft's label, the
|
|
1327
|
+
draft line its paragraph starts on, and where every candidate came from.
|
|
1328
|
+
|
|
1329
|
+
The verdict is `{ pick, confidence, reason }`: a label, a number from 0 to 1, and what decided it.
|
|
1330
|
+
It passes when the pick is not the draft's passage: the judge could not tell. A judge who picks at
|
|
1331
|
+
random also passes, three times in four with four candidates, so one lineup is weak evidence of a
|
|
1332
|
+
voice. A failing lineup is the stronger signal, and its reason says where to look.
|
|
1333
|
+
|
|
1334
|
+
| Id | Kind | Meaning |
|
|
1335
|
+
|---|---|---|
|
|
1336
|
+
| `judge-lineup-picked` | fail | the judge picked the draft's passage; gives its confidence and reason, at the paragraph's line |
|
|
1337
|
+
| `judge-lineup-pick-unknown` | invalid | the pick is not one of the lineup's labels |
|
|
1338
|
+
|
|
1339
|
+
### reader
|
|
1340
|
+
|
|
1341
|
+
Reads the draft as the audience block's reader. Applies when `writing.audience` is written and its
|
|
1342
|
+
check has a rubric. Inputs: `audience` (`who`, `funnel_now`, `knows`, `terms`, `believes_now`,
|
|
1343
|
+
`wants`, `reads_on`, `reader`) and the `draft`. The verdict is `{ lost_at: [{ evidence, why }],
|
|
1344
|
+
stopped_at, would_take_next_step, next_step }`, where `stopped_at` is `{ evidence, why }` or
|
|
1345
|
+
`null` when the reader read to the end, and must be present either way. The reader names its own
|
|
1346
|
+
next step: it is not shown `goal.next_if_worked`, so a reader that would act and a doctor that
|
|
1347
|
+
says the reader would not take the spec's next step can both be right. In the essay example they
|
|
1348
|
+
agree: both name the card. It passes when the reader read to the end and would take its next step
|
|
1349
|
+
now.
|
|
1350
|
+
|
|
1351
|
+
| Id | Kind | Meaning |
|
|
1352
|
+
|---|---|---|
|
|
1353
|
+
| `judge-reader-lost` | warn | the reader got lost here, and why |
|
|
1354
|
+
| `judge-reader-stopped` | fail | the reader stopped reading here, and why |
|
|
1355
|
+
| `judge-reader-next-step` | fail | the reader would not take its next step now |
|
|
1356
|
+
|
|
1357
|
+
### persona
|
|
1358
|
+
|
|
1359
|
+
Reads the draft as the persona block's speaker, against the claims ledger. Applies when
|
|
1360
|
+
`writing.persona` is written and its check has a rubric, and the claims ledger, if
|
|
1361
|
+
`sources.ledger` names one, can be read. Inputs: `persona` (`identity`, `stance`, `may_assert`,
|
|
1362
|
+
`will_not_say`), `claims` (the text of every well-formed line of the claims ledger, in order) and
|
|
1363
|
+
the `draft`. With no ledger declared, `claims` is `null`: facts cannot be checked against
|
|
1364
|
+
sources, so the instructions say not to report one, the schema leaves that kind out, and a break
|
|
1365
|
+
of that kind is invalid. The verdict is `{ breaks: [{ evidence, kind, why }] }`, where `kind` is
|
|
1366
|
+
`stance` (the voice leaves its stance), `assertion` (it asserts something outside `may_assert`),
|
|
1367
|
+
`will_not_say` or `unsourced_fact` (a fact no claim holds). It passes when there is no break.
|
|
1368
|
+
|
|
1369
|
+
| Id | Kind | Meaning |
|
|
1370
|
+
|---|---|---|
|
|
1371
|
+
| `judge-persona-break` | fail | the voice breaks here; the message names the kind and gives the judge's why |
|
|
1372
|
+
| `judge-persona-kind-unknown` | invalid | a break's kind is not one of the four |
|
|
1373
|
+
| `judge-persona-no-ledger` | invalid | a break is `unsourced_fact`, and the spec declares no claims ledger |
|
|
1374
|
+
|
|
1375
|
+
### attribution
|
|
1376
|
+
|
|
1377
|
+
Fiction only: a blind test of whether the characters' voices can be told apart. The packet holds
|
|
1378
|
+
the draft's dialogue lines with the narration and the speaker removed, and each speaking
|
|
1379
|
+
character's `speech` block, `golden_lines` and `rejected_lines`; it carries no draft. The judge
|
|
1380
|
+
names a speaker for every line, and `record` scores the answers against the true speakers, which
|
|
1381
|
+
only `attribution.key.json` holds. Applies when `fiction: true`, at least two characters have a
|
|
1382
|
+
speech block and one of them has a check rubric, and the draft has attributable lines from at
|
|
1383
|
+
least two speakers.
|
|
1384
|
+
|
|
1385
|
+
A dialogue line is a double-quoted span, straight or curly, inside one paragraph, outside code. A
|
|
1386
|
+
quote split by a speech tag (`"Twenty minutes," Ines said, "then we fold it."`) is one line when
|
|
1387
|
+
its first part ends in a comma and the narration between the parts is exactly one tag and its
|
|
1388
|
+
comma, nothing else. Narration that holds anything more joins nothing: in `"Leave it there," Ines
|
|
1389
|
+
said, and Theo muttered, "No chance at all."` the first part is Ines's, and the second is left
|
|
1390
|
+
out, since a second speaker brought in by a beat, a pronoun or a verb off the list cannot be read
|
|
1391
|
+
mechanically. The second part is left out whatever ends that narration, even another tag (`Ines
|
|
1392
|
+
said, and Theo said,`), and so is the second part of `"Twenty minutes," Ines said, wiping her
|
|
1393
|
+
hands, "then we fold it."`. A line's speaker comes only from a speech tag:
|
|
1394
|
+
narration in the same paragraph directly after the closing mark (`"...," Ines said`) or directly
|
|
1395
|
+
before the opening mark, ending in a comma or colon (`Ines said, "..."`). A tag is a subject and
|
|
1396
|
+
one of the verbs said, asked, told, replied, called, whispered, shouted, answered, added and went
|
|
1397
|
+
on, or their present tense (says, asks, goes on). The subject is:
|
|
1398
|
+
|
|
1399
|
+
- a character's id, or its `name` if it has one: whole words, any case, and an id's words may be
|
|
1400
|
+
joined by a space, a hyphen or an underscore, so `old-man` is named by "old man";
|
|
1401
|
+
- "I", when `persona.identity` is `character:<id>`: the narrator speaks;
|
|
1402
|
+
- "she" or "he", when exactly two characters have speech blocks and one of them narrates: the
|
|
1403
|
+
other one speaks. After a quote the pronoun must be lower case.
|
|
1404
|
+
|
|
1405
|
+
The verb may come first ("said Ines") only for said, replied, whispered, shouted and went on (and
|
|
1406
|
+
their present tense), since "Ines told Theo" names the person spoken to. Anything else leaves the
|
|
1407
|
+
line out, and every doubt does: no tag, a possessive ("Ines's") or an action beat ("Theo nodded"),
|
|
1408
|
+
a subject that is not a character with a speech block, tags that name two different speakers, or a
|
|
1409
|
+
line of three words or more that repeats, or is part of, a golden or rejected line, which would
|
|
1410
|
+
give its speaker away. So does the second quote in
|
|
1411
|
+
`"Not a bakery," she said. "They want the room."`, since no tag sits against it. The key is never wrong, and recall pays for it: the worked
|
|
1412
|
+
story has 36 dialogue lines, and 19 of them are in the test, 10 by Ines and 9 by Theo. The packet
|
|
1413
|
+
counts the lines left out, and the key lists each with its reason.
|
|
1414
|
+
|
|
1415
|
+
Inputs: `characters`, `lines` (each `{ id, text }`, numbered L1 up in draft order, whitespace
|
|
1416
|
+
reflowed) and `excluded`, the count left out. The verdict is `{ lines: [{ id, speaker }] }`,
|
|
1417
|
+
every line id exactly once, a speaker by id or name, in any case. Accuracy is taken per speaker
|
|
1418
|
+
and averaged, so naming one character for every line cannot pass: the station passes when that
|
|
1419
|
+
mean is 80 percent or more, compared exactly, with percentages rounded down. This is the worked
|
|
1420
|
+
story's sample:
|
|
1421
|
+
|
|
1422
|
+
```
|
|
1423
|
+
attribution: pass
|
|
1424
|
+
accuracy per speaker: ines 10/10 (100%), theo 9/9 (100%); mean 100%, passing at 80%; 17 dialogue lines left out (8 repeating a golden or rejected line, 9 with no speech tag)
|
|
1425
|
+
verdict: one-shot
|
|
1426
|
+
```
|
|
1427
|
+
|
|
1428
|
+
| Id | Kind | Meaning |
|
|
1429
|
+
|---|---|---|
|
|
1430
|
+
| `judge-attribution-accuracy` | fail | the mean accuracy over speakers is under 80%; gives each speaker's |
|
|
1431
|
+
| `judge-attribution-miss` | warn | a line was given to the wrong character, at its draft line |
|
|
1432
|
+
| `judge-attribution-speaker-unknown` | invalid | a speaker is not a character in the packet |
|
|
1433
|
+
| `judge-attribution-line-missing` | invalid | a line in the packet has no answer |
|
|
1434
|
+
| `judge-attribution-line-unknown` | invalid | the verdict answers an id that is not a line in the packet |
|
|
1435
|
+
| `judge-attribution-line-duplicate` | invalid | a line is answered more than once |
|
|
1436
|
+
|
|
1437
|
+
### knowledge
|
|
1438
|
+
|
|
1439
|
+
Fiction only: whether anyone knows something before their timeline gives it to them. Applies when
|
|
1440
|
+
`fiction: true` and at least one character has a knowledge timeline (an entry with both `by` and
|
|
1441
|
+
`knows`) and one of those characters has a check rubric. Inputs: `characters` (each with its
|
|
1442
|
+
`knowledge` entries) and the `draft`. The verdict is `{ leaks: [{ character, evidence,
|
|
1443
|
+
knows_too_early }] }`, and it passes when there is no leak.
|
|
1444
|
+
|
|
1445
|
+
hyperspec does not read the draft's structure: `by` goes into the packet as the spec wrote it,
|
|
1446
|
+
and the judge maps it onto the draft. The worked story's timelines say `scene-1` to `scene-4`,
|
|
1447
|
+
and its draft's four headings are times, 3:40 to 6:55, in the scene list's order, so a judge maps
|
|
1448
|
+
them by order. Write `by` as something a reader of the draft can find.
|
|
1449
|
+
|
|
1450
|
+
| Id | Kind | Meaning |
|
|
1451
|
+
|---|---|---|
|
|
1452
|
+
| `judge-knowledge-leak` | fail | a character knows something too early, at its evidence's line |
|
|
1453
|
+
| `judge-knowledge-character-unknown` | invalid | a leak names a character with no timeline in the packet |
|
|
1454
|
+
|
|
1455
|
+
### Any judgment station
|
|
1456
|
+
|
|
1457
|
+
| Id | Kind | Meaning |
|
|
1458
|
+
|---|---|---|
|
|
1459
|
+
| `judge-verdict-not-json` | invalid | the verdict file is not JSON |
|
|
1460
|
+
| `judge-verdict-shape` | invalid | a field is missing, of the wrong type or out of range, or a required note, reason or why is empty |
|
|
1461
|
+
| `judge-evidence-missing` | invalid | an evidence field is empty or not a string |
|
|
1462
|
+
| `judge-evidence-too-short` | invalid | an evidence span has fewer than three words |
|
|
1463
|
+
| `judge-evidence-not-found` | invalid | an evidence span is not in the draft, word for word |
|
|
1464
|
+
| `judge-stale` | stale | the spec, the draft, or a file the station reads changed since the packet was prepared, or the packet was edited |
|
|
1465
|
+
| `judge-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
|
|
1466
|
+
| `judge-<name>-crashed` | invalid | the station threw; at `prepare` its packet is not written, at `record` nothing is recorded |
|
|
1467
|
+
|
|
1468
|
+
### Judge lines in the runs ledger
|
|
1469
|
+
|
|
1470
|
+
Each recorded verdict appends one line to the spec's `improvement.ledger`:
|
|
1471
|
+
|
|
1472
|
+
```json
|
|
1473
|
+
{"at":"2026-09-29T16:27:36.883Z","kind":"judge","station":"lineup","draft":"essay/draft.md","draft_sha256":"<sha256 of the draft>","spec_sha256":"<sha256 of the spec>","packet_sha256":"<sha256 of the packet>","inputs_sha256":"<sha256 of its inputs>","status":"fail","verdict":"not-improved","reason":"failing stations: lineup"}
|
|
1474
|
+
```
|
|
1475
|
+
|
|
1476
|
+
`draft` is relative to the spec's folder, as in a check line. `packet_sha256` is the SHA-256 of
|
|
1477
|
+
the packet the judge was shown, taken with its two paths written as the ledger writes them
|
|
1478
|
+
(relative to the spec's folder), so the same packet prepared from another folder, or as
|
|
1479
|
+
`./draft.md`, hashes the same. `inputs_sha256` is the SHA-256 of the packet's `inputs` alone. The
|
|
1480
|
+
verdict follows the same rules as a check line, compared with
|
|
1481
|
+
the most recent earlier judge line for the same station and the same draft, with three more:
|
|
1482
|
+
|
|
1483
|
+
- **What changed** is judged by the packet. The draft or the spec is named when its bytes
|
|
1484
|
+
changed. When neither did and the packet's inputs did, the files the station reads besides them
|
|
1485
|
+
are named: `the DNA scope (<scope_dir>: scope.md and goldens)` for `lineup`, `the claims ledger
|
|
1486
|
+
(<path>)` for `persona`. When only the rest of the packet changed, which a later hyperspec
|
|
1487
|
+
release can do by rewording a station's instructions, the reason says `the packet's fixed text
|
|
1488
|
+
(hyperspec's instructions or format) changed`. So adding the golden a failing lineup asked for, or the claims a
|
|
1489
|
+
failing persona asked for, and passing on the new packet is `improved`, "the DNA scope
|
|
1490
|
+
(dna/essay-new-managers-teach: scope.md and goldens) changed; stations now pass: lineup". An
|
|
1491
|
+
improved judge line always names what changed, "draft changed; stations now pass: doctor".
|
|
1492
|
+
- **one-shot** also needs these draft bytes never to have been judged by this station before,
|
|
1493
|
+
under any name. A copy or a rename of a judged draft gets `not-improved`, with the reason
|
|
1494
|
+
"these draft bytes were judged before as <path>: <status>".
|
|
1495
|
+
- **improved** also needs the packet to have changed since the failing line. A judge can answer
|
|
1496
|
+
differently about an identical packet, and that is not the work improving: the verdict is
|
|
1497
|
+
`not-improved`, "the verdict changed; nothing the judge was shown changed".
|
|
1498
|
+
|
|
1499
|
+
Otherwise the reasons are check's, with `judgment` where check says `check`: `failing stations:
|
|
1500
|
+
doctor`, `no change since the last passing judgment`, `no change since the last judgment; still
|
|
1501
|
+
failing: doctor`, or `draft changed; every station still passes`. Check and judge each read only
|
|
1502
|
+
their own lines, and learn reads none, so judging a draft never changes what `check` says about
|
|
1503
|
+
it, or the reverse. Judge lines keep lint's test 9 passing, as check lines do.
|
|
1504
|
+
|
|
1505
|
+
### The worked examples
|
|
1506
|
+
|
|
1507
|
+
Both examples ship the packets `prepare` writes for their drafts, in `essay/judge/` and
|
|
1508
|
+
`story/judge/`, with no key beside them, and one sample verdict per packet in
|
|
1509
|
+
`essay/sample-verdicts/` and `story/sample-verdicts/`. The samples were filled in by hand, as
|
|
1510
|
+
one careful judge would, to show the shape and what `record` does with it; each says so in its
|
|
1511
|
+
`sample` field. Another judge may answer differently. A test records every sample on a fresh
|
|
1512
|
+
copy on every release.
|
|
1513
|
+
|
|
1514
|
+
| Example | Station | Sample | What the judge found |
|
|
1515
|
+
|---|---|---|---|
|
|
1516
|
+
| essay | doctor | pass | every condition holds, and the reader would write the three questions on a card, the goal's next step and the draft's close |
|
|
1517
|
+
| essay | lineup | fail | the draft's passage is the only one resting on a survey figure; the goldens hold no numbers, so they do not cover this kind of passage |
|
|
1518
|
+
| essay | reader | pass | read to the end, lost nowhere; the reader's own next step is the card |
|
|
1519
|
+
| essay | persona | pass | the mentor stance holds, and every figure is in the claims ledger |
|
|
1520
|
+
| story | doctor | pass | every condition holds, and a reader would look for the author's other stories |
|
|
1521
|
+
| story | reader | pass | read to the end; lost for a moment at "proving cabinet" and "peel", two warnings |
|
|
1522
|
+
| story | persona | fail | three process details (the deck oven's heat-up time, the rolls' bake time, how the starter is fed) are in no claim, and the rubric allows none outside the ledger |
|
|
1523
|
+
| story | attribution | pass | all 19 lines named right by voice alone |
|
|
1524
|
+
| story | knowledge | pass | neither character knows anything early |
|
|
1525
|
+
|
|
1526
|
+
Both examples pass every station of `check`. Each failure here is something no deterministic
|
|
1527
|
+
station can see.
|
|
1528
|
+
|
|
1529
|
+
## Learning from edits
|
|
1530
|
+
|
|
1531
|
+
A factory writes a first draft, and a person edits it into the draft they approve. Every edit is
|
|
1532
|
+
something the spec did not say, or did not say well enough. `learn prepare` diffs the two drafts
|
|
1533
|
+
sentence by sentence and writes a packet of the edits; an outside judge names, for each edit, the
|
|
1534
|
+
one block of the spec that would have prevented it; `learn record` validates that verdict, counts
|
|
1535
|
+
the edits by block, and names one next move for the block with the most. It never edits the spec:
|
|
1536
|
+
the move is a suggestion for whoever keeps it.
|
|
1537
|
+
|
|
1538
|
+
```bash
|
|
1539
|
+
npx @supersuit/hyperspec learn prepare essay.hyperspec.md --first essay/learn/first-draft.md --approved essay/draft.md --out essay/learn
|
|
1540
|
+
npx @supersuit/hyperspec learn record essay/learn/learn.packet.json --verdict essay/sample-verdicts/learn.verdict.json
|
|
1541
|
+
```
|
|
1542
|
+
|
|
1543
|
+
```
|
|
1544
|
+
essay/learn/learn.packet.json
|
|
1545
|
+
5 edits over 8 sentences: 3 replaced, 2 deleted
|
|
1546
|
+
```
|
|
1547
|
+
|
|
1548
|
+
```
|
|
1549
|
+
learn: 5 edits classified
|
|
1550
|
+
dna 2 edits, 4 sentences
|
|
1551
|
+
materials 1 edit, 2 sentences
|
|
1552
|
+
persona 1 edit, 1 sentence
|
|
1553
|
+
none 1 edit, 1 sentence
|
|
1554
|
+
next move: dna, 4 of 8 sentences (2 of 5 edits): add a golden or a style rule
|
|
1555
|
+
verdict: not-improved (edits by block: dna 2 edits (4 sentences), materials 1 edit (2 sentences), persona 1 edit (1 sentence), none 1 edit (1 sentence); not yet applied to the spec)
|
|
1556
|
+
```
|
|
1557
|
+
|
|
1558
|
+
`prepare` lints the spec first, exactly as `judge prepare` does, then writes
|
|
1559
|
+
`<out>/learn.packet.json` and prints how many edits it found, or `no edits: the first draft and the
|
|
1560
|
+
approved draft match; nothing to learn`. The rules on `--out`, `--force` and the paths are the
|
|
1561
|
+
judge's: the folder must exist, an existing packet needs `--force`, and `record` runs in the
|
|
1562
|
+
folder `prepare` ran in. The essay example ships this pair and its packet in `essay/learn/`, so
|
|
1563
|
+
add `--force` to write the packet again there.
|
|
1564
|
+
|
|
1565
|
+
Exit codes: `prepare` **0** written, **2** usage (no spec path, `--first`, `--approved` or
|
|
1566
|
+
`--out`, a folder that is missing, a draft that cannot be read, a spec without `profile: writing`,
|
|
1567
|
+
a packet that exists without `--force`, or drafts too large to diff), and lint's own **1** or
|
|
1568
|
+
**3** when the spec is not ready; `record` **0** recorded, **1** the verdict is invalid or the
|
|
1569
|
+
packet stale or altered (nothing recorded), **2** usage.
|
|
1570
|
+
|
|
1571
|
+
### Sentence units
|
|
1572
|
+
|
|
1573
|
+
The drafts are compared one unit at a time, and a unit is one sentence, one heading, or one list
|
|
1574
|
+
item:
|
|
1575
|
+
|
|
1576
|
+
- A blank line always ends a unit. When the first line of a paragraph is a Markdown heading, that
|
|
1577
|
+
line is a unit of its own; a later line that starts with `#` is a hard wrap in the text, and is
|
|
1578
|
+
read as text.
|
|
1579
|
+
- Other lines are split into sentences at `.`, `!` or `?` followed by whitespace, never inside a
|
|
1580
|
+
quotation that opens and closes on one line (straight or curly), so a quoted passage of several
|
|
1581
|
+
sentences on one line is one unit. Each list item starts a unit.
|
|
1582
|
+
- A unit that ends in a common abbreviation (Mr., Mrs., Ms., Dr., Prof., Sr., Jr., St., vs., cf.,
|
|
1583
|
+
e.g., i.e.) or a single capital initial ("J.") runs on into the next, and so does a unit
|
|
1584
|
+
followed by one that starts with a lower-case letter ("the U.S. economy").
|
|
1585
|
+
|
|
1586
|
+
### The edits
|
|
1587
|
+
|
|
1588
|
+
Units are compared with their whitespace collapsed and matched by a longest common subsequence.
|
|
1589
|
+
Every run of unmatched units is one edit, a hunk: `deleted` (only in the first draft), `inserted`
|
|
1590
|
+
(only in the approved one) or `replaced`. A hunk never crosses a paragraph break or a heading, so
|
|
1591
|
+
a paragraph rewritten from end to end is one hunk and a change in two paragraphs is two. A hunk's
|
|
1592
|
+
texts are the drafts' own words from its first unit to its last, and its `sentences` is the
|
|
1593
|
+
larger of its two unit counts. A run that reads the same once whitespace is collapsed is a
|
|
1594
|
+
reflow, not an edit, so rewrapping lines, or joining or splitting paragraphs without changing a
|
|
1595
|
+
word, makes no hunk. The drafts' common start and end are set aside first; if what is left would
|
|
1596
|
+
need more than 10,000,000 comparisons, `prepare` stops with a usage error naming both sentence
|
|
1597
|
+
counts. Learn from a chapter or a scene at a time.
|
|
1598
|
+
|
|
1599
|
+
### The learn packet
|
|
1600
|
+
|
|
1601
|
+
`{ "hyperspec_learn": "0.1", "spec", "spec_sha256", "first", "first_sha256", "approved",
|
|
1602
|
+
"approved_sha256", "blocks", "instructions", "hunks", "verdict_schema" }`. `blocks` lists the
|
|
1603
|
+
writing blocks the spec has written, in schema order, then `none`; a deferred block is not
|
|
1604
|
+
listed. Each hunk is `{ id, kind, first, approved, sentences }`, with ids E1 up, and `null` on
|
|
1605
|
+
the side that has no text. The packet carries no draft beyond its hunks: the instructions tell
|
|
1606
|
+
the judge to read the spec at its path. A learn packet has no answer key.
|
|
1607
|
+
|
|
1608
|
+
The verdict is `{ edits: [{ id, block, why }] }`: every hunk id exactly once, a `block` from the
|
|
1609
|
+
packet's `blocks`, and a `why` saying what that block should have said. `none` means no block of
|
|
1610
|
+
the spec could have prevented the edit, a typo say. As with a judge packet, `record` hashes the
|
|
1611
|
+
spec and both drafts again, rebuilds the packet, and requires the file to be exactly those bytes
|
|
1612
|
+
before it reads the verdict.
|
|
1613
|
+
|
|
1614
|
+
### The tally and the next move
|
|
1615
|
+
|
|
1616
|
+
The tally counts, for each block the verdict names, its edits and the sentences they touched.
|
|
1617
|
+
Blocks are ordered by sentences, then edits, then the order of the table below with `none` last,
|
|
1618
|
+
and the next move goes to the first block that is not `none`, so a paragraph rewritten from end to
|
|
1619
|
+
end weighs as the sentences it rewrote. There is one move per block:
|
|
1620
|
+
|
|
1621
|
+
| Block | Next move |
|
|
1622
|
+
|---|---|
|
|
1623
|
+
| `materials` | mark or add the material the edit drew on |
|
|
1624
|
+
| `dna` | add a golden or a style rule |
|
|
1625
|
+
| `persona` | tighten the persona's stance, may_assert or will_not_say |
|
|
1626
|
+
| `audience` | extend the audience's knows or terms |
|
|
1627
|
+
| `goal` | tighten a goal condition's fails_when |
|
|
1628
|
+
| `form` | adjust the form block's length or shape |
|
|
1629
|
+
| `spine` | restate the spine's claim or its order |
|
|
1630
|
+
| `sources` | add or cite a source in the claims ledger |
|
|
1631
|
+
| `characters` | extend a character's speech or knowledge |
|
|
1632
|
+
|
|
1633
|
+
When every edit is `none`, the move is `none`: no block could have prevented any edit, so the spec
|
|
1634
|
+
has nothing to learn from the pair.
|
|
1635
|
+
|
|
1636
|
+
`record` appends one line to the runs ledger, `{ at, kind: "learn", first, first_sha256,
|
|
1637
|
+
approved, approved_sha256, spec_sha256, verdict, reason, tally }`, with both drafts' paths
|
|
1638
|
+
relative to the spec's folder. Its verdict is always `not-improved`: the spec has not changed yet,
|
|
1639
|
+
and the reason gives the counts by block. Learn reads no earlier line, and check and judge ignore
|
|
1640
|
+
learn lines.
|
|
1641
|
+
|
|
1642
|
+
| Id | Kind | Meaning |
|
|
1643
|
+
|---|---|---|
|
|
1644
|
+
| `learn-verdict-not-json` | invalid | the verdict file is not JSON |
|
|
1645
|
+
| `learn-verdict-shape` | invalid | the verdict is not `{ edits: [...] }`, or an entry is not an object with an id |
|
|
1646
|
+
| `learn-edit-unknown` | invalid | an id is not a hunk in the packet |
|
|
1647
|
+
| `learn-edit-duplicate` | invalid | a hunk is classified more than once |
|
|
1648
|
+
| `learn-edit-missing` | invalid | a hunk is not classified |
|
|
1649
|
+
| `learn-block-unknown` | invalid | a block is not one of the nine blocks or `none` |
|
|
1650
|
+
| `learn-block-absent` | invalid | a block is one the spec has not written |
|
|
1651
|
+
| `learn-why-missing` | invalid | a why is empty or not a string |
|
|
1652
|
+
| `learn-stale` | stale | the spec or a draft no longer matches the hash the packet recorded: it changed since prepare, or the packet was edited |
|
|
1653
|
+
| `learn-packet-altered` | invalid | the packet is not the one `prepare` builds from the files on disk |
|
|
1654
|
+
|
|
1655
|
+
The essay example's pair is a first draft that differs from `essay/draft.md` by five edits: a
|
|
1656
|
+
paragraph of hedged advice and a hedged sentence, where the approved draft gives instructions; a
|
|
1657
|
+
claim about what most managers do; the walking aside from the voice memo; and a typo. The sample
|
|
1658
|
+
verdict puts the two hedges on `dna`, since no golden or style rule shows advice given flat, and
|
|
1659
|
+
the tally sends the next move there.
|
|
1660
|
+
|
|
1012
1661
|
## Deferring a block
|
|
1013
1662
|
|
|
1014
1663
|
A block can be deferred, never silently missing. A required block that is absent fails test 1
|
|
@@ -1057,7 +1706,7 @@ not exist.
|
|
|
1057
1706
|
|
|
1058
1707
|
## Worked examples
|
|
1059
1708
|
|
|
1060
|
-
|
|
1709
|
+
Three complete specs ship in [`examples/writing/`](examples/writing/), each with every file it
|
|
1061
1710
|
names:
|
|
1062
1711
|
|
|
1063
1712
|
- `essay.hyperspec.md`: an essay for new managers on running a first one-on-one. Three materials
|
|
@@ -1066,8 +1715,11 @@ names:
|
|
|
1066
1715
|
- `story.hyperspec.md`: a short story, `fiction: true`, narrated by one of its two characters.
|
|
1067
1716
|
Each character has speech rules, a knowledge timeline by scene, and golden and rejected lines
|
|
1068
1717
|
in a voice you can tell apart from the other's.
|
|
1718
|
+
- `course.hyperspec.md`: a four-lesson course in two part files, declared a sequence (see
|
|
1719
|
+
[Sequential works](#sequential-works)), with an outline promising each lesson's terms. `check`
|
|
1720
|
+
reads the parts with no `--draft` and passes them, with two forward-pointer warnings.
|
|
1069
1721
|
|
|
1070
|
-
Every material in
|
|
1722
|
+
Every material in all three is marked. Between them the essay and the story use all seven labels, each with
|
|
1071
1723
|
the field it needs, and every spine claim cites the segments that support it. Each segments file
|
|
1072
1724
|
keeps the boundaries `segments init` wrote, in paragraph mode for prose and sentence mode for
|
|
1073
1725
|
bulleted notes, so you can re-run it and compare.
|
|
@@ -1076,15 +1728,22 @@ Both lint `pass (9/9)` with `writing: 9/9 blocks complete` and no findings. Each
|
|
|
1076
1728
|
draft written to it, `essay/draft.md` and `story/draft.md`, with its claims ledger beside it, and
|
|
1077
1729
|
both drafts pass every station of `check`: the essay with one dna warning, described under
|
|
1078
1730
|
[dna](#dna), and the story with dna skipped, since it names no scope folder, and quotes skipped,
|
|
1079
|
-
since it is fiction.
|
|
1080
|
-
specs and checks
|
|
1731
|
+
since it is fiction. The course lints the same way and its parts pass every station. A test
|
|
1732
|
+
lints all three specs and checks their drafts on every release, so they cannot drift from the tool.
|
|
1733
|
+
|
|
1734
|
+
Each also ships the packets `judge prepare` writes for its draft, in `essay/judge/` and
|
|
1735
|
+
`story/judge/`, and one sample verdict per packet in `essay/sample-verdicts/` and
|
|
1736
|
+
`story/sample-verdicts/`, filled in by hand and marked as samples; what each found is under
|
|
1737
|
+
[The worked examples](#the-worked-examples). The essay adds a learn pair in `essay/learn/`: a
|
|
1738
|
+
first draft, the packet `learn prepare` writes comparing it with `essay/draft.md`, and a sample
|
|
1739
|
+
learn verdict (see [Learning from edits](#learning-from-edits)). A test checks that every packet
|
|
1740
|
+
is what `prepare` writes now and records every sample.
|
|
1081
1741
|
|
|
1082
1742
|
## What later versions add
|
|
1083
1743
|
|
|
1084
|
-
This release is the schema, its lint, marked materials, scoped DNA,
|
|
1085
|
-
deterministic stations
|
|
1086
|
-
|
|
1087
|
-
|
|
1088
|
-
|
|
1089
|
-
|
|
1090
|
-
lands in the spec or the skill that wrote the draft rather than in one draft.
|
|
1744
|
+
This release is the schema, its lint, marked materials, scoped DNA, `check` with eight
|
|
1745
|
+
deterministic stations (sequential works among them), six judgment stations written as packets for
|
|
1746
|
+
an outside judge, and learn.
|
|
1747
|
+
Next: lineups over several passages of one draft, so that one lucky pick carries less weight, and
|
|
1748
|
+
a learn step that reads the runs ledger across drafts for the stations that keep failing and the
|
|
1749
|
+
changes that made them pass, beside what one pair of drafts shows.
|