@supersuit/hyperspec 0.6.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +91 -0
- package/README.md +37 -0
- package/SPEC.md +2 -2
- package/WRITING.md +542 -10
- package/bin/hyperspec.mjs +183 -0
- package/examples/writing/essay/judge/doctor.packet.json +108 -0
- package/examples/writing/essay/judge/lineup.packet.json +64 -0
- package/examples/writing/essay/judge/persona.packet.json +73 -0
- package/examples/writing/essay/judge/reader.packet.json +93 -0
- package/examples/writing/essay/learn/first-draft.md +84 -0
- package/examples/writing/essay/learn/learn.packet.json +106 -0
- package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
- package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
- package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
- package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
- package/examples/writing/essay.hyperspec.md +6 -1
- package/examples/writing/story/judge/attribution.packet.json +194 -0
- package/examples/writing/story/judge/doctor.packet.json +108 -0
- package/examples/writing/story/judge/knowledge.packet.json +77 -0
- package/examples/writing/story/judge/persona.packet.json +73 -0
- package/examples/writing/story/judge/reader.packet.json +94 -0
- package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
- package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
- package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
- package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
- package/examples/writing/story.hyperspec.md +7 -3
- package/package.json +1 -1
- package/src/check.mjs +75 -129
- package/src/draft.mjs +26 -0
- package/src/judge.mjs +386 -0
- package/src/judges/attribution.mjs +360 -0
- package/src/judges/doctor.mjs +126 -0
- package/src/judges/index.mjs +31 -0
- package/src/judges/knowledge.mjs +111 -0
- package/src/judges/lineup.mjs +272 -0
- package/src/judges/persona.mjs +137 -0
- package/src/judges/reader.mjs +111 -0
- package/src/learn.mjs +422 -0
- package/src/ledger.mjs +108 -0
- package/src/sentences.mjs +81 -0
- package/src/stations/claims.mjs +44 -39
- package/src/stations/quotes.mjs +6 -4
- package/src/writing.mjs +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,96 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.7.0 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
A spec can now have its judgments made and recorded. `check` covers what a function of the spec
|
|
6
|
+
and the draft can decide; the rest of a spec's checks are rubrics, and until this release nothing
|
|
7
|
+
ran them. hyperspec still calls no model. `hyperspec judge prepare` writes one packet per judgment
|
|
8
|
+
station, holding the rubric, fixed instructions, the inputs and the exact shape of the answer, for
|
|
9
|
+
an outside judge to fill: your agent, any model, or a person. `hyperspec judge record` checks the
|
|
10
|
+
verdict, derives the station's status from it by a fixed rule, and records it in the runs ledger.
|
|
11
|
+
Every passage a judge quotes must be in the draft, and the two blind tests are scored
|
|
12
|
+
against answer keys the judge never sees. `hyperspec learn` works from the other end: given the
|
|
13
|
+
first draft a factory wrote and the draft a person approved, it lists the edits, has a judge name
|
|
14
|
+
the spec block that should have prevented each one, and names one next move.
|
|
15
|
+
|
|
16
|
+
**No behavior change for lint or check.** Every 0.6.0 test passes unchanged, no finding id changed,
|
|
17
|
+
and no schema field was added. The runs ledger gains two kinds of line, `judge` and `learn`; check
|
|
18
|
+
reads only its own, so a 0.6 ledger's check verdicts are exactly what they were, and the new lines
|
|
19
|
+
keep lint's test 9 passing.
|
|
20
|
+
|
|
21
|
+
- `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force] [--json]`
|
|
22
|
+
lints the spec first, as `check` does, then writes `<station>.packet.json` for each station that
|
|
23
|
+
applies, in the order `doctor, lineup, reader, persona, attribution, knowledge`, and prints a
|
|
24
|
+
skip line with the reason for each that does not. Packets keep the paths as given and are
|
|
25
|
+
byte-identical for the same files. It refuses to overwrite any file without `--force`, naming
|
|
26
|
+
each. Exit 0 written, 1 a station crashed, 2 usage, or lint's own 1 or 3.
|
|
27
|
+
- The two blind tests write an answer key beside their packets, `lineup.key.json` and
|
|
28
|
+
`attribution.key.json`, for a person to read. Hand a judge only the `*.packet.json` files, and
|
|
29
|
+
each of the two blind packets to its own fresh context, with no access to the draft and apart
|
|
30
|
+
from the other packets, which carry it: their
|
|
31
|
+
instructions say to decide from the packet's inputs alone and open no file it names, but a judge
|
|
32
|
+
that can open the draft can always cheat. `record` never reads a key: it builds it again.
|
|
33
|
+
- `hyperspec judge record <packet> --verdict <file> [--json]` hashes the spec and the draft again
|
|
34
|
+
and rebuilds the packet from them and from the files the station reads (the DNA goldens, the
|
|
35
|
+
claims ledger). A file that changed makes the packet stale (`judge-stale`), and a packet that is
|
|
36
|
+
not those exact bytes is refused (`judge-packet-altered`); either way nothing is recorded. A
|
|
37
|
+
stale packet has `learn record`'s shape, `{ "invalid": true, "stale": true }` with `--json` and
|
|
38
|
+
"stale packet, nothing recorded" on the terminal. It must run in the folder `prepare` ran in. Then the verdict is validated, every problem named. Exit
|
|
39
|
+
0 the station passed, 1 it failed or the verdict was refused, 2 usage.
|
|
40
|
+
- The evidence rule: every span a verdict cites is at least three words and appears in the draft,
|
|
41
|
+
on whole words, after whitespace runs become one space and curly, low and angle quotation marks
|
|
42
|
+
and apostrophes become straight ones (primes do not).
|
|
43
|
+
- Six stations. `doctor`: every goal condition, and whether the reader would take
|
|
44
|
+
`goal.next_if_worked` now. `lineup`: the draft's paragraph nearest the goldens' median length
|
|
45
|
+
beside up to three goldens' paragraphs, each reflowed to one line and shuffled with a seed
|
|
46
|
+
derived from the draft's full text, which the packet does not carry, so the packet cannot reveal
|
|
47
|
+
the order (a draft that is one paragraph and nothing else skips, since its hash would name its
|
|
48
|
+
candidate); it passes when the judge picks a golden, which a random pick does three times in four.
|
|
49
|
+
`reader`: where the audience's reader got lost (warnings) or stopped, and whether they would
|
|
50
|
+
take their own next step. `persona`: every break of stance, assertion, `will_not_say` or an
|
|
51
|
+
unsourced fact, against the claims ledger's texts (`null` with no ledger, and then no
|
|
52
|
+
unsourced fact may be reported). `attribution`, fiction only: the speaker of each dialogue line
|
|
53
|
+
whose speech tag names one, scored per speaker and averaged, passing at 80 percent; a line whose
|
|
54
|
+
speaker is in doubt (a split quote joins only across exactly one speech tag, and the part after
|
|
55
|
+
any longer narration is left out), or that repeats a golden or rejected line, is left out and
|
|
56
|
+
counted, so the key is never wrong (the worked story tests 19 of its 36 lines). `knowledge`, fiction only: every
|
|
57
|
+
place a character knows something before their timeline gives it to them.
|
|
58
|
+
- Judge ledger lines carry `packet_sha256`, the hash of what the judge was shown, and
|
|
59
|
+
`inputs_sha256`, the hash of its inputs, and follow check's rules against the last judge line
|
|
60
|
+
for the same station and draft, saying `judgment` where check says `check` ("no change since the
|
|
61
|
+
last passing judgment"), with three more: what changed is judged by the packet, naming the
|
|
62
|
+
draft, the spec, the file the station reads besides them when the inputs changed (the DNA scope,
|
|
63
|
+
the claims ledger), or the packet's fixed text when only that changed, so adding the golden or
|
|
64
|
+
the claims a failing station asked for and passing is improved; one-shot needs draft bytes never judged by that
|
|
65
|
+
station under any name; and improved needs the packet to have changed, since a judge answering
|
|
66
|
+
differently about the same packet is not the work improving ("the verdict changed; nothing the
|
|
67
|
+
judge was shown changed").
|
|
68
|
+
- `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]`
|
|
69
|
+
splits both drafts into sentence units (a paragraph's opening heading, each list item and each
|
|
70
|
+
sentence, with common abbreviations and lower-case continuations joined), diffs them, and
|
|
71
|
+
writes `learn.packet.json` with every edit as a hunk (`deleted`, `inserted` or `replaced`, with
|
|
72
|
+
both texts and a sentence count). A hunk never crosses a paragraph or a heading, and reflowed
|
|
73
|
+
text is no edit. Past 10,000,000 comparisons after the common start and end are set aside, it is
|
|
74
|
+
a usage error.
|
|
75
|
+
- `hyperspec learn record <packet> --verdict <file>` checks that every hunk is classified once by a
|
|
76
|
+
block the spec has written, or `none`, counts edits and sentences per block, and names one move
|
|
77
|
+
for the block with the most sentences, from a fixed table (`dna`: add a golden or a style rule).
|
|
78
|
+
It appends a `learn` line to the runs ledger, always `not-improved` since the spec has not
|
|
79
|
+
changed yet, and never edits the spec.
|
|
80
|
+
- The worked examples ship their packets (`essay/judge/`, `story/judge/`), with no key, and one
|
|
81
|
+
sample verdict per packet (`*/sample-verdicts/`), filled in by hand and marked `sample`. Two
|
|
82
|
+
of them fail, honestly: the essay's lineup (the draft's passage is the only one resting on a
|
|
83
|
+
figure) and the story's persona (three process details no claim holds). The essay's goal now
|
|
84
|
+
names the card its draft ends on as the reader's next step (`goal.next_if_worked`, matching the
|
|
85
|
+
spec's `agenda-card` decision), so its doctor sample passes. The essay adds a learn pair in `essay/learn/` and a sample verdict whose
|
|
86
|
+
tally sends the next move to `dna`. A test holds every packet to what `prepare` writes and records
|
|
87
|
+
every sample.
|
|
88
|
+
- WRITING.md gains "Judging a draft" and "Learning from edits": the commands and exit codes, the
|
|
89
|
+
packet, the evidence rule, what `record` refuses, each station's rules and findings table (held
|
|
90
|
+
by a test to the ids the code raises and their kind), the ledger rules, the samples and what they
|
|
91
|
+
found, sentence units, the diff, the tally and its move table. The quotes station's section
|
|
92
|
+
now points at `attribution` for fiction's dialogue.
|
|
93
|
+
|
|
3
94
|
## 0.6.0 (2026-09-29)
|
|
4
95
|
|
|
5
96
|
A spec can now check a draft. Until this release hyperspec could tell you whether a writing spec
|
package/README.md
CHANGED
|
@@ -29,6 +29,10 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
|
|
|
29
29
|
| `hyperspec dna init <scope-dir> --writer W --form F --audience A --purpose P` | Start a writer-DNA scope folder. Refuses to overwrite an existing `scope.md`. |
|
|
30
30
|
| `hyperspec dna measure <scope-dir>` | Check every golden in a scope and write its measured features. |
|
|
31
31
|
| `hyperspec check <spec> --draft <file> [--only a,b]` | Run a writing spec's deterministic stations against a draft. |
|
|
32
|
+
| `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force]` | Write one packet per judgment station, for an outside judge to fill. |
|
|
33
|
+
| `hyperspec judge record <packet> --verdict <file>` | Check a judge's verdict against its packet, derive the station's status, and record it. |
|
|
34
|
+
| `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]` | Write the edits between a first draft and the approved one, for a judge to classify by spec block. |
|
|
35
|
+
| `hyperspec learn record <packet> --verdict <file>` | Count the classified edits by block and name one next move. |
|
|
32
36
|
| `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
|
|
33
37
|
| `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
|
|
34
38
|
| `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
|
|
@@ -47,6 +51,12 @@ error or a file that cannot be read, has broken frontmatter, or is not a hypersp
|
|
|
47
51
|
error, a draft that cannot be read, or a spec without `profile: writing`. A spec that is not ready to check against exits with lint's
|
|
48
52
|
own code, 1 or 3, and no station runs.
|
|
49
53
|
|
|
54
|
+
`hyperspec judge record` exits 0 when the station passed and 1 when it failed or the verdict was
|
|
55
|
+
refused (invalid, or its packet stale or edited), and `hyperspec learn record` 0 when it recorded
|
|
56
|
+
and 1 when it refused. `judge prepare` and `learn prepare` exit 0 when they wrote their packets and
|
|
57
|
+
with lint's own code when the spec is not ready, and `judge prepare` exits 1 when a station could
|
|
58
|
+
not build its packet. All four exit 2 on a usage error.
|
|
59
|
+
|
|
50
60
|
The recipe commands use the same numbers: 0 ok, 1 a check failed or the child regressed, 2
|
|
51
61
|
usage or unreadable input, 3 pending, when `regenerate` has stages waiting for a runner.
|
|
52
62
|
`regenerate` also exits 1 when a stage it ran reported a failing verdict, or none. A stage it
|
|
@@ -200,6 +210,33 @@ spec's runs ledger with a verdict. Both examples ship a draft that passes. What
|
|
|
200
210
|
checks and cannot check, and every finding, are in
|
|
201
211
|
[WRITING.md](WRITING.md#checking-a-draft).
|
|
202
212
|
|
|
213
|
+
### Judging a draft and learning from edits
|
|
214
|
+
|
|
215
|
+
The rest of a spec's checks are judgments: whether each goal condition holds, where a reader gets
|
|
216
|
+
lost, whether a passage can be told from the writer's goldens, whether the persona holds, and in
|
|
217
|
+
fiction whether the characters' voices can be told apart and whether anyone knows something too
|
|
218
|
+
early. hyperspec never calls a model. `judge prepare` writes a packet per station for an outside
|
|
219
|
+
judge, your agent, any model or a person, and `judge record` checks the verdict (every quoted
|
|
220
|
+
passage must be in the draft, and the blind tests are scored against keys the judge never
|
|
221
|
+
sees), derives the station's status, and records it in the runs ledger.
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
|
|
225
|
+
npx @supersuit/hyperspec judge record essay/judge/doctor.packet.json --verdict doctor.verdict.json
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
Hand the judge the `*.packet.json` files only: the answer keys are written beside them. Give the
|
|
229
|
+
two blind packets (`lineup`, `attribution`) to a judge in a fresh context with no access to the
|
|
230
|
+
draft, such as a new conversation, never to an agent working in the draft's folder: a judge that
|
|
231
|
+
can open the draft can always find the answer, whatever the packet tells it. Give each blind
|
|
232
|
+
packet its own context, apart from the other packets too, since those carry the draft. `learn`
|
|
233
|
+
closes the loop from the other end. Given the first draft a factory wrote and the draft a person
|
|
234
|
+
approved, `learn prepare` lists the edits, a judge names the spec block that should have prevented
|
|
235
|
+
each, and `learn record` counts them by block and names one next move, such as "add a golden or a
|
|
236
|
+
style rule". It never edits the spec. Both examples ship their packets and hand-filled sample
|
|
237
|
+
verdicts. The packet shapes, every station's rules and findings, and the learn tally are in
|
|
238
|
+
[WRITING.md](WRITING.md#judging-a-draft).
|
|
239
|
+
|
|
203
240
|
## The format
|
|
204
241
|
|
|
205
242
|
A hyperspec is a markdown file with a YAML frontmatter block: `decisions`, `requirements`,
|
package/SPEC.md
CHANGED
|
@@ -97,7 +97,7 @@ examples:
|
|
|
97
97
|
- path: examples/minimal.hyperspec.md
|
|
98
98
|
why: the smallest spec that passes all nine tests
|
|
99
99
|
resume:
|
|
100
|
-
next_action: collect adopter issues on 0.
|
|
100
|
+
next_action: collect adopter issues on 0.7, the judge and learn commands included, and cut 0.8 from them
|
|
101
101
|
feedback:
|
|
102
102
|
issues: https://github.com/SupersuitUp/hyperspec/issues
|
|
103
103
|
fork: MIT; fork it for your own purposes and say so in your SPEC
|
|
@@ -109,7 +109,7 @@ improvement:
|
|
|
109
109
|
|
|
110
110
|
A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
|
|
111
111
|
|
|
112
|
-
**Version 0.
|
|
112
|
+
**Version 0.7.0** (2026-09-29)
|
|
113
113
|
|
|
114
114
|
## What makes a spec a hyperspec
|
|
115
115
|
|