@supersuit/hyperspec 0.5.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +166 -0
- package/README.md +62 -1
- package/SPEC.md +4 -4
- package/WRITING.md +770 -9
- package/bin/hyperspec.mjs +252 -0
- package/examples/writing/essay/claims.jsonl +9 -0
- package/examples/writing/essay/draft.md +82 -0
- package/examples/writing/essay/judge/doctor.packet.json +108 -0
- package/examples/writing/essay/judge/lineup.packet.json +64 -0
- package/examples/writing/essay/judge/persona.packet.json +73 -0
- package/examples/writing/essay/judge/reader.packet.json +93 -0
- package/examples/writing/essay/learn/first-draft.md +84 -0
- package/examples/writing/essay/learn/learn.packet.json +106 -0
- package/examples/writing/essay/materials/interview-notes.md.segments.jsonl +2 -2
- package/examples/writing/essay/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/essay/sample-verdicts/learn.verdict.json +30 -0
- package/examples/writing/essay/sample-verdicts/lineup.verdict.json +6 -0
- package/examples/writing/essay/sample-verdicts/persona.verdict.json +4 -0
- package/examples/writing/essay/sample-verdicts/reader.verdict.json +7 -0
- package/examples/writing/essay.hyperspec.md +23 -7
- package/examples/writing/story/claims.jsonl +9 -0
- package/examples/writing/story/draft.md +267 -0
- package/examples/writing/story/judge/attribution.packet.json +194 -0
- package/examples/writing/story/judge/doctor.packet.json +108 -0
- package/examples/writing/story/judge/knowledge.packet.json +77 -0
- package/examples/writing/story/judge/persona.packet.json +73 -0
- package/examples/writing/story/judge/reader.packet.json +94 -0
- package/examples/writing/story/sample-verdicts/attribution.verdict.json +81 -0
- package/examples/writing/story/sample-verdicts/doctor.verdict.json +43 -0
- package/examples/writing/story/sample-verdicts/knowledge.verdict.json +4 -0
- package/examples/writing/story/sample-verdicts/persona.verdict.json +20 -0
- package/examples/writing/story/sample-verdicts/reader.verdict.json +16 -0
- package/examples/writing/story.hyperspec.md +24 -5
- package/package.json +1 -1
- package/src/check.mjs +191 -0
- package/src/dna.mjs +4 -1
- package/src/draft.mjs +26 -0
- package/src/judge.mjs +386 -0
- package/src/judges/attribution.mjs +360 -0
- package/src/judges/doctor.mjs +126 -0
- package/src/judges/index.mjs +31 -0
- package/src/judges/knowledge.mjs +111 -0
- package/src/judges/lineup.mjs +272 -0
- package/src/judges/persona.mjs +137 -0
- package/src/judges/reader.mjs +111 -0
- package/src/learn.mjs +422 -0
- package/src/ledger.mjs +108 -0
- package/src/sentences.mjs +81 -0
- package/src/stations/claims.mjs +155 -0
- package/src/stations/dna.mjs +126 -0
- package/src/stations/form.mjs +115 -0
- package/src/stations/index.mjs +27 -0
- package/src/stations/links.mjs +275 -0
- package/src/stations/private.mjs +117 -0
- package/src/stations/quotes.mjs +170 -0
- package/src/stations/terms.mjs +131 -0
- package/src/stations/util.mjs +99 -0
- package/src/writing-fields.mjs +26 -2
- package/src/writing.mjs +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,171 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.7.0 (2026-09-29)
|
|
4
|
+
|
|
5
|
+
A spec can now have its judgments made and recorded. `check` covers what a function of the spec
|
|
6
|
+
and the draft can decide; the rest of a spec's checks are rubrics, and until this release nothing
|
|
7
|
+
ran them. hyperspec still calls no model. `hyperspec judge prepare` writes one packet per judgment
|
|
8
|
+
station, holding the rubric, fixed instructions, the inputs and the exact shape of the answer, for
|
|
9
|
+
an outside judge to fill: your agent, any model, or a person. `hyperspec judge record` checks the
|
|
10
|
+
verdict, derives the station's status from it by a fixed rule, and records it in the runs ledger.
|
|
11
|
+
Every passage a judge quotes must be in the draft, and the two blind tests are scored
|
|
12
|
+
against answer keys the judge never sees. `hyperspec learn` works from the other end: given the
|
|
13
|
+
first draft a factory wrote and the draft a person approved, it lists the edits, has a judge name
|
|
14
|
+
the spec block that should have prevented each one, and names one next move.
|
|
15
|
+
|
|
16
|
+
**No behavior change for lint or check.** Every 0.6.0 test passes unchanged, no finding id changed,
|
|
17
|
+
and no schema field was added. The runs ledger gains two kinds of line, `judge` and `learn`; check
|
|
18
|
+
reads only its own, so a 0.6 ledger's check verdicts are exactly what they were, and the new lines
|
|
19
|
+
keep lint's test 9 passing.
|
|
20
|
+
|
|
21
|
+
- `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force] [--json]`
|
|
22
|
+
lints the spec first, as `check` does, then writes `<station>.packet.json` for each station that
|
|
23
|
+
applies, in the order `doctor, lineup, reader, persona, attribution, knowledge`, and prints a
|
|
24
|
+
skip line with the reason for each that does not. Packets keep the paths as given and are
|
|
25
|
+
byte-identical for the same files. It refuses to overwrite any file without `--force`, naming
|
|
26
|
+
each. Exit 0 written, 1 a station crashed, 2 usage, or lint's own 1 or 3.
|
|
27
|
+
- The two blind tests write an answer key beside their packets, `lineup.key.json` and
|
|
28
|
+
`attribution.key.json`, for a person to read. Hand a judge only the `*.packet.json` files, and
|
|
29
|
+
each of the two blind packets to its own fresh context, with no access to the draft and apart
|
|
30
|
+
from the other packets, which carry it: their
|
|
31
|
+
instructions say to decide from the packet's inputs alone and open no file it names, but a judge
|
|
32
|
+
that can open the draft can always cheat. `record` never reads a key: it builds it again.
|
|
33
|
+
- `hyperspec judge record <packet> --verdict <file> [--json]` hashes the spec and the draft again
|
|
34
|
+
and rebuilds the packet from them and from the files the station reads (the DNA goldens, the
|
|
35
|
+
claims ledger). A file that changed makes the packet stale (`judge-stale`), and a packet that is
|
|
36
|
+
not those exact bytes is refused (`judge-packet-altered`); either way nothing is recorded. A
|
|
37
|
+
stale packet has `learn record`'s shape, `{ "invalid": true, "stale": true }` with `--json` and
|
|
38
|
+
"stale packet, nothing recorded" on the terminal. It must run in the folder `prepare` ran in. Then the verdict is validated, every problem named. Exit
|
|
39
|
+
0 the station passed, 1 it failed or the verdict was refused, 2 usage.
|
|
40
|
+
- The evidence rule: every span a verdict cites is at least three words and appears in the draft,
|
|
41
|
+
on whole words, after whitespace runs become one space and curly, low and angle quotation marks
|
|
42
|
+
and apostrophes become straight ones (primes do not).
|
|
43
|
+
- Six stations. `doctor`: every goal condition, and whether the reader would take
|
|
44
|
+
`goal.next_if_worked` now. `lineup`: the draft's paragraph nearest the goldens' median length
|
|
45
|
+
beside up to three goldens' paragraphs, each reflowed to one line and shuffled with a seed
|
|
46
|
+
derived from the draft's full text, which the packet does not carry, so the packet cannot reveal
|
|
47
|
+
the order (a draft that is one paragraph and nothing else skips, since its hash would name its
|
|
48
|
+
candidate); it passes when the judge picks a golden, which a random pick does three times in four.
|
|
49
|
+
`reader`: where the audience's reader got lost (warnings) or stopped, and whether they would
|
|
50
|
+
take their own next step. `persona`: every break of stance, assertion, `will_not_say` or an
|
|
51
|
+
unsourced fact, against the claims ledger's texts (`null` with no ledger, and then no
|
|
52
|
+
unsourced fact may be reported). `attribution`, fiction only: the speaker of each dialogue line
|
|
53
|
+
whose speech tag names one, scored per speaker and averaged, passing at 80 percent; a line whose
|
|
54
|
+
speaker is in doubt (a split quote joins only across exactly one speech tag, and the part after
|
|
55
|
+
any longer narration is left out), or that repeats a golden or rejected line, is left out and
|
|
56
|
+
counted, so the key is never wrong (the worked story tests 19 of its 36 lines). `knowledge`, fiction only: every
|
|
57
|
+
place a character knows something before their timeline gives it to them.
|
|
58
|
+
- Judge ledger lines carry `packet_sha256`, the hash of what the judge was shown, and
|
|
59
|
+
`inputs_sha256`, the hash of its inputs, and follow check's rules against the last judge line
|
|
60
|
+
for the same station and draft, saying `judgment` where check says `check` ("no change since the
|
|
61
|
+
last passing judgment"), with three more: what changed is judged by the packet, naming the
|
|
62
|
+
draft, the spec, the file the station reads besides them when the inputs changed (the DNA scope,
|
|
63
|
+
the claims ledger), or the packet's fixed text when only that changed, so adding the golden or
|
|
64
|
+
the claims a failing station asked for and passing is improved; one-shot needs draft bytes never judged by that
|
|
65
|
+
station under any name; and improved needs the packet to have changed, since a judge answering
|
|
66
|
+
differently about the same packet is not the work improving ("the verdict changed; nothing the
|
|
67
|
+
judge was shown changed").
|
|
68
|
+
- `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]`
|
|
69
|
+
splits both drafts into sentence units (a paragraph's opening heading, each list item and each
|
|
70
|
+
sentence, with common abbreviations and lower-case continuations joined), diffs them, and
|
|
71
|
+
writes `learn.packet.json` with every edit as a hunk (`deleted`, `inserted` or `replaced`, with
|
|
72
|
+
both texts and a sentence count). A hunk never crosses a paragraph or a heading, and reflowed
|
|
73
|
+
text is no edit. Past 10,000,000 comparisons after the common start and end are set aside, it is
|
|
74
|
+
a usage error.
|
|
75
|
+
- `hyperspec learn record <packet> --verdict <file>` checks that every hunk is classified once by a
|
|
76
|
+
block the spec has written, or `none`, counts edits and sentences per block, and names one move
|
|
77
|
+
for the block with the most sentences, from a fixed table (`dna`: add a golden or a style rule).
|
|
78
|
+
It appends a `learn` line to the runs ledger, always `not-improved` since the spec has not
|
|
79
|
+
changed yet, and never edits the spec.
|
|
80
|
+
- The worked examples ship their packets (`essay/judge/`, `story/judge/`), with no key, and one
|
|
81
|
+
sample verdict per packet (`*/sample-verdicts/`), filled in by hand and marked `sample`. Two
|
|
82
|
+
of them fail, honestly: the essay's lineup (the draft's passage is the only one resting on a
|
|
83
|
+
figure) and the story's persona (three process details no claim holds). The essay's goal now
|
|
84
|
+
names the card its draft ends on as the reader's next step (`goal.next_if_worked`, matching the
|
|
85
|
+
spec's `agenda-card` decision), so its doctor sample passes. The essay adds a learn pair in `essay/learn/` and a sample verdict whose
|
|
86
|
+
tally sends the next move to `dna`. A test holds every packet to what `prepare` writes and records
|
|
87
|
+
every sample.
|
|
88
|
+
- WRITING.md gains "Judging a draft" and "Learning from edits": the commands and exit codes, the
|
|
89
|
+
packet, the evidence rule, what `record` refuses, each station's rules and findings table (held
|
|
90
|
+
by a test to the ids the code raises and their kind), the ledger rules, the samples and what they
|
|
91
|
+
found, sentence units, the diff, the tally and its move table. The quotes station's section
|
|
92
|
+
now points at `attribution` for fiction's dialogue.
|
|
93
|
+
|
|
94
|
+
## 0.6.0 (2026-09-29)
|
|
95
|
+
|
|
96
|
+
A spec can now check a draft. Until this release hyperspec could tell you whether a writing spec
|
|
97
|
+
was ready; it could not tell you whether the piece written from it met the spec. `hyperspec
|
|
98
|
+
check` runs seven deterministic stations against a draft: the length and required parts, the
|
|
99
|
+
terms the reader needs defined, the claims ledger, quotations, private material, the writer's
|
|
100
|
+
measured style, and links. None of them calls a model or touches the network, so the same draft
|
|
101
|
+
and spec always give the same answer. Each run leaves one line in the spec's runs ledger saying
|
|
102
|
+
whether the draft passed first time, improved, or did not, and why.
|
|
103
|
+
|
|
104
|
+
**No behavior change for lint.** Every spec without `writing.audience.terms` passes and fails
|
|
105
|
+
exactly as it did in 0.5.0, and no lint finding id changed. That field is the one schema
|
|
106
|
+
addition, and it is optional.
|
|
107
|
+
|
|
108
|
+
- `hyperspec check <spec> --draft <file> [--json] [--only a,b]` needs a writing spec
|
|
109
|
+
(`profile: writing`) and lints it first: a spec that fails lint, or is blocked on an open
|
|
110
|
+
decision, runs no station and exits with lint's code. Then it runs every station in a fixed
|
|
111
|
+
order, `form, terms, claims, quotes, private, dna, links`, and prints each one's `pass`, `fail`
|
|
112
|
+
with findings, or `skip` with the reason. Warnings print under their station and never fail it.
|
|
113
|
+
Exit 0 when every station that ran passed, 1 when one failed, 2 on usage (no spec, no
|
|
114
|
+
`--draft`, an unreadable draft, a spec without the writing profile, an `--only` that names no
|
|
115
|
+
known station); under `--json` a usage error is one document, `{ spec, draft, error }`.
|
|
116
|
+
`--only` runs a subset, still in the fixed order. A finding names the draft line it points at
|
|
117
|
+
where there is one, quotes at most 80 characters of the draft, and never prints an absolute
|
|
118
|
+
path. A byte order mark at the start of the draft is ignored. A station that throws becomes one
|
|
119
|
+
failing finding, `station-<name>-crashed`, and the rest still run.
|
|
120
|
+
- `form`: word count against `writing.form.length` (only `unit: words` is measured; another unit
|
|
121
|
+
skips the station), and every `required_parts` entry present as an ATX heading of that text
|
|
122
|
+
(indented up to three spaces, closing `#`s allowed), or as a line starting `part:`, outside
|
|
123
|
+
code blocks.
|
|
124
|
+
- `terms`: every term in the new optional `writing.audience.terms`, other than those in `knows`,
|
|
125
|
+
is defined at its first appearance: in that sentence or the next, the term followed within six
|
|
126
|
+
words by `is`, `means`, `refers to` or a colon, or by a parenthesis. A mechanical proxy for a
|
|
127
|
+
definition, and documented as one. No `terms` list: skip.
|
|
128
|
+
- `claims`: every claim in the JSONL ledger at `writing.sources.ledger` (`text`, `source`,
|
|
129
|
+
optional `span`) still appears in the draft word for word, and has a real source; unsourced
|
|
130
|
+
claims warn instead under `unsourced_claim: warn`, and point at the draft line where the claim
|
|
131
|
+
appears. A missing ledger fails. The ledger is the list of claims: the station does not decide
|
|
132
|
+
what counts as one.
|
|
133
|
+
- `quotes`: every double-quoted span of four words or more appears word for word in a `quote` or
|
|
134
|
+
`story` segment of a marked material, never a private one; when the sentence names a quote
|
|
135
|
+
segment's speaker, by the full name or by its first word (when that word has two or more
|
|
136
|
+
letters and is not a common function word such as "the"), the span must come from that
|
|
137
|
+
speaker. A spec with `fiction: true` skips the station, since a character's dialogue is
|
|
138
|
+
invented rather than quoted.
|
|
139
|
+
- `private`: no run of eight words from a `private` segment appears in the draft. Segments of four
|
|
140
|
+
to seven words are checked whole; shorter ones are counted in one warning and never quoted.
|
|
141
|
+
- `dna`: with a current `writing.dna.scope_dir`, the draft is measured the way goldens are and
|
|
142
|
+
each feature compared with the scope's, from v ÷ 1.5 to the larger of v × 1.5 and v + 5; drift,
|
|
143
|
+
and an em dash where the scope has none (pointing at the first one), are warnings. No scope
|
|
144
|
+
folder: skip.
|
|
145
|
+
- `links`: inline, reference and bare links are well-formed http, https or mailto (a bare
|
|
146
|
+
`https://` with no host fails); relative links resolve to a file beside the draft; a `/` link
|
|
147
|
+
warns, since there is no site root to resolve it against; a full or collapsed reference needs
|
|
148
|
+
its definition. A bare `[label]` is a link only when its label is defined, so `[sic]`, `[x]` and
|
|
149
|
+
`[1]` are text. No network access.
|
|
150
|
+
- The runs ledger: each check appends `{ at, kind: "check", draft, draft_sha256, spec_sha256,
|
|
151
|
+
stations, verdict }` to `improvement.ledger`, with `draft` relative to the spec's folder. A run
|
|
152
|
+
with `--only` adds `partial: true`, is `not-improved` with the reason `partial run: <stations>`,
|
|
153
|
+
and is ignored by later verdicts. A full run is compared with the last full line for the same
|
|
154
|
+
draft: `one-shot` when there is none and every station passes; `improved` when that line failed
|
|
155
|
+
and every station passes now, with `change` naming exactly the stations that failed then and
|
|
156
|
+
pass now; otherwise `not-improved`, with a `reason` that says which case it is, such as "no
|
|
157
|
+
change since the last passing check" or "spec changed; failing stations: terms". The lines keep
|
|
158
|
+
lint's test 9 passing.
|
|
159
|
+
- New optional field `writing.audience.terms`: a list of real strings when present (test 1,
|
|
160
|
+
`writing-audience-terms`, naming a scalar that is not a list, an empty list, and each entry that
|
|
161
|
+
is not a string, is empty or is a placeholder). Absent, nothing changes.
|
|
162
|
+
- The worked examples each ship a draft that passes every station, with its claims ledger:
|
|
163
|
+
`examples/writing/essay/draft.md` and `examples/writing/story/draft.md`. Both specs now list
|
|
164
|
+
`audience.terms`, and their `required_parts` are the drafts' headings. The essay's interview
|
|
165
|
+
quotes now name their speaker `dana, an engineering manager`, and the essay links the survey
|
|
166
|
+
summary it cites. A test runs `check` on both. WRITING.md gains "Checking a draft": each station, what it cannot check, a findings table
|
|
167
|
+
per station held by a test to the ids the stations raise, and the ledger line.
|
|
168
|
+
|
|
3
169
|
## 0.5.0 (2026-09-29)
|
|
4
170
|
|
|
5
171
|
Scoped writer DNA. A writer does not have one voice: the same person writes differently for a
|
package/README.md
CHANGED
|
@@ -28,6 +28,11 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
|
|
|
28
28
|
| `hyperspec segments init <material> --id <mid> [--out F] [--by paragraph\|sentence]` | Split a material into segments to label. Refuses to overwrite an existing file. |
|
|
29
29
|
| `hyperspec dna init <scope-dir> --writer W --form F --audience A --purpose P` | Start a writer-DNA scope folder. Refuses to overwrite an existing `scope.md`. |
|
|
30
30
|
| `hyperspec dna measure <scope-dir>` | Check every golden in a scope and write its measured features. |
|
|
31
|
+
| `hyperspec check <spec> --draft <file> [--only a,b]` | Run a writing spec's deterministic stations against a draft. |
|
|
32
|
+
| `hyperspec judge prepare <spec> --draft <file> --out <dir> [--only a,b] [--force]` | Write one packet per judgment station, for an outside judge to fill. |
|
|
33
|
+
| `hyperspec judge record <packet> --verdict <file>` | Check a judge's verdict against its packet, derive the station's status, and record it. |
|
|
34
|
+
| `hyperspec learn prepare <spec> --first <draft> --approved <draft> --out <dir> [--force]` | Write the edits between a first draft and the approved one, for a judge to classify by spec block. |
|
|
35
|
+
| `hyperspec learn record <packet> --verdict <file>` | Count the classified edits by block and name one next move. |
|
|
31
36
|
| `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
|
|
32
37
|
| `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
|
|
33
38
|
| `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
|
|
@@ -42,6 +47,16 @@ Every command except `init`, `segments init` and `dna init` takes `--json`. `hyp
|
|
|
42
47
|
test fails, 3 when every test passes but a decision is still open (blocked), and 2 on a usage
|
|
43
48
|
error or a file that cannot be read, has broken frontmatter, or is not a hyperspec.
|
|
44
49
|
|
|
50
|
+
`hyperspec check` exits 0 when every station it ran passed, 1 when one failed, and 2 on a usage
|
|
51
|
+
error, a draft that cannot be read, or a spec without `profile: writing`. A spec that is not ready to check against exits with lint's
|
|
52
|
+
own code, 1 or 3, and no station runs.
|
|
53
|
+
|
|
54
|
+
`hyperspec judge record` exits 0 when the station passed and 1 when it failed or the verdict was
|
|
55
|
+
refused (invalid, or its packet stale or edited), and `hyperspec learn record` 0 when it recorded
|
|
56
|
+
and 1 when it refused. `judge prepare` and `learn prepare` exit 0 when they wrote their packets and
|
|
57
|
+
with lint's own code when the spec is not ready, and `judge prepare` exits 1 when a station could
|
|
58
|
+
not build its packet. All four exit 2 on a usage error.
|
|
59
|
+
|
|
45
60
|
The recipe commands use the same numbers: 0 ok, 1 a check failed or the child regressed, 2
|
|
46
61
|
usage or unreadable input, 3 pending, when `regenerate` has stages waiting for a runner.
|
|
47
62
|
`regenerate` also exits 1 when a stage it ran reported a failing verdict, or none. A stage it
|
|
@@ -132,7 +147,7 @@ The skeleton fails until every placeholder is real and its four open questions (
|
|
|
132
147
|
who speaks, who reads, what changes) are answered. The blocks, every field, and which test
|
|
133
148
|
each rule reports under are in [WRITING.md](WRITING.md). Two complete specs that pass with
|
|
134
149
|
nothing to warn ship in `examples/writing/`: an essay for new managers, and a short story with
|
|
135
|
-
two characters whose voices a judge can tell apart.
|
|
150
|
+
two characters whose voices a judge can tell apart. Each comes with a draft written to it.
|
|
136
151
|
|
|
137
152
|
### Marking materials
|
|
138
153
|
|
|
@@ -176,6 +191,52 @@ it did in 0.4. The folder shape, every feature, and every finding are in
|
|
|
176
191
|
[WRITING.md](WRITING.md#scoped-dna); `readScope` and `measureFeatures` are exported from
|
|
177
192
|
`@supersuit/hyperspec/writing`.
|
|
178
193
|
|
|
194
|
+
### Checking a draft
|
|
195
|
+
|
|
196
|
+
Once a draft exists, `check` holds it to its spec with seven stations, none of which calls a
|
|
197
|
+
model or touches the network: `form` (length and required parts), `terms` (every word in the new
|
|
198
|
+
optional `writing.audience.terms` is defined where it first appears), `claims` (the claims
|
|
199
|
+
ledger still matches the draft, and every claim has a source), `quotes` (in nonfiction, every quotation of four
|
|
200
|
+
words or more is word for word in a marked quote), `private` (no run of eight words from a
|
|
201
|
+
private segment), `dna` (the draft's measured style beside its scope's, as warnings) and `links`
|
|
202
|
+
(well-formed, and relative links resolve).
|
|
203
|
+
|
|
204
|
+
```bash
|
|
205
|
+
npx @supersuit/hyperspec check essay.hyperspec.md --draft essay/draft.md
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
It lints the spec first, prints each station's pass, fail or skip, and appends one line to the
|
|
209
|
+
spec's runs ledger with a verdict. Both examples ship a draft that passes. What each station
|
|
210
|
+
checks and cannot check, and every finding, are in
|
|
211
|
+
[WRITING.md](WRITING.md#checking-a-draft).
|
|
212
|
+
|
|
213
|
+
### Judging a draft and learning from edits
|
|
214
|
+
|
|
215
|
+
The rest of a spec's checks are judgments: whether each goal condition holds, where a reader gets
|
|
216
|
+
lost, whether a passage can be told from the writer's goldens, whether the persona holds, and in
|
|
217
|
+
fiction whether the characters' voices can be told apart and whether anyone knows something too
|
|
218
|
+
early. hyperspec never calls a model. `judge prepare` writes a packet per station for an outside
|
|
219
|
+
judge, your agent, any model or a person, and `judge record` checks the verdict (every quoted
|
|
220
|
+
passage must be in the draft, and the blind tests are scored against keys the judge never
|
|
221
|
+
sees), derives the station's status, and records it in the runs ledger.
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
npx @supersuit/hyperspec judge prepare essay.hyperspec.md --draft essay/draft.md --out essay/judge
|
|
225
|
+
npx @supersuit/hyperspec judge record essay/judge/doctor.packet.json --verdict doctor.verdict.json
|
|
226
|
+
```
|
|
227
|
+
|
|
228
|
+
Hand the judge the `*.packet.json` files only: the answer keys are written beside them. Give the
|
|
229
|
+
two blind packets (`lineup`, `attribution`) to a judge in a fresh context with no access to the
|
|
230
|
+
draft, such as a new conversation, never to an agent working in the draft's folder: a judge that
|
|
231
|
+
can open the draft can always find the answer, whatever the packet tells it. Give each blind
|
|
232
|
+
packet its own context, apart from the other packets too, since those carry the draft. `learn`
|
|
233
|
+
closes the loop from the other end. Given the first draft a factory wrote and the draft a person
|
|
234
|
+
approved, `learn prepare` lists the edits, a judge names the spec block that should have prevented
|
|
235
|
+
each, and `learn record` counts them by block and names one next move, such as "add a golden or a
|
|
236
|
+
style rule". It never edits the spec. Both examples ship their packets and hand-filled sample
|
|
237
|
+
verdicts. The packet shapes, every station's rules and findings, and the learn tally are in
|
|
238
|
+
[WRITING.md](WRITING.md#judging-a-draft).
|
|
239
|
+
|
|
179
240
|
## The format
|
|
180
241
|
|
|
181
242
|
A hyperspec is a markdown file with a YAML frontmatter block: `decisions`, `requirements`,
|
package/SPEC.md
CHANGED
|
@@ -97,7 +97,7 @@ examples:
|
|
|
97
97
|
- path: examples/minimal.hyperspec.md
|
|
98
98
|
why: the smallest spec that passes all nine tests
|
|
99
99
|
resume:
|
|
100
|
-
next_action: collect adopter issues on 0.
|
|
100
|
+
next_action: collect adopter issues on 0.7, the judge and learn commands included, and cut 0.8 from them
|
|
101
101
|
feedback:
|
|
102
102
|
issues: https://github.com/SupersuitUp/hyperspec/issues
|
|
103
103
|
fork: MIT; fork it for your own purposes and say so in your SPEC
|
|
@@ -109,7 +109,7 @@ improvement:
|
|
|
109
109
|
|
|
110
110
|
A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
|
|
111
111
|
|
|
112
|
-
**Version 0.
|
|
112
|
+
**Version 0.7.0** (2026-09-29)
|
|
113
113
|
|
|
114
114
|
## What makes a spec a hyperspec
|
|
115
115
|
|
|
@@ -220,7 +220,7 @@ Each row lists every condition under which `hyperspec lint` fails that test. A w
|
|
|
220
220
|
| 8 its adopters can push back on it | `feedback.issues` or `feedback.fork` missing |
|
|
221
221
|
| 9 it improves itself | `improvement.ledger` missing; a ledger path that exists and is not a readable file; if the ledger file exists, a line that is not a JSON object, a `verdict` outside one-shot, improved or not-improved, `improved` without `change`, `not-improved` without `reason`. A declared ledger that does not exist yet is a warning |
|
|
222
222
|
|
|
223
|
-
A profile adds its own conditions to these rows. The writing profile's are in [WRITING.md](WRITING.md#the-test-mapping), including the checks on each material's segments file: every material marked, every segment labeled from the closed set and matching its material word for word, the marking current (tests 1 and 4), and no spine claim citing a private or question segment (test 5). A writing spec that names a writer-DNA scope folder with `writing.dna.scope_dir` is also checked against it: the scope matches the spec (test 1), every golden has a person's approval and a source (test 4), no golden comes from outside the scope (test 5), and every golden has its why and the scope's measured features are current (test 6). `scope_dir` is optional, and without it nothing changes.
|
|
223
|
+
A profile adds its own conditions to these rows. The writing profile's are in [WRITING.md](WRITING.md#the-test-mapping), including the checks on each material's segments file: every material marked, every segment labeled from the closed set and matching its material word for word, the marking current (tests 1 and 4), and no spine claim citing a private or question segment (test 5). A writing spec that names a writer-DNA scope folder with `writing.dna.scope_dir` is also checked against it: the scope matches the spec (test 1), every golden has a person's approval and a source (test 4), no golden comes from outside the scope (test 5), and every golden has its why and the scope's measured features are current (test 6). `scope_dir` is optional, and without it nothing changes. The writing profile also has an optional `writing.audience.terms`, the words a piece uses that its reader may not know: when present it must list real terms (test 1), and `hyperspec check` reads it to require each term's definition where the draft first uses it. Without it nothing changes.
|
|
224
224
|
|
|
225
225
|
## Exit codes
|
|
226
226
|
|
|
@@ -236,7 +236,7 @@ A profile adds its own conditions to these rows. The writing profile's are in [W
|
|
|
236
236
|
Every run of a skill that works from a hyperspec writes one line to the ledger named in `improvement.ledger`, one JSON object per line, with a `verdict`:
|
|
237
237
|
|
|
238
238
|
- **one-shot**: no intervention, nothing to learn.
|
|
239
|
-
- **improved**: the skill, the spec template, or a component library changed, and the line carries `change`, naming what changed.
|
|
239
|
+
- **improved**: the skill, the spec template, or a component library changed, and the line carries `change`, naming what changed. On a `kind: "check"` line, which `hyperspec check` writes, it means the draft now passes every station after the last full check of it failed, and `change` names those stations.
|
|
240
240
|
- **not-improved**: nothing changed, and the line carries `reason`, a reason a later session can argue with, such as "the correction was about this piece only" or "the fix belongs to a shipped skill and was filed as an issue".
|
|
241
241
|
|
|
242
242
|
Silence is not a verdict. A run that learned nothing has to say so and why, and a ledger line with none of the three verdicts fails the ninth test.
|