@supersuit/hyperspec 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,64 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.4.0 (2026-09-29)
4
+
5
+ Marking materials. Before a writing spec can pass, every material it draws on (a brain dump, a
6
+ transcript, a set of interview notes) is split into segments, and each segment is labeled with
7
+ what a draft may use it as: a claim with its source, the author's own claim, a story with its
8
+ teller, a quote with its speaker, a stance, an open question, an aside, or something private.
9
+ The spine then cites segments rather than whole files, so every claim points at the exact words
10
+ behind it. hyperspec splits and checks; an agent or a person labels. It still calls no model.
11
+
12
+ **Behavior change for 0.3 writing specs:** a writing spec's materials must now be marked. A
13
+ material item with no `segments:` field fails test 1, so a writing spec that passed 0.3.0 fails
14
+ until each of its materials has a segments file, written with `hyperspec segments init` and
15
+ labeled. Specs with no profile are unaffected.
16
+
17
+ - `hyperspec segments init <material> --id <mid> [--out <file>] [--by paragraph|sentence]`
18
+ writes `<material>.segments.jsonl`: a header naming the material, its path and the SHA-256 of
19
+ its bytes, then one line per segment with character offsets, the verbatim text, and the label
20
+ `unlabeled`. Paragraph mode is the default. It refuses to overwrite a file, and exits 2 with a
21
+ plain message on a missing material, a material with nothing in it, an `--out` folder that
22
+ does not exist, or an unknown `--by`.
23
+ - In sentence mode, a new line that opens on a list marker (`-`, `*`, `+`, or a number followed
24
+ by `.` or `)`, then a space) starts a new segment, and a numbered item's own `1.` is not read
25
+ as a sentence ending. A bullet that is entirely a quotation ending in `."` used to run into the
26
+ next bullet; it now stands on its own. Paragraph mode is unchanged.
27
+ - Seven labels, a closed set: `claim` (needs `source`, or `own: true`), `story` (needs
28
+ `teller`), `quote` (needs `speaker`), `stance`, `question`, `aside` and `private`. `unlabeled`
29
+ is never accepted. A placeholder word (`TODO`, `n/a`, `tbd`, `...`, `???` and the rest) counts
30
+ as missing in these fields, in the header, and in segment ids, as it does everywhere else in
31
+ the linter.
32
+ - What `lint` checks on each segments file, under test 1: the file exists and parses, its header
33
+ names the right material, every label is from the set, ids are unique, and segments never
34
+ overlap and cover every character that is not whitespace (the finding quotes the first
35
+ uncovered text and gives its offset), and the material has some text to mark. Under test 4: every segment's text
36
+ matches the material word for word, each label carries the field it needs, and the material
37
+ has not changed since it was marked (its SHA-256 still matches). A changed material fails as
38
+ stale until it is marked again.
39
+ - A spine claim may cite `material#segment`. The segment must exist (test 4), and citing a
40
+ `private` or `question` segment fails test 5. A bare material id still cites the whole
41
+ material; `material#` with nothing after the `#` fails as an unknown segment.
42
+ - Every marking finding id starts `writing-materials-` or `writing-spine-`, and every message
43
+ names the material, and the segment where there is one. Paths in messages read as the spec
44
+ wrote them, never resolved to a folder on your machine, so `--json` output is the same
45
+ everywhere. WRITING.md lists every finding with its test.
46
+ - A new export, `@supersuit/hyperspec/writing`, gives your own tools `MATERIAL_LABELS` and
47
+ `readSegments`, the same parse-and-check lint runs. `readSegments` returns
48
+ `{ header, segments, findings }` and never throws; `displayPath` and `materialDisplayPath` set
49
+ how the files are named in its messages.
50
+ - `hyperspec init --profile writing` names a segments file for its placeholder material, and the
51
+ materials check reads "every segment of every material carries a label from the closed set,
52
+ matches its source verbatim, and the markings are current".
53
+ - Both writing examples ship with every material marked. Between them they use all seven labels,
54
+ and every spine claim cites segments.
55
+ - WRITING.md gains a Marking materials section: the file format with a worked sample, the labels
56
+ and what each needs, coverage, staleness (including a line-ending conversion, which changes the
57
+ hash), citing segments, every finding with its test, and the import. README and SPEC.md point
58
+ at it.
59
+ - The repository's `.gitattributes` keeps example materials and test fixtures LF on every
60
+ checkout, so the hashes their segments files pin still match on a Windows clone.
61
+
3
62
  ## 0.3.0 (2026-09-29)
4
63
 
5
64
  The writing profile. A piece of writing can now carry a hyperspec that names everything an
package/README.md CHANGED
@@ -25,13 +25,14 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
25
25
  | `hyperspec lint <file...> [--json]` | Score each hyperspec against the nine tests. |
26
26
  | `hyperspec init <file> [--title T] [--kind K]` | Write a new hyperspec skeleton. Refuses to overwrite an existing file. |
27
27
  | `hyperspec init <file> --profile writing [--title T] [--form F] [--fiction]` | Write a writing-spec skeleton, every block shown with placeholders. |
28
+ | `hyperspec segments init <material> --id <mid> [--out F] [--by paragraph\|sentence]` | Split a material into segments to label. Refuses to overwrite an existing file. |
28
29
  | `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
29
30
  | `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
30
31
  | `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
31
32
  | `hyperspec regenerate <recipe> --out <path> --clicker <slug> <one change> [--run cmd]` | Make a child recipe from a parent and one named change, rerunning only the stages it reaches. |
32
33
  | `hyperspec compare <child-recipe> --doctor cmd` | Grade a child and its parent through one doctor against one spec. |
33
34
 
34
- Every command except `init` takes `--json`. `hyperspec --help` prints every flag.
35
+ Every command except `init` and `segments init` takes `--json`. `hyperspec --help` prints every flag.
35
36
 
36
37
  ## Exit codes
37
38
 
@@ -131,6 +132,26 @@ each rule reports under are in [WRITING.md](WRITING.md). Two complete specs that
131
132
  nothing to warn ship in `examples/writing/`: an essay for new managers, and a short story with
132
133
  two characters whose voices a judge can tell apart.
133
134
 
135
+ ### Marking materials
136
+
137
+ Every material a writing spec draws on is marked before the spec can pass: split into segments,
138
+ and each segment labeled with what it may be used as (claim, story, quote, stance, question,
139
+ aside, private). hyperspec does the splitting and the checking; an agent or a person does the
140
+ labeling.
141
+
142
+ ```bash
143
+ npx @supersuit/hyperspec segments init materials/voice-memo.md --id voice-memo
144
+ npx @supersuit/hyperspec lint essay.hyperspec.md
145
+ ```
146
+
147
+ The first writes `materials/voice-memo.md.segments.jsonl`, every segment `unlabeled`. Name that
148
+ file as `segments:` on the material item, label every segment, and `lint` checks that each label
149
+ is from the set and carries what it needs, that each segment matches the material word for word,
150
+ that the material has not changed since, and that no spine claim cites a private or question
151
+ segment. The file format, the labels, and every finding are in
152
+ [WRITING.md](WRITING.md#marking-materials). A tool of your own can run the same check with
153
+ `import { readSegments } from "@supersuit/hyperspec/writing"`.
154
+
134
155
  ## The format
135
156
 
136
157
  A hyperspec is a markdown file with a YAML frontmatter block: `decisions`, `requirements`,
package/SPEC.md CHANGED
@@ -97,7 +97,7 @@ examples:
97
97
  - path: examples/minimal.hyperspec.md
98
98
  why: the smallest spec that passes all nine tests
99
99
  resume:
100
- next_action: collect adopter issues on 0.3, the writing profile included, and cut 0.4 from them
100
+ next_action: collect adopter issues on 0.4, materials marking included, and cut 0.5 from them
101
101
  feedback:
102
102
  issues: https://github.com/SupersuitUp/hyperspec/issues
103
103
  fork: MIT; fork it for your own purposes and say so in your SPEC
@@ -109,7 +109,7 @@ improvement:
109
109
 
110
110
  A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
111
111
 
112
- **Version 0.3.0** (2026-09-29)
112
+ **Version 0.4.0** (2026-09-29)
113
113
 
114
114
  ## What makes a spec a hyperspec
115
115
 
@@ -220,6 +220,8 @@ Each row lists every condition under which `hyperspec lint` fails that test. A w
220
220
  | 8 its adopters can push back on it | `feedback.issues` or `feedback.fork` missing |
221
221
  | 9 it improves itself | `improvement.ledger` missing; a ledger path that exists and is not a readable file; if the ledger file exists, a line that is not a JSON object, a `verdict` outside one-shot, improved or not-improved, `improved` without `change`, `not-improved` without `reason`. A declared ledger that does not exist yet is a warning |
222
222
 
223
+ A profile adds its own conditions to these rows. The writing profile's are in [WRITING.md](WRITING.md#the-test-mapping), including the checks on each material's segments file: every material marked, every segment labeled from the closed set and matching its material word for word, the marking current (tests 1 and 4), and no spine claim citing a private or question segment (test 5).
224
+
223
225
  ## Exit codes
224
226
 
225
227
  `hyperspec lint` reports the worst result across every file it is given:
@@ -243,7 +245,7 @@ Silence is not a verdict. A run that learned nothing has to say so and why, and
243
245
 
244
246
  A profile adds the rules for one kind of work on top of the nine tests. A spec opts in with a top-level `profile:` naming it. A profile never adds a tenth test: every finding it raises reports under one of the nine, with an id that starts with the profile's name, and the score stays out of nine. `lint` prints one more line for a profiled spec, how many of the profile's blocks are complete. A `profile` this linter does not know is a warning under test 7, and none of its rules are checked.
245
247
 
246
- One profile ships: `writing`, for essays, chapters, letters, stories and anything else an agent drafts for a person to read. Its blocks, its fields, which test each rule reports under, and `hyperspec init --profile writing` are in [WRITING.md](WRITING.md).
248
+ One profile ships: `writing`, for essays, chapters, letters, stories and anything else an agent drafts for a person to read. Its blocks, its fields, which test each rule reports under, `hyperspec init --profile writing`, and marking materials with `hyperspec segments init` are in [WRITING.md](WRITING.md).
247
249
 
248
250
  ## Recipes
249
251
 
package/WRITING.md CHANGED
@@ -46,7 +46,7 @@ thinking out loud, `considered` for something someone has reviewed, `verified` f
46
46
  checked against its source. A brain dump is the most valuable input and the least structured, so
47
47
  marking it is the step that turns thinking into something an agent can cite. Each material is
48
48
  split into segments and each segment gets a label saying what it may be used as (see
49
- [Materials labels](#materials-labels)).
49
+ [Marking materials](#marking-materials)).
50
50
 
51
51
  ### 2. Writer DNA: who is writing, and how they sound, for this purpose
52
52
 
@@ -104,8 +104,8 @@ chapters around it, a wiki article checks its links, a text message checks bubbl
104
104
  ### 7. Spine: what it argues
105
105
 
106
106
  The kind of argument (thesis, testimony, primer, letter, story; the set is open) and the claim
107
- chain: the three to seven claims the piece has to land, in order, each pointing at the materials
108
- that support it.
107
+ chain: the three to seven claims the piece has to land, in order, each pointing at the marked
108
+ segments that support it.
109
109
 
110
110
  ### 8. Sources and claims: what lets another agent pick it up
111
111
 
@@ -159,12 +159,13 @@ writing:
159
159
  items: # at least one
160
160
  - id: voice-memo # unique across items
161
161
  path: materials/voice-memo.md
162
+ segments: materials/voice-memo.md.segments.jsonl # written by hyperspec segments init, then labeled
162
163
  produced_by: example-author
163
164
  captured: "2026-09-12"
164
165
  how: voice memo, transcribed
165
166
  trust: raw # raw | considered | verified
166
167
  check:
167
- station: every segment of every material carries a label from the closed set
168
+ station: every segment of every material carries a label from the closed set, matches its source verbatim, and the markings are current
168
169
  source: capture step
169
170
  author: agent:claude
170
171
  dna:
@@ -246,16 +247,17 @@ writing:
246
247
  claims: # 3 to 7, in order, each with its own id
247
248
  - id: c1
248
249
  text: the first one-on-one is the one meeting the report should set the agenda for
249
- materials: # ids from materials.items; voice-memo#segment is accepted
250
- - voice-memo
250
+ materials: # material#segment cites one segment; never a private or question one
251
+ - voice-memo#s3
251
252
  - id: c2
252
253
  text: status belongs in the tracker
253
- materials:
254
+ materials: # a bare material id cites the whole material
254
255
  - voice-memo
255
256
  - id: c3
256
257
  text: three questions are enough to hand the meeting over
257
258
  materials:
258
- - voice-memo
259
+ - voice-memo#s2
260
+ - voice-memo#s5
259
261
  check:
260
262
  rubric: each claim lands, in order, and nothing is argued outside the chain
261
263
  source: spine interview
@@ -317,11 +319,11 @@ Each row lists what the writing profile adds to that test. The core conditions i
317
319
 
318
320
  | Test | A writing spec fails it when |
319
321
  |---|---|
320
- | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
322
+ | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
321
323
  | 2 every requirement can fail | `goal.conditions` lists fewer than five or more than ten distinct ids, lists an id twice, or names an id that is not a top-level requirement |
322
324
  | 3 every requirement names its check | a block or a character has no `check` with a `station` or a `rubric` |
323
- | 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items` |
324
- | 5 negative space is specified | `persona.will_not_say` is empty; `persona.facts_from` is anything other than `sources` |
325
+ | 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items`, or a segment that is not in that material's segments file; a segment's text does not match its material word for word; a material changed after it was marked; a claim segment has no `source` and no `own`, a story no `teller`, a quote no `speaker` |
326
+ | 5 negative space is specified | `persona.will_not_say` is empty; `persona.facts_from` is anything other than `sources`; a spine claim cites a `private` or a `question` segment |
325
327
  | 6 examples outrank adjectives | a golden has no `why`; a material, `dna.rules`, golden or character `entity` path does not exist or is not a file; a character has no golden lines or no rejected lines, or has the same line in both (compared trimmed and case-folded) |
326
328
  | 7 a stranger can resume it | `writing.progress` exists. An unknown `profile:` is a warning |
327
329
  | 8 its adopters can push back on it | nothing further; the core rule applies |
@@ -342,25 +344,188 @@ Each row lists what the writing profile adds to that test. The core conditions i
342
344
 
343
345
  `form.name` and `spine.kind` are open: name the form and the kind of argument in your own words.
344
346
 
345
- ## Materials labels
347
+ ## Marking materials
348
+
349
+ A brain dump mixes things a draft may use with things it may not: a checked fact, an opinion, a
350
+ story from the author's own week, a line said in confidence. Marking tells them apart before an
351
+ agent drafts anything. Each material is split into segments, each segment gets one label saying
352
+ what it may be used as, and the spine cites segments, so every claim in the piece points at the
353
+ exact words that support it.
354
+
355
+ hyperspec never decides a label. `segments init` splits a material the same way every time, an
356
+ agent or a person labels each segment by editing the file it wrote, and `lint` checks everything
357
+ a rule can check: every segment carries a label from the closed set and the field that label
358
+ needs, matches its material word for word, and was marked against the material as it reads now.
359
+
360
+ A material item with no `segments` field fails test 1. Marking comes before specifying, so a
361
+ writing spec cannot pass until every material it draws on is marked.
362
+
363
+ ### Marking a material
364
+
365
+ ```bash
366
+ npx @supersuit/hyperspec segments init materials/voice-memo.md --id voice-memo
367
+ ```
368
+
369
+ `hyperspec segments init <material> --id <mid> [--out <file>] [--by paragraph|sentence]` writes
370
+ `<material>.segments.jsonl`, or the path `--out` names. `--by paragraph`, the default, makes one
371
+ segment per paragraph. `--by sentence` makes one per sentence, and a new line that opens on a list
372
+ marker (`-`, `*`, `+`, `1.` or `1)`, then a space) also starts a segment, so each bullet in a set
373
+ of notes stands on its own. Every segment starts as `unlabeled`, which lint never accepts. `init`
374
+ refuses to overwrite a file that exists, and exits 2 on a material that does not exist or a
375
+ `--by` it does not know, a material with nothing in it, and an `--out` folder that does not
376
+ exist.
377
+
378
+ Then name the file on the material item, as `segments:` beside `path:`, and label every segment.
379
+ You may also move a boundary by hand, splitting one segment in two or joining two, as long as
380
+ the rules under [Coverage](#coverage) still hold.
381
+
382
+ ### The segments file
346
383
 
347
- Before a material is used, it is split into segments, and each segment gets one of seven labels.
348
- The label decides what the segment may become in the draft.
384
+ JSON Lines: one object per line. Line 1 is a header, and every later line is one segment. For this
385
+ material, `materials/voice-memo.md`:
349
386
 
350
- | Label | Means | May be used as |
387
+ ```text
388
+ Voice memo, recorded on a walk. Raw thinking.
389
+
390
+ My first one-on-one as a manager was a disaster. I ran it from my own list.
391
+
392
+ The first one-on-one is the one meeting the report should set the agenda for.
393
+
394
+ My first manager said this to me in my second week, and I wrote it down.
395
+
396
+ "Ask what they want to talk about, then stop talking."
397
+ ```
398
+
399
+ `segments init` writes five segments, and once they are labeled the file reads:
400
+
401
+ ```jsonl
402
+ {"material":"voice-memo","path":"materials/voice-memo.md","sha256":"66c62b995a6f29c72f2a9a20c6b27deaef5d9b195bb6bab96336c06f9a8f2bfc"}
403
+ {"id":"s1","start":0,"end":45,"label":"aside","text":"Voice memo, recorded on a walk. Raw thinking."}
404
+ {"id":"s2","start":47,"end":122,"label":"story","teller":"example-author","text":"My first one-on-one as a manager was a disaster. I ran it from my own list."}
405
+ {"id":"s3","start":124,"end":201,"label":"claim","own":true,"text":"The first one-on-one is the one meeting the report should set the agenda for."}
406
+ {"id":"s4","start":203,"end":275,"label":"story","teller":"example-author","text":"My first manager said this to me in my second week, and I wrote it down."}
407
+ {"id":"s5","start":277,"end":331,"label":"quote","speaker":"the author's first manager","text":"\"Ask what they want to talk about, then stop talking.\""}
408
+ ```
409
+
410
+ The author's framing, s4, is a segment of its own, so the quote, s5, holds only the manager's
411
+ words, which is all a `quote` may hold.
412
+
413
+ - **Header.** `material` is the item's id and must match it. `path` records the material path
414
+ given to `segments init`; lint reads the material from the item's own `path`. `sha256` is the
415
+ SHA-256 of the material file's bytes when it was marked.
416
+ - **`id`** is unique within the file. `init` writes `s1`, `s2` and so on; any id works, and it is
417
+ what the spine cites.
418
+ - **`start` and `end`** are character offsets into the material's text read as UTF-8, counted as
419
+ JavaScript string indices (UTF-16 code units), with `end` exclusive.
420
+ - **`text`** is exactly the material's characters from `start` to `end`.
421
+ - **`label`**, plus the one field some labels need (below). Those fields are strings, except `own`.
422
+
423
+ ### The labels
424
+
425
+ | Label | Means | May be used as | Needs |
426
+ |---|---|---|---|
427
+ | `claim` | a statement of fact about the world | only with a source, or as the author's own claim said as such | `source`, non-empty, or `own: true` |
428
+ | `story` | something that happened, told by someone who was there | testimony, with the teller named | `teller` |
429
+ | `quote` | words someone said, verbatim | quoted exactly, never paraphrased inside quotation marks | `speaker` |
430
+ | `stance` | an opinion or conviction | the author's position | nothing more |
431
+ | `question` | something open | a prompt for the interview, never an assertion | nothing more |
432
+ | `aside` | true but off the thread | held back unless the spine needs it | nothing more |
433
+ | `private` | not for this audience | never used; kept for context | nothing more |
434
+
435
+ `own` counts when it is `true` or the string `"true"`. Any other value, `false` included, leaves
436
+ it unset, and a claim with no `source` then fails. A placeholder word such as `TODO`, `n/a` or
437
+ `???` counts as missing here as it does everywhere in a hyperspec (see [The schema](#the-schema)),
438
+ so `source: "TODO"` fails like no source at all. The same holds for the header's fields and for
439
+ segment ids.
440
+
441
+ ### Coverage
442
+
443
+ Taken in order of `start`, whatever order the lines are in, segments never overlap, and between
444
+ them they cover every character of the material that is not whitespace. Whitespace between
445
+ segments may be left out, which is what `init` does. Segment ids are unique within a file. When
446
+ text is left uncovered, the finding gives the offset of the first uncovered stretch and quotes up
447
+ to 60 characters of it. A material with no text that is not whitespace has nothing to mark and
448
+ fails test 1.
449
+
450
+ ### When a material changes
451
+
452
+ The header's `sha256` pins the material as it was when it was marked. If the material changes,
453
+ lint fails the segments file as stale (test 4), because its offsets and labels describe text that
454
+ is no longer there. Mark it again: run `segments init` with `--out` to a new file, point the
455
+ material item at it, and label every segment, carrying labels over from the old file wherever the
456
+ text did not change.
457
+
458
+ The hash is over the file's bytes, so a change nobody would call an edit still counts. Converting
459
+ line endings is the common one: a material marked with LF endings reads as stale once an editor
460
+ or a checkout setting rewrites it with CRLF. Mark a material in the line endings it will be kept
461
+ in, and if it lives in git, pin them with a `.gitattributes` line such as
462
+ `materials/** text eol=lf`.
463
+
464
+ ### Citing segments in the spine
465
+
466
+ A spine claim cites a segment as `<material>#<segment>`, such as `voice-memo#s3`. The segment has
467
+ to exist in that material's segments file (test 4). A `private` segment is never used and a
468
+ `question` is never an assertion, so a claim citing either fails test 5. A bare material id, such
469
+ as `voice-memo`, still cites the whole material; `voice-memo#`, with nothing after the `#`, is not
470
+ a bare id and fails as an unknown segment. When a claim cites a segment of a material whose
471
+ segments cannot be read at all (the material is not marked, its file is missing, or the file
472
+ holds no segments), lint says so once for that material rather than once per citation.
473
+
474
+ ### Findings
475
+
476
+ Every marking finding fails the test in its row. `<segment>` is the segment's id, or its
477
+ position when it has none; `<line>` is a line number in the segments file; `<n>` is the claim's
478
+ position in `spine.claims`, counting from 0. Every message names the material, and the segment
479
+ where there is one, and prints paths as the spec wrote them, so the output is the same on every
480
+ machine.
481
+
482
+ | Id | Test | Fails when |
351
483
  |---|---|---|
352
- | `claim` | a statement of fact about the world | only with a source, or as the author's own claim said as such |
353
- | `story` | something that happened, told by someone who was there | testimony, with the teller named |
354
- | `quote` | words someone said, verbatim | quoted exactly, never paraphrased inside quotation marks |
355
- | `stance` | an opinion or conviction | the author's position |
356
- | `question` | something open | a prompt for the interview, never an assertion |
357
- | `aside` | true but off the thread | held back unless the spine needs it |
358
- | `private` | not for this audience | never used; kept for context |
359
-
360
- The linter defines the set once, as `MATERIAL_LABELS` in `src/writing.mjs`. This release checks
361
- the materials list itself; it does not yet read segment files or check their labels. A spine
362
- claim may point at a segment as `m1#segment`, and today only the material id before the `#` is
363
- checked.
484
+ | `writing-materials-unmarked` | 1 | a material item has no `segments` field |
485
+ | `writing-materials-segments-missing` | 1 | the segments file does not exist or cannot be read |
486
+ | `writing-materials-material-missing` | 1 | the material file cannot be read. Lint reports a missing material path under test 6 instead, so this comes only from `readSegments` |
487
+ | `writing-materials-empty` | 1 | the material has no text that is not whitespace |
488
+ | `writing-materials-header` | 1 | line 1 is not a JSON object, or has no `material`, `path` or `sha256` |
489
+ | `writing-materials-header-material` | 1 | the header names a different material from the item |
490
+ | `writing-materials-json-line-<line>` | 1 | a segment line is not a JSON object |
491
+ | `writing-materials-segment-id-<line>` | 1 | a segment has no id |
492
+ | `writing-materials-segment-id` | 1 | two segments share an id |
493
+ | `writing-materials-label-<segment>` | 1 | a label outside the seven, `unlabeled` included |
494
+ | `writing-materials-segment-shape-<segment>` | 1 | `start` and `end` are not whole numbers with `start` at least 0, `end` greater than `start`, and `end` no further than the material's length |
495
+ | `writing-materials-overlap` | 1 | two segments overlap |
496
+ | `writing-materials-coverage` | 1 | text that is not whitespace lies outside every segment |
497
+ | `writing-materials-text-<segment>` | 4 | `text` is not the material's characters from `start` to `end` |
498
+ | `writing-materials-stale` | 4 | the material's SHA-256 no longer matches the header |
499
+ | `writing-materials-claim-source-<segment>` | 4 | a claim has no `source` and no `own` |
500
+ | `writing-materials-story-teller-<segment>` | 4 | a story has no `teller` |
501
+ | `writing-materials-quote-speaker-<segment>` | 4 | a quote has no `speaker` |
502
+ | `writing-spine-materials-segments-unresolvable-<material>` | 4 | a claim cites a segment of a material whose segments cannot be read |
503
+ | `writing-spine-claim-<n>-materials-segment-unknown` | 4 | a claim cites a segment that is not in the file |
504
+ | `writing-spine-claim-<n>-materials-segment-private` | 5 | a claim cites a `private` segment |
505
+ | `writing-spine-claim-<n>-materials-segment-question` | 5 | a claim cites a `question` segment |
506
+
507
+ ### Reading segments from your own tool
508
+
509
+ A tool that labels materials, such as an agent's capture step or an editor, can import the label
510
+ set and the same parse-and-check lint runs:
511
+
512
+ ```js
513
+ import { MATERIAL_LABELS, readSegments } from "@supersuit/hyperspec/writing";
514
+
515
+ const { header, segments, findings } = readSegments("materials/voice-memo.md.segments.jsonl", {
516
+ materialPath: "materials/voice-memo.md",
517
+ materialId: "voice-memo",
518
+ });
519
+ ```
520
+
521
+ `readSegments` never throws. It returns the parsed header (or `null`), every segment line that
522
+ parsed as a JSON object, and findings in the shape lint reports: `test`, `id`, `severity`,
523
+ `message` and `fix`. Without `materialPath` it runs only the checks that need no material text
524
+ (the header, ids, labels and label fields); with it, it also checks verbatim text, coverage,
525
+ overlap and staleness. `materialId`, when given, has to match the header's `material`.
526
+ Two more options, `displayPath` and `materialDisplayPath`, set how the two files are named in
527
+ messages (lint passes the paths as the spec wrote them); by default the paths are printed as
528
+ given. `MATERIAL_LABELS` is the seven labels, in the order of the table above.
364
529
 
365
530
  ## Deferring a block
366
531
 
@@ -394,7 +559,9 @@ npx @supersuit/hyperspec init story.hyperspec.md --profile writing --form "short
394
559
  The skeleton shows every required block in schema order with every field present as a `TODO`
395
560
  placeholder. `dna`, `persona`, `audience` and `goal` also carry an open decision whose question
396
561
  says what you have to answer before the placeholder means anything. `--form` sets both `kind:`
397
- and `writing.form.name`, and defaults to `essay`. `--fiction` sets `fiction: true` and adds one
562
+ and `writing.form.name`, and defaults to `essay`. The material item names
563
+ `materials/TODO.md.segments.jsonl`, the file `segments init` writes for `materials/TODO.md`, so
564
+ materials keeps failing until a real material is marked. `--fiction` sets `fiction: true` and adds one
398
565
  character with the same treatment. The skeleton never passes: it lints `fail`, with
399
566
  `writing: 1/9 blocks complete` (or `0/9` with `--fiction`), until the placeholders and the open
400
567
  decisions are replaced with real content.
@@ -415,12 +582,16 @@ names:
415
582
  Each character has speech rules, a knowledge timeline by scene, and golden and rejected lines
416
583
  in a voice you can tell apart from the other's.
417
584
 
585
+ Every material in both is marked. Between them the two examples use all seven labels, each with
586
+ the field it needs, and every spine claim cites the segments that support it. Each segments file
587
+ keeps the boundaries `segments init` wrote, in paragraph mode for prose and sentence mode for
588
+ bulleted notes, so you can re-run it and compare.
589
+
418
590
  Both lint `pass (9/9)` with `writing: 9/9 blocks complete` and no findings. A test runs them on
419
591
  every release, so they cannot drift from the linter.
420
592
 
421
593
  ## What later versions add
422
594
 
423
- This release is the schema and its lint. Later versions build on it in order: marking materials
424
- (a brain dump or transcript in, labeled segments out, with the labels above enforced), scoped
425
- DNA with annotated goldens filed by form, audience and purpose, and the stations themselves,
426
- running the checks each block names and grading drafts against the goal.
595
+ This release is the schema, its lint, and marked materials. Later versions build on it in order:
596
+ scoped DNA with annotated goldens filed by form, audience and purpose, and the stations
597
+ themselves, running the checks each block names and grading drafts against the goal.
package/bin/hyperspec.mjs CHANGED
@@ -1,5 +1,5 @@
1
1
  #!/usr/bin/env node
2
- import { existsSync, statSync, writeFileSync } from "node:fs";
2
+ import { existsSync, statSync, writeFileSync, readFileSync } from "node:fs";
3
3
  import { dirname, resolve } from "node:path";
4
4
  import { loadSpec } from "../src/load.mjs";
5
5
  import { lintSpec } from "../src/rules.mjs";
@@ -12,6 +12,8 @@ import { approve } from "../src/writer.mjs";
12
12
  import { reproduce } from "../src/reproduce.mjs";
13
13
  import { regenerate } from "../src/regenerate.mjs";
14
14
  import { compare } from "../src/compare.mjs";
15
+ import { splitSegments } from "../src/segments.mjs";
16
+ import { sha256 } from "../src/hash.mjs";
15
17
 
16
18
  const HELP = `hyperspec <command> [options]
17
19
 
@@ -30,6 +32,17 @@ const HELP = `hyperspec <command> [options]
30
32
  writing, --kind with it (use --form), or a folder that does not
31
33
  exist
32
34
 
35
+ segments init <material> --id <mid> [--out <file>] [--by paragraph|sentence]
36
+ split a material into candidate segments, written as JSONL to
37
+ <material>.segments.jsonl by default; every segment starts label:
38
+ unlabeled, never valid in lint; label each one by hand (claim,
39
+ story, quote, stance, question, aside, private), then run hyperspec
40
+ lint on the spec; --by sentence also starts a segment at each list
41
+ item (-, *, +, 1. or 1) then a space); refuses to overwrite an
42
+ existing file (exit 2); exit 2 for a missing material, a material
43
+ with nothing in it, an --out folder that does not exist, or a --by
44
+ outside paragraph/sentence
45
+
33
46
  recipe check <output-or-recipe> [--json]
34
47
  check a recipe's completeness (a path not ending .recipe.json
35
48
  means <path>.recipe.json)
@@ -98,6 +111,41 @@ if (cmd === "init") {
98
111
  process.exit(0);
99
112
  }
100
113
 
114
+ if (cmd === "segments") {
115
+ const sub = argv[1];
116
+
117
+ if (sub === "init") {
118
+ const parsed = parseArgs(argv.slice(2), { valueFlags: ["--id", "--out", "--by"] });
119
+ if (parsed.error) { console.error(parsed.error); process.exit(2); }
120
+ const [material] = parsed.positionals;
121
+ if (!material) { console.error("segments init needs a material path"); process.exit(2); }
122
+ if (!parsed.values["--id"]) { console.error("segments init needs --id <material id>"); process.exit(2); }
123
+ const id = parsed.values["--id"];
124
+ const by = parsed.values["--by"] ?? "paragraph";
125
+ if (by !== "paragraph" && by !== "sentence") { console.error(`--by must be paragraph or sentence, not "${by}"`); process.exit(2); }
126
+ let materialStat;
127
+ try { materialStat = statSync(material); } catch { materialStat = null; }
128
+ if (!materialStat || !materialStat.isFile()) { console.error(`material not found: ${material}`); process.exit(2); }
129
+ const out = parsed.values["--out"] ?? `${material}.segments.jsonl`;
130
+ if (existsSync(out)) { console.error(`refusing to overwrite ${out}`); process.exit(2); }
131
+ const outFolder = dirname(resolve(out));
132
+ if (!existsSync(outFolder) || !statSync(outFolder).isDirectory()) { console.error(`folder does not exist: ${dirname(out)}; create it first`); process.exit(2); }
133
+
134
+ const buf = readFileSync(material);
135
+ const text = buf.toString("utf8");
136
+ if (!/\S/.test(text)) { console.error(`nothing to mark: ${material} has no text`); process.exit(2); }
137
+ const segments = splitSegments(text, { by });
138
+ const header = { material: id, path: material, sha256: sha256(buf) };
139
+ const lines = [JSON.stringify(header), ...segments.map((s) => JSON.stringify(s))];
140
+ writeFileSync(out, `${lines.join("\n")}\n`);
141
+ console.log(`${segments.length} segments written to ${out}. Label every segment (claim, story, quote, stance, question, aside, private), then run hyperspec lint on the spec.`);
142
+ process.exit(0);
143
+ }
144
+
145
+ console.error(`unknown segments subcommand: ${sub}\n\n${HELP}`);
146
+ process.exit(2);
147
+ }
148
+
101
149
  if (cmd === "lint") {
102
150
  const json = argv.includes("--json");
103
151
  // Lint takes one flag, --json, and no flag takes a value, so a flag is dropped on its own and
@@ -359,7 +407,7 @@ if (cmd === "compare") {
359
407
  if (json) console.log(JSON.stringify(result, null, 2));
360
408
 
361
409
  // Both a usage error and a check-failed error (missing output, escaping path, ...) map to
362
- // 2 here — compare's own ok:false without a usage flag is still "could not grade", i.e. an
410
+ // 2 here: compare's own ok:false without a usage flag is still "could not grade", i.e. an
363
411
  // unreadable-input class failure, not a graded-but-worse-1 class one.
364
412
  if (!result.ok) {
365
413
  if (!json) console.error(result.error);
@@ -0,0 +1,10 @@
1
+ {"material":"interview","path":"essay/materials/interview-notes.md","sha256":"3c4fd3ff19435027b1bb9584e648128d59feb3184f67eeaf6b09ed9fa8bcfaeb"}
2
+ {"id":"s1","start":0,"end":118,"label":"aside","text":"Interview notes, a thirty-minute call with an engineering manager of eight years, taken by the\nauthor during the call."}
3
+ {"id":"s2","start":119,"end":199,"label":"aside","text":"Considered: the manager reviewed these notes afterwards and corrected two\nlines."}
4
+ {"id":"s3","start":201,"end":240,"label":"claim","source":"interview with an engineering manager of eight years","text":"- Her rule: the report owns the agenda."}
5
+ {"id":"s4","start":241,"end":356,"label":"claim","source":"interview with an engineering manager of eight years","text":"She keeps a shared document per person; they add items\n before the meeting, and she adds hers last, at the bottom."}
6
+ {"id":"s5","start":357,"end":448,"label":"quote","speaker":"the engineering manager interviewed","text":"- \"If I have something urgent, it is not a one-on-one topic. I send it the day it happens.\""}
7
+ {"id":"s6","start":449,"end":617,"label":"claim","source":"interview with an engineering manager of eight years","text":"- The first one-on-one with a new report is always the same: she asks how they like to receive\n feedback, in writing or out loud, right away or at the end of the week."}
8
+ {"id":"s7","start":618,"end":673,"label":"claim","source":"interview with an engineering manager of eight years","text":"- Her warning: new managers treat silence as a problem."}
9
+ {"id":"s8","start":674,"end":733,"label":"quote","speaker":"the engineering manager interviewed","text":"\"Wait. Count to five. The real answer is\n the second one.\""}
10
+ {"id":"s9","start":734,"end":812,"label":"claim","source":"interview with an engineering manager of eight years","text":"- She cancels a one-on-one only when the report asks her to, never on her own."}
@@ -0,0 +1,6 @@
1
+ {"material":"survey","path":"essay/materials/team-survey.md","sha256":"ae2c821b1731ad02eb3280c28872aa9d2163728b868f262554d0cbfd3da19b97"}
2
+ {"id":"s1","start":0,"end":139,"label":"aside","text":"Summary of an internal survey the author's team ran in the spring, 41 responses, figures checked\nagainst the raw export by a second person."}
3
+ {"id":"s2","start":140,"end":149,"label":"aside","text":"Verified."}
4
+ {"id":"s3","start":151,"end":261,"label":"claim","source":"team survey, spring, 41 responses, figures checked against the raw export","text":"- 29 of 41 said their most useful one-on-one in the last quarter was one where they brought the\n first topic."}
5
+ {"id":"s4","start":262,"end":358,"label":"claim","source":"team survey, spring, 41 responses, figures checked against the raw export","text":"- 11 of 41 said at least one of their one-on-ones in the last quarter was mostly project status."}
6
+ {"id":"s5","start":359,"end":449,"label":"claim","source":"team survey, spring, 41 responses, figures checked against the raw export","text":"- The most common free-text request, in 9 responses: \"ask me what I want to work on next.\""}
@@ -0,0 +1,7 @@
1
+ {"material":"voice-memo","path":"essay/materials/voice-memo.md","sha256":"c1e07e7d6886f0f3abe638500b614a94883b66f259d272b213f7fe63e70afdc8"}
2
+ {"id":"s1","start":0,"end":87,"label":"aside","text":"Voice memo transcript, recorded by the author on a walk, lightly cleaned. Raw thinking."}
3
+ {"id":"s2","start":89,"end":385,"label":"story","teller":"example-author","text":"My first one-on-one as a manager was a disaster, and the reason was simple: I ran it. I had a\nlist. I went down the list. Project status, blockers, the thing from Tuesday. Thirty minutes\nlater my report said \"cool, thanks\" and left, and I had learned nothing I could not have read in\nthe tracker."}
4
+ {"id":"s3","start":387,"end":572,"label":"claim","own":true,"text":"What I wish someone had told me: the first one-on-one is the only meeting where they get to set\nthe agenda. If you set it, you have told them what the meeting is for, and it is for you."}
5
+ {"id":"s4","start":574,"end":799,"label":"claim","own":true,"text":"The three questions I use now. What is taking more of your energy than it should? What do you\nwant to be doing more of in six months? What should I stop doing, or start doing, that would\nmake your week easier? Then I shut up."}
6
+ {"id":"s5","start":801,"end":897,"label":"stance","text":"Status goes in the tracker. If a one-on-one is a status meeting, cancel it and read the tracker."}
7
+ {"id":"s6","start":899,"end":1037,"label":"aside","text":"Aside, probably not for this piece: my second manager used to walk the one-on-ones outside. I\nliked it but I do not think it is the point."}