@supersuit/hyperspec 0.3.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/CHANGELOG.md +127 -0
  2. package/README.md +46 -1
  3. package/SPEC.md +5 -3
  4. package/WRITING.md +475 -40
  5. package/bin/hyperspec.mjs +157 -3
  6. package/examples/writing/dna/essay-new-managers-teach/features.json +56 -0
  7. package/examples/writing/dna/essay-new-managers-teach/goldens/README.md +14 -0
  8. package/examples/writing/dna/essay-new-managers-teach/goldens/close.md +9 -0
  9. package/examples/writing/dna/essay-new-managers-teach/goldens/opening.md +9 -0
  10. package/examples/writing/dna/essay-new-managers-teach/goldens/status.md +10 -0
  11. package/examples/writing/dna/essay-new-managers-teach/scope.md +11 -0
  12. package/examples/writing/essay/materials/interview-notes.md.segments.jsonl +10 -0
  13. package/examples/writing/essay/materials/team-survey.md.segments.jsonl +6 -0
  14. package/examples/writing/essay/materials/voice-memo.md.segments.jsonl +7 -0
  15. package/examples/writing/essay.hyperspec.md +24 -13
  16. package/examples/writing/story/materials/bakery-visit.md +1 -0
  17. package/examples/writing/story/materials/bakery-visit.md.segments.jsonl +13 -0
  18. package/examples/writing/story/materials/notes.md +2 -0
  19. package/examples/writing/story/materials/notes.md.segments.jsonl +7 -0
  20. package/examples/writing/story/materials/scene-list.md.segments.jsonl +14 -0
  21. package/examples/writing/story.hyperspec.md +8 -5
  22. package/package.json +2 -1
  23. package/src/blobs.mjs +1 -1
  24. package/src/compare.mjs +6 -6
  25. package/src/dna.mjs +471 -0
  26. package/src/fsutil.mjs +1 -1
  27. package/src/labels.mjs +6 -0
  28. package/src/reproduce.mjs +5 -5
  29. package/src/segments.mjs +407 -0
  30. package/src/writing-exports.mjs +11 -0
  31. package/src/writing-fields.mjs +279 -6
  32. package/src/writing-template.mjs +16 -1
  33. package/src/writing.mjs +4 -4
  34. package/examples/writing/essay/goldens/close.md +0 -2
  35. package/examples/writing/essay/goldens/opening.md +0 -2
package/WRITING.md CHANGED
@@ -46,7 +46,7 @@ thinking out loud, `considered` for something someone has reviewed, `verified` f
46
46
  checked against its source. A brain dump is the most valuable input and the least structured, so
47
47
  marking it is the step that turns thinking into something an agent can cite. Each material is
48
48
  split into segments and each segment gets a label saying what it may be used as (see
49
- [Materials labels](#materials-labels)).
49
+ [Marking materials](#marking-materials)).
50
50
 
51
51
  ### 2. Writer DNA: who is writing, and how they sound, for this purpose
52
52
 
@@ -61,10 +61,14 @@ whole design of this block.
61
61
  email.
62
62
  - **Every golden carries a note on why it is golden.** A golden without its reason teaches the
63
63
  surface; the reason teaches the move.
64
+ - **Features** are measured per scope from its goldens: sentence and paragraph length,
65
+ punctuation habits, pronouns, signature words. Two scopes of one writer measure differently,
66
+ and each keeps its own numbers.
64
67
 
65
- DNA is proven by a blind lineup within its scope: a judge sees a generated passage beside real
66
- goldens of the same kind and tries to pick it out. Every writer has their own DNA, and nobody's
67
- scope feeds anybody else's.
68
+ A scope is a folder on disk, and [Scoped DNA](#scoped-dna) covers it: its shape, the golden file,
69
+ what is measured, and how a spec names it. A later release proves DNA with a blind lineup within
70
+ its scope: a judge sees a generated passage beside real goldens of the same kind and tries to
71
+ pick it out. Every writer has their own DNA, and nobody's scope feeds anybody else's.
68
72
 
69
73
  ### 3. Persona: who the piece speaks as
70
74
 
@@ -104,8 +108,8 @@ chapters around it, a wiki article checks its links, a text message checks bubbl
104
108
  ### 7. Spine: what it argues
105
109
 
106
110
  The kind of argument (thesis, testimony, primer, letter, story; the set is open) and the claim
107
- chain: the three to seven claims the piece has to land, in order, each pointing at the materials
108
- that support it.
111
+ chain: the three to seven claims the piece has to land, in order, each pointing at the marked
112
+ segments that support it.
109
113
 
110
114
  ### 8. Sources and claims: what lets another agent pick it up
111
115
 
@@ -159,24 +163,26 @@ writing:
159
163
  items: # at least one
160
164
  - id: voice-memo # unique across items
161
165
  path: materials/voice-memo.md
166
+ segments: materials/voice-memo.md.segments.jsonl # written by hyperspec segments init, then labeled
162
167
  produced_by: example-author
163
168
  captured: "2026-09-12"
164
169
  how: voice memo, transcribed
165
170
  trust: raw # raw | considered | verified
166
171
  check:
167
- station: every segment of every material carries a label from the closed set
172
+ station: every segment of every material carries a label from the closed set, matches its source verbatim, and the markings are current
168
173
  source: capture step
169
174
  author: agent:claude
170
175
  dna:
171
176
  writer: example-author
177
+ scope_dir: dna/essay-new-managers-teach # optional; a scope folder (see Scoped DNA), and every golden below then lives in its goldens/
172
178
  scope:
173
179
  form: essay
174
180
  audience: new managers
175
181
  purpose: teach
176
182
  rules: style-rules.md # your style rules file, the always-on layer
177
183
  goldens: # at least one
178
- - path: goldens/opening.md
179
- why: one plain claim, then a second sentence that turns it into something to do
184
+ - path: dna/essay-new-managers-teach/goldens/opening.md
185
+ why: one plain claim in the first sentence, then two short sentences that turn it into something to do
180
186
  check:
181
187
  rubric: blind lineup within this scope
182
188
  source: goldens marked on the review page
@@ -246,16 +252,17 @@ writing:
246
252
  claims: # 3 to 7, in order, each with its own id
247
253
  - id: c1
248
254
  text: the first one-on-one is the one meeting the report should set the agenda for
249
- materials: # ids from materials.items; voice-memo#segment is accepted
250
- - voice-memo
255
+ materials: # material#segment cites one segment; never a private or question one
256
+ - voice-memo#s3
251
257
  - id: c2
252
258
  text: status belongs in the tracker
253
- materials:
259
+ materials: # a bare material id cites the whole material
254
260
  - voice-memo
255
261
  - id: c3
256
262
  text: three questions are enough to hand the meeting over
257
263
  materials:
258
- - voice-memo
264
+ - voice-memo#s2
265
+ - voice-memo#s5
259
266
  check:
260
267
  rubric: each claim lands, in order, and nothing is argued outside the chain
261
268
  source: spine interview
@@ -317,12 +324,12 @@ Each row lists what the writing profile adds to that test. The core conditions i
317
324
 
318
325
  | Test | A writing spec fails it when |
319
326
  |---|---|
320
- | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
327
+ | 1 every decision is accounted for | a required block is missing and not deferred; a required field is missing; a closed-set value is outside its set (`trust`, `reader`, `change.kind`, the shape of `identity`, `unsourced_claim`); `identity: character:<id>` names a character that is not in `writing.characters`; `fiction` is present and is anything other than `true` or `false`; two materials, two spine claims or two characters share an id; `form.length.min` or `max` is not a whole number of at least 1, or `min` is greater than `max`; `spine.claims` has fewer than three or more than seven distinct claims; a character has no knowledge entry, or an entry lacks `by` or `knows`; a material has no text, or no `segments` field, or its segments file is missing, malformed, labels a segment outside the seven (`unlabeled` included), repeats a segment id, or has segments that overlap or leave text uncovered; `dna.scope_dir` is present and is a placeholder; with `dna.scope_dir`, its `scope.md` is missing, unreadable or lacks a field, its writer, form, audience or purpose differs from the spec's, or its `goldens/` folder is missing or empty, or holds a golden that cannot be read, whose frontmatter never closes, or that has no passage. A `stance` outside the four is a warning, and so is `unsourced_claim: warn` |
321
328
  | 2 every requirement can fail | `goal.conditions` lists fewer than five or more than ten distinct ids, lists an id twice, or names an id that is not a top-level requirement |
322
329
  | 3 every requirement names its check | a block or a character has no `check` with a `station` or a `rubric` |
323
- | 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items` |
324
- | 5 negative space is specified | `persona.will_not_say` is empty; `persona.facts_from` is anything other than `sources` |
325
- | 6 examples outrank adjectives | a golden has no `why`; a material, `dna.rules`, golden or character `entity` path does not exist or is not a file; a character has no golden lines or no rejected lines, or has the same line in both (compared trimmed and case-folded) |
330
+ | 4 every field says where it came from and who wrote it | a block or a character has no `source` or no `author`; a spine claim names no materials, or names a material id that is not in `materials.items`, or a segment that is not in that material's segments file; a segment's text does not match its material word for word; a material changed after it was marked; a claim segment has no `source` and no `own`, a story no `teller`, a quote no `speaker`; with `dna.scope_dir`, a golden in the scope has no `approved_by`, an approver that starts `agent:`, or no `source` |
331
+ | 5 negative space is specified | `persona.will_not_say` is empty; `persona.facts_from` is anything other than `sources`; a spine claim cites a `private` or a `question` segment; with `dna.scope_dir`, a golden the spec lists is not one of the scope's goldens (it lives outside the scope's `goldens/` folder, is a symlink that resolves outside it, or sits in a subfolder, is `README.md` or is not a `.md` file), or the scope's `goldens/` folder resolves outside the scope |
332
+ | 6 examples outrank adjectives | a golden has no `why`; a material, `dna.rules`, golden or character `entity` path does not exist or is not a file; a character has no golden lines or no rejected lines, or has the same line in both (compared trimmed and case-folded); with `dna.scope_dir`, a golden in the scope has no `why`, or the scope's `features.json` is missing or is not what `dna measure` would write now |
326
333
  | 7 a stranger can resume it | `writing.progress` exists. An unknown `profile:` is a warning |
327
334
  | 8 its adopters can push back on it | nothing further; the core rule applies |
328
335
  | 9 it improves itself | nothing further; the core rule applies |
@@ -342,25 +349,442 @@ Each row lists what the writing profile adds to that test. The core conditions i
342
349
 
343
350
  `form.name` and `spine.kind` are open: name the form and the kind of argument in your own words.
344
351
 
345
- ## Materials labels
352
+ ## Marking materials
353
+
354
+ A brain dump mixes things a draft may use with things it may not: a checked fact, an opinion, a
355
+ story from the author's own week, a line said in confidence. Marking tells them apart before an
356
+ agent drafts anything. Each material is split into segments, each segment gets one label saying
357
+ what it may be used as, and the spine cites segments, so every claim in the piece points at the
358
+ exact words that support it.
359
+
360
+ hyperspec never decides a label. `segments init` splits a material the same way every time, an
361
+ agent or a person labels each segment by editing the file it wrote, and `lint` checks everything
362
+ a rule can check: every segment carries a label from the closed set and the field that label
363
+ needs, matches its material word for word, and was marked against the material as it reads now.
364
+
365
+ A material item with no `segments` field fails test 1. Marking comes before specifying, so a
366
+ writing spec cannot pass until every material it draws on is marked.
367
+
368
+ ### Marking a material
369
+
370
+ ```bash
371
+ npx @supersuit/hyperspec segments init materials/voice-memo.md --id voice-memo
372
+ ```
373
+
374
+ `hyperspec segments init <material> --id <mid> [--out <file>] [--by paragraph|sentence]` writes
375
+ `<material>.segments.jsonl`, or the path `--out` names. `--by paragraph`, the default, makes one
376
+ segment per paragraph. `--by sentence` makes one per sentence, and a new line that opens on a list
377
+ marker (`-`, `*`, `+`, `1.` or `1)`, then a space) also starts a segment, so each bullet in a set
378
+ of notes stands on its own. Every segment starts as `unlabeled`, which lint never accepts. `init`
379
+ refuses to overwrite a file that exists, and exits 2 on a material that does not exist or a
380
+ `--by` it does not know, a material with nothing in it, and an `--out` folder that does not
381
+ exist.
382
+
383
+ Then name the file on the material item, as `segments:` beside `path:`, and label every segment.
384
+ You may also move a boundary by hand, splitting one segment in two or joining two, as long as
385
+ the rules under [Coverage](#coverage) still hold.
386
+
387
+ ### The segments file
388
+
389
+ JSON Lines: one object per line. Line 1 is a header, and every later line is one segment. For this
390
+ material, `materials/voice-memo.md`:
391
+
392
+ ```text
393
+ Voice memo, recorded on a walk. Raw thinking.
394
+
395
+ My first one-on-one as a manager was a disaster. I ran it from my own list.
396
+
397
+ The first one-on-one is the one meeting the report should set the agenda for.
398
+
399
+ My first manager said this to me in my second week, and I wrote it down.
400
+
401
+ "Ask what they want to talk about, then stop talking."
402
+ ```
403
+
404
+ `segments init` writes five segments, and once they are labeled the file reads:
405
+
406
+ ```jsonl
407
+ {"material":"voice-memo","path":"materials/voice-memo.md","sha256":"66c62b995a6f29c72f2a9a20c6b27deaef5d9b195bb6bab96336c06f9a8f2bfc"}
408
+ {"id":"s1","start":0,"end":45,"label":"aside","text":"Voice memo, recorded on a walk. Raw thinking."}
409
+ {"id":"s2","start":47,"end":122,"label":"story","teller":"example-author","text":"My first one-on-one as a manager was a disaster. I ran it from my own list."}
410
+ {"id":"s3","start":124,"end":201,"label":"claim","own":true,"text":"The first one-on-one is the one meeting the report should set the agenda for."}
411
+ {"id":"s4","start":203,"end":275,"label":"story","teller":"example-author","text":"My first manager said this to me in my second week, and I wrote it down."}
412
+ {"id":"s5","start":277,"end":331,"label":"quote","speaker":"the author's first manager","text":"\"Ask what they want to talk about, then stop talking.\""}
413
+ ```
414
+
415
+ The author's framing, s4, is a segment of its own, so the quote, s5, holds only the manager's
416
+ words, which is all a `quote` may hold.
417
+
418
+ - **Header.** `material` is the item's id and must match it. `path` records the material path
419
+ given to `segments init`; lint reads the material from the item's own `path`. `sha256` is the
420
+ SHA-256 of the material file's bytes when it was marked.
421
+ - **`id`** is unique within the file. `init` writes `s1`, `s2` and so on; any id works, and it is
422
+ what the spine cites.
423
+ - **`start` and `end`** are character offsets into the material's text read as UTF-8, counted as
424
+ JavaScript string indices (UTF-16 code units), with `end` exclusive.
425
+ - **`text`** is exactly the material's characters from `start` to `end`.
426
+ - **`label`**, plus the one field some labels need (below). Those fields are strings, except `own`.
427
+
428
+ ### The labels
429
+
430
+ | Label | Means | May be used as | Needs |
431
+ |---|---|---|---|
432
+ | `claim` | a statement of fact about the world | only with a source, or as the author's own claim said as such | `source`, non-empty, or `own: true` |
433
+ | `story` | something that happened, told by someone who was there | testimony, with the teller named | `teller` |
434
+ | `quote` | words someone said, verbatim | quoted exactly, never paraphrased inside quotation marks | `speaker` |
435
+ | `stance` | an opinion or conviction | the author's position | nothing more |
436
+ | `question` | something open | a prompt for the interview, never an assertion | nothing more |
437
+ | `aside` | true but off the thread | held back unless the spine needs it | nothing more |
438
+ | `private` | not for this audience | never used; kept for context | nothing more |
439
+
440
+ `own` counts when it is `true` or the string `"true"`. Any other value, `false` included, leaves
441
+ it unset, and a claim with no `source` then fails. A placeholder word such as `TODO`, `n/a` or
442
+ `???` counts as missing here as it does everywhere in a hyperspec (see [The schema](#the-schema)),
443
+ so `source: "TODO"` fails like no source at all. The same holds for the header's fields and for
444
+ segment ids.
445
+
446
+ ### Coverage
447
+
448
+ Taken in order of `start`, whatever order the lines are in, segments never overlap, and between
449
+ them they cover every character of the material that is not whitespace. Whitespace between
450
+ segments may be left out, which is what `init` does. Segment ids are unique within a file. When
451
+ text is left uncovered, the finding gives the offset of the first uncovered stretch and quotes up
452
+ to 60 characters of it. A material with no text that is not whitespace has nothing to mark and
453
+ fails test 1.
454
+
455
+ ### When a material changes
456
+
457
+ The header's `sha256` pins the material as it was when it was marked. If the material changes,
458
+ lint fails the segments file as stale (test 4), because its offsets and labels describe text that
459
+ is no longer there. Mark it again: run `segments init` with `--out` to a new file, point the
460
+ material item at it, and label every segment, carrying labels over from the old file wherever the
461
+ text did not change.
462
+
463
+ The hash is over the file's bytes, so a change nobody would call an edit still counts. Converting
464
+ line endings is the common one: a material marked with LF endings reads as stale once an editor
465
+ or a checkout setting rewrites it with CRLF. Mark a material in the line endings it will be kept
466
+ in, and if it lives in git, pin them with a `.gitattributes` line such as
467
+ `materials/** text eol=lf`.
468
+
469
+ ### Citing segments in the spine
470
+
471
+ A spine claim cites a segment as `<material>#<segment>`, such as `voice-memo#s3`. The segment has
472
+ to exist in that material's segments file (test 4). A `private` segment is never used and a
473
+ `question` is never an assertion, so a claim citing either fails test 5. A bare material id, such
474
+ as `voice-memo`, still cites the whole material; `voice-memo#`, with nothing after the `#`, is not
475
+ a bare id and fails as an unknown segment. When a claim cites a segment of a material whose
476
+ segments cannot be read at all (the material is not marked, its file is missing, or the file
477
+ holds no segments), lint says so once for that material rather than once per citation.
478
+
479
+ ### Findings
480
+
481
+ Every marking finding fails the test in its row. `<segment>` is the segment's id, or its
482
+ position when it has none; `<line>` is a line number in the segments file; `<n>` is the claim's
483
+ position in `spine.claims`, counting from 0. Every message names the material, and the segment
484
+ where there is one, and prints paths as the spec wrote them, so the output is the same on every
485
+ machine.
486
+
487
+ | Id | Test | Fails when |
488
+ |---|---|---|
489
+ | `writing-materials-unmarked` | 1 | a material item has no `segments` field |
490
+ | `writing-materials-segments-missing` | 1 | the segments file does not exist or cannot be read |
491
+ | `writing-materials-material-missing` | 1 | the material file cannot be read. Lint reports a missing material path under test 6 instead, so this comes only from `readSegments` |
492
+ | `writing-materials-empty` | 1 | the material has no text that is not whitespace |
493
+ | `writing-materials-header` | 1 | line 1 is not a JSON object, or has no `material`, `path` or `sha256` |
494
+ | `writing-materials-header-material` | 1 | the header names a different material from the item |
495
+ | `writing-materials-json-line-<line>` | 1 | a segment line is not a JSON object |
496
+ | `writing-materials-segment-id-<line>` | 1 | a segment has no id |
497
+ | `writing-materials-segment-id` | 1 | two segments share an id |
498
+ | `writing-materials-label-<segment>` | 1 | a label outside the seven, `unlabeled` included |
499
+ | `writing-materials-segment-shape-<segment>` | 1 | `start` and `end` are not whole numbers with `start` at least 0, `end` greater than `start`, and `end` no further than the material's length |
500
+ | `writing-materials-overlap` | 1 | two segments overlap |
501
+ | `writing-materials-coverage` | 1 | text that is not whitespace lies outside every segment |
502
+ | `writing-materials-text-<segment>` | 4 | `text` is not the material's characters from `start` to `end` |
503
+ | `writing-materials-stale` | 4 | the material's SHA-256 no longer matches the header |
504
+ | `writing-materials-claim-source-<segment>` | 4 | a claim has no `source` and no `own` |
505
+ | `writing-materials-story-teller-<segment>` | 4 | a story has no `teller` |
506
+ | `writing-materials-quote-speaker-<segment>` | 4 | a quote has no `speaker` |
507
+ | `writing-spine-materials-segments-unresolvable-<material>` | 4 | a claim cites a segment of a material whose segments cannot be read |
508
+ | `writing-spine-claim-<n>-materials-segment-unknown` | 4 | a claim cites a segment that is not in the file |
509
+ | `writing-spine-claim-<n>-materials-segment-private` | 5 | a claim cites a `private` segment |
510
+ | `writing-spine-claim-<n>-materials-segment-question` | 5 | a claim cites a `question` segment |
511
+
512
+ ### Reading segments from your own tool
513
+
514
+ A tool that labels materials, such as an agent's capture step or an editor, can import the label
515
+ set and the same parse-and-check lint runs:
516
+
517
+ ```js
518
+ import { MATERIAL_LABELS, readSegments } from "@supersuit/hyperspec/writing";
519
+
520
+ const { header, segments, findings } = readSegments("materials/voice-memo.md.segments.jsonl", {
521
+ materialPath: "materials/voice-memo.md",
522
+ materialId: "voice-memo",
523
+ });
524
+ ```
525
+
526
+ `readSegments` never throws. It returns the parsed header (or `null`), every segment line that
527
+ parsed as a JSON object, and findings in the shape lint reports: `test`, `id`, `severity`,
528
+ `message` and `fix`. Without `materialPath` it runs only the checks that need no material text
529
+ (the header, ids, labels and label fields); with it, it also checks verbatim text, coverage,
530
+ overlap and staleness. `materialId`, when given, has to match the header's `material`.
531
+ Two more options, `displayPath` and `materialDisplayPath`, set how the two files are named in
532
+ messages (lint passes the paths as the spec wrote them); by default the paths are printed as
533
+ given. `MATERIAL_LABELS` is the seven labels, in the order of the table above.
534
+
535
+ ## Scoped DNA
536
+
537
+ A writer does not have one voice, so hyperspec does not keep one. A writer's DNA is kept per
538
+ **scope**, a form, an audience and a purpose together, and each scope is a folder holding its own
539
+ goldens and its own measurements.
540
+
541
+ The reason is a leak. A passage can be exactly right for one kind of writing and wrong for
542
+ another. The short, warm sentences that make a text message land read as thin in a theology
543
+ essay, and the long, qualified sentences that make the essay careful read as evasive on a landing
544
+ page. Pool every golden a writer has into one set and an agent learns the moves of each kind of
545
+ writing and carries them into the others. Filed by scope, a golden feeds only work that shares
546
+ its scope, so the moves it teaches stay where they are right. Retrieval is by scope, never by
547
+ "best writing overall".
548
+
549
+ Scoped DNA is optional in this release. A spec that names no scope folder lints exactly as it did
550
+ in 0.4.
551
+
552
+ ### The scope folder
553
+
554
+ ```text
555
+ dna/essay-new-managers-teach/
556
+ scope.md writer, form, audience, purpose, optional notes
557
+ goldens/
558
+ README.md the golden file shape; never read as a golden
559
+ close.md one golden per file
560
+ opening.md
561
+ status.md
562
+ features.json written by dna measure, never by hand
563
+ ```
564
+
565
+ `scope.md` carries the scope in its frontmatter; its body is free text for people:
566
+
567
+ ```markdown
568
+ ---
569
+ writer: example-author
570
+ form: essay
571
+ audience: new managers
572
+ purpose: teach
573
+ ---
574
+ ```
575
+
576
+ `writer`, `form`, `audience` and `purpose` are required, and `notes` is optional. Name the folder
577
+ after its scope so a person can tell scopes apart at a glance. hyperspec reads the scope from
578
+ `scope.md`, never from the folder's name.
579
+
580
+ `goldens/` must be a real folder inside the scope. A `goldens/` that resolves somewhere else, such
581
+ as a symlink to another scope's goldens, would carry that scope's passages into this one under
582
+ this scope's name, so `dna measure` refuses it and lint fails it under test 5, both naming where it
583
+ leads. A whole scope folder reached through a symlink is fine, because `scope.md` travels with
584
+ it.
585
+
586
+ ### A golden
587
+
588
+ A golden is a real passage the writer marked as right, one per file in `goldens/`. Every `.md`
589
+ file directly in `goldens/` is a golden except `README.md`, which is for notes to people and is
590
+ never read as a golden. A subfolder, a file with another extension, and a symlink sitting in the
591
+ folder are not read either. This is the essay example's opening:
592
+
593
+ ```markdown
594
+ ---
595
+ why: one plain claim in the first sentence, then two short sentences that turn it into something to do
596
+ approved_by: example-author
597
+ source: first draft of this essay's opening paragraph, marked golden on the review page
598
+ approved_on: "2026-09-18"
599
+ ---
600
+
601
+ Your first one-on-one with a new report is the only meeting on your calendar where they should
602
+ set the agenda. Everything else you run. This one you hand over.
603
+ ```
604
+
605
+ - **`why`** (required) names the move the passage teaches. A golden without its reason teaches
606
+ the surface: an agent copies its length, its words and its rhythm. The reason teaches the move,
607
+ which carries over to a passage that shares none of those.
608
+ - **`approved_by`** (required) is the person who approved it, as a slug. Golden means a human
609
+ approved it, so an approver that starts `agent:` is refused. An agent may propose a golden;
610
+ only a person makes one.
611
+ - **`source`** (required) says where the passage came from: a draft, an earlier piece, a review
612
+ page. It lets someone check that the passage is real and find the context it was written in.
613
+ - **`approved_on`** (optional) is the date it was approved.
614
+ - **The body** is the passage, verbatim. Whitespace before and after it is dropped, and nothing
615
+ inside it is changed.
616
+
617
+ A placeholder counts as missing in every one of these fields, as it does everywhere in a
618
+ hyperspec (see [The schema](#the-schema)).
619
+
620
+ ### Starting a scope
621
+
622
+ ```bash
623
+ mkdir -p dna
624
+ npx @supersuit/hyperspec dna init dna/essay-new-managers-teach --writer example-author --form essay --audience "new managers" --purpose teach
625
+ ```
626
+
627
+ `hyperspec dna init <scope-dir> --writer W --form F --audience A --purpose P` writes `scope.md`
628
+ and a `goldens/` folder holding only a README on the golden file shape. All four flags are
629
+ required. It refuses to overwrite an existing `scope.md`, never replaces a `goldens/README.md`
630
+ that is already there, and exits 2 with a plain message on a missing flag, a flag whose value is a
631
+ placeholder, or a scope folder whose parent folder does not exist. Then add one file per golden.
632
+
633
+ ### Measuring a scope
634
+
635
+ ```bash
636
+ npx @supersuit/hyperspec dna measure dna/essay-new-managers-teach
637
+ ```
638
+
639
+ ```
640
+ dna/essay-new-managers-teach: measured 3 goldens
641
+ word_count 118, sentence length mean 13.111 median 15 p90 25
642
+ signature words: first, report, tracker
643
+ wrote dna/essay-new-managers-teach/features.json
644
+ ```
645
+
646
+ `hyperspec dna measure <scope-dir> [--json]` reads every golden, checks each one's own fields, and
647
+ writes `<scope-dir>/features.json`. If the scope or any golden fails a check (no `why`, an agent
648
+ approver, no passage, a `goldens/` folder that resolves outside the scope), it prints the findings,
649
+ writes nothing and exits 1, so a hollow or borrowed golden is never measured into the DNA. It exits 0 when it wrote the file and 2 on a usage error. `--json`
650
+ prints the same result as JSON. The same goldens always produce the same bytes.
651
+
652
+ `features.json` holds `dna` (the version of this format, `"0.1"`), `scope` (the four fields from
653
+ `scope.md`), `goldens` (each golden's path inside the folder and the SHA-256 of its file, sorted
654
+ by path) and `features`.
655
+
656
+ ### What is measured
346
657
 
347
- Before a material is used, it is split into segments, and each segment gets one of seven labels.
348
- The label decides what the segment may become in the draft.
658
+ Every feature is a count or a ratio computed from the goldens' text. None of them is a judgment:
659
+ a number says how the writer writes in this scope, never whether the writing is good, and
660
+ hyperspec calls no model to get it. Words are pooled across every golden in the scope, so their
661
+ order changes nothing. Paragraphs and sentences are split exactly as `segments init` splits them,
662
+ so the two never disagree about where a boundary falls. A word is a run of letters, digits and
663
+ apostrophes, lowercased. Every number is rounded to three decimal places, and a rate is per 1000
664
+ words.
349
665
 
350
- | Label | Means | May be used as |
666
+ | Feature | What it counts | What it is for |
351
667
  |---|---|---|
352
- | `claim` | a statement of fact about the world | only with a source, or as the author's own claim said as such |
353
- | `story` | something that happened, told by someone who was there | testimony, with the teller named |
354
- | `quote` | words someone said, verbatim | quoted exactly, never paraphrased inside quotation marks |
355
- | `stance` | an opinion or conviction | the author's position |
356
- | `question` | something open | a prompt for the interview, never an assertion |
357
- | `aside` | true but off the thread | held back unless the spine needs it |
358
- | `private` | not for this audience | never used; kept for context |
359
-
360
- The linter defines the set once, as `MATERIAL_LABELS` in `src/writing.mjs`. This release checks
361
- the materials list itself; it does not yet read segment files or check their labels. A spine
362
- claim may point at a segment as `m1#segment`, and today only the material id before the `#` is
363
- checked.
668
+ | `word_count` | words across every golden | how much text the other numbers rest on; a scope of a few dozen words measures loosely |
669
+ | `sentence_length` | words per sentence: `mean`, `median` and `p90` (nearest rank) | the writer's usual sentence, and how long their long ones run, which a mean hides |
670
+ | `paragraph_length` | per paragraph, the mean number of sentences (`mean_sentences`) and of words (`mean_words`) | how much the writer puts in one block before a break |
671
+ | `rates_per_1000_words` | commas, semicolons, colons, em dashes, en dashes, exclamation marks, question marks, parentheses (each one counted) and double quotation marks, straight or curly | punctuation habits, which carry much of how a voice sounds |
672
+ | `contraction_rate` | words with an apostrophe between two letters | how conversational the writer is in this scope |
673
+ | `first_person_singular_rate` | I, me, my, mine, myself | how much the writer speaks as themselves |
674
+ | `first_person_plural_rate` | we, us, our, ours, ourselves | how much the writer speaks as a group, or alongside the reader |
675
+ | `second_person_rate` | you, your, yours, yourself, yourselves | how directly the writer addresses the reader |
676
+ | `mean_word_length` | characters per word | plain words or long ones |
677
+ | `signature_words` | up to 15 words of four or more letters that are not common function words and appear at least twice, most frequent first, ties in alphabetical order | the vocabulary the writer returns to in this scope |
678
+
679
+ ### When the scope changes
680
+
681
+ `features.json` is current only when it is exactly what `dna measure` would write from the scope
682
+ as it reads now: the same goldens, pinned by the SHA-256 of each file; the same four fields as
683
+ `scope.md`; the format version `"0.1"`; and the same numbers. Add a golden, remove one, change any
684
+ byte of one (its passage or its frontmatter), edit `scope.md`, or edit a number by hand, and lint
685
+ fails the scope as stale under test 6. The finding names what differs: each golden added, removed
686
+ or changed, each scope field that changed, an unknown version, or each feature whose number no
687
+ longer matches a fresh measurement. Run `dna measure` again. The hash is over bytes, so a
688
+ line-ending conversion counts as a change, as it does for a segments file (see
689
+ [When a material changes](#when-a-material-changes)).
690
+
691
+ ### Naming the scope in a spec
692
+
693
+ `writing.dna.scope_dir` points a writing spec at its scope folder, relative to the spec like every
694
+ other path. The essay example's `dna` block:
695
+
696
+ ```yaml
697
+ dna:
698
+ writer: example-author
699
+ scope_dir: dna/essay-new-managers-teach
700
+ scope:
701
+ form: essay
702
+ audience: new managers
703
+ purpose: teach
704
+ rules: style-rules.md
705
+ goldens:
706
+ - path: dna/essay-new-managers-teach/goldens/opening.md
707
+ why: one plain claim in the first sentence, then two short sentences that turn it into something to do
708
+ ```
709
+
710
+ `scope_dir` is optional. Without it, `dna` lints exactly as it did in 0.4: each golden the spec
711
+ lists needs a path to a file and a `why`, and no scope folder is read. With it, lint also checks
712
+ that:
713
+
714
+ - `scope.md`'s writer equals `dna.writer`, and its form, audience and purpose equal `dna.scope`,
715
+ compared trimmed and ignoring case (test 1);
716
+ - every golden the spec lists is one of the scope's goldens: after following any symlink, a `.md`
717
+ file directly in the scope's own `goldens/` folder, other than `README.md` (test 5). A golden
718
+ from another scope is a leak, the exact thing a scope exists to prevent, and so is a passage in
719
+ a subfolder, in `README.md` or in another kind of file, which would feed the spec without ever
720
+ being checked or measured;
721
+ - every golden in the folder, listed in the spec or not, has a `why` (test 6), an `approved_by`
722
+ that names a person and a `source` (test 4), and a passage (test 1);
723
+ - `features.json` exists and is current, as [When the scope changes](#when-the-scope-changes)
724
+ defines it (test 6).
725
+
726
+ The spec still gives each golden it lists a `why`, as in 0.4; the essay example keeps it the same
727
+ as the golden file's own. A `scope_dir` that is present but a placeholder, such as `TODO`, fails
728
+ test 1 on its own, so a skeleton cannot pass by leaving it unfilled.
729
+
730
+ ### Findings
731
+
732
+ Every scoped-DNA finding id starts `writing-dna-` and fails the test in its row. Messages name the
733
+ scope folder as the spec wrote it (or as it was given to `dna measure`) and each golden by its
734
+ path inside the folder, so the output is the same on every machine. `<field>` is `writer`,
735
+ `form`, `audience` or `purpose`.
736
+
737
+ | Id | Test | Fails when |
738
+ |---|---|---|
739
+ | `writing-dna-scope-dir` | 1 | `writing.dna.scope_dir` is present and is a placeholder |
740
+ | `writing-dna-scope-<field>` | 1 | the spec's own `writing.dna.scope` has no `form`, `audience` or `purpose` (this check runs with or without `scope_dir`, and is the id 0.4 used) |
741
+ | `writing-dna-scope-missing` | 1 | `scope.md` does not exist, cannot be read, or its frontmatter does not parse |
742
+ | `writing-dna-scope-file-<field>` | 1 | `scope.md` has no such field |
743
+ | `writing-dna-scope-mismatch-<field>` | 1 | `scope.md` and the spec disagree on that field |
744
+ | `writing-dna-goldens-missing` | 1 | the `goldens/` folder does not exist or cannot be read |
745
+ | `writing-dna-goldens-empty` | 1 | `goldens/` holds no golden |
746
+ | `writing-dna-goldens-outside` | 5 | `goldens/` resolves to a folder outside the scope, such as a symlink to another scope's goldens |
747
+ | `writing-dna-golden-unreadable` | 1 | a golden file cannot be read |
748
+ | `writing-dna-golden-frontmatter` | 1 | a golden's frontmatter opens with `---` and never closes |
749
+ | `writing-dna-golden-empty` | 1 | a golden has no passage |
750
+ | `writing-dna-golden-approved-by` | 4 | a golden has no `approved_by` |
751
+ | `writing-dna-golden-approved-by-agent` | 4 | a golden's `approved_by` starts `agent:`, in any case |
752
+ | `writing-dna-golden-source` | 4 | a golden has no `source` |
753
+ | `writing-dna-golden-leak` | 5 | a golden the spec lists is not one of the scope's goldens: it lives outside the scope's `goldens/` folder, is a symlink that resolves outside it, or sits in a subfolder, is `README.md` or is not a `.md` file |
754
+ | `writing-dna-golden-why` | 6 | a golden has no `why` |
755
+ | `writing-dna-features-missing` | 6 | the scope has no `features.json`, or it is not valid JSON |
756
+ | `writing-dna-features-stale` | 6 | `features.json` is not what `dna measure` would write now: a golden was added, removed or changed, `scope.md` changed, the version is unknown, or a number differs from a fresh measurement |
757
+
758
+ `dna measure` raises the ids that come from the folder alone: every row except `scope-dir`,
759
+ `scope-<field>`, `scope-mismatch-<field>`, `golden-leak` and the two `features-` rows, which
760
+ need a spec to compare against. Lint raises all of them.
761
+
762
+ ### Reading a scope from your own tool
763
+
764
+ A tool of your own, such as a review page that files goldens, can read a scope and measure it the
765
+ way `dna measure` does:
766
+
767
+ ```js
768
+ import { readScope, measureFeatures } from "@supersuit/hyperspec/writing";
769
+
770
+ const { scope, goldens, findings } = readScope("dna/essay-new-managers-teach");
771
+ const features = measureFeatures(goldens.map((g) => g.text));
772
+ ```
773
+
774
+ `readScope` never throws for a folder path, whatever is or is not in the folder. It returns the scope's four fields and `notes` (or `null` when
775
+ `scope.md` cannot be read at all), every golden it could read, with its `path`, `why`,
776
+ `approved_by`, `source`, `approved_on`, `text` and `sha256`, and findings in the shape lint
777
+ reports. `displayDir` sets how the folder is named in messages. `measureFeatures` takes an array
778
+ of passages and returns the `features` object `dna measure` writes. It reads no file and returns
779
+ the same object for the same passages.
780
+
781
+ ### The worked example
782
+
783
+ The essay in [`examples/writing/`](examples/writing/) takes its voice from
784
+ `dna/essay-new-managers-teach/`: three goldens, each with its `why`, a person's approval and its
785
+ source, and a `features.json` that `dna measure` wrote. A test measures the folder again on every
786
+ release and requires the same bytes, so the example cannot drift from the tool. The short story
787
+ beside it lists its goldens in the spec with no scope folder, the 0.4 shape, which still passes.
364
788
 
365
789
  ## Deferring a block
366
790
 
@@ -394,8 +818,12 @@ npx @supersuit/hyperspec init story.hyperspec.md --profile writing --form "short
394
818
  The skeleton shows every required block in schema order with every field present as a `TODO`
395
819
  placeholder. `dna`, `persona`, `audience` and `goal` also carry an open decision whose question
396
820
  says what you have to answer before the placeholder means anything. `--form` sets both `kind:`
397
- and `writing.form.name`, and defaults to `essay`. `--fiction` sets `fiction: true` and adds one
398
- character with the same treatment. The skeleton never passes: it lints `fail`, with
821
+ and `writing.form.name`, and defaults to `essay`. The material item names
822
+ `materials/TODO.md.segments.jsonl`, the file `segments init` writes for `materials/TODO.md`, so
823
+ materials keeps failing until a real material is marked. `dna` shows `scope_dir: TODO`, which
824
+ fails until it names a scope folder (see [Scoped DNA](#scoped-dna)) or is deleted, since the
825
+ field is optional. `--fiction` sets `fiction: true` and adds one character with the same
826
+ treatment. The skeleton never passes: it lints `fail`, with
399
827
  `writing: 1/9 blocks complete` (or `0/9` with `--fiction`), until the placeholders and the open
400
828
  decisions are replaced with real content.
401
829
 
@@ -410,17 +838,24 @@ Two complete specs ship in [`examples/writing/`](examples/writing/), each with e
410
838
  names:
411
839
 
412
840
  - `essay.hyperspec.md`: an essay for new managers on running a first one-on-one. Three materials
413
- at three trust levels, scoped DNA with two annotated goldens, a four-claim spine.
841
+ at three trust levels, its voice from the scope folder `dna/essay-new-managers-teach/` with three
842
+ annotated goldens and their measured features, a four-claim spine.
414
843
  - `story.hyperspec.md`: a short story, `fiction: true`, narrated by one of its two characters.
415
844
  Each character has speech rules, a knowledge timeline by scene, and golden and rejected lines
416
845
  in a voice you can tell apart from the other's.
417
846
 
847
+ Every material in both is marked. Between them the two examples use all seven labels, each with
848
+ the field it needs, and every spine claim cites the segments that support it. Each segments file
849
+ keeps the boundaries `segments init` wrote, in paragraph mode for prose and sentence mode for
850
+ bulleted notes, so you can re-run it and compare.
851
+
418
852
  Both lint `pass (9/9)` with `writing: 9/9 blocks complete` and no findings. A test runs them on
419
853
  every release, so they cannot drift from the linter.
420
854
 
421
855
  ## What later versions add
422
856
 
423
- This release is the schema and its lint. Later versions build on it in order: marking materials
424
- (a brain dump or transcript in, labeled segments out, with the labels above enforced), scoped
425
- DNA with annotated goldens filed by form, audience and purpose, and the stations themselves,
426
- running the checks each block names and grading drafts against the goal.
857
+ This release is the schema, its lint, marked materials, and scoped DNA with measured features.
858
+ Later versions build on it in order. The first compares a draft against its scope: its features
859
+ beside the scope's features, and a blind lineup in which a judge sees a generated passage among
860
+ the scope's goldens and tries to pick it out. Then the stations themselves, running the checks
861
+ each block names and grading drafts against the goal.