@supersuit/hyperspec 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +42 -0
- package/README.md +74 -2
- package/SPEC.md +190 -10
- package/bin/hyperspec.mjs +277 -0
- package/examples/recipe/doctor.mjs +11 -0
- package/examples/recipe/essay.hyperspec.md +50 -0
- package/examples/recipe/factory.mjs +29 -0
- package/examples/recipe/materials/call-2.md +2 -0
- package/examples/recipe/materials/call.md +3 -0
- package/examples/recipe/materials/notes.md +3 -0
- package/examples/recipe/runner.mjs +20 -0
- package/examples/recipe/runs.jsonl +0 -0
- package/examples/recipe/stages.mjs +31 -0
- package/package.json +6 -1
- package/runs.jsonl +0 -0
- package/src/blobs.mjs +77 -0
- package/src/compare.mjs +189 -0
- package/src/fsutil.mjs +33 -0
- package/src/hash.mjs +26 -0
- package/src/recipe.mjs +164 -0
- package/src/regenerate.mjs +318 -0
- package/src/reproduce.mjs +156 -0
- package/src/rules.mjs +33 -6
- package/src/template.mjs +12 -3
- package/src/writer.mjs +125 -0
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,47 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.2.0 (2026-09-28)
|
|
4
|
+
|
|
5
|
+
Recipes. Every output a factory makes can now carry a recipe beside it: what made it, from
|
|
6
|
+
what, and who approved it, with every input and every intermediate step kept as its exact
|
|
7
|
+
bytes. From a recipe you can check that an output is still exactly what was made, make it
|
|
8
|
+
again with one more ingredient while reusing every step that ingredient does not reach, and
|
|
9
|
+
grade the new output against the old one, so "the new one is better" is a number. hyperspec
|
|
10
|
+
calls no model to do any of this: the steps that need one are commands you supply.
|
|
11
|
+
|
|
12
|
+
- `hyperspec recipe check <output-or-recipe>` reports what a recipe is missing: the factory
|
|
13
|
+
version, the spec hash and who wrote each of its fields, an input's hash, a stage's verdict,
|
|
14
|
+
the clicker, the approver. `hyperspec recipe approve <recipe> --by <slug>` records the
|
|
15
|
+
approver.
|
|
16
|
+
- `hyperspec reproduce <recipe> [--restore]` re-checks every recorded hash and names the
|
|
17
|
+
first that fails. `--restore` rewrites the output from its stored bytes.
|
|
18
|
+
- `hyperspec regenerate <recipe> --out <path> --clicker <slug>` takes one change,
|
|
19
|
+
`--add-input name=path [--reads stage]...`, `--swap-input name=path` or
|
|
20
|
+
`--factory-version v`, reruns only the stages it reaches through `--run <runner>`, reuses
|
|
21
|
+
the rest, and writes a child recipe naming its parent and the change. Without `--run` the
|
|
22
|
+
child is written with those stages pending, and the command exits 3. A runner that prints
|
|
23
|
+
no `VERDICT` line fails that stage. An added input that no stage reads is refused.
|
|
24
|
+
- `hyperspec compare <child-recipe> --doctor <command>` grades the child and its parent with
|
|
25
|
+
one doctor against one spec, exits 1 on a regression naming the change as the suspect, and
|
|
26
|
+
appends the result to the spec's improvement ledger. It refuses an output file that no
|
|
27
|
+
longer matches its recipe, so a hand edit is never scored as the change's doing.
|
|
28
|
+
- A recorded hash must be 64 lowercase hex characters. Anything else names no blob, so a
|
|
29
|
+
crafted recipe cannot point a read outside the store.
|
|
30
|
+
- `@supersuit/hyperspec/recipe` exports `startRecipe` and `approve`, which a factory calls as
|
|
31
|
+
it runs to record inputs and stages and write the recipe.
|
|
32
|
+
- The package now declares `exports`, so `@supersuit/hyperspec/recipe` and
|
|
33
|
+
`@supersuit/hyperspec/package.json` are the only paths you can import. Deep imports of
|
|
34
|
+
`src/` files, which resolved in 0.1.0, no longer do.
|
|
35
|
+
- `examples/recipe/` is a worked factory, runner and doctor. The README walks the full loop,
|
|
36
|
+
and a test runs that walkthrough exactly as written, so the two cannot drift apart.
|
|
37
|
+
- SPEC.md gains a Recipes section: the recipe file, the stage key, the blob store, each
|
|
38
|
+
command's contract, the runner and doctor contracts, and every exit code.
|
|
39
|
+
- `lint` fails an id used twice across decisions and requirements (test 1), and an example
|
|
40
|
+
that is a folder or is the spec itself (test 6). It warns on a `hyperspec` version it does
|
|
41
|
+
not know (test 7) and on a declared ledger that does not exist yet (test 9).
|
|
42
|
+
- `hyperspec init` quotes a title or kind that a YAML reader would read back differently.
|
|
43
|
+
- The README's exit-code sentence names every case that exits 2, matching SPEC.md.
|
|
44
|
+
|
|
3
45
|
## 0.1.0 (2026-09-28)
|
|
4
46
|
|
|
5
47
|
- The hyperspecification standard: a markdown file with a YAML frontmatter block, versioned
|
package/README.md
CHANGED
|
@@ -24,12 +24,84 @@ improvement ledger. Every test is defined in [SPEC.md](SPEC.md).
|
|
|
24
24
|
|---|---|
|
|
25
25
|
| `hyperspec lint <file...> [--json]` | Score each hyperspec against the nine tests. |
|
|
26
26
|
| `hyperspec init <file> [--title T] [--kind K]` | Write a new hyperspec skeleton. Refuses to overwrite an existing file. |
|
|
27
|
+
| `hyperspec recipe check <output-or-recipe>` | Check that a recipe records everything the standard asks for. |
|
|
28
|
+
| `hyperspec recipe approve <recipe> --by <slug>` | Record who approved the output. |
|
|
29
|
+
| `hyperspec reproduce <recipe> [--restore]` | Re-check every hash the recipe recorded. Never runs a model. |
|
|
30
|
+
| `hyperspec regenerate <recipe> --out <path> --clicker <slug> <one change> [--run cmd]` | Make a child recipe from a parent and one named change, rerunning only the stages it reaches. |
|
|
31
|
+
| `hyperspec compare <child-recipe> --doctor cmd` | Grade a child and its parent through one doctor against one spec. |
|
|
32
|
+
|
|
33
|
+
Every command except `init` takes `--json`. `hyperspec --help` prints every flag.
|
|
27
34
|
|
|
28
35
|
## Exit codes
|
|
29
36
|
|
|
30
37
|
`hyperspec lint` exits 0 when every test passes and nothing is open, 1 when at least one
|
|
31
|
-
test fails,
|
|
32
|
-
|
|
38
|
+
test fails, 3 when every test passes but a decision is still open (blocked), and 2 on a usage
|
|
39
|
+
error or a file that cannot be read, has broken frontmatter, or is not a hyperspec.
|
|
40
|
+
|
|
41
|
+
The recipe commands use the same numbers: 0 ok, 1 a check failed or the child regressed, 2
|
|
42
|
+
usage or unreadable input, 3 pending, when `regenerate` has stages waiting for a runner.
|
|
43
|
+
`regenerate` also exits 1 when a stage it ran reported a failing verdict, or none. A stage it
|
|
44
|
+
reused keeps its parent's verdict and does not change the exit code.
|
|
45
|
+
|
|
46
|
+
## Recipes
|
|
47
|
+
|
|
48
|
+
A hyperspec says what the work must be. A recipe says what one piece of work was made from.
|
|
49
|
+
Every output a factory makes can carry one beside it, `<output>.recipe.json`: the factory and
|
|
50
|
+
its version, the spec and who wrote each of its fields, every input by path and SHA-256 hash,
|
|
51
|
+
every stage with what it read and its verdict, and who pressed go and who approved. Inputs and
|
|
52
|
+
stage outputs are kept by content in `.hyperspec/blobs/`, so a file edited next month cannot
|
|
53
|
+
change what an old recipe reproduces.
|
|
54
|
+
|
|
55
|
+
A recipe runs again three ways. `reproduce` re-checks every recorded hash and never runs a
|
|
56
|
+
model. `regenerate` takes one named change, reruns only the stages that change reaches, and
|
|
57
|
+
reuses the rest. The runner you give it prints each stage's output and a `VERDICT` line; a
|
|
58
|
+
runner that prints no verdict fails the stage. `compare` grades the new output and its parent
|
|
59
|
+
through the same doctor against the same spec, so a claim that the new one is better is a
|
|
60
|
+
number. It grades only the bytes each recipe records, and refuses an output edited since.
|
|
61
|
+
|
|
62
|
+
### 30 seconds
|
|
63
|
+
|
|
64
|
+
The package ships a worked example in `examples/recipe/`: a factory with two inputs and three
|
|
65
|
+
stages, a runner and a doctor, all plain text transforms with no model in them.
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
cp -r node_modules/@supersuit/hyperspec/examples/recipe recipe-demo
|
|
69
|
+
cd recipe-demo
|
|
70
|
+
node factory.mjs
|
|
71
|
+
npx hyperspec recipe approve essay.md.recipe.json --by you
|
|
72
|
+
npx hyperspec reproduce essay.md.recipe.json
|
|
73
|
+
npx hyperspec regenerate essay.md.recipe.json --out essay-2.md --clicker you --add-input call-2=materials/call-2.md --reads claims --run "node runner.mjs"
|
|
74
|
+
npx hyperspec compare essay-2.md.recipe.json --doctor "node doctor.mjs"
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
The second call reaches the `claims` stage and the `draft` that reads it, so `terms` is reused.
|
|
78
|
+
The child carries two more claims, and the comparison lands in the spec's ledger:
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
claims: rerun
|
|
82
|
+
terms: reuse
|
|
83
|
+
draft: rerun
|
|
84
|
+
child scored 5 vs parent 3 (delta 2)
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
A factory writes its recipe with the writer API:
|
|
88
|
+
|
|
89
|
+
```js
|
|
90
|
+
import { startRecipe } from "@supersuit/hyperspec/recipe";
|
|
91
|
+
|
|
92
|
+
const recipe = startRecipe({
|
|
93
|
+
output: "essay.md",
|
|
94
|
+
factory: { name: "my-factory", version: "1.0.0" },
|
|
95
|
+
spec: "essay.hyperspec.md",
|
|
96
|
+
clicker: "you",
|
|
97
|
+
});
|
|
98
|
+
recipe.input("call", "materials/call.md");
|
|
99
|
+
recipe.stage({ id: "claims", reads: ["input:call"], output: claimsText, verdict: { station: "claims-not-empty", pass: true, note: "" } });
|
|
100
|
+
recipe.finish();
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
The schema, the stage key, the runner and doctor contracts, and every exit code are in
|
|
104
|
+
[SPEC.md](SPEC.md#recipes).
|
|
33
105
|
|
|
34
106
|
## The format
|
|
35
107
|
|
package/SPEC.md
CHANGED
|
@@ -17,8 +17,26 @@ decisions:
|
|
|
17
17
|
chosen_by: human
|
|
18
18
|
- id: exit-codes
|
|
19
19
|
state: decided
|
|
20
|
-
value: "0 every test passes and nothing is open; 1 at least one test fails; 3 every test passes but a decision is open; 2 usage or IO error"
|
|
21
|
-
source: SPEC.md,
|
|
20
|
+
value: "0 every test passes and nothing is open; 1 at least one test fails; 3 every test passes but a decision is open; 2 usage or IO error. The recipe commands use the same numbers, with 3 meaning a regeneration waits on a runner"
|
|
21
|
+
source: SPEC.md, sections "Exit codes" and "Recipe exit codes", and README.md, section "Exit codes"
|
|
22
|
+
author: agent:claude
|
|
23
|
+
chosen_by: agent
|
|
24
|
+
- id: recipe-verbs
|
|
25
|
+
state: decided
|
|
26
|
+
value: a recipe runs again three ways, reproduce (re-check every recorded hash), regenerate (one named change, rerunning only the stages it reaches) and compare (both outputs graded by one doctor against one spec)
|
|
27
|
+
source: "the recipe standard: \"Every output can be reproduced, and regenerated with one more ingredient\""
|
|
28
|
+
author: gary-sheng
|
|
29
|
+
chosen_by: human
|
|
30
|
+
- id: recipe-store
|
|
31
|
+
state: decided
|
|
32
|
+
value: a recipe records every input, the spec and every stage output by SHA-256 hash, and each distinct content is kept once in a store at .hyperspec/blobs/<first two hex characters>/<hash>
|
|
33
|
+
source: SPEC.md, section "Recipes", "The blob store"
|
|
34
|
+
author: agent:claude
|
|
35
|
+
chosen_by: agent
|
|
36
|
+
- id: no-model-calls
|
|
37
|
+
state: decided
|
|
38
|
+
value: hyperspec never calls a model; a runner command the caller supplies reruns stages, and a doctor command the caller supplies grades outputs
|
|
39
|
+
source: SPEC.md, section "Recipes"
|
|
22
40
|
author: agent:claude
|
|
23
41
|
chosen_by: agent
|
|
24
42
|
requirements:
|
|
@@ -36,14 +54,44 @@ requirements:
|
|
|
36
54
|
station: "test/rules.test.mjs, \"every finding names its test\""
|
|
37
55
|
source: CHANGELOG.md, 0.1.0, "every finding names its test, a severity, a message and a fix"
|
|
38
56
|
author: agent:claude
|
|
57
|
+
- id: r3
|
|
58
|
+
text: reproduce runs no model and no command; it only re-hashes what the recipe recorded
|
|
59
|
+
fails_when: reproduce reports success for a recipe whose recorded blob does not hash to its recorded value
|
|
60
|
+
check:
|
|
61
|
+
station: test/reproduce.test.mjs
|
|
62
|
+
source: SPEC.md, section "Recipes", "reproduce"
|
|
63
|
+
author: agent:claude
|
|
64
|
+
- id: r4
|
|
65
|
+
text: regenerate reuses every stage whose key is unchanged and records the parent and the change
|
|
66
|
+
fails_when: a stage whose recomputed key equals its parent's recorded key is rerun, or a child recipe has no parent or no change
|
|
67
|
+
check:
|
|
68
|
+
station: test/regenerate.test.mjs
|
|
69
|
+
source: SPEC.md, section "Recipes", "regenerate"
|
|
70
|
+
author: agent:claude
|
|
71
|
+
- id: r5
|
|
72
|
+
text: compare flags a child that scores lower than its parent under the same doctor and spec
|
|
73
|
+
fails_when: compare exits 0 when the child scores lower than its parent
|
|
74
|
+
check:
|
|
75
|
+
station: test/cli-recipe.test.mjs
|
|
76
|
+
source: SPEC.md, section "Recipes", "compare"
|
|
77
|
+
author: agent:claude
|
|
78
|
+
- id: r6
|
|
79
|
+
text: the README's recipe walkthrough runs exactly as written against the shipped example
|
|
80
|
+
fails_when: a command in the README walkthrough exits non-zero, or prints something other than the output the README shows
|
|
81
|
+
check:
|
|
82
|
+
station: test/example-recipe.test.mjs
|
|
83
|
+
source: README.md, section "Recipes"
|
|
84
|
+
author: agent:claude
|
|
39
85
|
rejects:
|
|
40
86
|
- prose advice where a field could be checked
|
|
41
87
|
- a second YAML parser
|
|
88
|
+
- a recipe that points at a path whose bytes can change
|
|
89
|
+
- reproducing an output by running a model again and hoping it says the same words
|
|
42
90
|
examples:
|
|
43
91
|
- path: examples/minimal.hyperspec.md
|
|
44
92
|
why: the smallest spec that passes all nine tests
|
|
45
93
|
resume:
|
|
46
|
-
next_action: collect adopter issues on 0.
|
|
94
|
+
next_action: collect adopter issues on 0.2, recipes included, and cut 0.3 from them
|
|
47
95
|
feedback:
|
|
48
96
|
issues: https://github.com/SupersuitUp/hyperspec/issues
|
|
49
97
|
fork: MIT; fork it for your own purposes and say so in your SPEC
|
|
@@ -55,7 +103,7 @@ improvement:
|
|
|
55
103
|
|
|
56
104
|
A person writing for another person leaves most of the specification unsaid, because the other person fills the gaps from shared context. An agent has none of that context, so it fills every gap with the average, and the average is what reads as middling. Hyperspecification is writing down the gaps. It is a level of detail that would feel like overkill between two people and is exactly enough for an agent: every decision the agent would otherwise guess is either decided, delegated with the rule for deciding it, or marked open, so the work stops instead of guessing.
|
|
57
105
|
|
|
58
|
-
**Version 0.
|
|
106
|
+
**Version 0.2.0** (2026-09-28)
|
|
59
107
|
|
|
60
108
|
## What makes a spec a hyperspec
|
|
61
109
|
|
|
@@ -101,7 +149,7 @@ Every run leaves a verdict: it went through clean, or it did not and the spec or
|
|
|
101
149
|
|
|
102
150
|
## The format (what `lint` reads)
|
|
103
151
|
|
|
104
|
-
A hyperspec is a markdown file with a YAML frontmatter block. This is the shape:
|
|
152
|
+
A hyperspec is a markdown file with a YAML frontmatter block. `hyperspec` names the version of this format the spec was written against; this linter knows `"0.1"`. This is the shape:
|
|
105
153
|
|
|
106
154
|
```yaml
|
|
107
155
|
---
|
|
@@ -156,15 +204,15 @@ Each row lists every condition under which `hyperspec lint` fails that test. A w
|
|
|
156
204
|
|
|
157
205
|
| Test | Fails when |
|
|
158
206
|
|---|---|
|
|
159
|
-
| 1 every decision is accounted for | no `decisions`; a decision with no or duplicate `id`; `state` not decided, delegated or open; decided without `value`; delegated without `rule`; open without `question` |
|
|
207
|
+
| 1 every decision is accounted for | no `decisions`; a decision with no or duplicate `id`; an `id` used by both a decision and a requirement, or by two requirements, since the two lists share one set of ids; `state` not decided, delegated or open; decided without `value`; delegated without `rule`; open without `question` |
|
|
160
208
|
| 2 every requirement can fail | no `requirements`; a requirement without `text` or without `fails_when`. A vague word in `fails_when` is a warning |
|
|
161
209
|
| 3 every requirement names its check | a requirement whose `check` has neither `station` nor `rubric` |
|
|
162
210
|
| 4 every field says where it came from and who wrote it | a decision or requirement without `source` or `author`; a decision whose `chosen_by` is not human or agent |
|
|
163
211
|
| 5 negative space is specified | `rejects` missing or empty; a `rejects` item that is not a plain string |
|
|
164
|
-
| 6 examples outrank adjectives | `examples` missing or empty; an example without `path` or `why`; a `path` that is not an http(s) URL and does not exist, read relative to the spec or as an absolute path |
|
|
165
|
-
| 7 a stranger can resume it | `resume.next_action` missing; a `next_action` that is only a no-action word (`continue`, `follow up`, `tbd`, `todo`, `keep going`, `pick it back up`, `n/a`, `none`); a `next_action` that says `as discussed` or `as mentioned earlier` or `above`. Those pointers in the body are a warning |
|
|
212
|
+
| 6 examples outrank adjectives | `examples` missing or empty; an example without `path` or `why`; a `path` that is not an http(s) URL and does not exist, is a folder, or is the spec itself, read relative to the spec or as an absolute path |
|
|
213
|
+
| 7 a stranger can resume it | `resume.next_action` missing; a `next_action` that is only a no-action word (`continue`, `follow up`, `tbd`, `todo`, `keep going`, `pick it back up`, `n/a`, `none`); a `next_action` that says `as discussed` or `as mentioned earlier` or `above`. Those pointers in the body are a warning, and so is a `hyperspec` version this linter does not know |
|
|
166
214
|
| 8 its adopters can push back on it | `feedback.issues` or `feedback.fork` missing |
|
|
167
|
-
| 9 it improves itself | `improvement.ledger` missing; a ledger path that exists and is not a readable file; if the ledger file exists, a line that is not a JSON object, a `verdict` outside one-shot, improved or not-improved, `improved` without `change`, `not-improved` without `reason
|
|
215
|
+
| 9 it improves itself | `improvement.ledger` missing; a ledger path that exists and is not a readable file; if the ledger file exists, a line that is not a JSON object, a `verdict` outside one-shot, improved or not-improved, `improved` without `change`, `not-improved` without `reason`. A declared ledger that does not exist yet is a warning |
|
|
168
216
|
|
|
169
217
|
## Exit codes
|
|
170
218
|
|
|
@@ -185,6 +233,138 @@ Every run of a skill that works from a hyperspec writes one line to the ledger n
|
|
|
185
233
|
|
|
186
234
|
Silence is not a verdict. A run that learned nothing has to say so and why, and a ledger line with none of the three verdicts fails the ninth test.
|
|
187
235
|
|
|
236
|
+
## Recipes
|
|
237
|
+
|
|
238
|
+
A hyperspec says what the work must be. A recipe records what one output was made from, so the output can be checked, made again with one change, and graded against the version before it. hyperspec never calls a model. Every step that needs one is a command the caller supplies: a runner for stages and a doctor for grading.
|
|
239
|
+
|
|
240
|
+
### The recipe file
|
|
241
|
+
|
|
242
|
+
A recipe sits beside its output as `<output>.recipe.json`: JSON with a two-space indent and a trailing newline. Every path inside it is relative to the recipe file's own directory.
|
|
243
|
+
|
|
244
|
+
```json
|
|
245
|
+
{
|
|
246
|
+
"recipe": "0.1",
|
|
247
|
+
"created": "2026-09-28T18:00:00.000Z",
|
|
248
|
+
"output": { "path": "essay.md", "sha256": "<hex>" },
|
|
249
|
+
"factory": { "name": "compose-a-piece", "version": "0.3.0" },
|
|
250
|
+
"spec": { "path": "essay.hyperspec.md", "sha256": "<hex>", "authors": { "audience": "gary-sheng", "length": "agent:claude" } },
|
|
251
|
+
"inputs": [
|
|
252
|
+
{ "name": "call", "path": "materials/call.md", "sha256": "<hex>", "order": 1 }
|
|
253
|
+
],
|
|
254
|
+
"stages": [
|
|
255
|
+
{
|
|
256
|
+
"id": "outline",
|
|
257
|
+
"reads": ["input:call", "spec"],
|
|
258
|
+
"model": { "name": "some-model", "temperature": 0.7 },
|
|
259
|
+
"key": "<hex>",
|
|
260
|
+
"output": { "sha256": "<hex>" },
|
|
261
|
+
"verdict": { "station": "outline-has-claim-chain", "pass": true, "note": "" }
|
|
262
|
+
}
|
|
263
|
+
],
|
|
264
|
+
"clicker": "gary-sheng",
|
|
265
|
+
"approver": "gary-sheng",
|
|
266
|
+
"parent": null,
|
|
267
|
+
"change": null
|
|
268
|
+
}
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
- `recipe` is the version of this schema, `"0.1"`.
|
|
272
|
+
- `output` is the final output's path and hash. It equals the last stage's output.
|
|
273
|
+
- `factory` names what made the output, and at which version.
|
|
274
|
+
- `spec` is the hyperspec the output was made from: its path, the hash of its bytes, and `authors`, which maps every decision and requirement id to that entry's `author`. Decisions and requirements share one set of ids, so the writer refuses a spec that uses an id twice, and `lint` fails it under test 1.
|
|
275
|
+
- `inputs` lists every input by name, path and hash, with `order` counting from 1 in the order each was taken in.
|
|
276
|
+
- `stages` lists every step in the order it ran. `model` holds the settings the stage ran with, and is left out when it had none. `verdict` is what the stage's station reported.
|
|
277
|
+
- `reads` names what a stage read: `input:<name>`, `stage:<id>` for an earlier stage, or `spec`. An empty `reads` means the stage read everything: every input in order, every earlier stage, then the spec. Such a stage reruns on every regeneration, and `recipe check` warns about it.
|
|
278
|
+
- `clicker` is who pressed go. `approver` is who approved the output, and stays `null` until someone does.
|
|
279
|
+
- `parent` and `change` are both `null` on a first recipe. On a regenerated one, `parent` is `{ "path", "sha256" }`, the parent recipe's path and the hash of its bytes, and `change` is one line naming what changed.
|
|
280
|
+
|
|
281
|
+
### The stage key
|
|
282
|
+
|
|
283
|
+
```
|
|
284
|
+
key = sha256(canonical({ reads, spec, factory, model }))
|
|
285
|
+
```
|
|
286
|
+
|
|
287
|
+
- `reads` is the list of `[ref, hash]` pairs in declared order, or in the everything order above for an empty `reads`. `input:<name>` resolves to that input's hash, `stage:<id>` to that stage's output hash, `spec` to the spec's hash.
|
|
288
|
+
- `spec` is the spec's hash, `factory` is the factory version, and `model` is the stage's model settings or `null`.
|
|
289
|
+
- Canonical JSON sorts keys at every depth, drops keys whose value is undefined, keeps array order, and carries no whitespace. Every hash is SHA-256, lowercase hex, over exact bytes.
|
|
290
|
+
|
|
291
|
+
A recorded hash is exactly 64 lowercase hex characters. Anything else names no blob: no path is ever built from it, `reproduce` fails the step that holds it, and `recipe check` fails it.
|
|
292
|
+
|
|
293
|
+
Two stages with the same key read the same bytes under the same spec, factory version and settings. A matching key is the only thing that lets `regenerate` reuse a stage.
|
|
294
|
+
|
|
295
|
+
### The blob store
|
|
296
|
+
|
|
297
|
+
Every input, the spec, and every stage output is kept once, by content, at `<root>/.hyperspec/blobs/<first two hex characters>/<hash>`, holding the exact bytes. A blob is written once, through a temporary file and a rename, and never overwritten. `<root>` is the `--store` flag, else the `HYPERSPEC_STORE` environment variable, else the nearest folder above the recipe holding `.hyperspec/` or `.git`, else the recipe's own folder.
|
|
298
|
+
|
|
299
|
+
This is what makes a recipe reproducible: a file edited next month does not change the bytes an old recipe points at.
|
|
300
|
+
|
|
301
|
+
### Writing a recipe
|
|
302
|
+
|
|
303
|
+
A factory records its recipe as it runs, through `@supersuit/hyperspec/recipe`:
|
|
304
|
+
|
|
305
|
+
- `startRecipe({ output, factory, spec, clicker, store })` loads the spec, stores its bytes, and fills `spec.authors`, recording `null` for an entry with no author. Paths resolve against the working directory.
|
|
306
|
+
- `input(name, path)` stores the file's bytes and records it. A repeated name is refused.
|
|
307
|
+
- `stage({ id, reads, model, output, verdict })` stores the output and computes the key. `model` and `verdict` are kept as JSON writes them, so the key is computed from exactly what the recipe file holds. An unknown read, a read of a later stage, and a repeated id are refused.
|
|
308
|
+
- `finish({ approver })` refuses a recipe with no stages. Otherwise it writes the output file from the last stage's blob, or checks that an output already on disk matches it, writes the recipe, and returns the completeness findings.
|
|
309
|
+
- `approve(recipePath, by)` sets the approver and returns the findings again.
|
|
310
|
+
|
|
311
|
+
`hyperspec recipe check <output-or-recipe>` runs the completeness check. It fails a recipe missing the factory name or version, the spec hash, `spec.authors`, the clicker or the approver; an input with no hash; any recorded hash that is not 64 lowercase hex characters; no stages; a stage with no verdict, a verdict whose `pass` is not true or false, or a stage still pending; a stored key that differs from the key recomputed from the recipe; a last stage whose output is not `output.sha256`; and a `parent` without a `change`, or the reverse. It warns on a stage that declares no reads, and on a spec id with no author, naming the id. It reads the recipe only: whether the blobs are still in the store is what `reproduce` checks. `hyperspec recipe approve <recipe> --by <slug>` records the approver.
|
|
312
|
+
|
|
313
|
+
### reproduce
|
|
314
|
+
|
|
315
|
+
`hyperspec reproduce <recipe> [--restore] [--store <dir>]` checks every hash the recipe recorded. It runs no model and no command. In order, it checks each input's blob, the spec's blob, each stage's output blob and the stage's key recomputed from the recipe, the final output's blob, and the output file on disk when there is one. It reports every step and names the first that fails. A pending stage fails.
|
|
316
|
+
|
|
317
|
+
`--restore` rewrites the output file from its blob, and only when that blob checks out. A recipe whose output path points outside its own folder is refused, and nothing is written there.
|
|
318
|
+
|
|
319
|
+
### regenerate
|
|
320
|
+
|
|
321
|
+
`hyperspec regenerate <recipe> --out <path> --clicker <slug>` takes exactly one change:
|
|
322
|
+
|
|
323
|
+
- `--add-input <name>=<path>` adds an input, and each `--reads <stage-id>` adds it to that stage's reads. `--reads` may be given more than once. A stage whose `reads` is empty already reads every input. An input that no stage would read is refused, since the child would be its parent with an unused input recorded.
|
|
324
|
+
- `--swap-input <name>=<path>` replaces an input. A file with the same bytes is refused, since swapping it would change nothing.
|
|
325
|
+
- `--factory-version <version>` names a factory version other than the parent's.
|
|
326
|
+
|
|
327
|
+
The child is the parent with that change applied. Each stage, in order, is reused only when two things hold: the key recomputed from the child's hashes equals the key the parent recorded, and the parent's recorded key matches the key recomputed from the parent's own record. A stage whose key differs reruns. A stage whose key matches but whose recorded output blob is missing or altered stops the regeneration, and nothing is written, since rerunning a stage whose key has not changed is exactly what regenerate must never do. With `--run <command>`, each stage is decided once everything it reads has run, so a stage whose upstream reran and came out byte for byte the same is still reused. Without `--run`, nothing runs: a stage that must rerun is written as pending, and every stage that reads it is pending too, with no key.
|
|
328
|
+
|
|
329
|
+
Nothing is written until the outcome is known. Then the new blobs, the child's output file, and `<out>.recipe.json` are written; while any stage is pending there is no output yet, so only the blobs and the recipe are. The output is written through a temporary file and read back against its hash. The child names its parent and the change, one line such as `added input call-2 (materials/call-2.md)`, `swapped input call to materials/call-v2.md`, or `factory 0.3.0 to 0.4.0`, which `--change <text>` replaces. Its approver is `null`. The parent recipe, its output and its blobs are only read. `--out` must not exist and must not be the parent's output, its folder must exist, and it is checked again just before anything is written.
|
|
330
|
+
|
|
331
|
+
**The runner contract.** hyperspec runs `/bin/sh -c <command>` once for each stage that must rerun.
|
|
332
|
+
|
|
333
|
+
- stdin is `{ "stage": "<id>", "reads": [{ "ref": "<ref>", "sha256": "<hex>", "path": "<file>" }], "model": <settings or null> }`. Each `path` is a temporary file holding that read's exact bytes, removed after the stage runs.
|
|
334
|
+
- stdout, byte for byte, is the stage's output.
|
|
335
|
+
- The last stderr line that begins `VERDICT ` carries the verdict as a JSON object whose `pass` is true or false. `station` defaults to `runner` and `note` to an empty string. With no such line, the verdict is `{ "station": "runner", "pass": false, "note": "runner reported no verdict" }`: silence is not a verdict, so a runner has to say something.
|
|
336
|
+
- A runner that exits non-zero, or a `VERDICT` line that is not such an object, stops the regeneration and names the stage, and nothing is written.
|
|
337
|
+
- A failing verdict from a stage that ran in this regeneration, including a runner that reported none, still produces a child, written in full, and the command exits 1 naming the stage.
|
|
338
|
+
- A reused stage keeps the verdict its parent recorded, and that verdict does not affect the exit code, even when it failed. Only the stages this regeneration ran are counted.
|
|
339
|
+
|
|
340
|
+
### compare
|
|
341
|
+
|
|
342
|
+
`hyperspec compare <child-recipe> --doctor <command> [--parent <recipe>] [--spec <file>]` grades the child's output and its parent's output with the same doctor command against the same spec. The parent defaults to the one the child names, and the spec to the child's spec. It warns when the parent recipe's bytes have changed since the child was made, and when the spec being graded against differs from the one the parent was made from; both outputs are still graded against that one file.
|
|
343
|
+
|
|
344
|
+
A score is attributed to a recipe, so the bytes graded must be the bytes the recipe records. Before the doctor runs, each output file is hashed against its recipe's `output.sha256`. A file that does not match, edited by hand or replaced since, is refused with `parent output does not match its recipe; run hyperspec reproduce --restore` (or `child`), and nothing is graded or appended to the ledger.
|
|
345
|
+
|
|
346
|
+
**The doctor contract.** hyperspec runs `/bin/sh -c <command>` once per output, with the same command both times. stdin is `{ "output": "<file>", "spec": "<file>" }`, both absolute paths. The last non-empty stdout line is a JSON object with a finite numeric `score` and an optional `notes`. Nothing from a recipe is placed into the command itself.
|
|
347
|
+
|
|
348
|
+
A child that scores lower than its parent has regressed. `compare` exits 1 and names the suspect, the child's `change`. When the spec declares `improvement.ledger`, `compare` appends one line to it, in the ledger's own vocabulary so the line passes the ninth test: `improved` with `change` when the child scores higher, and otherwise `not-improved` with a `reason` naming both scores, which begins `regression: ` when the child scored lower. The line also carries `kind: "compare"`, both recipe paths, both scores, and `regressed`. A ledger path outside the spec's folder is not written, and a warning says so.
|
|
349
|
+
|
|
350
|
+
### Recipe exit codes
|
|
351
|
+
|
|
352
|
+
The recipe commands use the same numbers as `lint`. Every one takes `--json`, which prints the result as one JSON document and keeps the same exit code.
|
|
353
|
+
|
|
354
|
+
| Command | 0 | 1 | 2 | 3 |
|
|
355
|
+
|---|---|---|---|---|
|
|
356
|
+
| `recipe check` | complete, warnings allowed | no recipe beside the path, or a check failed | a recipe that cannot be read or is not JSON | |
|
|
357
|
+
| `recipe approve` | approver recorded, remaining findings printed | | `--by` missing, or a recipe that cannot be read | |
|
|
358
|
+
| `reproduce` | every hash checks out | a check failed | a usage error, or a recipe that cannot be read | |
|
|
359
|
+
| `regenerate` | child written, and every stage that ran reported a passing verdict | a runner failed or a blob it needs is missing, and nothing was written; or the child was written with a failing verdict from a stage that ran | a usage error, such as no change or more than one, a missing `--out` or `--clicker`, an `--out` that exists, an unknown stage in `--reads`, an added input no stage reads, a parent or input that cannot be read | child written with stages waiting for a runner |
|
|
360
|
+
| `compare` | the child did not regress | the child regressed | a usage error, a recipe or spec that cannot be read, a child that names no parent when none is given, a missing output file or one that does not match its recipe, or a doctor that failed | |
|
|
361
|
+
|
|
362
|
+
### Known limits in 0.2
|
|
363
|
+
|
|
364
|
+
- A pending child cannot yet be finished in place. Rerunning `regenerate` with `--run` is refused because the child recipe exists; to produce the output, rerun `regenerate` on the parent with a runner and a new `--out`.
|
|
365
|
+
- Approve a recipe before regenerating from it. Approval rewrites the recipe's bytes, and a child records its parent's hash, so approving a parent after a child exists makes `compare` warn that the parent recipe changed.
|
|
366
|
+
- `reproduce`, `regenerate` and `compare` are commands, not yet library functions. `@supersuit/hyperspec/recipe` exports the writer (`startRecipe`, `approve`); a factory runs the other three through the `hyperspec` command.
|
|
367
|
+
|
|
188
368
|
## Why now
|
|
189
369
|
|
|
190
|
-
A human reader treats a thousand-line spec as a burden, so specs were written short and the gaps were filled from shared context. A model reads all of it at almost no cost and uses every line. Detail that would have been waste between two people is now the cheapest input there is. That is
|
|
370
|
+
A human reader treats a thousand-line spec as a burden, so specs were written short and the gaps were filled from shared context. A model reads all of it at almost no cost and uses every line. Detail that would have been waste between two people is now the cheapest input there is. That is why hyperspecification exists now and could not have before: the reader changed, so the economics of writing everything down changed with it.
|