task-pipeline-skill 1.71.1 → 1.73.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,119 @@
1
1
  # Changelog
2
2
 
3
+ ## v1.73.0 — 2026-08-20 — the registry could not see the templates, the scripts, or its own registers
4
+
5
+
6
+ **The claim registry had the class for this exact incident and it fired on nothing.** Three
7
+ shipped surfaces said *34 reference files* over a directory of 35 — `scripts/graph.py`'s
8
+ `doctrine` docstring and `templates/run.md` twice — and `npm test` printed
9
+ `reference files: dormant (truth 35)`. Two holes, either of which was enough: the pattern
10
+ knew only the word order ``N files under `references/` ``, and the corpus was eight named
11
+ files plus `references/**`, so `templates/` and `scripts/` were never opened. The class now
12
+ reads any phrasing of the count, the corpus reads the templates and the shipped script, and
13
+ the class is armed at three agreeing sites instead of dormant.
14
+
15
+ **And it never read this repository's own registers.** `docs/DOCMAP.md` names decisions,
16
+ open questions, the board, the ledger and the retro as the registers, and its propagation
17
+ matrix sends *a number stated in a living document* to this registry — which could not see
18
+ any of them. `docs/OPEN_QUESTIONS.md` said *the 250 guards* against a workflow defining 390,
19
+ in a phrasing the guard class already knew. Bringing them in refused 26 statements and every
20
+ one was narration, so a number inside a **dated item** is a record and exempt; on the board
21
+ the discriminator is the State cell rather than the date, because every row names the day it
22
+ was filed and B-001 — open — was stating its description budget and its reference-file count
23
+ as facts about now.
24
+
25
+ **Fourteen consecutive releases carry no run stamp**, `v1.60.1` through `v1.72.0`, and the
26
+ retro's honest-gap section named only `v1.16.0`–`v1.23.0`. Its *Measured, not recalled*
27
+ receipt could not produce the measurement: it grepped `docs/superpowers/retro.md`, removed at
28
+ v1.53.0, and grepped for a tag's own commit when a stamp names the commit the *run* ended on.
29
+ Rewritten as a tag-range walk that reads the archive too, and a guard now requires every
30
+ release after the newest stamp to be named in that section.
31
+
32
+ Also: `## Unreleased` is where the guard count lives between a tag and the next bump (B-104);
33
+ `Environment` is a required cell on every verification row with the vocabulary read out of
34
+ the shipped template (B-099's other half); the graph schema states its three node rules
35
+ behind `$ref`s and the checker follows them (B-079, proved by moving them, not by a fixture);
36
+ an open board row's `file:N-M` must quote the phrase it points at, which caught five stale
37
+ citations and then caught this change's own edits four more times; the acceptance ladder is
38
+ policy **`AP-1`** with an owner and an in-force date, and `gates.md`'s *the framework fixes
39
+ no stage count* is scoped to the pipeline's shape; `read:` and `gate:` are reported
40
+ **unattested** instead of claimed *never agent-written*, because the ledger is the file the
41
+ agent appends to at every stage; `validate.yml` no longer re-validates a SHA on its own tag
42
+ push; and every documented `npm` equation is compared against `package.json` — `CLAUDE.md`
43
+ had `npm test` as `validate.py` alone, dropping 129 graph cases.
44
+
45
+ **What the suite found that no reading did.** This change broke **thirteen** existing
46
+ plants and `npm run test:all` named every one: the two verification-header probes spell
47
+ the whole header and it gained a column; two CHANGELOG probes scoped themselves to
48
+ `^## v` and the count now lives in `## Unreleased`; **seven** graph-schema probes walk
49
+ `node.allOf` inline and the rules moved behind `$ref`. All thirteen were repaired by
50
+ deriving the guard's own scope rather than restating it, which is what `learned.md`'s
51
+ *sweep the class* asks for — and the sweep needed two rounds: six schema probes were
52
+ repaired together and the seventh surfaced on the next full run, having died on
53
+ `KeyError: 'if'` where the others had died on their own asserts. A class fixed in six of
54
+ seven places is the shape standing instruction R-003 exists for.
55
+
56
+ One of the twelve was sharper than the rest, and it was self-inflicted twice over: the
57
+ honesty note explaining that `OQ-0002` no longer restates a total put an ISO date in the
58
+ row, which made the row a dated record and **disarmed the register plant** — green over
59
+ exactly the stale total it had been written to catch. `OQ-####` now uses its `Status`
60
+ cell, the discriminator the board already uses, and the register plant sits on an open
61
+ board row where no prose edit beside it can turn it off.
62
+
63
+ Guards: 390 → **412** · property checks 9 → **14**
64
+
65
+
66
+ ## v1.72.0 — a node says how it will be closed
67
+
68
+ **B-080 closed, and with it the last of four requirements this pack's own manifesto named
69
+ as unbuilt.** The Proof of Done manifesto cites this repository as its reference
70
+ implementation and filed four gaps against it by id; three closed on 2026-08-17, and this
71
+ is the fourth. The public document now says all four are built and names the commit for
72
+ each.
73
+
74
+ ### `check` — the per-node completion test
75
+
76
+ A node could not say how it would be closed, while shipped doctrine told the verifier to
77
+ "run the checks the task named". The doctrine read a field the schema did not have.
78
+
79
+ `check` is a string, **required on every node whose status is not `parked`** — the parked
80
+ node being the one nobody will close, where a placeholder would be confidence without
81
+ correctness. Never a list: a node needing two unrelated checks is a node doing two jobs, and
82
+ one gate made of two commands is `a && b`, which is still one gate. Stated twice on purpose,
83
+ in the schema and in `violations()`, because the schema never runs against a live graph.
84
+ `add --check` is required like `--why`, and a replan entry naming no check is refused before
85
+ the close writes anything.
86
+
87
+ ### The record the manifesto opens with now reads as one event
88
+
89
+ `residue.md`'s opening record is quoted by an external document and permalinked from a public
90
+ site. It said the inventory was queried "one minute later" beside a `ps` line showing the
91
+ process at 3:12 — two accurate observations of different things, framed so a reader had to
92
+ reconcile them by hand. Both moments are now stated, the interval between them named, and
93
+ **what the record never carried is named too**: the wall-clock times, which `git log --follow`
94
+ proves were never in it. A disambiguation, not a correction — no number changed.
95
+
96
+ The E4 sentence claimed more than its own block: "that **they** have been up for three days"
97
+ from a measurement of the oldest one. Now "the oldest of them".
98
+
99
+ ### Three plants that could not run, and one row that could not be parsed
100
+
101
+ CI failed twice on this release's work, and neither failure was noise.
102
+
103
+ A ledger row wrote a pytest filter as a backticked `-k a|b|c`; markdown does not care about
104
+ backticks, so two pipes split one cell into three and that row carried ten cells against a
105
+ header of eight. Locally the damage sat in a column nothing read, so every gate run passed.
106
+ Every `REQ` row must now match the header's cell count, with the remedy named in the refusal.
107
+
108
+ Then three schema plants died on `KeyError` before reaching their own `PLANT DID NOT LAND`
109
+ assert, because the new `check` rule added an `allOf` branch keyed on `not: {const: parked}`
110
+ and each plant indexed the const directly. All three traverse tolerantly now. **The assert
111
+ catches a plant that does not land; nothing catches a plant that cannot run** — the third
112
+ instance of that class in three days, and it is filed family-wide rather than here.
113
+
114
+ Guards: 389 → **390**. The new rule arrived without a plant, which this repository's own standard refuses; that is the plant. `graph.py` 114 → 129 cases.
115
+
116
+
3
117
 
4
118
  ## v1.71.1 — four shapes of a green that measured nothing
5
119
 
@@ -55,8 +169,20 @@ only reader that checks.
55
169
  and what makes each self-checking. Where a claim genuinely cannot be derived, **say so in the
56
170
  row**; correcting it silently a second time teaches readers that somebody else is checking.
57
171
 
58
- Guards: 376 → **376**. No guard was added or removed this release is doctrine, and each of the
59
- four rules is cited from the reference that owns it rather than restated.
172
+ > **After the tag, on the same tree.** `v1.71.1` is cut, and the guard-count claim
173
+ > `test/validate.py` reads is the newest `## vX.Y.Z` section's which post-tag is this one,
174
+ > with no version heading yet open for the work that follows. So the live figure is stated
175
+ > here rather than inside a released sentence, and the released sentence below is not edited:
176
+ > B-080 closed on 2026-08-19 with eight plants and TP-02 with five on the same day, so
177
+ > Guards: 376 → **389**. The guard has no home for a count between a tag and the next bump;
178
+ > filed as `B-104`, and the note is what keeps the number true meanwhile. **The bold belongs
179
+ > on the number alone**: the plant that corrupts this count reads `Guards: N → **M**` out of
180
+ > the raw section, so bolding the whole phrase left the validator reading a count no plant
181
+ > could still corrupt — which is the form this note was introduced with, and no second
182
+ > count-shaped figure belongs in this section for the same reason.
183
+
184
+ Guards at the tag: 376 → 376. No guard was added or removed — that release is doctrine, and
185
+ each of the four rules is cited from the reference that owns it rather than restated.
60
186
 
61
187
  Closes #39, #40, #48, #49.
62
188
 
package/CONTRIBUTING.md CHANGED
@@ -518,6 +518,80 @@ runs.
518
518
  *(guard: `does not say` and `rule 1 of the list it is printed beside`
519
519
  and `no stage names`)*
520
520
 
521
+ **56. A node says how it will be closed, and the doctrine that reads that field names
522
+ one the node has.** `agents/verifier.md` ordered the verifier to run *the checks the task
523
+ named* while `graph.schema.json`'s node carried no field in which a task could name one —
524
+ two files shipped on one day, and the instruction pointed at an absence that left the
525
+ verifier the two options the same paragraph forbids: invent a check, or run everything.
526
+ `check` is now a node property, **required on every node except a `parked` one** — the one
527
+ node nobody will close, where a placeholder would be worse than the gap. The rule is
528
+ stated twice on purpose, in the schema and in `violations()`, because the schema is never
529
+ applied to a live graph; both are RUN against a planted graph rather than inspected. And
530
+ the two homes are compared directly: whatever the verifier is told to read off the node
531
+ must be a property the schema declares.
532
+ *(guard: `node declares no` and `rule that can fire` and `off the node, and`)*
533
+
534
+ **57. The claim registry reads every surface that states a number, including the ones this
535
+ repository writes about itself.** Its corpus was eight named files plus `references/**`, so
536
+ `templates/` and `scripts/` were invisible — three shipped surfaces said "34 reference files"
537
+ over a directory of 35 while the class printed `dormant` — and so were the registers
538
+ `docs/DOCMAP.md` names, where `docs/OPEN_QUESTIONS.md` said "the 250 guards" against a
539
+ workflow defining 390. The corpus now covers both, and a class recognises every phrasing of
540
+ its count rather than one word order. A number inside a **dated item** in a register is a
541
+ record and exempt; on the board the discriminator is the **State** cell, because every row
542
+ names the day it was filed and an open row is a claim about now.
543
+ *(guard: `reference files` and `— derive the number or delete it`)*
544
+
545
+ **58. Every documented `npm` command means what `package.json` runs.** `CLAUDE.md` glossed
546
+ `npm test` as `python3 test/validate.py`, dropping `graph_test.py` and its 129 cases — the
547
+ suite-outside-the-run class stated the other way round — and called `npm run test:all` "both"
548
+ where it runs eight scripts. An equation is compared against the script body after one level
549
+ of `npm run` resolution, and a bare `npm run X` in a document about this repository must be a
550
+ script that exists. Portable doctrine under `plugins/` and `cursor/` is out of scope: it names
551
+ a host project's commands.
552
+ *(guard: `is glossed as` and `declares no such script`)*
553
+
554
+ **59. A release either carries a run stamp or is recorded as a gap, and the guard-count claim
555
+ has a home before the bump.** Fourteen consecutive releases had no stamp while the retro named
556
+ only `v1.16.0`–`v1.23.0`, and the receipt that was supposed to prove it grepped a path removed
557
+ at v1.53.0 for a tag's own commit — a stamp names the commit the *run* ended on. Scoped to the
558
+ trailing stretch: 84 of 117 tags predate the register and backfilling is forbidden. Separately,
559
+ the count guard reads the topmost `## ` section, so `## Unreleased` is where the number lives
560
+ between a tag and the next bump, and it must sit above every version heading.
561
+ *(guard: `release(s) after the newest run stamp` and `section sits below a released version`)*
562
+
563
+ **60. A `file:line` range in an open board row quotes the phrase it points at.** Five
564
+ citations resolved to real lines and pointed at other text; a line number is the most fragile
565
+ address a document carries, because every edit above it moves it and nothing notices. Closed
566
+ rows are records and are left alone; single-line citations cannot be quoted and are disclosed
567
+ as unanchored rather than failed.
568
+ *(guard: `quotes no phrase from it`)*
569
+
570
+ **61. Every evidence row records the environment it ran in, and a claim of provenance the
571
+ format cannot check is marked unattested instead.** `Observed at` said which tree a check saw
572
+ and nothing said where it ran, so a preview smoke test and a production one entered the record
573
+ in the same shape — in a pack whose own `learned.md` records a suite green on every author's
574
+ machine and 1039 failures on a clean runner. And `read:`/`gate:` were declared *hook-written,
575
+ never agent-written* while both land in the file the agent appends to at every stage: no writer
576
+ field, no provenance check, so the count is reported `unattested` and the claim is not made.
577
+ *(guard: `has no `Environment` cell` and `prints a count and never says`)*
578
+
579
+ **62. The acceptance ladder is a versioned policy with an owner, and `gates.md`'s
580
+ fixes-nothing sentence is scoped to the pipeline's shape.** Both rules stood unscoped side by
581
+ side for seventy releases — *the framework fixes no stage count and no gate assignment* beside
582
+ twelve fixed criteria — so a reader could take either as the whole rule, and a table accepted
583
+ under v1.20 doctrine was indistinguishable from one accepted under v1.70. The block carries
584
+ `AP-1`, an owner and an in-force date; an amendment moves the version and lands with a decision
585
+ row.
586
+ *(guard: `the acceptance policy carries no` and `stands unscoped beside`)*
587
+
588
+ **63. A heading may not declare a bound nothing enforces.** The retro's *Recent log* read
589
+ *entries from the last five run stamps* over 25 entries reaching back nine days, borrowing the
590
+ stamp section's wording without its cap — filed twice as B-060 and B-069 and disclosed by the
591
+ file about itself. Checked in the live retro and in `templates/retro.md`, which seeded the
592
+ false bound into every host project.
593
+ *(guard: `declares the bound` and `and nothing enforces it`)*
594
+
521
595
  ## Adding or changing doctrine
522
596
 
523
597
  - **Change one idea per PR.** These files are read by agents under load; a PR that
package/SKILL-CARD.md CHANGED
@@ -12,7 +12,7 @@ harmless.
12
12
  |---|---|
13
13
  | **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
14
14
  | **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
15
- | **Version** | 1.71.1 |
15
+ | **Version** | 1.73.0 |
16
16
  | **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
17
17
  | **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki, and **one of two browser channels** — `playwright` (CLI or MCP) or `chrome-devtools` (MCP); either satisfies the browser step and neither is required. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
18
18
  | **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "task-pipeline-skill",
3
- "version": "1.71.1",
3
+ "version": "1.73.0",
4
4
  "description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
5
5
  "bin": {
6
6
  "task-pipeline": "bin/task-pipeline.js"
@@ -2,7 +2,7 @@
2
2
  "name": "task-pipeline",
3
3
  "displayName": "Task Pipeline",
4
4
  "description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
5
- "version": "1.71.1",
5
+ "version": "1.73.0",
6
6
  "author": {
7
7
  "name": "ssheleg",
8
8
  "url": "https://x.com/sshlg93"
@@ -34,7 +34,10 @@ than ending in a question nobody will see.
34
34
  "not_done": ["what was asked and is not"],
35
35
  "not_verified": ["what was BUILT and no check touched"],
36
36
  "blockers": [{ "what": "…", "blocks": ["N-009"], "can_continue_around": true }],
37
- "replan": { "possible": true, "add": [], "park": ["N-009"], "why": "…" },
37
+ "replan": { "possible": true,
38
+ "add": [{ "title": "…", "owner": "implementer",
39
+ "serves": "REQ-004", "check": "…" }],
40
+ "park": ["N-009"], "why": "…" },
38
41
  "evidence": ["the command and the output that proves each `done` row"]
39
42
  }
40
43
  ```
@@ -58,12 +61,20 @@ ledger exists to catch.
58
61
 
59
62
  1. **Read the node's `serves`** — the REQ or the goal clause. That is the standard.
60
63
  Not what the diff does; what was asked.
61
- 2. **Run the checks the task named.** Not a check you invented, and not `npm test`
62
- alone if the task named something narrowera green from a check nobody watched
63
- fail against a planted defect is not evidence.
64
+ 2. **Run the node's `check`.** It is a field on the node — one command, or a named
65
+ judgement where no command can decide itand it is what closes this node. Not a
66
+ check you invented, and not `npm test` alone when the node named something narrower:
67
+ a green from a check nobody watched fail against a planted defect is not evidence.
68
+ Where the `check` is a judgement, record the verdict **as** judgement
69
+ (`references/gates.md`), never as an exit code.
70
+ **A node with no `check` is a refusal, not a judgement call.** `graph.py validate`
71
+ exits 1 and names it, because inventing a check and running everything are the two
72
+ things this step forbids — and until B-080 closed, this paragraph asked for a field
73
+ the schema did not have.
64
74
  3. **`done` takes one row per claim, and each needs a line in `evidence`** — the
65
75
  command and what it printed. Paraphrase is not evidence. If you cannot produce
66
- the output, the row belongs in `not_done`.
76
+ the output, the row belongs in `not_done`. The `check`'s own output is the first
77
+ row: the node said how it would be closed, so the proof it closed is that output.
67
78
  4. **`not_done` is not a failure report.** It is what the next iteration picks up,
68
79
  so write it as work rather than as blame.
69
80
  5. **Every blocker says what it `blocks` and whether the run `can_continue_around`
@@ -72,6 +83,9 @@ ledger exists to catch.
72
83
  6. **`replan.possible: false` needs a `why` a person can act on.** A stop with no
73
84
  reason is indistinguishable from a stall, and the operator is the one who has to
74
85
  tell them apart.
86
+ 7. **Every `replan.add` entry names its own `check`.** A node you create is a node the
87
+ next verifier has to close, and handing it the absence you were handed is how the
88
+ defect returns one iteration later. `close` refuses the verdict and names the key.
75
89
 
76
90
  ## Three ways this goes wrong
77
91
 
@@ -80,6 +94,7 @@ ledger exists to catch.
80
94
  | «The tests pass, so it is done» | The node serves a REQ, not a suite. A green suite that never exercised the requirement proves the suite ran |
81
95
  | «Close it and note the gap» | A `done` with a caveat is a `not_done` somebody will read as finished. Split the row |
82
96
  | «This blocker stops everything» | Say whether it does. `can_continue_around: true` is what keeps a run moving past one bad node, and guessing it wrong costs either the run or the correctness |
97
+ | «The node names no check, so I'll run the suite» | That is the choice this agent may not make. `validate` refuses the node before you are dispatched; a node whose completion test nobody wrote is a planning defect, and reporting it is worth more than a green suite |
83
98
 
84
99
  ## Where the doctrine is
85
100
 
@@ -14,6 +14,7 @@
14
14
  "status": "done",
15
15
  "blocked_by": [],
16
16
  "serves": "REQ-001",
17
+ "check": "npm test — graph.example.json validates against graph.schema.json",
17
18
  "evidence": [
18
19
  "npm test → PASS: task-pipeline structure valid, with graph.example.json validated against graph.schema.json"
19
20
  ]
@@ -27,6 +28,7 @@
27
28
  "N-001"
28
29
  ],
29
30
  "serves": "REQ-002",
31
+ "check": "python3 test/graph_test.py — the validate fixtures",
30
32
  "evidence": null
31
33
  },
32
34
  {
@@ -38,6 +40,7 @@
38
40
  "N-001"
39
41
  ],
40
42
  "serves": "REQ-005",
43
+ "check": "judgment: the reviewer rules the seven keys sufficient — recorded AS judgement, references/gates.md",
41
44
  "evidence": null
42
45
  },
43
46
  {
@@ -110,6 +110,15 @@
110
110
  "minLength": 1,
111
111
  "description": "A REQ id, or a clause of the goal. Required, because a node that serves neither is work nobody asked for — and the pipeline's answer to that is to park it WITH THAT AS THE REASON rather than to do it quietly."
112
112
  },
113
+ "check": {
114
+ "type": "string",
115
+ "minLength": 1,
116
+ "pattern": "\\S",
117
+ "not": {
118
+ "pattern": "[\\n\\r]"
119
+ },
120
+ "description": "**How this node will be closed** \u2014 the command a verifier runs, or the named judgement that stands in where no command can (`references/gates.md`'s third gate type: name the judge, record the verdict AS judgement). Its output becomes the node's `evidence` row.\n\nRequired on every node except a `parked` one, and that is the whole point of the field: `agents/verifier.md` tells the verifier to run *the check this node names*, so a node with none leaves it the two options that same paragraph forbids \u2014 invent a check, or run everything. Shipped doctrine read a field the schema did not have (B-080); the field the doctrine reads now exists.\n\nOne string, never a list, because the requirement it serves is singular throughout \u2014 one input, one job, one output, one owner, **one** completion test. A node needing two unrelated checks is a node doing two jobs, and the answer is to split it rather than to widen this field. One gate made of two commands is `a && b`, which is still one gate.\n\nA `parked` node is exempt because it is the one node nobody will close, and a placeholder there \u2014 *n/a, parked* \u2014 is the confidence-without-correctness this schema refuses everywhere else. The field survives a park: `park` sets the status and never removes what the node said it would run."
121
+ },
113
122
  "evidence": {
114
123
  "type": [
115
124
  "array",
@@ -140,57 +149,13 @@
140
149
  },
141
150
  "allOf": [
142
151
  {
143
- "if": {
144
- "properties": {
145
- "status": {
146
- "const": "done"
147
- }
148
- },
149
- "required": [
150
- "status"
151
- ]
152
- },
153
- "then": {
154
- "required": [
155
- "evidence"
156
- ],
157
- "properties": {
158
- "evidence": {
159
- "type": "array",
160
- "minItems": 1,
161
- "items": {
162
- "type": "string",
163
- "minLength": 1,
164
- "pattern": "\\S",
165
- "description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
166
- }
167
- }
168
- }
169
- }
152
+ "$ref": "#/definitions/rule_check_unless_parked"
170
153
  },
171
154
  {
172
- "if": {
173
- "properties": {
174
- "status": {
175
- "const": "parked"
176
- }
177
- },
178
- "required": [
179
- "status"
180
- ]
181
- },
182
- "then": {
183
- "required": [
184
- "parked_reason"
185
- ],
186
- "properties": {
187
- "parked_reason": {
188
- "type": "string",
189
- "minLength": 1,
190
- "pattern": "\\S"
191
- }
192
- }
193
- }
155
+ "$ref": "#/definitions/rule_done_needs_evidence"
156
+ },
157
+ {
158
+ "$ref": "#/definitions/rule_parked_needs_reason"
194
159
  }
195
160
  ]
196
161
  },
@@ -248,6 +213,88 @@
248
213
  "description": "The reason, written for a person reading it later. The non-whitespace pattern is required for the same reason `parked_reason` needs one: `minLength: 1` counts a space."
249
214
  }
250
215
  }
216
+ },
217
+ "rule_check_unless_parked": {
218
+ "description": "B-080's rule, factored out of `node.allOf` so it is stated once and referenced. Factoring it exposed B-079: `test/validate.py`'s `_conditionals()` recursed into `allOf` and never followed a `$ref`, so a schema that shares a rule this way read as a schema with NO conditionals and both REQ-006 and REQ-012 were reported missing from a schema that states them. The shipped schema now uses the `$ref` form, so the deref branch is exercised by every run of the gate rather than by one fixture.",
219
+ "if": {
220
+ "properties": {
221
+ "status": {
222
+ "not": {
223
+ "const": "parked"
224
+ }
225
+ }
226
+ },
227
+ "required": [
228
+ "status"
229
+ ]
230
+ },
231
+ "then": {
232
+ "required": [
233
+ "check"
234
+ ],
235
+ "properties": {
236
+ "check": {
237
+ "type": "string",
238
+ "minLength": 1,
239
+ "pattern": "\\S"
240
+ }
241
+ }
242
+ }
243
+ },
244
+ "rule_done_needs_evidence": {
245
+ "description": "REQ-006 as draft-07 states it: `done` implies at least one non-whitespace evidence string. Referenced from `node.allOf`, never inlined, so the rule has one home.",
246
+ "if": {
247
+ "properties": {
248
+ "status": {
249
+ "const": "done"
250
+ }
251
+ },
252
+ "required": [
253
+ "status"
254
+ ]
255
+ },
256
+ "then": {
257
+ "required": [
258
+ "evidence"
259
+ ],
260
+ "properties": {
261
+ "evidence": {
262
+ "type": "array",
263
+ "minItems": 1,
264
+ "items": {
265
+ "type": "string",
266
+ "minLength": 1,
267
+ "pattern": "\\S",
268
+ "description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
269
+ }
270
+ }
271
+ }
272
+ }
273
+ },
274
+ "rule_parked_needs_reason": {
275
+ "description": "REQ-012: a parked node names why it is parked. Referenced from `node.allOf`.",
276
+ "if": {
277
+ "properties": {
278
+ "status": {
279
+ "const": "parked"
280
+ }
281
+ },
282
+ "required": [
283
+ "status"
284
+ ]
285
+ },
286
+ "then": {
287
+ "required": [
288
+ "parked_reason"
289
+ ],
290
+ "properties": {
291
+ "parked_reason": {
292
+ "type": "string",
293
+ "minLength": 1,
294
+ "pattern": "\\S"
295
+ }
296
+ }
297
+ }
251
298
  }
252
299
  }
253
300
  }
@@ -319,6 +319,29 @@ spends its length refusing.
319
319
 
320
320
  ## GATE (manual)
321
321
 
322
+ > **Acceptance policy `AP-1` · owner: the operator of the project running this pipeline ·
323
+ > in force since 2026-08-20 · supersedes: the unversioned ladder that stood before it.**
324
+ >
325
+ > This block is the policy, and naming it that closes a contradiction that stood in the
326
+ > shipped doctrine. [`gates.md`](gates.md) → *Axis A — the stage gate type* says which stages are
327
+ > manual is the operator's decision and that **the framework fixes no stage count and no
328
+ > gate assignment** — while the ladder below fixes twelve criteria and the statuses they
329
+ > may take. Both were true and neither was scoped, so a reader could take either as the
330
+ > rule.
331
+ >
332
+ > **The scope, stated once:** `gates.md`'s sentence governs the **pipeline's shape** — how
333
+ > many stages a project runs and which of them are `auto`, `judgment` or `manual`, all of
334
+ > it in the project's own `pipeline.json`. `AP-1` governs **what stage 10's manual gate
335
+ > asks when a project runs one.** A project may drop stage 10, or make it `auto` for a
336
+ > class of work, and `AP-1` then does not apply to it; a project that keeps it manual gets
337
+ > these criteria and this vocabulary, not a per-run selection from them.
338
+ >
339
+ > **Changing it is a decision, not an edit.** An amended criterion moves the version to
340
+ > `AP-2` and lands with a row in the project's decision register, because a policy that can
341
+ > be edited silently is the *«recorded in the slot reserved for what a machine
342
+ > established»* failure one level up: an acceptance standard nobody can cite by version is
343
+ > one every run re-negotiates.
344
+
322
345
  All of:
323
346
 
324
347
  1. **Every shipped REQ has a verification row, and every row names a REQ its own run
@@ -47,12 +47,21 @@ The method most audits use is **horizontal**: compare the documents against each
47
47
  other, then do it again. It works, and then it fails in a way that is invisible
48
48
  from inside it. Measured over seven passes on a production repository:
49
49
 
50
- | Pass | Findings | …of which the previous pass's own fixes caused |
51
- |---|---|---|
52
- | 4 | 12 | 5 |
53
- | 5 | 17 | 9 |
54
- | 6 | 13 | 10 |
55
- | 7 | 19 | 4 |
50
+ | Pass | Findings | …of which the previous pass's own fixes caused | Self-inflicted share |
51
+ |---|---|---|---|
52
+ | 4 | 12 | 5 | 42% |
53
+ | 5 | 17 | 9 | 53% |
54
+ | 6 | 13 | 10 | 77% |
55
+ | 7 | 19 | 4 | 21% — see the note below |
56
+
57
+ **The trend above is measured over passes four to six, and pass seven is not part of
58
+ it.** 42% → 53% → 77% is the decay this section teaches; the seventh pass fell back to
59
+ 4 of 19 and **the record does not say what that pass did differently.** The row stays,
60
+ with the gap named, for two reasons. Deleting a measured row to protect a claim is the
61
+ opposite of what this file asks of every gate it describes — and the vertical pass
62
+ described below cannot be credited for the drop, because it is a separate pass over the
63
+ same repository, not the seventh. Whatever pass seven did, it is unrecorded, and an
64
+ unrecorded cause is what this table has to say about it.
56
65
 
57
66
  By pass six the audit was **mostly repairing itself**. Not fatigue — arithmetic.
58
67
  Each pass edits the corpus the next pass reads, so the newest edits are always the
@@ -58,6 +58,15 @@ not the operator confirming it is what they asked for, and no amount of checking
58
58
  makes it one. Which stages are manual is the **operator's** decision, recorded in
59
59
  their `pipeline.json`; the framework fixes no stage count and no gate assignment.
60
60
 
61
+ **That sentence is about the pipeline's SHAPE, and nothing else.** It says a project
62
+ chooses how many stages it runs and which of them wait for a person. It does not say the
63
+ criteria inside a gate are per-run negotiable: where a project keeps stage 10 manual, what
64
+ that gate asks is [`acceptance.md`](acceptance.md)'s policy **`AP-1`**, which is versioned
65
+ and has an owner. The two rules stood side by side unscoped until 2026-08-20 (`B-091`), and
66
+ a reader could take either as the whole rule — *the framework fixes nothing* and *the ladder
67
+ is fixed* are both in the shipped doctrine, which is how an acceptance standard becomes
68
+ something every run re-argues.
69
+
61
70
  ## The judgment gate — a ruling is not a measurement
62
71
 
63
72
  Two types were not enough, and the gap was not cosmetic. A reviewer's ruling, a check that
@@ -369,12 +369,18 @@ deduplicated, and always exiting 0 — a hook that can fail a `Read` breaks ever
369
369
  every session. `scripts/graph.py doctrine` reports how many of the bundle's reference files
370
370
  a run opened and lists the rest.
371
371
 
372
- Why it is hook-written is the same reason as above: a claim about what somebody read,
373
- written by the party the claim is about, is not evidence. And why the verb prints
372
+ Why a hook appends it is the same reason as above: a claim about what somebody read,
373
+ written by the party the claim is about, is not evidence. **And that is an intent rather
374
+ than a proof, so the verb reports `unattested`:** the ledger is the file the agent appends
375
+ to at every stage and carries no writer field, so nothing in it separates a hook-written
376
+ line from one an agent typed. Saying *never agent-written* was a provenance claim the
377
+ format cannot support — B-014's class, in the mechanism built to close it.
378
+
379
+ And why the verb prints
374
380
  `unmeasured` rather than `0` when there are no such lines is the same reason again — the
375
381
  hook being absent and the run reading nothing are **opposite facts** the ledger cannot
376
382
  separate, so it claims neither. A `0` there would be the reassuring answer to a question
377
- nobody asked, over 34 files nobody checked.
383
+ nobody asked, over 35 files nobody checked.
378
384
 
379
385
  It is a disclosure: no floor, no direction, never a target. The moment the number becomes
380
386
  something to raise, a run will open files to raise it.
@@ -27,15 +27,29 @@ Both are checked the same way and this file covers both.
27
27
 
28
28
  ## The measured reason this file exists
29
29
 
30
- On 2026-08-11, mid-run, a monitor was armed to watch CI. One minute later the
31
- harness task inventory was queried:
30
+ On 2026-08-11, mid-run, a monitor was armed to watch CI. **What follows is two
31
+ observations of two different things, taken at two different moments**, and each
32
+ one is stamped so that a reader is not left reconciling them. One minute after
33
+ the monitor was armed, the harness task inventory was queried; the process table
34
+ was read after that:
32
35
 
33
36
  ```
34
37
  TaskList → "No tasks found"
35
38
  ps -eo pid,etime,cmd → 52693 03:12 /bin/zsh -c … gh pr checks …
36
39
  ```
37
40
 
38
- The monitor was **alive and polling**, and the inventory tool reported nothing.
41
+ `etime` is elapsed life, so the `03:12` in that line dates its own reading:
42
+ three minutes and twelve seconds after the monitor was armed, a little over two
43
+ minutes after the inventory had come back empty. The monitor was **alive and
44
+ polling**, and the inventory tool reported nothing. Both readings were accurate.
45
+ They were readings of different things, made about two minutes apart.
46
+
47
+ **What this record does not carry, said plainly: the wall-clock times.** Neither
48
+ observation was stamped with the hour at which it was taken, so what was
49
+ observed is the two offsets above — one minute to the inventory, `03:12` of
50
+ process life at the `ps` reading — and the interval between them is derived from
51
+ those two. The absolute times are not recoverable from anything this record
52
+ kept, and nothing here reconstructs them.
39
53
 
40
54
  This is the class this whole doctrine exists to catch: **a check answered by
41
55
  something that is not its subject.** An inventory that does not enumerate the
@@ -201,9 +215,9 @@ judgement, the item is reported, not ended** — which puts it back under the ru
201
215
  above rather than creating an exception to it.
202
216
 
203
217
  **A foreign item never becomes spent.** The 18 containers above belong to other
204
- projects; that they have been up for three days is information for whoever owns them,
205
- not permission. The third owner state widens what a run may clean **inside its own
206
- project** and widens nothing at all outside it.
218
+ projects; that the oldest of them has been up for three days is information for
219
+ whoever owns them, not permission. The third owner state widens what a run may
220
+ clean **inside its own project** and widens nothing at all outside it.
207
221
 
208
222
  ## Rationalizations
209
223
 
@@ -19,7 +19,7 @@ justifies reading it protects one section while the file below it doubles.
19
19
  | Artifact | Parts | How it is read |
20
20
  |---|---|---|
21
21
  | `docs/evidence/retro.md` — **one per project** | **Standing instructions** (max **10**) · **Run stamps** (max **10**, oldest rotate out) | stage 0, **in full** — both are bounded by a **cap**, which *one line each* never was |
22
- | the same file's **Recent log** | entries from the last five run stamps narrative, and capped by nothing | stage 0, **queried** by the task's nouns. It said *in full* until 2026-08-10, when it measured **74%** of the file: an uncapped section inside a binding source is what makes the capped part get skimmed |
22
+ | the same file's **Recent log** | narrative entries, **uncapped by design** — the heading said *entries from the last five run stamps* until 2026-08-20 while the section held 25 going back nine days, a bound in a heading that nothing enforced | stage 0, **queried** by the task's nouns. It said *in full* until 2026-08-10, when it measured **74%** of the file: an uncapped section inside a binding source is what makes the capped part get skimmed |
23
23
  | `docs/evidence/retro/YYYY-QN.md` — the archive | every entry and every retirement ever written, append-only | **queried** by the task's nouns; never read end to end |
24
24
 
25
25
  Seed the archive from [`../templates/retro-archive.md`](../templates/retro-archive.md).
@@ -264,8 +264,11 @@ never that the work was skipped quietly.
264
264
  decomposition is a decision, never an omission.
265
265
  - **Where the queue is a work graph, it is WRITTEN here**
266
266
  ([`work-graph.md`](work-graph.md)): `.task-pipeline/graph.json`, carrying the frozen REQ
267
- ids so `serves` resolves, one node per unit of work with its owner and what it touches,
268
- and an edge per dependency **naming what it hands over**. Then
267
+ ids so `serves` resolves, one node per unit of work with its owner, what it touches and
268
+ **the check that will close it**, and an edge per dependency **naming what it hands
269
+ over**. The check is the planner's job and not the verifier's: a node whose completion
270
+ test is decided by whoever happens to close it is a node two closes will disagree about,
271
+ and `validate` refuses one that names none. Then
269
272
  `python3 scripts/graph.py validate` — a graph that does not validate is not a queue, and
270
273
  `next` refuses to walk one. The reason to prefer it over a prose plan is measured, not
271
274
  aesthetic: a 400-node graph and a 4-node graph produce the same 27-byte frontier, so the
@@ -40,6 +40,7 @@ turn of every loop.
40
40
  | `nodes[].serves` | the REQ or goal clause it exists for. A node serving neither is **parked with that as the reason** |
41
41
  | `nodes[].blocked_by` | what must close first. This is what the frontier obeys |
42
42
  | `nodes[].touches` | what it **mutates**. Two runnable nodes writing one file is the false parallelism [`planning.md`](planning.md) refuses — *distinct is not the same as independent, and the check is what they touch, never what they are called* |
43
+ | `nodes[].check` | **how this node will be closed** — one command, or the named judgement where no command can decide it. Required on every node except a `parked` one. `agents/verifier.md` runs it and reports its output as the evidence row; before this field existed that instruction pointed at an absence, leaving the verifier the two things it forbids — invent a check, or run everything (B-080) |
43
44
  | `nodes[].evidence` | required when `status` is `done`. A node called done by assertion is what evidence exists to prevent |
44
45
  | `nodes[].parked_reason` | required when `status` is `parked`. A park with no reason is indistinguishable, a week later, from work quietly dropped |
45
46
  | `edges[].payload` | what the dependency hands over. **An edge carrying no named artifact is chronology drawn as architecture** |
@@ -68,6 +69,14 @@ the graph does. Ties break on declaration order, so the frontier is stable betwe
68
69
  an unstable one costs more than it looks, because an agent calling `next` twice starts the
69
70
  other node.
70
71
 
72
+ **One node, one completion test.** `check` is a string rather than a list, because the
73
+ requirement is singular throughout — one input, one job, one output, one owner, one
74
+ completion test. A node needing two unrelated checks is a node doing two jobs, and the
75
+ answer is to split it; one gate made of two commands is `a && b`, which is still one gate.
76
+ A **`parked`** node is the single exemption: it is the one node nobody will close, and
77
+ *n/a — parked* in that field is confidence without correctness. `park` never removes what
78
+ the node said it would run.
79
+
71
80
  **`close` stamps the commit; the verifier never supplies it.** A verdict written after the
72
81
  tree moved is evidence about a different tree, and an agent cannot name the wrong commit if
73
82
  it is never the one naming one.
@@ -114,6 +123,7 @@ does not have.
114
123
 
115
124
  | The excuse | Why it is wrong |
116
125
  |---|---|
126
+ | *"The verifier will work out what to run."* | Then the node's completion test is decided by whoever happens to close it, and two closes of one node disagree. The field is where the planner says it, and `validate` refuses a node that does not |
117
127
  | *"I'll just read the graph, it's only forty nodes."* | Forty is the number today. The property being protected is that four hundred costs the same, and it stops being true the first time the model reads the file |
118
128
  | *"The frontier is short; one extra line won't matter."* | It is paid on every iteration of every loop. That is what *the frontier and nothing else* means |
119
129
  | *"`blocked_by` already says the dependency, the edge is bookkeeping."* | The edge carries the **payload**. A dependency handing over nothing named is the fake edge with a field around it |
@@ -19,7 +19,10 @@ channel ships it. A dependency here would make the graph Claude-Code-shaped.
19
19
  states everything JSON Schema can — the required fields, `owner` non-empty, an edge's
20
20
  payload, and `done` implying non-empty evidence. What a schema cannot reach is
21
21
  cross-document and cross-node: whether an `owner` names a role that EXISTS, whether
22
- `serves` resolves, and whether the edges cycle. Those three are here. The split is
22
+ `serves` resolves, and whether the edges cycle. Those three are here. So is every rule
23
+ the format states, because the schema is never applied to a LIVE graph — including
24
+ B-080's `check`, without which `agents/verifier.md` instructs an agent to read a field
25
+ that does not exist. The split is
23
26
  where the format actually puts it, which is not where the first draft drew it.
24
27
 
25
28
  Exit codes are the contract (standing instruction R-004 — the next command is
@@ -238,6 +241,25 @@ def violations(graph):
238
241
  "nobody asked for is work the brief cannot account for, and the "
239
242
  "coverage relation cannot reach it")
240
243
 
244
+ # B-080 — the node says HOW it will be closed, and shipped doctrine reads it.
245
+ # `agents/verifier.md` told the verifier to run *the check the task named* while
246
+ # the schema had no field to name one, so the instruction pointed at an absence
247
+ # and the verifier's only options were the two that paragraph forbids: invent a
248
+ # check, or run everything. A `parked` node is the one exemption — it is the one
249
+ # node nobody will close, and a placeholder there would be worse than the gap.
250
+ if n.get("status") != "parked":
251
+ chk = n.get("check")
252
+ if not isinstance(chk, str) or not chk.strip():
253
+ out.append(f"{nid}: `check` is {chk!r} — B-080: a node that cannot say how "
254
+ "it will be closed leaves the verifier inventing a check or "
255
+ "running everything, and `agents/verifier.md` forbids both. Name "
256
+ "the command, or the judge where no command can decide it. Only a "
257
+ "`parked` node is exempt")
258
+ elif any(c in chk for c in "\n\r"):
259
+ out.append(f"{nid}: `check` contains a line break. A completion test that is "
260
+ "two commands cannot say which one closed the node, and the "
261
+ "verifier reports its output as one evidence row")
262
+
241
263
  if n.get("status") == "done":
242
264
  ev = n.get("evidence")
243
265
  if not isinstance(ev, list) or not [e for e in ev
@@ -480,6 +502,16 @@ def verdict_violations(v):
480
502
  for nid in rp.get("add") or []:
481
503
  if not isinstance(nid, dict) or "title" not in nid:
482
504
  out.append("replan.add entries must be nodes with at least a `title`")
505
+ continue
506
+ # B-080 again, at the one place a node is created by a verdict rather than by
507
+ # a person. Checked HERE so the refusal names the key before anything is
508
+ # written: leaving it to `violations()` after the close would abort the whole
509
+ # close over a field the verdict could have been asked for.
510
+ _c = nid.get("check")
511
+ if not isinstance(_c, str) or not _c.strip():
512
+ out.append(f"replan.add entry {nid.get('title')!r} names no `check` — the "
513
+ "node it creates has to say how IT will be closed, or the next "
514
+ "verifier faces the absence this one was told to read")
483
515
  return out
484
516
 
485
517
 
@@ -679,7 +711,14 @@ def cmd_add(graph, args):
679
711
  die("owner %r is not a role this pipeline ships%s — nothing was written"
680
712
  % (args.owner, (" — did you mean %s?" % near[0]) if near else ""))
681
713
 
682
- for name, val in (("--title", title), ("--serves", serves)):
714
+ check = (args.check or "").strip()
715
+ if not check:
716
+ die("`add` needs a --check with something in it: the command that closes this node, "
717
+ "or the judge where no command can decide it. A node that cannot say how it "
718
+ "will be closed leaves the verifier inventing a check or running everything — "
719
+ "B-080, and `agents/verifier.md` forbids both. Nothing was written")
720
+
721
+ for name, val in (("--title", title), ("--serves", serves), ("--check", check)):
683
722
  if any(c in val for c in "\n\r"):
684
723
  die("%s contains a line break. `next` prints one row per node and the loop "
685
724
  "reads those rows, so a break here forges a row for a node that does not "
@@ -729,7 +768,7 @@ def cmd_add(graph, args):
729
768
  "allocated from the highest in use" % nid)
730
769
 
731
770
  new = {"id": nid, "title": title, "owner": args.owner, "status": "pending",
732
- "blocked_by": blocked, "serves": serves, "evidence": None}
771
+ "blocked_by": blocked, "serves": serves, "check": check, "evidence": None}
733
772
  touches = list(dict.fromkeys(t.strip() for t in (args.touches or []) if t.strip()))
734
773
  if touches:
735
774
  new["touches"] = touches
@@ -870,13 +909,21 @@ def cmd_producer(graph, args):
870
909
  def cmd_doctrine(graph, args):
871
910
  """Which doctrine this run actually read — B-061.
872
911
 
873
- The bundle is 34 reference files. A run reads some subset and nothing recorded which,
912
+ The bundle is 35 reference files. A run reads some subset and nothing recorded which,
874
913
  so **a skipped file and a read one were indistinguishable** — the class every guard in
875
914
  this repository exists to catch, left standing over the doctrine itself.
876
915
 
877
- `read:` lines in the run ledger are written by a hook, never by the agent, for the same
878
- reason `gate:` is: a claim about what somebody read, written by the party the claim is
879
- about, is not evidence.
916
+ `read:` lines are written by a hook rather than by the agent, for the same reason `gate:`
917
+ is: a claim about what somebody read, written by the party the claim is about, is not
918
+ evidence.
919
+
920
+ **And that is an intent, not a proof, so every line here is reported as UNATTESTED.**
921
+ The ledger is `.task-pipeline/run.md` — the file the agent appends to at every stage —
922
+ so nothing in it distinguishes a hook-written line from one an agent typed. The doctrine
923
+ said *hook-written, never agent-written*, which is a provenance claim this script cannot
924
+ check and no format here carries; B-014's class, committed by the mechanism built to
925
+ close it. Until the ledger can attest a writer, the honest output is the count plus the
926
+ word: read as *this is what the ledger says, and the ledger cannot say who wrote it*.
880
927
 
881
928
  **The one rule that matters here: no `read:` lines means UNMEASURED, never «read
882
929
  nothing».** Zero would be the reassuring answer to a question nobody asked, and this
@@ -919,7 +966,10 @@ def cmd_doctrine(graph, args):
919
966
  return 0
920
967
 
921
968
  unread = [r for r in refs if r not in read]
922
- print(f"doctrine: {len(read)} of {len(refs)} reference files read")
969
+ print(f"doctrine: {len(read)} of {len(refs)} reference files read — unattested")
970
+ print(" unattested: the ledger is the file the agent appends to at every stage, "
971
+ "so nothing in it proves the hook wrote these lines rather than an agent. The "
972
+ "count is what the ledger says; who wrote it is not recorded.")
923
973
  print(" a disclosure: no floor, no direction, never a target. A run that needs "
924
974
  "four files and reads four is not worse than one that reads thirty.")
925
975
  for r in unread:
@@ -993,14 +1043,17 @@ def cmd_close(graph, args):
993
1043
  title = str(spec.get("title", "")).strip()
994
1044
  owner = str(spec.get("owner", "implementer")).strip()
995
1045
  serves = str(spec.get("serves", node.get("serves"))).strip()
1046
+ check = str(spec.get("check", "")).strip()
996
1047
  if not title:
997
1048
  die("replan.add carries an entry with no title — nothing was written")
1049
+ if not check:
1050
+ die("replan.add carries an entry with no check — nothing was written")
998
1051
  if owner not in ROLES:
999
1052
  die("replan.add names owner %r, which is not a role this pipeline ships" % owner)
1000
1053
  new_id = next_id({n.get("id") for n in graph["nodes"]})
1001
1054
  graph["nodes"].append({"id": new_id, "title": title, "owner": owner,
1002
1055
  "status": "pending", "blocked_by": [], "serves": serves,
1003
- "evidence": None})
1056
+ "check": check, "evidence": None})
1004
1057
  revise(graph, "add", new_id, why or ("re-planned by the verdict on " + nid))
1005
1058
  added.append(new_id)
1006
1059
  for pid in rp.get("park") or []:
@@ -1074,6 +1127,10 @@ def main(argv=None):
1074
1127
  p_add.add_argument("--title", required=True)
1075
1128
  p_add.add_argument("--owner", required=True)
1076
1129
  p_add.add_argument("--serves", required=True)
1130
+ p_add.add_argument("--check", required=True,
1131
+ help="the command that closes this node, or the judge where no "
1132
+ "command can decide it — the verifier runs it and reports its "
1133
+ "output as the evidence row")
1077
1134
  p_add.add_argument("--blocked-by", dest="blocked_by", action="append", default=[])
1078
1135
  p_add.add_argument("--carries", action="append", default=[],
1079
1136
  help="what each --blocked-by hands over; pairs in the order written")
@@ -31,7 +31,7 @@ has not fired in the last five run stamps, or in the last sixty days. At eleven
31
31
  rows, the oldest never-fired
32
32
  row goes — the cap is not negotiable, ranking is.
33
33
 
34
- ## Recent log — entries from the last five run stamps (newest first)
34
+ ## Recent log — narrative entries, uncapped and queried rather than read (newest first)
35
35
 
36
36
  Older entries and every retirement **move** to `docs/evidence/retro/YYYY-QN.md`
37
37
  at the prune. Moving is not deleting: the archive is append-only and holds the
@@ -20,23 +20,35 @@ Run: `<topic>` · started `<YYYY-MM-DD>` · module map: `<path or "none">`
20
20
 
21
21
  ## `read:` — which doctrine this run actually opened
22
22
 
23
- The bundle is 34 reference files and nothing recorded which of them a run read, so **a
23
+ The bundle is 35 reference files and nothing recorded which of them a run read, so **a
24
24
  skipped file and a read one were indistinguishable** — the class every guard in this
25
25
  pipeline exists to catch, left standing over the doctrine itself.
26
26
 
27
- `read:` is **hook-written, never agent-written**, for the same reason `gate:` is: a claim
28
- about what somebody read, written by the party the claim is about, is not evidence. The
27
+ `read:` is **written by a hook rather than by the agent**, for the same reason `gate:` is: a
28
+ claim about what somebody read, written by the party the claim is about, is not evidence. The
29
29
  hook is in `hooks.example.json`, matches `Read`, records the path only
30
30
  when it is inside the bundle's `references/`, deduplicates, and **always exits 0** — a hook
31
31
  that can fail a `Read` would break every turn in every session.
32
32
 
33
+ **That is an intent, not a proof, and the line says so.** This file is the one the agent
34
+ appends to at every stage, so nothing in it distinguishes a hook-written line from one an
35
+ agent typed: there is no writer field, and `scripts/graph.py` has no provenance check
36
+ because the format gives it nothing to check. The doctrine here said *hook-written, never
37
+ agent-written* until 2026-08-20 — a provenance claim in the file whose whole subject is that
38
+ a claim by the interested party is not evidence, which is B-014's class committed by the
39
+ mechanism built to close it. So `doctrine` and the `gate:` reader both report **`unattested`**
40
+ beside their counts, and the two claims that survive are the ones the ledger can support:
41
+ *this line is in the ledger*, and *nobody recorded who wrote it*. Attesting a writer needs a
42
+ field this format does not have; inventing one that an agent can also fill would restate the
43
+ same claim one level down.
44
+
33
45
  `scripts/graph.py doctrine` reads these lines and prints one of three things:
34
46
 
35
47
  | It prints | When | Why not just a number |
36
48
  |---|---|---|
37
49
  | `unmeasured — no run ledger` | there is no ledger | nothing to read from |
38
50
  | `unmeasured — the ledger carries no read: lines` | the hook is absent, **or** the run opened no doctrine | two opposite facts, and the ledger cannot separate them, so neither is claimed |
39
- | `N of 34 reference files read`, then each unread one | the hook is installed and fired | the count alone says there is a gap, not where |
51
+ | `N of 35 reference files read — unattested`, then each unread one | the hook is installed and fired | the count alone says there is a gap, not where — and `unattested` says the ledger cannot name who wrote the lines |
40
52
 
41
53
  **It is a disclosure: no floor, no direction, never a target.** A run that needs four files
42
54
  and reads four is not worse than one that reads thirty — and the moment the number becomes
@@ -58,15 +70,17 @@ hand: <N|10> — task "<quoted>" — done <n> — surfaced <n> — decisions <n
58
70
  holds: <stage id> — <n> (<class: what, owner>; … or "none") — enumerated <n>/8 classes, <unlooked: classes not enumerable>
59
71
  gate: <stage id> — command "<cmd>" — exit <N> — <ISO-8601>
60
72
  event: <compact|session-end|subagent> — <detail> — <ISO-8601>
61
- read: references/<file>.md # hook-written, deduplicated, one per file opened
73
+ read: references/<file>.md # hook-appended, deduped, UNATTESTED (no writer field)
62
74
  ```
63
75
 
64
76
  - **`stage:`** — written when a gate **returns**, not when the stage is entered. The
65
77
  rail's `✓` is derived from this line and from nothing else; a glyph set from memory
66
78
  is a summary that is confidently wrong exactly when it matters.
67
- - **`gate:`** — written by `hooks/gate-observer.sh`, never by an agent. It is the
68
- only line here that records what a command **did** rather than what somebody
69
- concluded: the exit code of the stage's declared `gate.command`, observed. The
79
+ - **`gate:`** — appended by `hooks/gate-observer.sh` rather than by an agent, and
80
+ **unattested** for the reason `read:` is: this file is agent-written at every stage
81
+ and carries no writer field, so the line's provenance is an intent the format cannot
82
+ prove. It is the only line here that records what a command **did** rather than what
83
+ somebody concluded: the exit code of the stage's declared `gate.command`, observed. The
70
84
  `stage:` line above it is the agent's claim, and the release gate requires the
71
85
  two to agree — without this, a gate reads a claim written by the party it
72
86
  constrains and confirms an assertion with itself. Absent where the project
@@ -41,6 +41,7 @@ worse than saying nothing.
41
41
  ## Contents
42
42
 
43
43
  - [Staleness — a row is true about the tree it OBSERVED](#staleness--a-row-is-true-about-the-tree-it-observed)
44
+ - [Environment — a proof is only valid where it ran](#environment--a-proof-is-only-valid-where-it-ran)
44
45
  - the ledger itself — one row per REQ, appended by stage 8
45
46
  - [What `Human` means, and what it does not](#what-human-means-and-what-it-does-not)
46
47
 
@@ -71,11 +72,39 @@ one. Four things overtake a row, and naming which one applies is the note's job:
71
72
  change in what it covers, a **dependency** change, an **environment** change, and a
72
73
  **policy** change — the last being the rule under which the evidence was accepted.
73
74
 
74
- | REQ | What | Run | Shipped in | Observed at | Auto | Human | Note |
75
- |---|---|---|---|---|---|---|---|
76
- | REQ-001 | CSV export from a report | `2026-07-28-export` | v1.4.0 | `5f21ac3` | pass | 2026-07-30 | opened the deployed page, exported, opened the file |
77
- | REQ-004 | XLSX export | `2026-07-28-export` | v1.4.0 | `5f21ac3` | pass | **never** | — |
78
- | REQ-007 | Export respects active filters | `2026-07-28-export` | v1.4.0 | `5f21ac3` | partial | **never** | CSV path only |
75
+ ## Environment a proof is only valid where it ran
76
+
77
+ `Observed at` says which tree the check saw. It does not say **where**, and without that a
78
+ smoke test against a preview URL enters the record in a shape indistinguishable from one
79
+ against production, and a suite green on a laptop with accumulated state is
80
+ indistinguishable from one green on a runner that started clean. Those are the two halves
81
+ of the same question, and only one was instrumented.
82
+
83
+ **`Environment` is a required cell on every row.** Its vocabulary is the project's own,
84
+ declared in the runbook and not invented per row — this file ships with four, and a project
85
+ that deploys differently declares different ones:
86
+
87
+ | Value | What it means |
88
+ |---|---|
89
+ | `production` | the deployed target real users reach |
90
+ | `preview` | a per-branch or per-PR deployment; proves the build, never the release |
91
+ | `ci` | a clean runner, no accumulated state |
92
+ | `local` | a developer machine, with whatever state it has |
93
+ | `—` | **recorded absence** — a row written before this column existed, or one whose environment nobody recorded. Not a value, and never a default to reach for |
94
+
95
+ A missing cell is not the same as `—`: the first is a row that forgot the question, and it is
96
+ **refused**. The second is an answer.
97
+
98
+ **A REQ claiming production behaviour may not be closed on a non-production
99
+ observation.** Stage 8 writes the row and refuses that pairing rather than recording it —
100
+ `pass` in `ci` against a requirement about the deployed product is the exact substitution
101
+ this column exists to make visible.
102
+
103
+ | REQ | What | Run | Shipped in | Observed at | Environment | Auto | Human | Note |
104
+ |---|---|---|---|---|---|---|---|---|
105
+ | REQ-001 | CSV export from a report | `2026-07-28-export` | v1.4.0 | `5f21ac3` | production | pass | 2026-07-30 | opened the deployed page, exported, opened the file |
106
+ | REQ-004 | XLSX export | `2026-07-28-export` | v1.4.0 | `5f21ac3` | ci | pass | **never** | — |
107
+ | REQ-007 | Export respects active filters | `2026-07-28-export` | v1.4.0 | `5f21ac3` | preview | partial | **never** | CSV path only |
79
108
 
80
109
  ## Columns
81
110
 
@@ -86,6 +115,10 @@ change in what it covers, a **dependency** change, an **environment** change, an
86
115
  - **Run** — the brief's topic slug, so the context is one file away.
87
116
  - **Shipped in** — the tag or commit that carried it. Where a project does not tag,
88
117
  the commit, and the same value every row of that run carries.
118
+ - **Observed at** — the commit the check ran against. `—` where nobody recorded one.
119
+ - **Environment** — where it ran, from the project's declared vocabulary. Required; `—` is
120
+ the recorded absence and an omitted cell is refused. See *Environment — a proof is only
121
+ valid where it ran*.
89
122
  - **Auto** — what the run's own gate said: `pass` · `partial` · `none`. Copied from the
90
123
  coverage table rather than re-derived; where the two disagree the coverage table wins
91
124
  and the disagreement is a finding. A coverage verdict of **`review`** — *no check can