task-pipeline-skill 1.71.1 → 1.73.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +128 -2
- package/CONTRIBUTING.md +74 -0
- package/SKILL-CARD.md +1 -1
- package/package.json +1 -1
- package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
- package/plugins/task-pipeline/agents/verifier.md +20 -5
- package/plugins/task-pipeline/skills/task-pipeline/graph.example.json +3 -0
- package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +96 -49
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +23 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +15 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/gates.md +9 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +9 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/residue.md +20 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/retrospective.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +5 -2
- package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +10 -0
- package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +66 -9
- package/plugins/task-pipeline/skills/task-pipeline/templates/retro.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +22 -8
- package/plugins/task-pipeline/skills/task-pipeline/templates/verification.md +38 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,119 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v1.73.0 — 2026-08-20 — the registry could not see the templates, the scripts, or its own registers
|
|
4
|
+
|
|
5
|
+
|
|
6
|
+
**The claim registry had the class for this exact incident and it fired on nothing.** Three
|
|
7
|
+
shipped surfaces said *34 reference files* over a directory of 35 — `scripts/graph.py`'s
|
|
8
|
+
`doctrine` docstring and `templates/run.md` twice — and `npm test` printed
|
|
9
|
+
`reference files: dormant (truth 35)`. Two holes, either of which was enough: the pattern
|
|
10
|
+
knew only the word order ``N files under `references/` ``, and the corpus was eight named
|
|
11
|
+
files plus `references/**`, so `templates/` and `scripts/` were never opened. The class now
|
|
12
|
+
reads any phrasing of the count, the corpus reads the templates and the shipped script, and
|
|
13
|
+
the class is armed at three agreeing sites instead of dormant.
|
|
14
|
+
|
|
15
|
+
**And it never read this repository's own registers.** `docs/DOCMAP.md` names decisions,
|
|
16
|
+
open questions, the board, the ledger and the retro as the registers, and its propagation
|
|
17
|
+
matrix sends *a number stated in a living document* to this registry — which could not see
|
|
18
|
+
any of them. `docs/OPEN_QUESTIONS.md` said *the 250 guards* against a workflow defining 390,
|
|
19
|
+
in a phrasing the guard class already knew. Bringing them in refused 26 statements and every
|
|
20
|
+
one was narration, so a number inside a **dated item** is a record and exempt; on the board
|
|
21
|
+
the discriminator is the State cell rather than the date, because every row names the day it
|
|
22
|
+
was filed and B-001 — open — was stating its description budget and its reference-file count
|
|
23
|
+
as facts about now.
|
|
24
|
+
|
|
25
|
+
**Fourteen consecutive releases carry no run stamp**, `v1.60.1` through `v1.72.0`, and the
|
|
26
|
+
retro's honest-gap section named only `v1.16.0`–`v1.23.0`. Its *Measured, not recalled*
|
|
27
|
+
receipt could not produce the measurement: it grepped `docs/superpowers/retro.md`, removed at
|
|
28
|
+
v1.53.0, and grepped for a tag's own commit when a stamp names the commit the *run* ended on.
|
|
29
|
+
Rewritten as a tag-range walk that reads the archive too, and a guard now requires every
|
|
30
|
+
release after the newest stamp to be named in that section.
|
|
31
|
+
|
|
32
|
+
Also: `## Unreleased` is where the guard count lives between a tag and the next bump (B-104);
|
|
33
|
+
`Environment` is a required cell on every verification row with the vocabulary read out of
|
|
34
|
+
the shipped template (B-099's other half); the graph schema states its three node rules
|
|
35
|
+
behind `$ref`s and the checker follows them (B-079, proved by moving them, not by a fixture);
|
|
36
|
+
an open board row's `file:N-M` must quote the phrase it points at, which caught five stale
|
|
37
|
+
citations and then caught this change's own edits four more times; the acceptance ladder is
|
|
38
|
+
policy **`AP-1`** with an owner and an in-force date, and `gates.md`'s *the framework fixes
|
|
39
|
+
no stage count* is scoped to the pipeline's shape; `read:` and `gate:` are reported
|
|
40
|
+
**unattested** instead of claimed *never agent-written*, because the ledger is the file the
|
|
41
|
+
agent appends to at every stage; `validate.yml` no longer re-validates a SHA on its own tag
|
|
42
|
+
push; and every documented `npm` equation is compared against `package.json` — `CLAUDE.md`
|
|
43
|
+
had `npm test` as `validate.py` alone, dropping 129 graph cases.
|
|
44
|
+
|
|
45
|
+
**What the suite found that no reading did.** This change broke **thirteen** existing
|
|
46
|
+
plants and `npm run test:all` named every one: the two verification-header probes spell
|
|
47
|
+
the whole header and it gained a column; two CHANGELOG probes scoped themselves to
|
|
48
|
+
`^## v` and the count now lives in `## Unreleased`; **seven** graph-schema probes walk
|
|
49
|
+
`node.allOf` inline and the rules moved behind `$ref`. All thirteen were repaired by
|
|
50
|
+
deriving the guard's own scope rather than restating it, which is what `learned.md`'s
|
|
51
|
+
*sweep the class* asks for — and the sweep needed two rounds: six schema probes were
|
|
52
|
+
repaired together and the seventh surfaced on the next full run, having died on
|
|
53
|
+
`KeyError: 'if'` where the others had died on their own asserts. A class fixed in six of
|
|
54
|
+
seven places is the shape standing instruction R-003 exists for.
|
|
55
|
+
|
|
56
|
+
One of the twelve was sharper than the rest, and it was self-inflicted twice over: the
|
|
57
|
+
honesty note explaining that `OQ-0002` no longer restates a total put an ISO date in the
|
|
58
|
+
row, which made the row a dated record and **disarmed the register plant** — green over
|
|
59
|
+
exactly the stale total it had been written to catch. `OQ-####` now uses its `Status`
|
|
60
|
+
cell, the discriminator the board already uses, and the register plant sits on an open
|
|
61
|
+
board row where no prose edit beside it can turn it off.
|
|
62
|
+
|
|
63
|
+
Guards: 390 → **412** · property checks 9 → **14**
|
|
64
|
+
|
|
65
|
+
|
|
66
|
+
## v1.72.0 — a node says how it will be closed
|
|
67
|
+
|
|
68
|
+
**B-080 closed, and with it the last of four requirements this pack's own manifesto named
|
|
69
|
+
as unbuilt.** The Proof of Done manifesto cites this repository as its reference
|
|
70
|
+
implementation and filed four gaps against it by id; three closed on 2026-08-17, and this
|
|
71
|
+
is the fourth. The public document now says all four are built and names the commit for
|
|
72
|
+
each.
|
|
73
|
+
|
|
74
|
+
### `check` — the per-node completion test
|
|
75
|
+
|
|
76
|
+
A node could not say how it would be closed, while shipped doctrine told the verifier to
|
|
77
|
+
"run the checks the task named". The doctrine read a field the schema did not have.
|
|
78
|
+
|
|
79
|
+
`check` is a string, **required on every node whose status is not `parked`** — the parked
|
|
80
|
+
node being the one nobody will close, where a placeholder would be confidence without
|
|
81
|
+
correctness. Never a list: a node needing two unrelated checks is a node doing two jobs, and
|
|
82
|
+
one gate made of two commands is `a && b`, which is still one gate. Stated twice on purpose,
|
|
83
|
+
in the schema and in `violations()`, because the schema never runs against a live graph.
|
|
84
|
+
`add --check` is required like `--why`, and a replan entry naming no check is refused before
|
|
85
|
+
the close writes anything.
|
|
86
|
+
|
|
87
|
+
### The record the manifesto opens with now reads as one event
|
|
88
|
+
|
|
89
|
+
`residue.md`'s opening record is quoted by an external document and permalinked from a public
|
|
90
|
+
site. It said the inventory was queried "one minute later" beside a `ps` line showing the
|
|
91
|
+
process at 3:12 — two accurate observations of different things, framed so a reader had to
|
|
92
|
+
reconcile them by hand. Both moments are now stated, the interval between them named, and
|
|
93
|
+
**what the record never carried is named too**: the wall-clock times, which `git log --follow`
|
|
94
|
+
proves were never in it. A disambiguation, not a correction — no number changed.
|
|
95
|
+
|
|
96
|
+
The E4 sentence claimed more than its own block: "that **they** have been up for three days"
|
|
97
|
+
from a measurement of the oldest one. Now "the oldest of them".
|
|
98
|
+
|
|
99
|
+
### Three plants that could not run, and one row that could not be parsed
|
|
100
|
+
|
|
101
|
+
CI failed twice on this release's work, and neither failure was noise.
|
|
102
|
+
|
|
103
|
+
A ledger row wrote a pytest filter as a backticked `-k a|b|c`; markdown does not care about
|
|
104
|
+
backticks, so two pipes split one cell into three and that row carried ten cells against a
|
|
105
|
+
header of eight. Locally the damage sat in a column nothing read, so every gate run passed.
|
|
106
|
+
Every `REQ` row must now match the header's cell count, with the remedy named in the refusal.
|
|
107
|
+
|
|
108
|
+
Then three schema plants died on `KeyError` before reaching their own `PLANT DID NOT LAND`
|
|
109
|
+
assert, because the new `check` rule added an `allOf` branch keyed on `not: {const: parked}`
|
|
110
|
+
and each plant indexed the const directly. All three traverse tolerantly now. **The assert
|
|
111
|
+
catches a plant that does not land; nothing catches a plant that cannot run** — the third
|
|
112
|
+
instance of that class in three days, and it is filed family-wide rather than here.
|
|
113
|
+
|
|
114
|
+
Guards: 389 → **390**. The new rule arrived without a plant, which this repository's own standard refuses; that is the plant. `graph.py` 114 → 129 cases.
|
|
115
|
+
|
|
116
|
+
|
|
3
117
|
|
|
4
118
|
## v1.71.1 — four shapes of a green that measured nothing
|
|
5
119
|
|
|
@@ -55,8 +169,20 @@ only reader that checks.
|
|
|
55
169
|
and what makes each self-checking. Where a claim genuinely cannot be derived, **say so in the
|
|
56
170
|
row**; correcting it silently a second time teaches readers that somebody else is checking.
|
|
57
171
|
|
|
58
|
-
|
|
59
|
-
|
|
172
|
+
> **After the tag, on the same tree.** `v1.71.1` is cut, and the guard-count claim
|
|
173
|
+
> `test/validate.py` reads is the newest `## vX.Y.Z` section's — which post-tag is this one,
|
|
174
|
+
> with no version heading yet open for the work that follows. So the live figure is stated
|
|
175
|
+
> here rather than inside a released sentence, and the released sentence below is not edited:
|
|
176
|
+
> B-080 closed on 2026-08-19 with eight plants and TP-02 with five on the same day, so
|
|
177
|
+
> Guards: 376 → **389**. The guard has no home for a count between a tag and the next bump;
|
|
178
|
+
> filed as `B-104`, and the note is what keeps the number true meanwhile. **The bold belongs
|
|
179
|
+
> on the number alone**: the plant that corrupts this count reads `Guards: N → **M**` out of
|
|
180
|
+
> the raw section, so bolding the whole phrase left the validator reading a count no plant
|
|
181
|
+
> could still corrupt — which is the form this note was introduced with, and no second
|
|
182
|
+
> count-shaped figure belongs in this section for the same reason.
|
|
183
|
+
|
|
184
|
+
Guards at the tag: 376 → 376. No guard was added or removed — that release is doctrine, and
|
|
185
|
+
each of the four rules is cited from the reference that owns it rather than restated.
|
|
60
186
|
|
|
61
187
|
Closes #39, #40, #48, #49.
|
|
62
188
|
|
package/CONTRIBUTING.md
CHANGED
|
@@ -518,6 +518,80 @@ runs.
|
|
|
518
518
|
*(guard: `does not say` and `rule 1 of the list it is printed beside`
|
|
519
519
|
and `no stage names`)*
|
|
520
520
|
|
|
521
|
+
**56. A node says how it will be closed, and the doctrine that reads that field names
|
|
522
|
+
one the node has.** `agents/verifier.md` ordered the verifier to run *the checks the task
|
|
523
|
+
named* while `graph.schema.json`'s node carried no field in which a task could name one —
|
|
524
|
+
two files shipped on one day, and the instruction pointed at an absence that left the
|
|
525
|
+
verifier the two options the same paragraph forbids: invent a check, or run everything.
|
|
526
|
+
`check` is now a node property, **required on every node except a `parked` one** — the one
|
|
527
|
+
node nobody will close, where a placeholder would be worse than the gap. The rule is
|
|
528
|
+
stated twice on purpose, in the schema and in `violations()`, because the schema is never
|
|
529
|
+
applied to a live graph; both are RUN against a planted graph rather than inspected. And
|
|
530
|
+
the two homes are compared directly: whatever the verifier is told to read off the node
|
|
531
|
+
must be a property the schema declares.
|
|
532
|
+
*(guard: `node declares no` and `rule that can fire` and `off the node, and`)*
|
|
533
|
+
|
|
534
|
+
**57. The claim registry reads every surface that states a number, including the ones this
|
|
535
|
+
repository writes about itself.** Its corpus was eight named files plus `references/**`, so
|
|
536
|
+
`templates/` and `scripts/` were invisible — three shipped surfaces said "34 reference files"
|
|
537
|
+
over a directory of 35 while the class printed `dormant` — and so were the registers
|
|
538
|
+
`docs/DOCMAP.md` names, where `docs/OPEN_QUESTIONS.md` said "the 250 guards" against a
|
|
539
|
+
workflow defining 390. The corpus now covers both, and a class recognises every phrasing of
|
|
540
|
+
its count rather than one word order. A number inside a **dated item** in a register is a
|
|
541
|
+
record and exempt; on the board the discriminator is the **State** cell, because every row
|
|
542
|
+
names the day it was filed and an open row is a claim about now.
|
|
543
|
+
*(guard: `reference files` and `— derive the number or delete it`)*
|
|
544
|
+
|
|
545
|
+
**58. Every documented `npm` command means what `package.json` runs.** `CLAUDE.md` glossed
|
|
546
|
+
`npm test` as `python3 test/validate.py`, dropping `graph_test.py` and its 129 cases — the
|
|
547
|
+
suite-outside-the-run class stated the other way round — and called `npm run test:all` "both"
|
|
548
|
+
where it runs eight scripts. An equation is compared against the script body after one level
|
|
549
|
+
of `npm run` resolution, and a bare `npm run X` in a document about this repository must be a
|
|
550
|
+
script that exists. Portable doctrine under `plugins/` and `cursor/` is out of scope: it names
|
|
551
|
+
a host project's commands.
|
|
552
|
+
*(guard: `is glossed as` and `declares no such script`)*
|
|
553
|
+
|
|
554
|
+
**59. A release either carries a run stamp or is recorded as a gap, and the guard-count claim
|
|
555
|
+
has a home before the bump.** Fourteen consecutive releases had no stamp while the retro named
|
|
556
|
+
only `v1.16.0`–`v1.23.0`, and the receipt that was supposed to prove it grepped a path removed
|
|
557
|
+
at v1.53.0 for a tag's own commit — a stamp names the commit the *run* ended on. Scoped to the
|
|
558
|
+
trailing stretch: 84 of 117 tags predate the register and backfilling is forbidden. Separately,
|
|
559
|
+
the count guard reads the topmost `## ` section, so `## Unreleased` is where the number lives
|
|
560
|
+
between a tag and the next bump, and it must sit above every version heading.
|
|
561
|
+
*(guard: `release(s) after the newest run stamp` and `section sits below a released version`)*
|
|
562
|
+
|
|
563
|
+
**60. A `file:line` range in an open board row quotes the phrase it points at.** Five
|
|
564
|
+
citations resolved to real lines and pointed at other text; a line number is the most fragile
|
|
565
|
+
address a document carries, because every edit above it moves it and nothing notices. Closed
|
|
566
|
+
rows are records and are left alone; single-line citations cannot be quoted and are disclosed
|
|
567
|
+
as unanchored rather than failed.
|
|
568
|
+
*(guard: `quotes no phrase from it`)*
|
|
569
|
+
|
|
570
|
+
**61. Every evidence row records the environment it ran in, and a claim of provenance the
|
|
571
|
+
format cannot check is marked unattested instead.** `Observed at` said which tree a check saw
|
|
572
|
+
and nothing said where it ran, so a preview smoke test and a production one entered the record
|
|
573
|
+
in the same shape — in a pack whose own `learned.md` records a suite green on every author's
|
|
574
|
+
machine and 1039 failures on a clean runner. And `read:`/`gate:` were declared *hook-written,
|
|
575
|
+
never agent-written* while both land in the file the agent appends to at every stage: no writer
|
|
576
|
+
field, no provenance check, so the count is reported `unattested` and the claim is not made.
|
|
577
|
+
*(guard: `has no `Environment` cell` and `prints a count and never says`)*
|
|
578
|
+
|
|
579
|
+
**62. The acceptance ladder is a versioned policy with an owner, and `gates.md`'s
|
|
580
|
+
fixes-nothing sentence is scoped to the pipeline's shape.** Both rules stood unscoped side by
|
|
581
|
+
side for seventy releases — *the framework fixes no stage count and no gate assignment* beside
|
|
582
|
+
twelve fixed criteria — so a reader could take either as the whole rule, and a table accepted
|
|
583
|
+
under v1.20 doctrine was indistinguishable from one accepted under v1.70. The block carries
|
|
584
|
+
`AP-1`, an owner and an in-force date; an amendment moves the version and lands with a decision
|
|
585
|
+
row.
|
|
586
|
+
*(guard: `the acceptance policy carries no` and `stands unscoped beside`)*
|
|
587
|
+
|
|
588
|
+
**63. A heading may not declare a bound nothing enforces.** The retro's *Recent log* read
|
|
589
|
+
*entries from the last five run stamps* over 25 entries reaching back nine days, borrowing the
|
|
590
|
+
stamp section's wording without its cap — filed twice as B-060 and B-069 and disclosed by the
|
|
591
|
+
file about itself. Checked in the live retro and in `templates/retro.md`, which seeded the
|
|
592
|
+
false bound into every host project.
|
|
593
|
+
*(guard: `declares the bound` and `and nothing enforces it`)*
|
|
594
|
+
|
|
521
595
|
## Adding or changing doctrine
|
|
522
596
|
|
|
523
597
|
- **Change one idea per PR.** These files are read by agents under load; a PR that
|
package/SKILL-CARD.md
CHANGED
|
@@ -12,7 +12,7 @@ harmless.
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
|
|
14
14
|
| **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
|
|
15
|
-
| **Version** | 1.
|
|
15
|
+
| **Version** | 1.73.0 |
|
|
16
16
|
| **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
|
|
17
17
|
| **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki, and **one of two browser channels** — `playwright` (CLI or MCP) or `chrome-devtools` (MCP); either satisfies the browser step and neither is required. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
|
|
18
18
|
| **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline-skill",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.73.0",
|
|
4
4
|
"description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"task-pipeline": "bin/task-pipeline.js"
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"name": "task-pipeline",
|
|
3
3
|
"displayName": "Task Pipeline",
|
|
4
4
|
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
|
|
5
|
-
"version": "1.
|
|
5
|
+
"version": "1.73.0",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "ssheleg",
|
|
8
8
|
"url": "https://x.com/sshlg93"
|
|
@@ -34,7 +34,10 @@ than ending in a question nobody will see.
|
|
|
34
34
|
"not_done": ["what was asked and is not"],
|
|
35
35
|
"not_verified": ["what was BUILT and no check touched"],
|
|
36
36
|
"blockers": [{ "what": "…", "blocks": ["N-009"], "can_continue_around": true }],
|
|
37
|
-
"replan": { "possible": true,
|
|
37
|
+
"replan": { "possible": true,
|
|
38
|
+
"add": [{ "title": "…", "owner": "implementer",
|
|
39
|
+
"serves": "REQ-004", "check": "…" }],
|
|
40
|
+
"park": ["N-009"], "why": "…" },
|
|
38
41
|
"evidence": ["the command and the output that proves each `done` row"]
|
|
39
42
|
}
|
|
40
43
|
```
|
|
@@ -58,12 +61,20 @@ ledger exists to catch.
|
|
|
58
61
|
|
|
59
62
|
1. **Read the node's `serves`** — the REQ or the goal clause. That is the standard.
|
|
60
63
|
Not what the diff does; what was asked.
|
|
61
|
-
2. **Run the
|
|
62
|
-
|
|
63
|
-
|
|
64
|
+
2. **Run the node's `check`.** It is a field on the node — one command, or a named
|
|
65
|
+
judgement where no command can decide it — and it is what closes this node. Not a
|
|
66
|
+
check you invented, and not `npm test` alone when the node named something narrower:
|
|
67
|
+
a green from a check nobody watched fail against a planted defect is not evidence.
|
|
68
|
+
Where the `check` is a judgement, record the verdict **as** judgement
|
|
69
|
+
(`references/gates.md`), never as an exit code.
|
|
70
|
+
**A node with no `check` is a refusal, not a judgement call.** `graph.py validate`
|
|
71
|
+
exits 1 and names it, because inventing a check and running everything are the two
|
|
72
|
+
things this step forbids — and until B-080 closed, this paragraph asked for a field
|
|
73
|
+
the schema did not have.
|
|
64
74
|
3. **`done` takes one row per claim, and each needs a line in `evidence`** — the
|
|
65
75
|
command and what it printed. Paraphrase is not evidence. If you cannot produce
|
|
66
|
-
the output, the row belongs in `not_done`.
|
|
76
|
+
the output, the row belongs in `not_done`. The `check`'s own output is the first
|
|
77
|
+
row: the node said how it would be closed, so the proof it closed is that output.
|
|
67
78
|
4. **`not_done` is not a failure report.** It is what the next iteration picks up,
|
|
68
79
|
so write it as work rather than as blame.
|
|
69
80
|
5. **Every blocker says what it `blocks` and whether the run `can_continue_around`
|
|
@@ -72,6 +83,9 @@ ledger exists to catch.
|
|
|
72
83
|
6. **`replan.possible: false` needs a `why` a person can act on.** A stop with no
|
|
73
84
|
reason is indistinguishable from a stall, and the operator is the one who has to
|
|
74
85
|
tell them apart.
|
|
86
|
+
7. **Every `replan.add` entry names its own `check`.** A node you create is a node the
|
|
87
|
+
next verifier has to close, and handing it the absence you were handed is how the
|
|
88
|
+
defect returns one iteration later. `close` refuses the verdict and names the key.
|
|
75
89
|
|
|
76
90
|
## Three ways this goes wrong
|
|
77
91
|
|
|
@@ -80,6 +94,7 @@ ledger exists to catch.
|
|
|
80
94
|
| «The tests pass, so it is done» | The node serves a REQ, not a suite. A green suite that never exercised the requirement proves the suite ran |
|
|
81
95
|
| «Close it and note the gap» | A `done` with a caveat is a `not_done` somebody will read as finished. Split the row |
|
|
82
96
|
| «This blocker stops everything» | Say whether it does. `can_continue_around: true` is what keeps a run moving past one bad node, and guessing it wrong costs either the run or the correctness |
|
|
97
|
+
| «The node names no check, so I'll run the suite» | That is the choice this agent may not make. `validate` refuses the node before you are dispatched; a node whose completion test nobody wrote is a planning defect, and reporting it is worth more than a green suite |
|
|
83
98
|
|
|
84
99
|
## Where the doctrine is
|
|
85
100
|
|
|
@@ -14,6 +14,7 @@
|
|
|
14
14
|
"status": "done",
|
|
15
15
|
"blocked_by": [],
|
|
16
16
|
"serves": "REQ-001",
|
|
17
|
+
"check": "npm test — graph.example.json validates against graph.schema.json",
|
|
17
18
|
"evidence": [
|
|
18
19
|
"npm test → PASS: task-pipeline structure valid, with graph.example.json validated against graph.schema.json"
|
|
19
20
|
]
|
|
@@ -27,6 +28,7 @@
|
|
|
27
28
|
"N-001"
|
|
28
29
|
],
|
|
29
30
|
"serves": "REQ-002",
|
|
31
|
+
"check": "python3 test/graph_test.py — the validate fixtures",
|
|
30
32
|
"evidence": null
|
|
31
33
|
},
|
|
32
34
|
{
|
|
@@ -38,6 +40,7 @@
|
|
|
38
40
|
"N-001"
|
|
39
41
|
],
|
|
40
42
|
"serves": "REQ-005",
|
|
43
|
+
"check": "judgment: the reviewer rules the seven keys sufficient — recorded AS judgement, references/gates.md",
|
|
41
44
|
"evidence": null
|
|
42
45
|
},
|
|
43
46
|
{
|
|
@@ -110,6 +110,15 @@
|
|
|
110
110
|
"minLength": 1,
|
|
111
111
|
"description": "A REQ id, or a clause of the goal. Required, because a node that serves neither is work nobody asked for — and the pipeline's answer to that is to park it WITH THAT AS THE REASON rather than to do it quietly."
|
|
112
112
|
},
|
|
113
|
+
"check": {
|
|
114
|
+
"type": "string",
|
|
115
|
+
"minLength": 1,
|
|
116
|
+
"pattern": "\\S",
|
|
117
|
+
"not": {
|
|
118
|
+
"pattern": "[\\n\\r]"
|
|
119
|
+
},
|
|
120
|
+
"description": "**How this node will be closed** \u2014 the command a verifier runs, or the named judgement that stands in where no command can (`references/gates.md`'s third gate type: name the judge, record the verdict AS judgement). Its output becomes the node's `evidence` row.\n\nRequired on every node except a `parked` one, and that is the whole point of the field: `agents/verifier.md` tells the verifier to run *the check this node names*, so a node with none leaves it the two options that same paragraph forbids \u2014 invent a check, or run everything. Shipped doctrine read a field the schema did not have (B-080); the field the doctrine reads now exists.\n\nOne string, never a list, because the requirement it serves is singular throughout \u2014 one input, one job, one output, one owner, **one** completion test. A node needing two unrelated checks is a node doing two jobs, and the answer is to split it rather than to widen this field. One gate made of two commands is `a && b`, which is still one gate.\n\nA `parked` node is exempt because it is the one node nobody will close, and a placeholder there \u2014 *n/a, parked* \u2014 is the confidence-without-correctness this schema refuses everywhere else. The field survives a park: `park` sets the status and never removes what the node said it would run."
|
|
121
|
+
},
|
|
113
122
|
"evidence": {
|
|
114
123
|
"type": [
|
|
115
124
|
"array",
|
|
@@ -140,57 +149,13 @@
|
|
|
140
149
|
},
|
|
141
150
|
"allOf": [
|
|
142
151
|
{
|
|
143
|
-
"
|
|
144
|
-
"properties": {
|
|
145
|
-
"status": {
|
|
146
|
-
"const": "done"
|
|
147
|
-
}
|
|
148
|
-
},
|
|
149
|
-
"required": [
|
|
150
|
-
"status"
|
|
151
|
-
]
|
|
152
|
-
},
|
|
153
|
-
"then": {
|
|
154
|
-
"required": [
|
|
155
|
-
"evidence"
|
|
156
|
-
],
|
|
157
|
-
"properties": {
|
|
158
|
-
"evidence": {
|
|
159
|
-
"type": "array",
|
|
160
|
-
"minItems": 1,
|
|
161
|
-
"items": {
|
|
162
|
-
"type": "string",
|
|
163
|
-
"minLength": 1,
|
|
164
|
-
"pattern": "\\S",
|
|
165
|
-
"description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
|
|
166
|
-
}
|
|
167
|
-
}
|
|
168
|
-
}
|
|
169
|
-
}
|
|
152
|
+
"$ref": "#/definitions/rule_check_unless_parked"
|
|
170
153
|
},
|
|
171
154
|
{
|
|
172
|
-
"
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
}
|
|
177
|
-
},
|
|
178
|
-
"required": [
|
|
179
|
-
"status"
|
|
180
|
-
]
|
|
181
|
-
},
|
|
182
|
-
"then": {
|
|
183
|
-
"required": [
|
|
184
|
-
"parked_reason"
|
|
185
|
-
],
|
|
186
|
-
"properties": {
|
|
187
|
-
"parked_reason": {
|
|
188
|
-
"type": "string",
|
|
189
|
-
"minLength": 1,
|
|
190
|
-
"pattern": "\\S"
|
|
191
|
-
}
|
|
192
|
-
}
|
|
193
|
-
}
|
|
155
|
+
"$ref": "#/definitions/rule_done_needs_evidence"
|
|
156
|
+
},
|
|
157
|
+
{
|
|
158
|
+
"$ref": "#/definitions/rule_parked_needs_reason"
|
|
194
159
|
}
|
|
195
160
|
]
|
|
196
161
|
},
|
|
@@ -248,6 +213,88 @@
|
|
|
248
213
|
"description": "The reason, written for a person reading it later. The non-whitespace pattern is required for the same reason `parked_reason` needs one: `minLength: 1` counts a space."
|
|
249
214
|
}
|
|
250
215
|
}
|
|
216
|
+
},
|
|
217
|
+
"rule_check_unless_parked": {
|
|
218
|
+
"description": "B-080's rule, factored out of `node.allOf` so it is stated once and referenced. Factoring it exposed B-079: `test/validate.py`'s `_conditionals()` recursed into `allOf` and never followed a `$ref`, so a schema that shares a rule this way read as a schema with NO conditionals and both REQ-006 and REQ-012 were reported missing from a schema that states them. The shipped schema now uses the `$ref` form, so the deref branch is exercised by every run of the gate rather than by one fixture.",
|
|
219
|
+
"if": {
|
|
220
|
+
"properties": {
|
|
221
|
+
"status": {
|
|
222
|
+
"not": {
|
|
223
|
+
"const": "parked"
|
|
224
|
+
}
|
|
225
|
+
}
|
|
226
|
+
},
|
|
227
|
+
"required": [
|
|
228
|
+
"status"
|
|
229
|
+
]
|
|
230
|
+
},
|
|
231
|
+
"then": {
|
|
232
|
+
"required": [
|
|
233
|
+
"check"
|
|
234
|
+
],
|
|
235
|
+
"properties": {
|
|
236
|
+
"check": {
|
|
237
|
+
"type": "string",
|
|
238
|
+
"minLength": 1,
|
|
239
|
+
"pattern": "\\S"
|
|
240
|
+
}
|
|
241
|
+
}
|
|
242
|
+
}
|
|
243
|
+
},
|
|
244
|
+
"rule_done_needs_evidence": {
|
|
245
|
+
"description": "REQ-006 as draft-07 states it: `done` implies at least one non-whitespace evidence string. Referenced from `node.allOf`, never inlined, so the rule has one home.",
|
|
246
|
+
"if": {
|
|
247
|
+
"properties": {
|
|
248
|
+
"status": {
|
|
249
|
+
"const": "done"
|
|
250
|
+
}
|
|
251
|
+
},
|
|
252
|
+
"required": [
|
|
253
|
+
"status"
|
|
254
|
+
]
|
|
255
|
+
},
|
|
256
|
+
"then": {
|
|
257
|
+
"required": [
|
|
258
|
+
"evidence"
|
|
259
|
+
],
|
|
260
|
+
"properties": {
|
|
261
|
+
"evidence": {
|
|
262
|
+
"type": "array",
|
|
263
|
+
"minItems": 1,
|
|
264
|
+
"items": {
|
|
265
|
+
"type": "string",
|
|
266
|
+
"minLength": 1,
|
|
267
|
+
"pattern": "\\S",
|
|
268
|
+
"description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
|
|
269
|
+
}
|
|
270
|
+
}
|
|
271
|
+
}
|
|
272
|
+
}
|
|
273
|
+
},
|
|
274
|
+
"rule_parked_needs_reason": {
|
|
275
|
+
"description": "REQ-012: a parked node names why it is parked. Referenced from `node.allOf`.",
|
|
276
|
+
"if": {
|
|
277
|
+
"properties": {
|
|
278
|
+
"status": {
|
|
279
|
+
"const": "parked"
|
|
280
|
+
}
|
|
281
|
+
},
|
|
282
|
+
"required": [
|
|
283
|
+
"status"
|
|
284
|
+
]
|
|
285
|
+
},
|
|
286
|
+
"then": {
|
|
287
|
+
"required": [
|
|
288
|
+
"parked_reason"
|
|
289
|
+
],
|
|
290
|
+
"properties": {
|
|
291
|
+
"parked_reason": {
|
|
292
|
+
"type": "string",
|
|
293
|
+
"minLength": 1,
|
|
294
|
+
"pattern": "\\S"
|
|
295
|
+
}
|
|
296
|
+
}
|
|
297
|
+
}
|
|
251
298
|
}
|
|
252
299
|
}
|
|
253
300
|
}
|
|
@@ -319,6 +319,29 @@ spends its length refusing.
|
|
|
319
319
|
|
|
320
320
|
## GATE (manual)
|
|
321
321
|
|
|
322
|
+
> **Acceptance policy `AP-1` · owner: the operator of the project running this pipeline ·
|
|
323
|
+
> in force since 2026-08-20 · supersedes: the unversioned ladder that stood before it.**
|
|
324
|
+
>
|
|
325
|
+
> This block is the policy, and naming it that closes a contradiction that stood in the
|
|
326
|
+
> shipped doctrine. [`gates.md`](gates.md) → *Axis A — the stage gate type* says which stages are
|
|
327
|
+
> manual is the operator's decision and that **the framework fixes no stage count and no
|
|
328
|
+
> gate assignment** — while the ladder below fixes twelve criteria and the statuses they
|
|
329
|
+
> may take. Both were true and neither was scoped, so a reader could take either as the
|
|
330
|
+
> rule.
|
|
331
|
+
>
|
|
332
|
+
> **The scope, stated once:** `gates.md`'s sentence governs the **pipeline's shape** — how
|
|
333
|
+
> many stages a project runs and which of them are `auto`, `judgment` or `manual`, all of
|
|
334
|
+
> it in the project's own `pipeline.json`. `AP-1` governs **what stage 10's manual gate
|
|
335
|
+
> asks when a project runs one.** A project may drop stage 10, or make it `auto` for a
|
|
336
|
+
> class of work, and `AP-1` then does not apply to it; a project that keeps it manual gets
|
|
337
|
+
> these criteria and this vocabulary, not a per-run selection from them.
|
|
338
|
+
>
|
|
339
|
+
> **Changing it is a decision, not an edit.** An amended criterion moves the version to
|
|
340
|
+
> `AP-2` and lands with a row in the project's decision register, because a policy that can
|
|
341
|
+
> be edited silently is the *«recorded in the slot reserved for what a machine
|
|
342
|
+
> established»* failure one level up: an acceptance standard nobody can cite by version is
|
|
343
|
+
> one every run re-negotiates.
|
|
344
|
+
|
|
322
345
|
All of:
|
|
323
346
|
|
|
324
347
|
1. **Every shipped REQ has a verification row, and every row names a REQ its own run
|
|
@@ -47,12 +47,21 @@ The method most audits use is **horizontal**: compare the documents against each
|
|
|
47
47
|
other, then do it again. It works, and then it fails in a way that is invisible
|
|
48
48
|
from inside it. Measured over seven passes on a production repository:
|
|
49
49
|
|
|
50
|
-
| Pass | Findings | …of which the previous pass's own fixes caused |
|
|
51
|
-
|
|
52
|
-
| 4 | 12 | 5 |
|
|
53
|
-
| 5 | 17 | 9 |
|
|
54
|
-
| 6 | 13 | 10 |
|
|
55
|
-
| 7 | 19 | 4 |
|
|
50
|
+
| Pass | Findings | …of which the previous pass's own fixes caused | Self-inflicted share |
|
|
51
|
+
|---|---|---|---|
|
|
52
|
+
| 4 | 12 | 5 | 42% |
|
|
53
|
+
| 5 | 17 | 9 | 53% |
|
|
54
|
+
| 6 | 13 | 10 | 77% |
|
|
55
|
+
| 7 | 19 | 4 | 21% — see the note below |
|
|
56
|
+
|
|
57
|
+
**The trend above is measured over passes four to six, and pass seven is not part of
|
|
58
|
+
it.** 42% → 53% → 77% is the decay this section teaches; the seventh pass fell back to
|
|
59
|
+
4 of 19 and **the record does not say what that pass did differently.** The row stays,
|
|
60
|
+
with the gap named, for two reasons. Deleting a measured row to protect a claim is the
|
|
61
|
+
opposite of what this file asks of every gate it describes — and the vertical pass
|
|
62
|
+
described below cannot be credited for the drop, because it is a separate pass over the
|
|
63
|
+
same repository, not the seventh. Whatever pass seven did, it is unrecorded, and an
|
|
64
|
+
unrecorded cause is what this table has to say about it.
|
|
56
65
|
|
|
57
66
|
By pass six the audit was **mostly repairing itself**. Not fatigue — arithmetic.
|
|
58
67
|
Each pass edits the corpus the next pass reads, so the newest edits are always the
|
|
@@ -58,6 +58,15 @@ not the operator confirming it is what they asked for, and no amount of checking
|
|
|
58
58
|
makes it one. Which stages are manual is the **operator's** decision, recorded in
|
|
59
59
|
their `pipeline.json`; the framework fixes no stage count and no gate assignment.
|
|
60
60
|
|
|
61
|
+
**That sentence is about the pipeline's SHAPE, and nothing else.** It says a project
|
|
62
|
+
chooses how many stages it runs and which of them wait for a person. It does not say the
|
|
63
|
+
criteria inside a gate are per-run negotiable: where a project keeps stage 10 manual, what
|
|
64
|
+
that gate asks is [`acceptance.md`](acceptance.md)'s policy **`AP-1`**, which is versioned
|
|
65
|
+
and has an owner. The two rules stood side by side unscoped until 2026-08-20 (`B-091`), and
|
|
66
|
+
a reader could take either as the whole rule — *the framework fixes nothing* and *the ladder
|
|
67
|
+
is fixed* are both in the shipped doctrine, which is how an acceptance standard becomes
|
|
68
|
+
something every run re-argues.
|
|
69
|
+
|
|
61
70
|
## The judgment gate — a ruling is not a measurement
|
|
62
71
|
|
|
63
72
|
Two types were not enough, and the gap was not cosmetic. A reviewer's ruling, a check that
|
|
@@ -369,12 +369,18 @@ deduplicated, and always exiting 0 — a hook that can fail a `Read` breaks ever
|
|
|
369
369
|
every session. `scripts/graph.py doctrine` reports how many of the bundle's reference files
|
|
370
370
|
a run opened and lists the rest.
|
|
371
371
|
|
|
372
|
-
Why
|
|
373
|
-
written by the party the claim is about, is not evidence. And
|
|
372
|
+
Why a hook appends it is the same reason as above: a claim about what somebody read,
|
|
373
|
+
written by the party the claim is about, is not evidence. **And that is an intent rather
|
|
374
|
+
than a proof, so the verb reports `unattested`:** the ledger is the file the agent appends
|
|
375
|
+
to at every stage and carries no writer field, so nothing in it separates a hook-written
|
|
376
|
+
line from one an agent typed. Saying *never agent-written* was a provenance claim the
|
|
377
|
+
format cannot support — B-014's class, in the mechanism built to close it.
|
|
378
|
+
|
|
379
|
+
And why the verb prints
|
|
374
380
|
`unmeasured` rather than `0` when there are no such lines is the same reason again — the
|
|
375
381
|
hook being absent and the run reading nothing are **opposite facts** the ledger cannot
|
|
376
382
|
separate, so it claims neither. A `0` there would be the reassuring answer to a question
|
|
377
|
-
nobody asked, over
|
|
383
|
+
nobody asked, over 35 files nobody checked.
|
|
378
384
|
|
|
379
385
|
It is a disclosure: no floor, no direction, never a target. The moment the number becomes
|
|
380
386
|
something to raise, a run will open files to raise it.
|
|
@@ -27,15 +27,29 @@ Both are checked the same way and this file covers both.
|
|
|
27
27
|
|
|
28
28
|
## The measured reason this file exists
|
|
29
29
|
|
|
30
|
-
On 2026-08-11, mid-run, a monitor was armed to watch CI.
|
|
31
|
-
|
|
30
|
+
On 2026-08-11, mid-run, a monitor was armed to watch CI. **What follows is two
|
|
31
|
+
observations of two different things, taken at two different moments**, and each
|
|
32
|
+
one is stamped so that a reader is not left reconciling them. One minute after
|
|
33
|
+
the monitor was armed, the harness task inventory was queried; the process table
|
|
34
|
+
was read after that:
|
|
32
35
|
|
|
33
36
|
```
|
|
34
37
|
TaskList → "No tasks found"
|
|
35
38
|
ps -eo pid,etime,cmd → 52693 03:12 /bin/zsh -c … gh pr checks …
|
|
36
39
|
```
|
|
37
40
|
|
|
38
|
-
|
|
41
|
+
`etime` is elapsed life, so the `03:12` in that line dates its own reading:
|
|
42
|
+
three minutes and twelve seconds after the monitor was armed, a little over two
|
|
43
|
+
minutes after the inventory had come back empty. The monitor was **alive and
|
|
44
|
+
polling**, and the inventory tool reported nothing. Both readings were accurate.
|
|
45
|
+
They were readings of different things, made about two minutes apart.
|
|
46
|
+
|
|
47
|
+
**What this record does not carry, said plainly: the wall-clock times.** Neither
|
|
48
|
+
observation was stamped with the hour at which it was taken, so what was
|
|
49
|
+
observed is the two offsets above — one minute to the inventory, `03:12` of
|
|
50
|
+
process life at the `ps` reading — and the interval between them is derived from
|
|
51
|
+
those two. The absolute times are not recoverable from anything this record
|
|
52
|
+
kept, and nothing here reconstructs them.
|
|
39
53
|
|
|
40
54
|
This is the class this whole doctrine exists to catch: **a check answered by
|
|
41
55
|
something that is not its subject.** An inventory that does not enumerate the
|
|
@@ -201,9 +215,9 @@ judgement, the item is reported, not ended** — which puts it back under the ru
|
|
|
201
215
|
above rather than creating an exception to it.
|
|
202
216
|
|
|
203
217
|
**A foreign item never becomes spent.** The 18 containers above belong to other
|
|
204
|
-
projects; that
|
|
205
|
-
not permission. The third owner state widens what a run may
|
|
206
|
-
project** and widens nothing at all outside it.
|
|
218
|
+
projects; that the oldest of them has been up for three days is information for
|
|
219
|
+
whoever owns them, not permission. The third owner state widens what a run may
|
|
220
|
+
clean **inside its own project** and widens nothing at all outside it.
|
|
207
221
|
|
|
208
222
|
## Rationalizations
|
|
209
223
|
|
|
@@ -19,7 +19,7 @@ justifies reading it protects one section while the file below it doubles.
|
|
|
19
19
|
| Artifact | Parts | How it is read |
|
|
20
20
|
|---|---|---|
|
|
21
21
|
| `docs/evidence/retro.md` — **one per project** | **Standing instructions** (max **10**) · **Run stamps** (max **10**, oldest rotate out) | stage 0, **in full** — both are bounded by a **cap**, which *one line each* never was |
|
|
22
|
-
| the same file's **Recent log** | entries from the last five run stamps
|
|
22
|
+
| the same file's **Recent log** | narrative entries, **uncapped by design** — the heading said *entries from the last five run stamps* until 2026-08-20 while the section held 25 going back nine days, a bound in a heading that nothing enforced | stage 0, **queried** by the task's nouns. It said *in full* until 2026-08-10, when it measured **74%** of the file: an uncapped section inside a binding source is what makes the capped part get skimmed |
|
|
23
23
|
| `docs/evidence/retro/YYYY-QN.md` — the archive | every entry and every retirement ever written, append-only | **queried** by the task's nouns; never read end to end |
|
|
24
24
|
|
|
25
25
|
Seed the archive from [`../templates/retro-archive.md`](../templates/retro-archive.md).
|
|
@@ -264,8 +264,11 @@ never that the work was skipped quietly.
|
|
|
264
264
|
decomposition is a decision, never an omission.
|
|
265
265
|
- **Where the queue is a work graph, it is WRITTEN here**
|
|
266
266
|
([`work-graph.md`](work-graph.md)): `.task-pipeline/graph.json`, carrying the frozen REQ
|
|
267
|
-
ids so `serves` resolves, one node per unit of work with its owner
|
|
268
|
-
and an edge per dependency **naming what it hands
|
|
267
|
+
ids so `serves` resolves, one node per unit of work with its owner, what it touches and
|
|
268
|
+
**the check that will close it**, and an edge per dependency **naming what it hands
|
|
269
|
+
over**. The check is the planner's job and not the verifier's: a node whose completion
|
|
270
|
+
test is decided by whoever happens to close it is a node two closes will disagree about,
|
|
271
|
+
and `validate` refuses one that names none. Then
|
|
269
272
|
`python3 scripts/graph.py validate` — a graph that does not validate is not a queue, and
|
|
270
273
|
`next` refuses to walk one. The reason to prefer it over a prose plan is measured, not
|
|
271
274
|
aesthetic: a 400-node graph and a 4-node graph produce the same 27-byte frontier, so the
|
|
@@ -40,6 +40,7 @@ turn of every loop.
|
|
|
40
40
|
| `nodes[].serves` | the REQ or goal clause it exists for. A node serving neither is **parked with that as the reason** |
|
|
41
41
|
| `nodes[].blocked_by` | what must close first. This is what the frontier obeys |
|
|
42
42
|
| `nodes[].touches` | what it **mutates**. Two runnable nodes writing one file is the false parallelism [`planning.md`](planning.md) refuses — *distinct is not the same as independent, and the check is what they touch, never what they are called* |
|
|
43
|
+
| `nodes[].check` | **how this node will be closed** — one command, or the named judgement where no command can decide it. Required on every node except a `parked` one. `agents/verifier.md` runs it and reports its output as the evidence row; before this field existed that instruction pointed at an absence, leaving the verifier the two things it forbids — invent a check, or run everything (B-080) |
|
|
43
44
|
| `nodes[].evidence` | required when `status` is `done`. A node called done by assertion is what evidence exists to prevent |
|
|
44
45
|
| `nodes[].parked_reason` | required when `status` is `parked`. A park with no reason is indistinguishable, a week later, from work quietly dropped |
|
|
45
46
|
| `edges[].payload` | what the dependency hands over. **An edge carrying no named artifact is chronology drawn as architecture** |
|
|
@@ -68,6 +69,14 @@ the graph does. Ties break on declaration order, so the frontier is stable betwe
|
|
|
68
69
|
an unstable one costs more than it looks, because an agent calling `next` twice starts the
|
|
69
70
|
other node.
|
|
70
71
|
|
|
72
|
+
**One node, one completion test.** `check` is a string rather than a list, because the
|
|
73
|
+
requirement is singular throughout — one input, one job, one output, one owner, one
|
|
74
|
+
completion test. A node needing two unrelated checks is a node doing two jobs, and the
|
|
75
|
+
answer is to split it; one gate made of two commands is `a && b`, which is still one gate.
|
|
76
|
+
A **`parked`** node is the single exemption: it is the one node nobody will close, and
|
|
77
|
+
*n/a — parked* in that field is confidence without correctness. `park` never removes what
|
|
78
|
+
the node said it would run.
|
|
79
|
+
|
|
71
80
|
**`close` stamps the commit; the verifier never supplies it.** A verdict written after the
|
|
72
81
|
tree moved is evidence about a different tree, and an agent cannot name the wrong commit if
|
|
73
82
|
it is never the one naming one.
|
|
@@ -114,6 +123,7 @@ does not have.
|
|
|
114
123
|
|
|
115
124
|
| The excuse | Why it is wrong |
|
|
116
125
|
|---|---|
|
|
126
|
+
| *"The verifier will work out what to run."* | Then the node's completion test is decided by whoever happens to close it, and two closes of one node disagree. The field is where the planner says it, and `validate` refuses a node that does not |
|
|
117
127
|
| *"I'll just read the graph, it's only forty nodes."* | Forty is the number today. The property being protected is that four hundred costs the same, and it stops being true the first time the model reads the file |
|
|
118
128
|
| *"The frontier is short; one extra line won't matter."* | It is paid on every iteration of every loop. That is what *the frontier and nothing else* means |
|
|
119
129
|
| *"`blocked_by` already says the dependency, the edge is bookkeeping."* | The edge carries the **payload**. A dependency handing over nothing named is the fake edge with a field around it |
|
|
@@ -19,7 +19,10 @@ channel ships it. A dependency here would make the graph Claude-Code-shaped.
|
|
|
19
19
|
states everything JSON Schema can — the required fields, `owner` non-empty, an edge's
|
|
20
20
|
payload, and `done` implying non-empty evidence. What a schema cannot reach is
|
|
21
21
|
cross-document and cross-node: whether an `owner` names a role that EXISTS, whether
|
|
22
|
-
`serves` resolves, and whether the edges cycle. Those three are here.
|
|
22
|
+
`serves` resolves, and whether the edges cycle. Those three are here. So is every rule
|
|
23
|
+
the format states, because the schema is never applied to a LIVE graph — including
|
|
24
|
+
B-080's `check`, without which `agents/verifier.md` instructs an agent to read a field
|
|
25
|
+
that does not exist. The split is
|
|
23
26
|
where the format actually puts it, which is not where the first draft drew it.
|
|
24
27
|
|
|
25
28
|
Exit codes are the contract (standing instruction R-004 — the next command is
|
|
@@ -238,6 +241,25 @@ def violations(graph):
|
|
|
238
241
|
"nobody asked for is work the brief cannot account for, and the "
|
|
239
242
|
"coverage relation cannot reach it")
|
|
240
243
|
|
|
244
|
+
# B-080 — the node says HOW it will be closed, and shipped doctrine reads it.
|
|
245
|
+
# `agents/verifier.md` told the verifier to run *the check the task named* while
|
|
246
|
+
# the schema had no field to name one, so the instruction pointed at an absence
|
|
247
|
+
# and the verifier's only options were the two that paragraph forbids: invent a
|
|
248
|
+
# check, or run everything. A `parked` node is the one exemption — it is the one
|
|
249
|
+
# node nobody will close, and a placeholder there would be worse than the gap.
|
|
250
|
+
if n.get("status") != "parked":
|
|
251
|
+
chk = n.get("check")
|
|
252
|
+
if not isinstance(chk, str) or not chk.strip():
|
|
253
|
+
out.append(f"{nid}: `check` is {chk!r} — B-080: a node that cannot say how "
|
|
254
|
+
"it will be closed leaves the verifier inventing a check or "
|
|
255
|
+
"running everything, and `agents/verifier.md` forbids both. Name "
|
|
256
|
+
"the command, or the judge where no command can decide it. Only a "
|
|
257
|
+
"`parked` node is exempt")
|
|
258
|
+
elif any(c in chk for c in "\n\r"):
|
|
259
|
+
out.append(f"{nid}: `check` contains a line break. A completion test that is "
|
|
260
|
+
"two commands cannot say which one closed the node, and the "
|
|
261
|
+
"verifier reports its output as one evidence row")
|
|
262
|
+
|
|
241
263
|
if n.get("status") == "done":
|
|
242
264
|
ev = n.get("evidence")
|
|
243
265
|
if not isinstance(ev, list) or not [e for e in ev
|
|
@@ -480,6 +502,16 @@ def verdict_violations(v):
|
|
|
480
502
|
for nid in rp.get("add") or []:
|
|
481
503
|
if not isinstance(nid, dict) or "title" not in nid:
|
|
482
504
|
out.append("replan.add entries must be nodes with at least a `title`")
|
|
505
|
+
continue
|
|
506
|
+
# B-080 again, at the one place a node is created by a verdict rather than by
|
|
507
|
+
# a person. Checked HERE so the refusal names the key before anything is
|
|
508
|
+
# written: leaving it to `violations()` after the close would abort the whole
|
|
509
|
+
# close over a field the verdict could have been asked for.
|
|
510
|
+
_c = nid.get("check")
|
|
511
|
+
if not isinstance(_c, str) or not _c.strip():
|
|
512
|
+
out.append(f"replan.add entry {nid.get('title')!r} names no `check` — the "
|
|
513
|
+
"node it creates has to say how IT will be closed, or the next "
|
|
514
|
+
"verifier faces the absence this one was told to read")
|
|
483
515
|
return out
|
|
484
516
|
|
|
485
517
|
|
|
@@ -679,7 +711,14 @@ def cmd_add(graph, args):
|
|
|
679
711
|
die("owner %r is not a role this pipeline ships%s — nothing was written"
|
|
680
712
|
% (args.owner, (" — did you mean %s?" % near[0]) if near else ""))
|
|
681
713
|
|
|
682
|
-
|
|
714
|
+
check = (args.check or "").strip()
|
|
715
|
+
if not check:
|
|
716
|
+
die("`add` needs a --check with something in it: the command that closes this node, "
|
|
717
|
+
"or the judge where no command can decide it. A node that cannot say how it "
|
|
718
|
+
"will be closed leaves the verifier inventing a check or running everything — "
|
|
719
|
+
"B-080, and `agents/verifier.md` forbids both. Nothing was written")
|
|
720
|
+
|
|
721
|
+
for name, val in (("--title", title), ("--serves", serves), ("--check", check)):
|
|
683
722
|
if any(c in val for c in "\n\r"):
|
|
684
723
|
die("%s contains a line break. `next` prints one row per node and the loop "
|
|
685
724
|
"reads those rows, so a break here forges a row for a node that does not "
|
|
@@ -729,7 +768,7 @@ def cmd_add(graph, args):
|
|
|
729
768
|
"allocated from the highest in use" % nid)
|
|
730
769
|
|
|
731
770
|
new = {"id": nid, "title": title, "owner": args.owner, "status": "pending",
|
|
732
|
-
"blocked_by": blocked, "serves": serves, "evidence": None}
|
|
771
|
+
"blocked_by": blocked, "serves": serves, "check": check, "evidence": None}
|
|
733
772
|
touches = list(dict.fromkeys(t.strip() for t in (args.touches or []) if t.strip()))
|
|
734
773
|
if touches:
|
|
735
774
|
new["touches"] = touches
|
|
@@ -870,13 +909,21 @@ def cmd_producer(graph, args):
|
|
|
870
909
|
def cmd_doctrine(graph, args):
|
|
871
910
|
"""Which doctrine this run actually read — B-061.
|
|
872
911
|
|
|
873
|
-
The bundle is
|
|
912
|
+
The bundle is 35 reference files. A run reads some subset and nothing recorded which,
|
|
874
913
|
so **a skipped file and a read one were indistinguishable** — the class every guard in
|
|
875
914
|
this repository exists to catch, left standing over the doctrine itself.
|
|
876
915
|
|
|
877
|
-
`read:` lines
|
|
878
|
-
|
|
879
|
-
|
|
916
|
+
`read:` lines are written by a hook rather than by the agent, for the same reason `gate:`
|
|
917
|
+
is: a claim about what somebody read, written by the party the claim is about, is not
|
|
918
|
+
evidence.
|
|
919
|
+
|
|
920
|
+
**And that is an intent, not a proof, so every line here is reported as UNATTESTED.**
|
|
921
|
+
The ledger is `.task-pipeline/run.md` — the file the agent appends to at every stage —
|
|
922
|
+
so nothing in it distinguishes a hook-written line from one an agent typed. The doctrine
|
|
923
|
+
said *hook-written, never agent-written*, which is a provenance claim this script cannot
|
|
924
|
+
check and no format here carries; B-014's class, committed by the mechanism built to
|
|
925
|
+
close it. Until the ledger can attest a writer, the honest output is the count plus the
|
|
926
|
+
word: read as *this is what the ledger says, and the ledger cannot say who wrote it*.
|
|
880
927
|
|
|
881
928
|
**The one rule that matters here: no `read:` lines means UNMEASURED, never «read
|
|
882
929
|
nothing».** Zero would be the reassuring answer to a question nobody asked, and this
|
|
@@ -919,7 +966,10 @@ def cmd_doctrine(graph, args):
|
|
|
919
966
|
return 0
|
|
920
967
|
|
|
921
968
|
unread = [r for r in refs if r not in read]
|
|
922
|
-
print(f"doctrine: {len(read)} of {len(refs)} reference files read")
|
|
969
|
+
print(f"doctrine: {len(read)} of {len(refs)} reference files read — unattested")
|
|
970
|
+
print(" unattested: the ledger is the file the agent appends to at every stage, "
|
|
971
|
+
"so nothing in it proves the hook wrote these lines rather than an agent. The "
|
|
972
|
+
"count is what the ledger says; who wrote it is not recorded.")
|
|
923
973
|
print(" a disclosure: no floor, no direction, never a target. A run that needs "
|
|
924
974
|
"four files and reads four is not worse than one that reads thirty.")
|
|
925
975
|
for r in unread:
|
|
@@ -993,14 +1043,17 @@ def cmd_close(graph, args):
|
|
|
993
1043
|
title = str(spec.get("title", "")).strip()
|
|
994
1044
|
owner = str(spec.get("owner", "implementer")).strip()
|
|
995
1045
|
serves = str(spec.get("serves", node.get("serves"))).strip()
|
|
1046
|
+
check = str(spec.get("check", "")).strip()
|
|
996
1047
|
if not title:
|
|
997
1048
|
die("replan.add carries an entry with no title — nothing was written")
|
|
1049
|
+
if not check:
|
|
1050
|
+
die("replan.add carries an entry with no check — nothing was written")
|
|
998
1051
|
if owner not in ROLES:
|
|
999
1052
|
die("replan.add names owner %r, which is not a role this pipeline ships" % owner)
|
|
1000
1053
|
new_id = next_id({n.get("id") for n in graph["nodes"]})
|
|
1001
1054
|
graph["nodes"].append({"id": new_id, "title": title, "owner": owner,
|
|
1002
1055
|
"status": "pending", "blocked_by": [], "serves": serves,
|
|
1003
|
-
"evidence": None})
|
|
1056
|
+
"check": check, "evidence": None})
|
|
1004
1057
|
revise(graph, "add", new_id, why or ("re-planned by the verdict on " + nid))
|
|
1005
1058
|
added.append(new_id)
|
|
1006
1059
|
for pid in rp.get("park") or []:
|
|
@@ -1074,6 +1127,10 @@ def main(argv=None):
|
|
|
1074
1127
|
p_add.add_argument("--title", required=True)
|
|
1075
1128
|
p_add.add_argument("--owner", required=True)
|
|
1076
1129
|
p_add.add_argument("--serves", required=True)
|
|
1130
|
+
p_add.add_argument("--check", required=True,
|
|
1131
|
+
help="the command that closes this node, or the judge where no "
|
|
1132
|
+
"command can decide it — the verifier runs it and reports its "
|
|
1133
|
+
"output as the evidence row")
|
|
1077
1134
|
p_add.add_argument("--blocked-by", dest="blocked_by", action="append", default=[])
|
|
1078
1135
|
p_add.add_argument("--carries", action="append", default=[],
|
|
1079
1136
|
help="what each --blocked-by hands over; pairs in the order written")
|
|
@@ -31,7 +31,7 @@ has not fired in the last five run stamps, or in the last sixty days. At eleven
|
|
|
31
31
|
rows, the oldest never-fired
|
|
32
32
|
row goes — the cap is not negotiable, ranking is.
|
|
33
33
|
|
|
34
|
-
## Recent log — entries
|
|
34
|
+
## Recent log — narrative entries, uncapped and queried rather than read (newest first)
|
|
35
35
|
|
|
36
36
|
Older entries and every retirement **move** to `docs/evidence/retro/YYYY-QN.md`
|
|
37
37
|
at the prune. Moving is not deleting: the archive is append-only and holds the
|
|
@@ -20,23 +20,35 @@ Run: `<topic>` · started `<YYYY-MM-DD>` · module map: `<path or "none">`
|
|
|
20
20
|
|
|
21
21
|
## `read:` — which doctrine this run actually opened
|
|
22
22
|
|
|
23
|
-
The bundle is
|
|
23
|
+
The bundle is 35 reference files and nothing recorded which of them a run read, so **a
|
|
24
24
|
skipped file and a read one were indistinguishable** — the class every guard in this
|
|
25
25
|
pipeline exists to catch, left standing over the doctrine itself.
|
|
26
26
|
|
|
27
|
-
`read:` is **
|
|
28
|
-
about what somebody read, written by the party the claim is about, is not evidence. The
|
|
27
|
+
`read:` is **written by a hook rather than by the agent**, for the same reason `gate:` is: a
|
|
28
|
+
claim about what somebody read, written by the party the claim is about, is not evidence. The
|
|
29
29
|
hook is in `hooks.example.json`, matches `Read`, records the path only
|
|
30
30
|
when it is inside the bundle's `references/`, deduplicates, and **always exits 0** — a hook
|
|
31
31
|
that can fail a `Read` would break every turn in every session.
|
|
32
32
|
|
|
33
|
+
**That is an intent, not a proof, and the line says so.** This file is the one the agent
|
|
34
|
+
appends to at every stage, so nothing in it distinguishes a hook-written line from one an
|
|
35
|
+
agent typed: there is no writer field, and `scripts/graph.py` has no provenance check
|
|
36
|
+
because the format gives it nothing to check. The doctrine here said *hook-written, never
|
|
37
|
+
agent-written* until 2026-08-20 — a provenance claim in the file whose whole subject is that
|
|
38
|
+
a claim by the interested party is not evidence, which is B-014's class committed by the
|
|
39
|
+
mechanism built to close it. So `doctrine` and the `gate:` reader both report **`unattested`**
|
|
40
|
+
beside their counts, and the two claims that survive are the ones the ledger can support:
|
|
41
|
+
*this line is in the ledger*, and *nobody recorded who wrote it*. Attesting a writer needs a
|
|
42
|
+
field this format does not have; inventing one that an agent can also fill would restate the
|
|
43
|
+
same claim one level down.
|
|
44
|
+
|
|
33
45
|
`scripts/graph.py doctrine` reads these lines and prints one of three things:
|
|
34
46
|
|
|
35
47
|
| It prints | When | Why not just a number |
|
|
36
48
|
|---|---|---|
|
|
37
49
|
| `unmeasured — no run ledger` | there is no ledger | nothing to read from |
|
|
38
50
|
| `unmeasured — the ledger carries no read: lines` | the hook is absent, **or** the run opened no doctrine | two opposite facts, and the ledger cannot separate them, so neither is claimed |
|
|
39
|
-
| `N of
|
|
51
|
+
| `N of 35 reference files read — unattested`, then each unread one | the hook is installed and fired | the count alone says there is a gap, not where — and `unattested` says the ledger cannot name who wrote the lines |
|
|
40
52
|
|
|
41
53
|
**It is a disclosure: no floor, no direction, never a target.** A run that needs four files
|
|
42
54
|
and reads four is not worse than one that reads thirty — and the moment the number becomes
|
|
@@ -58,15 +70,17 @@ hand: <N|10> — task "<quoted>" — done <n> — surfaced <n> — decisions <n
|
|
|
58
70
|
holds: <stage id> — <n> (<class: what, owner>; … or "none") — enumerated <n>/8 classes, <unlooked: classes not enumerable>
|
|
59
71
|
gate: <stage id> — command "<cmd>" — exit <N> — <ISO-8601>
|
|
60
72
|
event: <compact|session-end|subagent> — <detail> — <ISO-8601>
|
|
61
|
-
read: references/<file>.md # hook-
|
|
73
|
+
read: references/<file>.md # hook-appended, deduped, UNATTESTED (no writer field)
|
|
62
74
|
```
|
|
63
75
|
|
|
64
76
|
- **`stage:`** — written when a gate **returns**, not when the stage is entered. The
|
|
65
77
|
rail's `✓` is derived from this line and from nothing else; a glyph set from memory
|
|
66
78
|
is a summary that is confidently wrong exactly when it matters.
|
|
67
|
-
- **`gate:`** —
|
|
68
|
-
|
|
69
|
-
|
|
79
|
+
- **`gate:`** — appended by `hooks/gate-observer.sh` rather than by an agent, and
|
|
80
|
+
**unattested** for the reason `read:` is: this file is agent-written at every stage
|
|
81
|
+
and carries no writer field, so the line's provenance is an intent the format cannot
|
|
82
|
+
prove. It is the only line here that records what a command **did** rather than what
|
|
83
|
+
somebody concluded: the exit code of the stage's declared `gate.command`, observed. The
|
|
70
84
|
`stage:` line above it is the agent's claim, and the release gate requires the
|
|
71
85
|
two to agree — without this, a gate reads a claim written by the party it
|
|
72
86
|
constrains and confirms an assertion with itself. Absent where the project
|
|
@@ -41,6 +41,7 @@ worse than saying nothing.
|
|
|
41
41
|
## Contents
|
|
42
42
|
|
|
43
43
|
- [Staleness — a row is true about the tree it OBSERVED](#staleness--a-row-is-true-about-the-tree-it-observed)
|
|
44
|
+
- [Environment — a proof is only valid where it ran](#environment--a-proof-is-only-valid-where-it-ran)
|
|
44
45
|
- the ledger itself — one row per REQ, appended by stage 8
|
|
45
46
|
- [What `Human` means, and what it does not](#what-human-means-and-what-it-does-not)
|
|
46
47
|
|
|
@@ -71,11 +72,39 @@ one. Four things overtake a row, and naming which one applies is the note's job:
|
|
|
71
72
|
change in what it covers, a **dependency** change, an **environment** change, and a
|
|
72
73
|
**policy** change — the last being the rule under which the evidence was accepted.
|
|
73
74
|
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
75
|
+
## Environment — a proof is only valid where it ran
|
|
76
|
+
|
|
77
|
+
`Observed at` says which tree the check saw. It does not say **where**, and without that a
|
|
78
|
+
smoke test against a preview URL enters the record in a shape indistinguishable from one
|
|
79
|
+
against production, and a suite green on a laptop with accumulated state is
|
|
80
|
+
indistinguishable from one green on a runner that started clean. Those are the two halves
|
|
81
|
+
of the same question, and only one was instrumented.
|
|
82
|
+
|
|
83
|
+
**`Environment` is a required cell on every row.** Its vocabulary is the project's own,
|
|
84
|
+
declared in the runbook and not invented per row — this file ships with four, and a project
|
|
85
|
+
that deploys differently declares different ones:
|
|
86
|
+
|
|
87
|
+
| Value | What it means |
|
|
88
|
+
|---|---|
|
|
89
|
+
| `production` | the deployed target real users reach |
|
|
90
|
+
| `preview` | a per-branch or per-PR deployment; proves the build, never the release |
|
|
91
|
+
| `ci` | a clean runner, no accumulated state |
|
|
92
|
+
| `local` | a developer machine, with whatever state it has |
|
|
93
|
+
| `—` | **recorded absence** — a row written before this column existed, or one whose environment nobody recorded. Not a value, and never a default to reach for |
|
|
94
|
+
|
|
95
|
+
A missing cell is not the same as `—`: the first is a row that forgot the question, and it is
|
|
96
|
+
**refused**. The second is an answer.
|
|
97
|
+
|
|
98
|
+
**A REQ claiming production behaviour may not be closed on a non-production
|
|
99
|
+
observation.** Stage 8 writes the row and refuses that pairing rather than recording it —
|
|
100
|
+
`pass` in `ci` against a requirement about the deployed product is the exact substitution
|
|
101
|
+
this column exists to make visible.
|
|
102
|
+
|
|
103
|
+
| REQ | What | Run | Shipped in | Observed at | Environment | Auto | Human | Note |
|
|
104
|
+
|---|---|---|---|---|---|---|---|---|
|
|
105
|
+
| REQ-001 | CSV export from a report | `2026-07-28-export` | v1.4.0 | `5f21ac3` | production | pass | 2026-07-30 | opened the deployed page, exported, opened the file |
|
|
106
|
+
| REQ-004 | XLSX export | `2026-07-28-export` | v1.4.0 | `5f21ac3` | ci | pass | **never** | — |
|
|
107
|
+
| REQ-007 | Export respects active filters | `2026-07-28-export` | v1.4.0 | `5f21ac3` | preview | partial | **never** | CSV path only |
|
|
79
108
|
|
|
80
109
|
## Columns
|
|
81
110
|
|
|
@@ -86,6 +115,10 @@ change in what it covers, a **dependency** change, an **environment** change, an
|
|
|
86
115
|
- **Run** — the brief's topic slug, so the context is one file away.
|
|
87
116
|
- **Shipped in** — the tag or commit that carried it. Where a project does not tag,
|
|
88
117
|
the commit, and the same value every row of that run carries.
|
|
118
|
+
- **Observed at** — the commit the check ran against. `—` where nobody recorded one.
|
|
119
|
+
- **Environment** — where it ran, from the project's declared vocabulary. Required; `—` is
|
|
120
|
+
the recorded absence and an omitted cell is refused. See *Environment — a proof is only
|
|
121
|
+
valid where it ran*.
|
|
89
122
|
- **Auto** — what the run's own gate said: `pass` · `partial` · `none`. Copied from the
|
|
90
123
|
coverage table rather than re-derived; where the two disagree the coverage table wins
|
|
91
124
|
and the disagreement is a finding. A coverage verdict of **`review`** — *no check can
|