task-pipeline-skill 1.72.0 → 1.73.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +63 -0
- package/CONTRIBUTING.md +61 -0
- package/SKILL-CARD.md +1 -1
- package/package.json +1 -1
- package/plugins/task-pipeline/.claude-plugin/plugin.json +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/graph.schema.json +85 -73
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +23 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +15 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/gates.md +9 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +9 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/retrospective.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +16 -5
- package/plugins/task-pipeline/skills/task-pipeline/templates/retro.md +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +22 -8
- package/plugins/task-pipeline/skills/task-pipeline/templates/verification.md +38 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,68 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v1.73.0 — 2026-08-20 — the registry could not see the templates, the scripts, or its own registers
|
|
4
|
+
|
|
5
|
+
|
|
6
|
+
**The claim registry had the class for this exact incident and it fired on nothing.** Three
|
|
7
|
+
shipped surfaces said *34 reference files* over a directory of 35 — `scripts/graph.py`'s
|
|
8
|
+
`doctrine` docstring and `templates/run.md` twice — and `npm test` printed
|
|
9
|
+
`reference files: dormant (truth 35)`. Two holes, either of which was enough: the pattern
|
|
10
|
+
knew only the word order ``N files under `references/` ``, and the corpus was eight named
|
|
11
|
+
files plus `references/**`, so `templates/` and `scripts/` were never opened. The class now
|
|
12
|
+
reads any phrasing of the count, the corpus reads the templates and the shipped script, and
|
|
13
|
+
the class is armed at three agreeing sites instead of dormant.
|
|
14
|
+
|
|
15
|
+
**And it never read this repository's own registers.** `docs/DOCMAP.md` names decisions,
|
|
16
|
+
open questions, the board, the ledger and the retro as the registers, and its propagation
|
|
17
|
+
matrix sends *a number stated in a living document* to this registry — which could not see
|
|
18
|
+
any of them. `docs/OPEN_QUESTIONS.md` said *the 250 guards* against a workflow defining 390,
|
|
19
|
+
in a phrasing the guard class already knew. Bringing them in refused 26 statements and every
|
|
20
|
+
one was narration, so a number inside a **dated item** is a record and exempt; on the board
|
|
21
|
+
the discriminator is the State cell rather than the date, because every row names the day it
|
|
22
|
+
was filed and B-001 — open — was stating its description budget and its reference-file count
|
|
23
|
+
as facts about now.
|
|
24
|
+
|
|
25
|
+
**Fourteen consecutive releases carry no run stamp**, `v1.60.1` through `v1.72.0`, and the
|
|
26
|
+
retro's honest-gap section named only `v1.16.0`–`v1.23.0`. Its *Measured, not recalled*
|
|
27
|
+
receipt could not produce the measurement: it grepped `docs/superpowers/retro.md`, removed at
|
|
28
|
+
v1.53.0, and grepped for a tag's own commit when a stamp names the commit the *run* ended on.
|
|
29
|
+
Rewritten as a tag-range walk that reads the archive too, and a guard now requires every
|
|
30
|
+
release after the newest stamp to be named in that section.
|
|
31
|
+
|
|
32
|
+
Also: `## Unreleased` is where the guard count lives between a tag and the next bump (B-104);
|
|
33
|
+
`Environment` is a required cell on every verification row with the vocabulary read out of
|
|
34
|
+
the shipped template (B-099's other half); the graph schema states its three node rules
|
|
35
|
+
behind `$ref`s and the checker follows them (B-079, proved by moving them, not by a fixture);
|
|
36
|
+
an open board row's `file:N-M` must quote the phrase it points at, which caught five stale
|
|
37
|
+
citations and then caught this change's own edits four more times; the acceptance ladder is
|
|
38
|
+
policy **`AP-1`** with an owner and an in-force date, and `gates.md`'s *the framework fixes
|
|
39
|
+
no stage count* is scoped to the pipeline's shape; `read:` and `gate:` are reported
|
|
40
|
+
**unattested** instead of claimed *never agent-written*, because the ledger is the file the
|
|
41
|
+
agent appends to at every stage; `validate.yml` no longer re-validates a SHA on its own tag
|
|
42
|
+
push; and every documented `npm` equation is compared against `package.json` — `CLAUDE.md`
|
|
43
|
+
had `npm test` as `validate.py` alone, dropping 129 graph cases.
|
|
44
|
+
|
|
45
|
+
**What the suite found that no reading did.** This change broke **thirteen** existing
|
|
46
|
+
plants and `npm run test:all` named every one: the two verification-header probes spell
|
|
47
|
+
the whole header and it gained a column; two CHANGELOG probes scoped themselves to
|
|
48
|
+
`^## v` and the count now lives in `## Unreleased`; **seven** graph-schema probes walk
|
|
49
|
+
`node.allOf` inline and the rules moved behind `$ref`. All thirteen were repaired by
|
|
50
|
+
deriving the guard's own scope rather than restating it, which is what `learned.md`'s
|
|
51
|
+
*sweep the class* asks for — and the sweep needed two rounds: six schema probes were
|
|
52
|
+
repaired together and the seventh surfaced on the next full run, having died on
|
|
53
|
+
`KeyError: 'if'` where the others had died on their own asserts. A class fixed in six of
|
|
54
|
+
seven places is the shape standing instruction R-003 exists for.
|
|
55
|
+
|
|
56
|
+
One of the twelve was sharper than the rest, and it was self-inflicted twice over: the
|
|
57
|
+
honesty note explaining that `OQ-0002` no longer restates a total put an ISO date in the
|
|
58
|
+
row, which made the row a dated record and **disarmed the register plant** — green over
|
|
59
|
+
exactly the stale total it had been written to catch. `OQ-####` now uses its `Status`
|
|
60
|
+
cell, the discriminator the board already uses, and the register plant sits on an open
|
|
61
|
+
board row where no prose edit beside it can turn it off.
|
|
62
|
+
|
|
63
|
+
Guards: 390 → **412** · property checks 9 → **14**
|
|
64
|
+
|
|
65
|
+
|
|
3
66
|
## v1.72.0 — a node says how it will be closed
|
|
4
67
|
|
|
5
68
|
**B-080 closed, and with it the last of four requirements this pack's own manifesto named
|
package/CONTRIBUTING.md
CHANGED
|
@@ -531,6 +531,67 @@ the two homes are compared directly: whatever the verifier is told to read off t
|
|
|
531
531
|
must be a property the schema declares.
|
|
532
532
|
*(guard: `node declares no` and `rule that can fire` and `off the node, and`)*
|
|
533
533
|
|
|
534
|
+
**57. The claim registry reads every surface that states a number, including the ones this
|
|
535
|
+
repository writes about itself.** Its corpus was eight named files plus `references/**`, so
|
|
536
|
+
`templates/` and `scripts/` were invisible — three shipped surfaces said "34 reference files"
|
|
537
|
+
over a directory of 35 while the class printed `dormant` — and so were the registers
|
|
538
|
+
`docs/DOCMAP.md` names, where `docs/OPEN_QUESTIONS.md` said "the 250 guards" against a
|
|
539
|
+
workflow defining 390. The corpus now covers both, and a class recognises every phrasing of
|
|
540
|
+
its count rather than one word order. A number inside a **dated item** in a register is a
|
|
541
|
+
record and exempt; on the board the discriminator is the **State** cell, because every row
|
|
542
|
+
names the day it was filed and an open row is a claim about now.
|
|
543
|
+
*(guard: `reference files` and `— derive the number or delete it`)*
|
|
544
|
+
|
|
545
|
+
**58. Every documented `npm` command means what `package.json` runs.** `CLAUDE.md` glossed
|
|
546
|
+
`npm test` as `python3 test/validate.py`, dropping `graph_test.py` and its 129 cases — the
|
|
547
|
+
suite-outside-the-run class stated the other way round — and called `npm run test:all` "both"
|
|
548
|
+
where it runs eight scripts. An equation is compared against the script body after one level
|
|
549
|
+
of `npm run` resolution, and a bare `npm run X` in a document about this repository must be a
|
|
550
|
+
script that exists. Portable doctrine under `plugins/` and `cursor/` is out of scope: it names
|
|
551
|
+
a host project's commands.
|
|
552
|
+
*(guard: `is glossed as` and `declares no such script`)*
|
|
553
|
+
|
|
554
|
+
**59. A release either carries a run stamp or is recorded as a gap, and the guard-count claim
|
|
555
|
+
has a home before the bump.** Fourteen consecutive releases had no stamp while the retro named
|
|
556
|
+
only `v1.16.0`–`v1.23.0`, and the receipt that was supposed to prove it grepped a path removed
|
|
557
|
+
at v1.53.0 for a tag's own commit — a stamp names the commit the *run* ended on. Scoped to the
|
|
558
|
+
trailing stretch: 84 of 117 tags predate the register and backfilling is forbidden. Separately,
|
|
559
|
+
the count guard reads the topmost `## ` section, so `## Unreleased` is where the number lives
|
|
560
|
+
between a tag and the next bump, and it must sit above every version heading.
|
|
561
|
+
*(guard: `release(s) after the newest run stamp` and `section sits below a released version`)*
|
|
562
|
+
|
|
563
|
+
**60. A `file:line` range in an open board row quotes the phrase it points at.** Five
|
|
564
|
+
citations resolved to real lines and pointed at other text; a line number is the most fragile
|
|
565
|
+
address a document carries, because every edit above it moves it and nothing notices. Closed
|
|
566
|
+
rows are records and are left alone; single-line citations cannot be quoted and are disclosed
|
|
567
|
+
as unanchored rather than failed.
|
|
568
|
+
*(guard: `quotes no phrase from it`)*
|
|
569
|
+
|
|
570
|
+
**61. Every evidence row records the environment it ran in, and a claim of provenance the
|
|
571
|
+
format cannot check is marked unattested instead.** `Observed at` said which tree a check saw
|
|
572
|
+
and nothing said where it ran, so a preview smoke test and a production one entered the record
|
|
573
|
+
in the same shape — in a pack whose own `learned.md` records a suite green on every author's
|
|
574
|
+
machine and 1039 failures on a clean runner. And `read:`/`gate:` were declared *hook-written,
|
|
575
|
+
never agent-written* while both land in the file the agent appends to at every stage: no writer
|
|
576
|
+
field, no provenance check, so the count is reported `unattested` and the claim is not made.
|
|
577
|
+
*(guard: `has no `Environment` cell` and `prints a count and never says`)*
|
|
578
|
+
|
|
579
|
+
**62. The acceptance ladder is a versioned policy with an owner, and `gates.md`'s
|
|
580
|
+
fixes-nothing sentence is scoped to the pipeline's shape.** Both rules stood unscoped side by
|
|
581
|
+
side for seventy releases — *the framework fixes no stage count and no gate assignment* beside
|
|
582
|
+
twelve fixed criteria — so a reader could take either as the whole rule, and a table accepted
|
|
583
|
+
under v1.20 doctrine was indistinguishable from one accepted under v1.70. The block carries
|
|
584
|
+
`AP-1`, an owner and an in-force date; an amendment moves the version and lands with a decision
|
|
585
|
+
row.
|
|
586
|
+
*(guard: `the acceptance policy carries no` and `stands unscoped beside`)*
|
|
587
|
+
|
|
588
|
+
**63. A heading may not declare a bound nothing enforces.** The retro's *Recent log* read
|
|
589
|
+
*entries from the last five run stamps* over 25 entries reaching back nine days, borrowing the
|
|
590
|
+
stamp section's wording without its cap — filed twice as B-060 and B-069 and disclosed by the
|
|
591
|
+
file about itself. Checked in the live retro and in `templates/retro.md`, which seeded the
|
|
592
|
+
false bound into every host project.
|
|
593
|
+
*(guard: `declares the bound` and `and nothing enforces it`)*
|
|
594
|
+
|
|
534
595
|
## Adding or changing doctrine
|
|
535
596
|
|
|
536
597
|
- **Change one idea per PR.** These files are read by agents under load; a PR that
|
package/SKILL-CARD.md
CHANGED
|
@@ -12,7 +12,7 @@ harmless.
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
|
|
14
14
|
| **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
|
|
15
|
-
| **Version** | 1.
|
|
15
|
+
| **Version** | 1.73.0 |
|
|
16
16
|
| **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
|
|
17
17
|
| **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki, and **one of two browser channels** — `playwright` (CLI or MCP) or `chrome-devtools` (MCP); either satisfies the browser step and neither is required. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
|
|
18
18
|
| **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline-skill",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.73.0",
|
|
4
4
|
"description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"task-pipeline": "bin/task-pipeline.js"
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"name": "task-pipeline",
|
|
3
3
|
"displayName": "Task Pipeline",
|
|
4
4
|
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
|
|
5
|
-
"version": "1.
|
|
5
|
+
"version": "1.73.0",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "ssheleg",
|
|
8
8
|
"url": "https://x.com/sshlg93"
|
|
@@ -149,83 +149,13 @@
|
|
|
149
149
|
},
|
|
150
150
|
"allOf": [
|
|
151
151
|
{
|
|
152
|
-
"
|
|
153
|
-
"properties": {
|
|
154
|
-
"status": {
|
|
155
|
-
"not": {
|
|
156
|
-
"const": "parked"
|
|
157
|
-
}
|
|
158
|
-
}
|
|
159
|
-
},
|
|
160
|
-
"required": [
|
|
161
|
-
"status"
|
|
162
|
-
]
|
|
163
|
-
},
|
|
164
|
-
"then": {
|
|
165
|
-
"required": [
|
|
166
|
-
"check"
|
|
167
|
-
],
|
|
168
|
-
"properties": {
|
|
169
|
-
"check": {
|
|
170
|
-
"type": "string",
|
|
171
|
-
"minLength": 1,
|
|
172
|
-
"pattern": "\\S"
|
|
173
|
-
}
|
|
174
|
-
}
|
|
175
|
-
}
|
|
152
|
+
"$ref": "#/definitions/rule_check_unless_parked"
|
|
176
153
|
},
|
|
177
154
|
{
|
|
178
|
-
"
|
|
179
|
-
"properties": {
|
|
180
|
-
"status": {
|
|
181
|
-
"const": "done"
|
|
182
|
-
}
|
|
183
|
-
},
|
|
184
|
-
"required": [
|
|
185
|
-
"status"
|
|
186
|
-
]
|
|
187
|
-
},
|
|
188
|
-
"then": {
|
|
189
|
-
"required": [
|
|
190
|
-
"evidence"
|
|
191
|
-
],
|
|
192
|
-
"properties": {
|
|
193
|
-
"evidence": {
|
|
194
|
-
"type": "array",
|
|
195
|
-
"minItems": 1,
|
|
196
|
-
"items": {
|
|
197
|
-
"type": "string",
|
|
198
|
-
"minLength": 1,
|
|
199
|
-
"pattern": "\\S",
|
|
200
|
-
"description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
|
|
201
|
-
}
|
|
202
|
-
}
|
|
203
|
-
}
|
|
204
|
-
}
|
|
155
|
+
"$ref": "#/definitions/rule_done_needs_evidence"
|
|
205
156
|
},
|
|
206
157
|
{
|
|
207
|
-
"
|
|
208
|
-
"properties": {
|
|
209
|
-
"status": {
|
|
210
|
-
"const": "parked"
|
|
211
|
-
}
|
|
212
|
-
},
|
|
213
|
-
"required": [
|
|
214
|
-
"status"
|
|
215
|
-
]
|
|
216
|
-
},
|
|
217
|
-
"then": {
|
|
218
|
-
"required": [
|
|
219
|
-
"parked_reason"
|
|
220
|
-
],
|
|
221
|
-
"properties": {
|
|
222
|
-
"parked_reason": {
|
|
223
|
-
"type": "string",
|
|
224
|
-
"minLength": 1,
|
|
225
|
-
"pattern": "\\S"
|
|
226
|
-
}
|
|
227
|
-
}
|
|
228
|
-
}
|
|
158
|
+
"$ref": "#/definitions/rule_parked_needs_reason"
|
|
229
159
|
}
|
|
230
160
|
]
|
|
231
161
|
},
|
|
@@ -283,6 +213,88 @@
|
|
|
283
213
|
"description": "The reason, written for a person reading it later. The non-whitespace pattern is required for the same reason `parked_reason` needs one: `minLength: 1` counts a space."
|
|
284
214
|
}
|
|
285
215
|
}
|
|
216
|
+
},
|
|
217
|
+
"rule_check_unless_parked": {
|
|
218
|
+
"description": "B-080's rule, factored out of `node.allOf` so it is stated once and referenced. Factoring it exposed B-079: `test/validate.py`'s `_conditionals()` recursed into `allOf` and never followed a `$ref`, so a schema that shares a rule this way read as a schema with NO conditionals and both REQ-006 and REQ-012 were reported missing from a schema that states them. The shipped schema now uses the `$ref` form, so the deref branch is exercised by every run of the gate rather than by one fixture.",
|
|
219
|
+
"if": {
|
|
220
|
+
"properties": {
|
|
221
|
+
"status": {
|
|
222
|
+
"not": {
|
|
223
|
+
"const": "parked"
|
|
224
|
+
}
|
|
225
|
+
}
|
|
226
|
+
},
|
|
227
|
+
"required": [
|
|
228
|
+
"status"
|
|
229
|
+
]
|
|
230
|
+
},
|
|
231
|
+
"then": {
|
|
232
|
+
"required": [
|
|
233
|
+
"check"
|
|
234
|
+
],
|
|
235
|
+
"properties": {
|
|
236
|
+
"check": {
|
|
237
|
+
"type": "string",
|
|
238
|
+
"minLength": 1,
|
|
239
|
+
"pattern": "\\S"
|
|
240
|
+
}
|
|
241
|
+
}
|
|
242
|
+
}
|
|
243
|
+
},
|
|
244
|
+
"rule_done_needs_evidence": {
|
|
245
|
+
"description": "REQ-006 as draft-07 states it: `done` implies at least one non-whitespace evidence string. Referenced from `node.allOf`, never inlined, so the rule has one home.",
|
|
246
|
+
"if": {
|
|
247
|
+
"properties": {
|
|
248
|
+
"status": {
|
|
249
|
+
"const": "done"
|
|
250
|
+
}
|
|
251
|
+
},
|
|
252
|
+
"required": [
|
|
253
|
+
"status"
|
|
254
|
+
]
|
|
255
|
+
},
|
|
256
|
+
"then": {
|
|
257
|
+
"required": [
|
|
258
|
+
"evidence"
|
|
259
|
+
],
|
|
260
|
+
"properties": {
|
|
261
|
+
"evidence": {
|
|
262
|
+
"type": "array",
|
|
263
|
+
"minItems": 1,
|
|
264
|
+
"items": {
|
|
265
|
+
"type": "string",
|
|
266
|
+
"minLength": 1,
|
|
267
|
+
"pattern": "\\S",
|
|
268
|
+
"description": "At least one non-whitespace character. `minLength: 1` alone accepts a single space, which is the one shape where this schema and `graph.py`'s verdict gate disagreed — the gate strips, the schema counted. A cross-check fixture comparing the two found it."
|
|
269
|
+
}
|
|
270
|
+
}
|
|
271
|
+
}
|
|
272
|
+
}
|
|
273
|
+
},
|
|
274
|
+
"rule_parked_needs_reason": {
|
|
275
|
+
"description": "REQ-012: a parked node names why it is parked. Referenced from `node.allOf`.",
|
|
276
|
+
"if": {
|
|
277
|
+
"properties": {
|
|
278
|
+
"status": {
|
|
279
|
+
"const": "parked"
|
|
280
|
+
}
|
|
281
|
+
},
|
|
282
|
+
"required": [
|
|
283
|
+
"status"
|
|
284
|
+
]
|
|
285
|
+
},
|
|
286
|
+
"then": {
|
|
287
|
+
"required": [
|
|
288
|
+
"parked_reason"
|
|
289
|
+
],
|
|
290
|
+
"properties": {
|
|
291
|
+
"parked_reason": {
|
|
292
|
+
"type": "string",
|
|
293
|
+
"minLength": 1,
|
|
294
|
+
"pattern": "\\S"
|
|
295
|
+
}
|
|
296
|
+
}
|
|
297
|
+
}
|
|
286
298
|
}
|
|
287
299
|
}
|
|
288
300
|
}
|
|
@@ -319,6 +319,29 @@ spends its length refusing.
|
|
|
319
319
|
|
|
320
320
|
## GATE (manual)
|
|
321
321
|
|
|
322
|
+
> **Acceptance policy `AP-1` · owner: the operator of the project running this pipeline ·
|
|
323
|
+
> in force since 2026-08-20 · supersedes: the unversioned ladder that stood before it.**
|
|
324
|
+
>
|
|
325
|
+
> This block is the policy, and naming it that closes a contradiction that stood in the
|
|
326
|
+
> shipped doctrine. [`gates.md`](gates.md) → *Axis A — the stage gate type* says which stages are
|
|
327
|
+
> manual is the operator's decision and that **the framework fixes no stage count and no
|
|
328
|
+
> gate assignment** — while the ladder below fixes twelve criteria and the statuses they
|
|
329
|
+
> may take. Both were true and neither was scoped, so a reader could take either as the
|
|
330
|
+
> rule.
|
|
331
|
+
>
|
|
332
|
+
> **The scope, stated once:** `gates.md`'s sentence governs the **pipeline's shape** — how
|
|
333
|
+
> many stages a project runs and which of them are `auto`, `judgment` or `manual`, all of
|
|
334
|
+
> it in the project's own `pipeline.json`. `AP-1` governs **what stage 10's manual gate
|
|
335
|
+
> asks when a project runs one.** A project may drop stage 10, or make it `auto` for a
|
|
336
|
+
> class of work, and `AP-1` then does not apply to it; a project that keeps it manual gets
|
|
337
|
+
> these criteria and this vocabulary, not a per-run selection from them.
|
|
338
|
+
>
|
|
339
|
+
> **Changing it is a decision, not an edit.** An amended criterion moves the version to
|
|
340
|
+
> `AP-2` and lands with a row in the project's decision register, because a policy that can
|
|
341
|
+
> be edited silently is the *«recorded in the slot reserved for what a machine
|
|
342
|
+
> established»* failure one level up: an acceptance standard nobody can cite by version is
|
|
343
|
+
> one every run re-negotiates.
|
|
344
|
+
|
|
322
345
|
All of:
|
|
323
346
|
|
|
324
347
|
1. **Every shipped REQ has a verification row, and every row names a REQ its own run
|
|
@@ -47,12 +47,21 @@ The method most audits use is **horizontal**: compare the documents against each
|
|
|
47
47
|
other, then do it again. It works, and then it fails in a way that is invisible
|
|
48
48
|
from inside it. Measured over seven passes on a production repository:
|
|
49
49
|
|
|
50
|
-
| Pass | Findings | …of which the previous pass's own fixes caused |
|
|
51
|
-
|
|
52
|
-
| 4 | 12 | 5 |
|
|
53
|
-
| 5 | 17 | 9 |
|
|
54
|
-
| 6 | 13 | 10 |
|
|
55
|
-
| 7 | 19 | 4 |
|
|
50
|
+
| Pass | Findings | …of which the previous pass's own fixes caused | Self-inflicted share |
|
|
51
|
+
|---|---|---|---|
|
|
52
|
+
| 4 | 12 | 5 | 42% |
|
|
53
|
+
| 5 | 17 | 9 | 53% |
|
|
54
|
+
| 6 | 13 | 10 | 77% |
|
|
55
|
+
| 7 | 19 | 4 | 21% — see the note below |
|
|
56
|
+
|
|
57
|
+
**The trend above is measured over passes four to six, and pass seven is not part of
|
|
58
|
+
it.** 42% → 53% → 77% is the decay this section teaches; the seventh pass fell back to
|
|
59
|
+
4 of 19 and **the record does not say what that pass did differently.** The row stays,
|
|
60
|
+
with the gap named, for two reasons. Deleting a measured row to protect a claim is the
|
|
61
|
+
opposite of what this file asks of every gate it describes — and the vertical pass
|
|
62
|
+
described below cannot be credited for the drop, because it is a separate pass over the
|
|
63
|
+
same repository, not the seventh. Whatever pass seven did, it is unrecorded, and an
|
|
64
|
+
unrecorded cause is what this table has to say about it.
|
|
56
65
|
|
|
57
66
|
By pass six the audit was **mostly repairing itself**. Not fatigue — arithmetic.
|
|
58
67
|
Each pass edits the corpus the next pass reads, so the newest edits are always the
|
|
@@ -58,6 +58,15 @@ not the operator confirming it is what they asked for, and no amount of checking
|
|
|
58
58
|
makes it one. Which stages are manual is the **operator's** decision, recorded in
|
|
59
59
|
their `pipeline.json`; the framework fixes no stage count and no gate assignment.
|
|
60
60
|
|
|
61
|
+
**That sentence is about the pipeline's SHAPE, and nothing else.** It says a project
|
|
62
|
+
chooses how many stages it runs and which of them wait for a person. It does not say the
|
|
63
|
+
criteria inside a gate are per-run negotiable: where a project keeps stage 10 manual, what
|
|
64
|
+
that gate asks is [`acceptance.md`](acceptance.md)'s policy **`AP-1`**, which is versioned
|
|
65
|
+
and has an owner. The two rules stood side by side unscoped until 2026-08-20 (`B-091`), and
|
|
66
|
+
a reader could take either as the whole rule — *the framework fixes nothing* and *the ladder
|
|
67
|
+
is fixed* are both in the shipped doctrine, which is how an acceptance standard becomes
|
|
68
|
+
something every run re-argues.
|
|
69
|
+
|
|
61
70
|
## The judgment gate — a ruling is not a measurement
|
|
62
71
|
|
|
63
72
|
Two types were not enough, and the gap was not cosmetic. A reviewer's ruling, a check that
|
|
@@ -369,12 +369,18 @@ deduplicated, and always exiting 0 — a hook that can fail a `Read` breaks ever
|
|
|
369
369
|
every session. `scripts/graph.py doctrine` reports how many of the bundle's reference files
|
|
370
370
|
a run opened and lists the rest.
|
|
371
371
|
|
|
372
|
-
Why
|
|
373
|
-
written by the party the claim is about, is not evidence. And
|
|
372
|
+
Why a hook appends it is the same reason as above: a claim about what somebody read,
|
|
373
|
+
written by the party the claim is about, is not evidence. **And that is an intent rather
|
|
374
|
+
than a proof, so the verb reports `unattested`:** the ledger is the file the agent appends
|
|
375
|
+
to at every stage and carries no writer field, so nothing in it separates a hook-written
|
|
376
|
+
line from one an agent typed. Saying *never agent-written* was a provenance claim the
|
|
377
|
+
format cannot support — B-014's class, in the mechanism built to close it.
|
|
378
|
+
|
|
379
|
+
And why the verb prints
|
|
374
380
|
`unmeasured` rather than `0` when there are no such lines is the same reason again — the
|
|
375
381
|
hook being absent and the run reading nothing are **opposite facts** the ledger cannot
|
|
376
382
|
separate, so it claims neither. A `0` there would be the reassuring answer to a question
|
|
377
|
-
nobody asked, over
|
|
383
|
+
nobody asked, over 35 files nobody checked.
|
|
378
384
|
|
|
379
385
|
It is a disclosure: no floor, no direction, never a target. The moment the number becomes
|
|
380
386
|
something to raise, a run will open files to raise it.
|
|
@@ -19,7 +19,7 @@ justifies reading it protects one section while the file below it doubles.
|
|
|
19
19
|
| Artifact | Parts | How it is read |
|
|
20
20
|
|---|---|---|
|
|
21
21
|
| `docs/evidence/retro.md` — **one per project** | **Standing instructions** (max **10**) · **Run stamps** (max **10**, oldest rotate out) | stage 0, **in full** — both are bounded by a **cap**, which *one line each* never was |
|
|
22
|
-
| the same file's **Recent log** | entries from the last five run stamps
|
|
22
|
+
| the same file's **Recent log** | narrative entries, **uncapped by design** — the heading said *entries from the last five run stamps* until 2026-08-20 while the section held 25 going back nine days, a bound in a heading that nothing enforced | stage 0, **queried** by the task's nouns. It said *in full* until 2026-08-10, when it measured **74%** of the file: an uncapped section inside a binding source is what makes the capped part get skimmed |
|
|
23
23
|
| `docs/evidence/retro/YYYY-QN.md` — the archive | every entry and every retirement ever written, append-only | **queried** by the task's nouns; never read end to end |
|
|
24
24
|
|
|
25
25
|
Seed the archive from [`../templates/retro-archive.md`](../templates/retro-archive.md).
|
|
@@ -909,13 +909,21 @@ def cmd_producer(graph, args):
|
|
|
909
909
|
def cmd_doctrine(graph, args):
|
|
910
910
|
"""Which doctrine this run actually read — B-061.
|
|
911
911
|
|
|
912
|
-
The bundle is
|
|
912
|
+
The bundle is 35 reference files. A run reads some subset and nothing recorded which,
|
|
913
913
|
so **a skipped file and a read one were indistinguishable** — the class every guard in
|
|
914
914
|
this repository exists to catch, left standing over the doctrine itself.
|
|
915
915
|
|
|
916
|
-
`read:` lines
|
|
917
|
-
|
|
918
|
-
|
|
916
|
+
`read:` lines are written by a hook rather than by the agent, for the same reason `gate:`
|
|
917
|
+
is: a claim about what somebody read, written by the party the claim is about, is not
|
|
918
|
+
evidence.
|
|
919
|
+
|
|
920
|
+
**And that is an intent, not a proof, so every line here is reported as UNATTESTED.**
|
|
921
|
+
The ledger is `.task-pipeline/run.md` — the file the agent appends to at every stage —
|
|
922
|
+
so nothing in it distinguishes a hook-written line from one an agent typed. The doctrine
|
|
923
|
+
said *hook-written, never agent-written*, which is a provenance claim this script cannot
|
|
924
|
+
check and no format here carries; B-014's class, committed by the mechanism built to
|
|
925
|
+
close it. Until the ledger can attest a writer, the honest output is the count plus the
|
|
926
|
+
word: read as *this is what the ledger says, and the ledger cannot say who wrote it*.
|
|
919
927
|
|
|
920
928
|
**The one rule that matters here: no `read:` lines means UNMEASURED, never «read
|
|
921
929
|
nothing».** Zero would be the reassuring answer to a question nobody asked, and this
|
|
@@ -958,7 +966,10 @@ def cmd_doctrine(graph, args):
|
|
|
958
966
|
return 0
|
|
959
967
|
|
|
960
968
|
unread = [r for r in refs if r not in read]
|
|
961
|
-
print(f"doctrine: {len(read)} of {len(refs)} reference files read")
|
|
969
|
+
print(f"doctrine: {len(read)} of {len(refs)} reference files read — unattested")
|
|
970
|
+
print(" unattested: the ledger is the file the agent appends to at every stage, "
|
|
971
|
+
"so nothing in it proves the hook wrote these lines rather than an agent. The "
|
|
972
|
+
"count is what the ledger says; who wrote it is not recorded.")
|
|
962
973
|
print(" a disclosure: no floor, no direction, never a target. A run that needs "
|
|
963
974
|
"four files and reads four is not worse than one that reads thirty.")
|
|
964
975
|
for r in unread:
|
|
@@ -31,7 +31,7 @@ has not fired in the last five run stamps, or in the last sixty days. At eleven
|
|
|
31
31
|
rows, the oldest never-fired
|
|
32
32
|
row goes — the cap is not negotiable, ranking is.
|
|
33
33
|
|
|
34
|
-
## Recent log — entries
|
|
34
|
+
## Recent log — narrative entries, uncapped and queried rather than read (newest first)
|
|
35
35
|
|
|
36
36
|
Older entries and every retirement **move** to `docs/evidence/retro/YYYY-QN.md`
|
|
37
37
|
at the prune. Moving is not deleting: the archive is append-only and holds the
|
|
@@ -20,23 +20,35 @@ Run: `<topic>` · started `<YYYY-MM-DD>` · module map: `<path or "none">`
|
|
|
20
20
|
|
|
21
21
|
## `read:` — which doctrine this run actually opened
|
|
22
22
|
|
|
23
|
-
The bundle is
|
|
23
|
+
The bundle is 35 reference files and nothing recorded which of them a run read, so **a
|
|
24
24
|
skipped file and a read one were indistinguishable** — the class every guard in this
|
|
25
25
|
pipeline exists to catch, left standing over the doctrine itself.
|
|
26
26
|
|
|
27
|
-
`read:` is **
|
|
28
|
-
about what somebody read, written by the party the claim is about, is not evidence. The
|
|
27
|
+
`read:` is **written by a hook rather than by the agent**, for the same reason `gate:` is: a
|
|
28
|
+
claim about what somebody read, written by the party the claim is about, is not evidence. The
|
|
29
29
|
hook is in `hooks.example.json`, matches `Read`, records the path only
|
|
30
30
|
when it is inside the bundle's `references/`, deduplicates, and **always exits 0** — a hook
|
|
31
31
|
that can fail a `Read` would break every turn in every session.
|
|
32
32
|
|
|
33
|
+
**That is an intent, not a proof, and the line says so.** This file is the one the agent
|
|
34
|
+
appends to at every stage, so nothing in it distinguishes a hook-written line from one an
|
|
35
|
+
agent typed: there is no writer field, and `scripts/graph.py` has no provenance check
|
|
36
|
+
because the format gives it nothing to check. The doctrine here said *hook-written, never
|
|
37
|
+
agent-written* until 2026-08-20 — a provenance claim in the file whose whole subject is that
|
|
38
|
+
a claim by the interested party is not evidence, which is B-014's class committed by the
|
|
39
|
+
mechanism built to close it. So `doctrine` and the `gate:` reader both report **`unattested`**
|
|
40
|
+
beside their counts, and the two claims that survive are the ones the ledger can support:
|
|
41
|
+
*this line is in the ledger*, and *nobody recorded who wrote it*. Attesting a writer needs a
|
|
42
|
+
field this format does not have; inventing one that an agent can also fill would restate the
|
|
43
|
+
same claim one level down.
|
|
44
|
+
|
|
33
45
|
`scripts/graph.py doctrine` reads these lines and prints one of three things:
|
|
34
46
|
|
|
35
47
|
| It prints | When | Why not just a number |
|
|
36
48
|
|---|---|---|
|
|
37
49
|
| `unmeasured — no run ledger` | there is no ledger | nothing to read from |
|
|
38
50
|
| `unmeasured — the ledger carries no read: lines` | the hook is absent, **or** the run opened no doctrine | two opposite facts, and the ledger cannot separate them, so neither is claimed |
|
|
39
|
-
| `N of
|
|
51
|
+
| `N of 35 reference files read — unattested`, then each unread one | the hook is installed and fired | the count alone says there is a gap, not where — and `unattested` says the ledger cannot name who wrote the lines |
|
|
40
52
|
|
|
41
53
|
**It is a disclosure: no floor, no direction, never a target.** A run that needs four files
|
|
42
54
|
and reads four is not worse than one that reads thirty — and the moment the number becomes
|
|
@@ -58,15 +70,17 @@ hand: <N|10> — task "<quoted>" — done <n> — surfaced <n> — decisions <n
|
|
|
58
70
|
holds: <stage id> — <n> (<class: what, owner>; … or "none") — enumerated <n>/8 classes, <unlooked: classes not enumerable>
|
|
59
71
|
gate: <stage id> — command "<cmd>" — exit <N> — <ISO-8601>
|
|
60
72
|
event: <compact|session-end|subagent> — <detail> — <ISO-8601>
|
|
61
|
-
read: references/<file>.md # hook-
|
|
73
|
+
read: references/<file>.md # hook-appended, deduped, UNATTESTED (no writer field)
|
|
62
74
|
```
|
|
63
75
|
|
|
64
76
|
- **`stage:`** — written when a gate **returns**, not when the stage is entered. The
|
|
65
77
|
rail's `✓` is derived from this line and from nothing else; a glyph set from memory
|
|
66
78
|
is a summary that is confidently wrong exactly when it matters.
|
|
67
|
-
- **`gate:`** —
|
|
68
|
-
|
|
69
|
-
|
|
79
|
+
- **`gate:`** — appended by `hooks/gate-observer.sh` rather than by an agent, and
|
|
80
|
+
**unattested** for the reason `read:` is: this file is agent-written at every stage
|
|
81
|
+
and carries no writer field, so the line's provenance is an intent the format cannot
|
|
82
|
+
prove. It is the only line here that records what a command **did** rather than what
|
|
83
|
+
somebody concluded: the exit code of the stage's declared `gate.command`, observed. The
|
|
70
84
|
`stage:` line above it is the agent's claim, and the release gate requires the
|
|
71
85
|
two to agree — without this, a gate reads a claim written by the party it
|
|
72
86
|
constrains and confirms an assertion with itself. Absent where the project
|
|
@@ -41,6 +41,7 @@ worse than saying nothing.
|
|
|
41
41
|
## Contents
|
|
42
42
|
|
|
43
43
|
- [Staleness — a row is true about the tree it OBSERVED](#staleness--a-row-is-true-about-the-tree-it-observed)
|
|
44
|
+
- [Environment — a proof is only valid where it ran](#environment--a-proof-is-only-valid-where-it-ran)
|
|
44
45
|
- the ledger itself — one row per REQ, appended by stage 8
|
|
45
46
|
- [What `Human` means, and what it does not](#what-human-means-and-what-it-does-not)
|
|
46
47
|
|
|
@@ -71,11 +72,39 @@ one. Four things overtake a row, and naming which one applies is the note's job:
|
|
|
71
72
|
change in what it covers, a **dependency** change, an **environment** change, and a
|
|
72
73
|
**policy** change — the last being the rule under which the evidence was accepted.
|
|
73
74
|
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
75
|
+
## Environment — a proof is only valid where it ran
|
|
76
|
+
|
|
77
|
+
`Observed at` says which tree the check saw. It does not say **where**, and without that a
|
|
78
|
+
smoke test against a preview URL enters the record in a shape indistinguishable from one
|
|
79
|
+
against production, and a suite green on a laptop with accumulated state is
|
|
80
|
+
indistinguishable from one green on a runner that started clean. Those are the two halves
|
|
81
|
+
of the same question, and only one was instrumented.
|
|
82
|
+
|
|
83
|
+
**`Environment` is a required cell on every row.** Its vocabulary is the project's own,
|
|
84
|
+
declared in the runbook and not invented per row — this file ships with four, and a project
|
|
85
|
+
that deploys differently declares different ones:
|
|
86
|
+
|
|
87
|
+
| Value | What it means |
|
|
88
|
+
|---|---|
|
|
89
|
+
| `production` | the deployed target real users reach |
|
|
90
|
+
| `preview` | a per-branch or per-PR deployment; proves the build, never the release |
|
|
91
|
+
| `ci` | a clean runner, no accumulated state |
|
|
92
|
+
| `local` | a developer machine, with whatever state it has |
|
|
93
|
+
| `—` | **recorded absence** — a row written before this column existed, or one whose environment nobody recorded. Not a value, and never a default to reach for |
|
|
94
|
+
|
|
95
|
+
A missing cell is not the same as `—`: the first is a row that forgot the question, and it is
|
|
96
|
+
**refused**. The second is an answer.
|
|
97
|
+
|
|
98
|
+
**A REQ claiming production behaviour may not be closed on a non-production
|
|
99
|
+
observation.** Stage 8 writes the row and refuses that pairing rather than recording it —
|
|
100
|
+
`pass` in `ci` against a requirement about the deployed product is the exact substitution
|
|
101
|
+
this column exists to make visible.
|
|
102
|
+
|
|
103
|
+
| REQ | What | Run | Shipped in | Observed at | Environment | Auto | Human | Note |
|
|
104
|
+
|---|---|---|---|---|---|---|---|---|
|
|
105
|
+
| REQ-001 | CSV export from a report | `2026-07-28-export` | v1.4.0 | `5f21ac3` | production | pass | 2026-07-30 | opened the deployed page, exported, opened the file |
|
|
106
|
+
| REQ-004 | XLSX export | `2026-07-28-export` | v1.4.0 | `5f21ac3` | ci | pass | **never** | — |
|
|
107
|
+
| REQ-007 | Export respects active filters | `2026-07-28-export` | v1.4.0 | `5f21ac3` | preview | partial | **never** | CSV path only |
|
|
79
108
|
|
|
80
109
|
## Columns
|
|
81
110
|
|
|
@@ -86,6 +115,10 @@ change in what it covers, a **dependency** change, an **environment** change, an
|
|
|
86
115
|
- **Run** — the brief's topic slug, so the context is one file away.
|
|
87
116
|
- **Shipped in** — the tag or commit that carried it. Where a project does not tag,
|
|
88
117
|
the commit, and the same value every row of that run carries.
|
|
118
|
+
- **Observed at** — the commit the check ran against. `—` where nobody recorded one.
|
|
119
|
+
- **Environment** — where it ran, from the project's declared vocabulary. Required; `—` is
|
|
120
|
+
the recorded absence and an omitted cell is refused. See *Environment — a proof is only
|
|
121
|
+
valid where it ran*.
|
|
89
122
|
- **Auto** — what the run's own gate said: `pass` · `partial` · `none`. Copied from the
|
|
90
123
|
coverage table rather than re-derived; where the two disagree the coverage table wins
|
|
91
124
|
and the disagreement is a finding. A coverage verdict of **`review`** — *no check can
|