task-pipeline-skill 1.73.0 → 1.75.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -117,7 +117,7 @@
117
117
  "not": {
118
118
  "pattern": "[\\n\\r]"
119
119
  },
120
- "description": "**How this node will be closed** \u2014 the command a verifier runs, or the named judgement that stands in where no command can (`references/gates.md`'s third gate type: name the judge, record the verdict AS judgement). Its output becomes the node's `evidence` row.\n\nRequired on every node except a `parked` one, and that is the whole point of the field: `agents/verifier.md` tells the verifier to run *the check this node names*, so a node with none leaves it the two options that same paragraph forbids \u2014 invent a check, or run everything. Shipped doctrine read a field the schema did not have (B-080); the field the doctrine reads now exists.\n\nOne string, never a list, because the requirement it serves is singular throughout \u2014 one input, one job, one output, one owner, **one** completion test. A node needing two unrelated checks is a node doing two jobs, and the answer is to split it rather than to widen this field. One gate made of two commands is `a && b`, which is still one gate.\n\nA `parked` node is exempt because it is the one node nobody will close, and a placeholder there \u2014 *n/a, parked* \u2014 is the confidence-without-correctness this schema refuses everywhere else. The field survives a park: `park` sets the status and never removes what the node said it would run."
120
+ "description": "**How this node will be closed** — the command a verifier runs, or the named judgement that stands in where no command can (`references/gates.md`'s third gate type: name the judge, record the verdict AS judgement). Its output becomes the node's `evidence` row.\n\nRequired on every node except a `parked` one, and that is the whole point of the field: `agents/verifier.md` tells the verifier to run *the check this node names*, so a node with none leaves it the two options that same paragraph forbids — invent a check, or run everything. Shipped doctrine read a field the schema did not have (B-080); the field the doctrine reads now exists.\n\nOne string, never a list, because the requirement it serves is singular throughout — one input, one job, one output, one owner, **one** completion test. A node needing two unrelated checks is a node doing two jobs, and the answer is to split it rather than to widen this field. One gate made of two commands is `a && b`, which is still one gate.\n\nA `parked` node is exempt because it is the one node nobody will close, and a placeholder there — *n/a, parked* — is the confidence-without-correctness this schema refuses everywhere else. The field survives a park: `park` sets the status and never removes what the node said it would run."
121
121
  },
122
122
  "evidence": {
123
123
  "type": [
@@ -145,6 +145,89 @@
145
145
  "pattern": "\\S"
146
146
  },
147
147
  "description": "What this node MUTATES — paths, register names, remote resource ids. `references/planning.md` states the rule the frontier needs: *distinct is not the same as independent, and the check is what they touch, never what they are called.* That rule lived in the markdown plan, and the graph replaced the plan as the thing deciding what runs next — so `next` could hand two agents two runnable nodes that write the same file, with nothing able to report it.\\n\\nOptional, and its absence is DISCLOSED rather than treated as «touches nothing»: `next` prints how many frontier nodes declared no targets, because a quiet run and a checked one must not look alike."
148
+ },
149
+ "certification": {
150
+ "type": "object",
151
+ "description": "What the three-tier certification recorded for this node. Written by `graph.py certify` on every round, pass or fail: a failing round that wrote nothing would erase the only evidence that a node is churning, which is the number the loop-guard ceiling reads.",
152
+ "additionalProperties": false,
153
+ "required": [
154
+ "round",
155
+ "tiers",
156
+ "at",
157
+ "history"
158
+ ],
159
+ "properties": {
160
+ "round": {
161
+ "type": "integer",
162
+ "minimum": 1
163
+ },
164
+ "tiers": {
165
+ "type": "object",
166
+ "additionalProperties": false,
167
+ "required": [
168
+ "unit",
169
+ "seam",
170
+ "product"
171
+ ],
172
+ "properties": {
173
+ "unit": {
174
+ "enum": [
175
+ "pass",
176
+ "fail"
177
+ ]
178
+ },
179
+ "seam": {
180
+ "enum": [
181
+ "pass",
182
+ "fail"
183
+ ]
184
+ },
185
+ "product": {
186
+ "enum": [
187
+ "pass",
188
+ "fail"
189
+ ]
190
+ }
191
+ }
192
+ },
193
+ "at": {
194
+ "type": "string",
195
+ "minLength": 1
196
+ },
197
+ "history": {
198
+ "type": "array",
199
+ "minItems": 1,
200
+ "items": {
201
+ "type": "object",
202
+ "additionalProperties": false,
203
+ "required": [
204
+ "unit",
205
+ "seam",
206
+ "product"
207
+ ],
208
+ "properties": {
209
+ "unit": {
210
+ "enum": [
211
+ "pass",
212
+ "fail"
213
+ ]
214
+ },
215
+ "seam": {
216
+ "enum": [
217
+ "pass",
218
+ "fail"
219
+ ]
220
+ },
221
+ "product": {
222
+ "enum": [
223
+ "pass",
224
+ "fail"
225
+ ]
226
+ }
227
+ }
228
+ }
229
+ }
230
+ }
148
231
  }
149
232
  },
150
233
  "allOf": [
@@ -0,0 +1,146 @@
1
+ # Certification — three readings at three distances, and all three must pass
2
+
3
+ A node is not closed by one agent's opinion. It is closed by **three independent
4
+ readings at escalating visibility**, and the run may not advance until all three
5
+ pass.
6
+
7
+ ## Contents
8
+
9
+ - Why one verifier is not enough, stated as the failure it produces
10
+ - The three tiers
11
+ - Blind, and it is the whole design
12
+ - A pass has to mean something, so two rules have teeth
13
+ - The report, and where each field lands
14
+ - The commands
15
+ - The fix cycle, and its ceiling
16
+ - What this costs, said out loud
17
+ - Rationalizations
18
+
19
+ ## Why one verifier is not enough, stated as the failure it produces
20
+
21
+ A verifier reads the diff it was handed. That is not a shortcoming of the agent; it
22
+ is the definition of its context. And it means a whole class of defect is invisible
23
+ to it by construction:
24
+
25
+ - the change is correct where it was made, and a **caller's contract moved** under it
26
+ - a **second implementation of the same rule** did not get the fix
27
+ - a **documented behaviour** is now false, and the document still reads as true
28
+ - **another feature** reaches the same path and nobody considered the interaction
29
+
30
+ None of these is a bug in the changed lines. All of them ship. Each is found by
31
+ looking one level further out than the change — which is a different reading, not a
32
+ longer one, because the context that finds it is the context that excludes the diff.
33
+
34
+ ## The three tiers
35
+
36
+ | Tier | Subject | Characteristic finding |
37
+ |---|---|---|
38
+ | `unit` | the changed functions, classes and branches, plus the node's own `check` | a branch nothing exercises; a boundary that moved |
39
+ | `seam` | everything that can reach the change — callers, callees, implementors, shared state, the neighbours' tests | a contract that moved under a dependent; the duplicate that did not get the fix |
40
+ | `product` | documentation, scenarios, user-visible strings, and the neighbouring features that share this path | a documented behaviour that is now false; an interaction nobody listed |
41
+
42
+ Agents: [`../../../agents/verifier-unit.md`](../../../agents/verifier-unit.md),
43
+ [`verifier-seam.md`](../../../agents/verifier-seam.md),
44
+ [`verifier-product.md`](../../../agents/verifier-product.md).
45
+
46
+ ## Blind, and it is the whole design
47
+
48
+ **The three run in parallel and no tier reads another's report.** Three readings
49
+ that inform each other are one opinion with three signatures — and the failure mode
50
+ is specific: an agent that has just read a convincing account of the implementation
51
+ will paraphrase it back as product truth. The disagreement between blind readings is
52
+ the instrument, so `graph.py certify` refuses a report whose prose cites another
53
+ tier's verdict.
54
+
55
+ Dispatch all three in one message so they run concurrently. Give each the node id,
56
+ its `serves`, and the diff — nothing else, and never another tier's output.
57
+
58
+ ## A pass has to mean something, so two rules have teeth
59
+
60
+ **A tier cannot pass on an empty `scope`.** `scope` is what the tier actually
61
+ opened, with `file:line`. A report that names nothing it read is a rubber stamp, and
62
+ three rubber stamps cost three times one verifier while reading as three times the
63
+ assurance — strictly worse than the single verdict it replaced. `certify` refuses it
64
+ by name.
65
+
66
+ **A tier cannot pass while carrying a `breaks` finding.** There are two severities
67
+ and no third, because a certification that admits a maybe admits everything:
68
+
69
+ - **`breaks`** — the node is not done. Carries a `check`: the command or judgement
70
+ that will prove the fix, for the same reason `replan.add` does. The finding
71
+ becomes a node the next round has to close, and handing that node the absence is
72
+ how the defect returns one round later.
73
+ - **`risk`** — found, judged survivable, and named. It ships, and it reaches the
74
+ closing verdict as a blocker with `can_continue_around: true`, which is exactly
75
+ what a named survivable finding is.
76
+
77
+ ## The report, and where each field lands
78
+
79
+ Eight keys, all required, `[]` a valid answer and silence not one. On a pass
80
+ `certify` assembles the canonical seven-key verdict that
81
+ [`work-graph.md`](work-graph.md) already specifies, and no field is used for
82
+ something it does not mean:
83
+
84
+ | Tier field | Becomes | Because |
85
+ |---|---|---|
86
+ | `confirms` | `done` | asked for, and now true |
87
+ | `not_examined` | `not_verified` | present, and no check touched it |
88
+ | `findings` at `risk` | `blockers`, `can_continue_around: true` | found, judged survivable, named |
89
+ | `evidence` | `evidence`, prefixed with the tier | the command and what it printed |
90
+
91
+ `certify` runs the assembled verdict through the same `verdict_violations` gate
92
+ `close` will apply, so a certification cannot hand the run a verdict its own
93
+ consumer refuses.
94
+
95
+ ## The commands
96
+
97
+ ```bash
98
+ # three reports in, one verdict out — exits 1 if any tier failed
99
+ graph.py certify --node N-007 \
100
+ --tier unit.json --tier seam.json --tier product.json
101
+
102
+ # unchanged, and still the only thing that moves the graph
103
+ graph.py close --verdict .task-pipeline/verdict-N-007.json
104
+ ```
105
+
106
+ ## The fix cycle, and its ceiling
107
+
108
+ A failing round **records itself and leaves the node open.** `certification` on the
109
+ node carries the round number, this round's three verdicts and the history of every
110
+ round — written on failure too, because a failing round that wrote nothing would
111
+ erase the only evidence that a node is churning.
112
+
113
+ The cycle is: `certify` → fail → the `breaks` findings become nodes (each already
114
+ carrying its `check`) → fix → `certify` again, round `N+1`. Same three tiers, same
115
+ blind dispatch. A tier that passed in an earlier round is **re-run**, because the
116
+ fix is a new change and the level it passed on is not the level it now faces.
117
+
118
+ **At the ceiling the gate measures rather than stops** —
119
+ [`loop-guard.md`](loop-guard.md). `--ceiling` defaults to 3. At or over it, `certify`
120
+ still runs and still reports the tiers; what it adds is the name of the tier that
121
+ has failed **every** round. A run spinning on one level needs the operator to see
122
+ *which* level:
123
+
124
+ - the same tier every round → the level is being misread, or the node is the wrong
125
+ shape. Not one fix away. Re-plan the node, do not attempt round four.
126
+ - different tiers each round → churn across levels. Usually one requirement that
127
+ was never decided, surfacing at whichever distance looks at it.
128
+
129
+ ## What this costs, said out loud
130
+
131
+ Three agents per node instead of one. That is the price of the visibility, and it is
132
+ paid per node rather than per run. The three are dispatched in parallel, so the
133
+ wall-clock cost is roughly one reading; the token cost is three. A node whose
134
+ `check` is mechanical and whose blast radius is genuinely nil still pays it — and a
135
+ tier with nothing to find says so in `scope` and `not_examined` rather than being
136
+ skipped, because **a skipped level and a clean level are indistinguishable
137
+ afterwards**, and only one of them is evidence.
138
+
139
+ ## Rationalizations
140
+
141
+ | Temptation | Why it is wrong |
142
+ |---|---|
143
+ | «All three would say the same thing» | Then all three say it, at a cost you already know, and the run has three signatures instead of one guess about what the other two would have found |
144
+ | «The seam tier can read the unit report first — it saves tokens» | It saves tokens by removing the second opinion. The reports are cheap; the independence is the product |
145
+ | «Tier 3 passed last round, skip it» | The fix is a new change. A tier's pass is about the tree it read, and that tree moved |
146
+ | «Round 4 will get it» | Read the ceiling's output. The same tier failing three times is a planning defect wearing a verification failure's clothes |
@@ -49,9 +49,23 @@ the same run and a description of the behaviour is not the behaviour. → the re
49
49
  SHA-resolution guard; the finding shape in [`setup.md`](setup.md); [`tdd.md`](tdd.md) →
50
50
  *When the thing under test is an agent*.
51
51
 
52
- **2. Numbers are computed, never restated.** A count in prose is a number that was true
53
- once. Derive it at check time and compare the stated one against the computed one as the
54
- same object. → [`learned.md`](learned.md) rule 8.
52
+ **2. Numbers are computed, never restated — and an example that instantiates a number
53
+ IS one.** A count in prose is a number that was true once. Derive it at check time and
54
+ compare the stated one against the computed one as the same object. → [`learned.md`](learned.md)
55
+ rule 8.
56
+
57
+ The half that costs more, because it is written by the person who understands the rule:
58
+ **a document quoting a form as an example is indistinguishable from the form.** A release
59
+ note explaining *"the count was written wrongly as `Guards: 412 → 412`"* has just placed a
60
+ second readable count in a section whose count a gate reads — so when the probe removes the
61
+ real one, the narrative still matches and the guard reports green over a section that states
62
+ nothing. Measured three times in one hour on 2026-08-22, each time inside prose *about* this
63
+ very failure: the release note, the board row filed against it, and the repair to the probe.
64
+
65
+ The rule the umbrella already states for commands — **name a dead command, never claim it**
66
+ — holds for a number, a version and a shape. Describe the wrong form; do not write it.
67
+ *"the bold sat around the whole phrase instead of around the second number"* is checkable by
68
+ a reader and invisible to a pattern; the same sentence with the digits in it is a live claim.
55
69
 
56
70
  **3. Every fact has exactly one home.** Other documents link to it; they never restate
57
71
  it. Two homes do not disagree on the day they are written — they disagree on the day one
@@ -31,6 +31,9 @@ elsewhere and is not restated here:
31
31
  - Anatomy of a project gate
32
32
  - Writing the check itself
33
33
  - Probing — plant, run, restore
34
+ - A probe rots, and every way it rots reports green
35
+ - A ratchet prices the rule, not the exception
36
+ - Run the whole suite locally before you push the tag
34
37
  - The neighbour probe — plant the evidence outside the subject
35
38
  - A ratchet's matcher is itself a check, and it needs a near-miss
36
39
  - A green probe is evidence only if the mutation is known to have landed
@@ -385,6 +388,95 @@ Otherwise the next reader has to redo it to know whether it was ever done.
385
388
 
386
389
  ---
387
390
 
391
+ ## A probe rots, and every way it rots reports green
392
+
393
+ The three assertions above prove a probe works **today**. They say nothing about the
394
+ day after, and a probe is uniquely exposed: the thing it guards is the thing that moves
395
+ it. Four rotted in one release on 2026-08-22, and none of them failed loudly — two
396
+ reported `caught`, two reported nothing at all.
397
+
398
+ **1. The anchor is a literal the guarded thing moves.** A probe pinned to *"the bundle
399
+ is N reference files"* stops landing the day a reference file is added — on the release
400
+ that changes the very number it guards. Same for a version, a count word, a phrase a
401
+ release rewrites. **Derive the anchor**: read whatever the text currently says and make
402
+ *that* wrong. A probe that computes `wrong = stated - 1` never needs maintaining.
403
+
404
+ **2. The precondition is inherited from the tree rather than created.** This is the
405
+ subtle one, because it is triggered by the system working correctly. A probe that
406
+ narrows a declared gap to expose *releases after the newest run stamp* has nothing to
407
+ expose the moment a release writes an honest stamp — the newest release is now the
408
+ newest stamp. A probe requiring an `## Unreleased` section has none the moment a release
409
+ absorbs it. **Both landed. Both proved nothing.** A probe must construct the state it
410
+ needs: remove the stamp, write the section, then plant the defect.
411
+
412
+ **3. The document quotes the form the probe removes.** Covered as
413
+ [`documentation.md`](documentation.md) canon 2's second half, and it belongs here too
414
+ because the probe is where it surfaces: prose describing the wrong shape *with real
415
+ values in it* is a second instance of the shape. The probe deletes the real one, the
416
+ narrative still matches, the guard is silent.
417
+
418
+ **How to see it before CI does.** Ask one question per probe: *what does this probe
419
+ LOOK FOR, and who is allowed to change it?* Where the answer is "the thing it guards",
420
+ it is rotting already. A repository-wide sweep is one grep — a two-digit literal inside
421
+ a needle, an `assert` or a `replace` — and the triage is: the number a probe **writes**
422
+ is correct, the number it **looks for** is the defect.
423
+
424
+ ## A ratchet prices the rule, not the exception
425
+
426
+ A floor that counts assertions is a floor that can be lowered by improving the code, if
427
+ the assertions are attached to the wrong things.
428
+
429
+ Measured: a coverage check called `check()` once per *exception* — a silent value, a kept
430
+ value, a promise — and simply `continue`d on the ordinary case. Collapsing four kept
431
+ values, **the remediation the requirement names first**, therefore dropped the count by
432
+ four against its floor and turned the suite red on the stricter answer. The only way
433
+ through was lowering a ratchet whose own reason says a falling count is how a deleted
434
+ requirement hides — so a legitimate lowering and the failure the floor exists to catch
435
+ became indistinguishable.
436
+
437
+ **One assertion per subject examined, whatever its verdict.** The ordinary case asserts
438
+ too. Then the floor is a function of how large the corpus is, not of how many exceptions
439
+ it happens to contain, and doing the right thing can never lower it.
440
+
441
+ **And measure the floor after the last edit, not before it.** Read the count, keep
442
+ editing, restate the count you read — every floor set that way sits below the true one,
443
+ and a floor is a minimum, so nothing ever says so. Where the gate cannot enforce
444
+ equality — it must not, or the ratchet stops allowing growth — it can still **print the
445
+ gap**: `floor 4026, ran 4027: 1 check is not pinned` turns a silent difference into a
446
+ visible one for the cost of one line.
447
+
448
+ ## Run the whole suite locally before you push the tag
449
+
450
+ Not a preference: an arithmetic. A release workflow that runs the full suite takes
451
+ twenty-five to forty minutes per round, and it reports one failure at a time. The same
452
+ suite on the machine that wrote the change takes twelve and reports all of them at once.
453
+
454
+ On 2026-08-22 one tag took **five CI rounds** — a stray key in a `run:` block, a missing
455
+ run stamp, a stamp cap, and then four rotted probes — where a single local `test:all`
456
+ before the first push would have found the last four together. Every refusal was correct.
457
+ The cost was entirely in asking the wrong machine.
458
+
459
+ **A green local suite is not evidence until you know which checks LOOKED.** This
460
+ instruction failed on its own release: the local run was green, CI was not, and the
461
+ difference was a precondition asking `os.path.isdir(".git")` — false in a submodule
462
+ checkout, where `.git` is a *file* holding a gitdir pointer. One check switched itself
463
+ off in the only checkout the family is developed in, silently, and had been doing so
464
+ since it was written. The repository had recorded that class **twice** already, in two
465
+ other files, and this instance was missed both times: knowing a class is not sweeping
466
+ it. Two consequences, and the second is the general one:
467
+
468
+ - ask `exists`, never `isdir`, of anything named `.git`;
469
+ - **a precondition that fails must disclose, not skip.** Where a check cannot run, it
470
+ appends to the unlooked list and the run prints it. A check guarded by a bare `and`
471
+ evaporates without a line of output, which is the one thing this file's own canon
472
+ forbids — and it evaporates most reliably in the environment its authors use.
473
+
474
+ The rule has a second half, and it is the one that makes it stick: **a tag is the only
475
+ thing that runs some checks.** A branch push cannot see a tag that does not exist yet, so
476
+ the tag-ancestry check, the version-sync check and the run-stamp check have no earlier
477
+ opportunity to fire. Locally, run them the way the release does — against the tree you are
478
+ about to tag, with the suite the release claims.
479
+
388
480
  ## The neighbour probe — plant the evidence outside the subject
389
481
 
390
482
  A probe proves a guard rejects **the phrasing its author had in mind**. That is less than
@@ -71,6 +71,7 @@ a row pointing outside the bundle is the defect this file exists to catch.
71
71
  | What a spec must lock, the UX-track order, the module dossier | `references/spec.md` |
72
72
  | The zero-context plan format, parallel groups, set equality | `references/planning.md` |
73
73
  | The work graph: its fields, the verbs and their exit codes, and the three invariants a schema cannot state | `references/work-graph.md`, `scripts/graph.py`, `graph.schema.json` |
74
+ | How a node is CLOSED: three blind readings at escalating visibility, all three required, and the round ledger the ceiling reads | `references/certification.md`, `agents/verifier-{unit,seam,product}.md`, `scripts/graph.py certify` |
74
75
  | Workspace isolation, the subagent loop, who may write the register | `references/build.md` |
75
76
  | The review rubric, diff packages, the three verdicts | `references/review.md` |
76
77
  | **False success** — the class, its known shapes and its two rules | `references/gates.md` |
@@ -273,6 +273,12 @@ never that the work was skipped quietly.
273
273
  `next` refuses to walk one. The reason to prefer it over a prose plan is measured, not
274
274
  aesthetic: a 400-node graph and a 4-node graph produce the same 27-byte frontier, so the
275
275
  cost of knowing what is next does not grow with the programme.
276
+ - **And the check is what THREE readings will run, not one**
277
+ ([`certification.md`](certification.md)). A node is closed by `unit`, `seam` and `product`
278
+ reports, dispatched blind and in parallel, and `certify` refuses the close until all three
279
+ pass. That is a planning fact, not only a verification one: a node whose blast radius nobody
280
+ can name is a node the seam and product tiers cannot scope, so `touches` and `serves` are
281
+ what make the two outer readings possible at all.
276
282
  - **The queue exists here, so the loop arms here** ([`continuity.md`](continuity.md) →
277
283
  *Part 1a*). Where `run.loop.arm` is `after-decomposition` and the map holds more than
278
284
  one module, arm the mode at the close of this stage and print one line: the mode, and
@@ -59,6 +59,7 @@ conditional on the code, never merely sequenced after it.**
59
59
  | `goal` | the release goal | `0` · `3` unstated |
60
60
  | `add` | the id it allocated | `0` · `1` refused |
61
61
  | `park` | the id and the reason | `0` · `1` refused |
62
+ | `certify` | the round, and on a failure every `breaks` finding with its fix and its check | `0` all three tiers passed · `1` a tier failed, or a report is malformed |
62
63
  | `close` | the goal, the new frontier count, and what was not verified | `0` · `1` refused **or the verdict stops the run** |
63
64
  | `producer` | what produced this proof — actor, model, runtime, skill, config, commit, trace | `0` |
64
65
  | `doctrine` | how many of the bundle's reference files this run opened | `0` |
@@ -77,6 +78,8 @@ A **`parked`** node is the single exemption: it is the one node nobody will clos
77
78
  *n/a — parked* in that field is confidence without correctness. `park` never removes what
78
79
  the node said it would run.
79
80
 
81
+ **A node is closed by three readings, not one.** `certify` takes one tier report from each of `unit`, `seam` and `product` — dispatched blind and in parallel — requires all three to pass, and assembles the seven-key verdict `close` consumes. `close`'s contract is unchanged; what changed is that the verdict is now built from three readings at different distances instead of written from one, because a change can be correct where it was made and wrong one level out. A failing round records itself and leaves the node open. Doctrine: [`certification.md`](certification.md).
82
+
80
83
  **`close` stamps the commit; the verifier never supplies it.** A verdict written after the
81
84
  tree moved is evidence about a different tree, and an agent cannot name the wrong commit if
82
85
  it is never the one naming one.