task-pipeline-skill 1.79.1 → 1.80.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/CHANGELOG.md +79 -0
  2. package/README.md +4 -3
  3. package/SKILL-CARD.md +1 -1
  4. package/cursor/rules/task-pipeline.mdc +3 -1
  5. package/evals/RESULTS.md +212 -9
  6. package/evals/evidence-docs.evals.json +109 -0
  7. package/evals/project-audit.evals.json +108 -0
  8. package/evals/run.py +47 -19
  9. package/package.json +1 -1
  10. package/plugins/task-pipeline/.claude-plugin/plugin.json +3 -2
  11. package/plugins/task-pipeline/commands/task-pipeline.md +2 -1
  12. package/plugins/task-pipeline/hooks/gate-observer.sh +19 -2
  13. package/plugins/task-pipeline/skills/evidence-docs/SKILL.md +1 -0
  14. package/plugins/task-pipeline/skills/project-audit/SKILL.md +7 -2
  15. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +49 -56
  16. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +1 -1
  17. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +1 -1
  18. package/plugins/task-pipeline/skills/task-pipeline/references/adoption.md +15 -4
  19. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +4 -2
  20. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -1
  21. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +11 -1
  22. package/plugins/task-pipeline/skills/task-pipeline/references/certification.md +7 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +33 -2
  24. package/plugins/task-pipeline/skills/task-pipeline/references/documentation.md +2 -2
  25. package/plugins/task-pipeline/skills/task-pipeline/references/exposure.md +7 -3
  26. package/plugins/task-pipeline/skills/task-pipeline/references/gates.md +17 -185
  27. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +13 -0
  28. package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/references/probing.md +202 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +8 -4
  31. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +60 -1
  32. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +32 -54
  33. package/plugins/task-pipeline/skills/task-pipeline/references/work-graph.md +1 -1
  34. package/plugins/task-pipeline/skills/task-pipeline/scripts/graph.py +28 -5
  35. package/plugins/task-pipeline/skills/task-pipeline/templates/backlog.md +6 -2
  36. package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +2 -2
@@ -58,6 +58,10 @@ better, plus one that is required only for user-facing work.
58
58
  | **playwright** (CLI — `playwright-cli open`, `snapshot`, `click`, `type`, and the two this pipeline actually asks for: `console` and `requests`; or MCP — `browser_navigate`, `browser_snapshot`, `browser_click`, `browser_console_messages`, `browser_network_requests`) | **stages 5–6 on any project with a web front end**, and **stage 8** on a deployed target — the same job as the row below: open the surface, snapshot it, read the console and the network log. Its own difference is **the channel a look arrives on**: the CLI is a shell command, so it does not put a tool schema in the context window the way an MCP channel does — that is upstream's own comparison and it is *CLI against MCP*, which makes it a claim about this row's own two halves before it is a claim about the row below. Both channels default to an **accessibility-tree snapshot rather than pixels**, so the ordinary look costs a page of text and no vision model; `screenshot` exists in both and costs one when you ask for it | **Recommended** — never a gate; absent → say the surface was verified **by reading the diff** and treat that as the weaker claim it is | CLI: `npm install -D @playwright/cli@latest` → `npx playwright-cli --help` (or `npm i -g` for the global binary; `playwright-cli install --skills` adds its agent skills, `--global` to the home directory). MCP: `claude mcp add playwright npx @playwright/mcp@latest` |
59
59
  | **chrome-devtools** (MCP — `list_pages`, `navigate_page`, `take_snapshot`, `take_screenshot`, `evaluate_script`, `list_console_messages`, `list_network_requests`, `lighthouse_audit`, `performance_start_trace`, `take_heapsnapshot`) | **stages 5–6 on any project with a web front end** — verify the **rendered** surface rather than the diff: computed layout, console errors, failed requests. **Stage 8** on a deployed web target: load the page and read what the browser did, not what the deploy said. Its own difference is **reach past the page**: a Lighthouse category (`lighthouse_audit`, which `seo-aeo-audit` builds on) and a heap snapshot have no equivalent in the row above. A performance trace has one — `playwright-cli tracing-start` records one — but only this row's arrives with `performance_analyze_insight` over it | **Recommended** — never a gate; absent → say the surface was verified **by reading the diff** and treat that as the weaker claim it is | `/plugin install chrome-devtools-mcp@claude-plugins-official` (or connect the MCP server directly) |
60
60
  | **[agent-sync](https://github.com/ssheleg/agent-sync)** (`/agent-sync`, **≥ 1.3.0** — `finish` did not exist before it, so an older install turns the stage-10 close-out into a command that is not there) | **guarded registers** — a lease before writing one, `reserve` before minting an id, `reconcile`/`record` for intent vs as-built, and `finish` for the stage-10 multi-repository close-out ([`documentation.md`](documentation.md)) | **Recommended** — never a gate. Absent → the run is **`ungated`** and must say so out loud; the discipline still applies, only the arbitration is missing | `npx sshlg-skills install` |
61
+ | **sheleg-dev** (`stripe-billing`, `crypto-payments`, `error-tracking`, `ad-tracking`, `google-signin`, `google-auth`, `frontend-performance`) | **stage 5 when the task wires money, tracking, sign-in or page speed** — the seams a generated integration gets wrong in ways no screen shows: the webhook is the payment, a thank-you-page event cannot know the charge cleared, a duplicated `event_id` counts revenue twice | **Recommended** — never a gate; absent → the integration ships on the host's own doctrine and the close-out names the seam nobody checked | `/plugin marketplace add ssheleg/sheleg-dev` → `/plugin install sheleg-dev@sheleg-dev` |
62
+ | **agent-stack** (`agent-orchestrator`, `agent-evals`, `agent-interop`, `agent-harness`) | **stage 5 when the thing being BUILT is an agent system** — the orchestrator's loop, evals that say whether it got better, the protocols it speaks (MCP/A2A), the wallet under LLM resale. Not for coordinating the agents editing this repository (that is agent-sync, above) | **Recommended** — never a gate; absent → the agent layer ships unevaluated and the close-out says so | `/plugin marketplace add ssheleg/agent-stack` → `/plugin install agent-stack@agent-stack` |
63
+ | **telegram-dev** (`telegram-bots`, `telegram-userbots`, `telegram-miniapps`) | **stage 5 when Telegram is the platform, not the transport** — `update_id` as the only idempotency key, a session file that is a logged-in person, a Mini App authenticated by one signed query string | **Recommended** — never a gate; absent → the surface ships on the Bot API docs alone and the close-out names the dedup and auth seams unverified | `/plugin marketplace add ssheleg/telegram-dev` → `/plugin install telegram-dev@telegram-dev` |
64
+ | **seo-aeo-audit** (`/seo-aeo-audit`) | **stage 8 when a logged-out reader or a crawler will see the shipped surface** — the check that a machine will find it; the design-time rule lives with the host, this is the audit after | **Recommended** — never a gate; absent → visibility ships undesigned and unaudited, said in those words | `/plugin marketplace add ssheleg/seo-aeo-audit` → `/plugin install seo-aeo-audit@seo-aeo-audit` |
61
65
  | ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
62
66
  | ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
63
67
 
@@ -127,6 +131,23 @@ Pipeline companions (stage doctrine is built in — nothing to install for it):
127
131
  claude mcp add playwright npx @playwright/mcp@latest
128
132
  (running without it — see the line below; without BOTH,
129
133
  the surface is verified by reading the diff)
134
+ ✗ sheleg-dev — this task wires money, tracking, sign-in or page speed;
135
+ it owns the seams no screen shows (the webhook IS the
136
+ payment, a thank-you event cannot know the charge cleared):
137
+ /plugin marketplace add ssheleg/sheleg-dev
138
+ /plugin install sheleg-dev@sheleg-dev
139
+ ✗ agent-stack — the thing being BUILT here is an agent system; it owns
140
+ the loop, the evals, the protocols and the wallet:
141
+ /plugin marketplace add ssheleg/agent-stack
142
+ /plugin install agent-stack@agent-stack
143
+ ✗ telegram-dev — Telegram is this task's platform, not its transport;
144
+ it owns update dedup, session files and Mini App auth:
145
+ /plugin marketplace add ssheleg/telegram-dev
146
+ /plugin install telegram-dev@telegram-dev
147
+ ✗ seo-aeo-audit — a logged-out reader will see the shipped surface;
148
+ stage 8 can audit what a machine will find:
149
+ /plugin marketplace add ssheleg/seo-aeo-audit
150
+ /plugin install seo-aeo-audit@seo-aeo-audit
130
151
  ✗ chrome-devtools — recommended when this project has a web front end:
131
152
  stages 5-6 check the RENDERED surface instead of the
132
153
  diff, stage 8 reads what the browser did after a deploy.
@@ -204,14 +225,24 @@ Rules:
204
225
  the MCP and then *continues text-only on its own*, so without a recorded answer
205
226
  the run silently narrows from "designed" to "described". The sweep row is
206
227
  `3 Design surface` ([`grill.md`](grill.md) → *The autonomy sweep*).
228
+ - **The four domain companions print only when stage-0's own reading warrants
229
+ them** — sheleg-dev when the task wires money, tracking, sign-in or speed;
230
+ agent-stack when the thing being built is an agent system; telegram-dev when
231
+ Telegram is the platform; seo-aeo-audit when a logged-out reader will see the
232
+ surface. A block that offers all four on a CSV parser teaches the operator to
233
+ stop reading the block.
207
234
  - **agent-sync**: detect via a resolving `/agent-sync` or a `.claude/agent-sync.json`
208
235
  in the project. Present → take a lease before writing a guarded register and
209
236
  reserve ids before minting them. Absent → print the line **once**, continue, and
210
237
  **record the run as `ungated`** — never describe the project as protected
211
238
  ([`documentation.md`](documentation.md) → *Registers are shared state*).
212
- - **The behaviour line prints whenever `evals/RESULTS.md` records no blind run**, and
239
+ - **The behaviour line prints whenever the eval ledger records no blind run**, and
213
240
  disappears the moment one is recorded — it is a state of the evidence, not a warning
214
- and not a ratchet. This bundle asks every project it touches for evidence rather than
241
+ and not a ratchet. The ledger is `evals/RESULTS.md` **in this skill's own
242
+ repository** (github.com/ssheleg/task-pipeline) — it does not ship in the bundle,
243
+ so an installed copy that cannot resolve it prints the line with `unknown` in
244
+ place of the counts rather than dropping it: an unreadable ledger and a clean one
245
+ are different facts. This bundle asks every project it touches for evidence rather than
215
246
  assertion, and until 2026-08-10 it made the opposite claim about itself by saying
216
247
  nothing: **one self-observed run by the author, zero blind runs on zero models**, with
217
248
  a preflight that reported companion availability in detail and its own confidence not
@@ -80,8 +80,8 @@ under every link checker, because the checker resolves from the file's home. →
80
80
  planted defect before its pass means anything — and the plant must be proven to have
81
81
  landed in the text the check actually parses. **A model used as a judge is the same
82
82
  object**: until it has been seen disagreeing with a human label on a case known to be
83
- bad, its pass is an opinion with a number attached. → [`gates.md`](gates.md) →
84
- *Probing*; [`learned.md`](learned.md) rules 4 and 5; [`tdd.md`](tdd.md) → *When the
83
+ bad, its pass is an opinion with a number attached. → [`probing.md`](probing.md) →
84
+ *Probing — plant, run, restore*; [`learned.md`](learned.md) rules 4 and 5; [`tdd.md`](tdd.md) → *When the
85
85
  thing under test is an agent*.
86
86
 
87
87
  **6. A check proves its scope and nothing beyond it.** Every gate carries what it does
@@ -82,9 +82,13 @@ invisible *precisely because nobody is running a pipeline*; a check that only ex
82
82
  inside a run can never say *"stop, fourteen things are unconfirmed."* So it takes no
83
83
  brief, opens no grill, and writes nothing on its own.
84
84
 
85
- It prints four sections, each read from a file this pipeline already keeps: the exposure
86
- line and its check-list, the board's open rows by computed priority, the carry-over
87
- ledgers' unresolved count, and the code graph's staleness where one exists.
85
+ It prints one section per source, each read from a file this pipeline already keeps: the
86
+ exposure line and its check-list, the board's open rows by computed priority, the
87
+ carry-over ledgers' unresolved count, the code graph's staleness where one exists — and
88
+ **abandoned runs**: any run ledger whose last `event:` line is `session-end` short of
89
+ acceptance ([`progress.md`](progress.md) → *The run's own lifecycle*). Before that
90
+ last one, a ledger that simply stopped was indistinguishable from a run still in
91
+ progress, which is exactly the invisibility this mode exists to end.
88
92
 
89
93
  **Where the operator asks it to file findings**, it appends board rows whose `Source`
90
94
  names the checkup and its date — so a row a machine created is distinguishable from one a
@@ -30,13 +30,9 @@ elsewhere and is not restated here:
30
30
  - False success — when a mechanism reports a win it never checked
31
31
  - Anatomy of a project gate
32
32
  - Writing the check itself
33
- - Probing — plant, run, restore
34
- - A probe rots, and every way it rots reports green
35
33
  - A ratchet prices the rule, not the exception
36
34
  - Run the whole suite locally before you push the tag
37
- - The neighbour probe — plant the evidence outside the subject
38
35
  - A ratchet's matcher is itself a check, and it needs a near-miss
39
- - A green probe is evidence only if the mutation is known to have landed
40
36
  - The false-positive budget
41
37
  - Ratchets
42
38
  - The result is the goal; the check is how you know
@@ -107,6 +103,15 @@ Gate assignment is the operator's call and the framework fixes none — so shipp
107
103
  reclassified stage list would contradict the sentence above it. The type exists; the
108
104
  project chooses where it applies.
109
105
 
106
+ **The rubric steers the builder before it judges the work.** A judgment gate's
107
+ criteria are read twice — by the judge at the gate, and earlier by whoever builds
108
+ toward it — and the second reading is absorbed as direction: Anthropic reports
109
+ agents converging on one aesthetic because a rubric's adjective ("museum
110
+ quality") was obeyed as a target rather than weighed as a bar (read 2026-08-30).
111
+ So write criteria as the direction you want taken, not only as the bar to clear;
112
+ the words will be obeyed either way, and a rubric written carelessly is an
113
+ instruction issued accidentally.
114
+
110
115
  | Rationalization | Why it is wrong |
111
116
  |---|---|
112
117
  | *"The reviewer approved it, so the gate passed."* | It did — as a judgement. Type it as one, or the table claims a machine agreed |
@@ -191,7 +196,10 @@ Four preconditions. Skipping any of them turns a run into a claim.
191
196
  added to an already-red base passes for the wrong reason and proves nothing.
192
197
  2. **The check has been probed.** Green from a check nobody has watched fail is
193
198
  worth nothing. If you did not plant the defect, you do not know what the green
194
- means.
199
+ means. How to plant, what a probe owes, and every way one rots is
200
+ [`probing.md`](probing.md) — the harness is `test/probe.py`
201
+ (`npm run test:probe`), and hand-rolling a fourth copy of its loop is how
202
+ three probes in one day proved the wrong guard.
195
203
  3. **You have read its scope header** and know what it does **not** cover. A gate
196
204
  is evidence for exactly the surface it walks; quoting it beyond that is how
197
205
  "the gate is green" becomes a false statement made in good faith.
@@ -338,89 +346,6 @@ exit 0
338
346
 
339
347
  ---
340
348
 
341
- ## Probing — plant, run, restore
342
-
343
- The law is [`audit.md`](audit.md)'s exit criterion. The procedure is this, and it is
344
- not optional for a check you intend to trust:
345
-
346
- ```bash
347
- cp -R . /tmp/probe && cd /tmp/probe
348
- python3 - <<'PY' # plant IN PYTHON, never sed -i
349
- p = "docs/DECISIONS.md"
350
- s = open(p).read().replace("DEC-0001", "DEC-9999", 1)
351
- open(p, "w").write(s)
352
- PY
353
- bash scripts/check-docs.sh; echo "planted -> exit=$?" # MUST be non-zero
354
- cd - && bash scripts/check-docs.sh; echo "clean -> exit=$?" # MUST be 0
355
- ```
356
-
357
- **Assert on `$?`, not on a `FAIL` line in the output.** A line on stdout is a
358
- decoration; the exit code is what CI reads.
359
-
360
- **Doubt the probe before you doubt the check.** On the project this comes from,
361
- **four of five** silent probes were the probe's fault: one added a *definition*
362
- where the check looks for an unresolved *reference*; one edited a string whose
363
- whitespace did not match; one hit the first prose mention instead of the table row;
364
- one flipped a row whose cell was already empty, so nothing was planted. A silent
365
- check is a claim about two things, and the probe is the one to doubt first — prove
366
- your edit landed in the text the check actually parses.
367
-
368
- **Three assertions, and the third is the one hand-rolled probes keep missing.** A
369
- plant that trips some *other* check has proved that other check works. Three probes
370
- in one day passed that way: one removed 1 of 3 identical lines and left the shape
371
- intact; one decremented a number inside an **already-released** section; one deleted
372
- the shouted spelling of a phrase and left the lowercase one. Each landed somewhere
373
- real and demonstrated nothing about the guard it was written for. So:
374
-
375
- 1. **the substitution landed** — `replace` matching nothing returns the string
376
- unchanged and raises nothing;
377
- 2. **the exit code is non-zero** — not a `FAIL` line on stdout;
378
- 3. **the message that fired belongs to the guard under test** — named up front, not
379
- recognised afterwards.
380
-
381
- `test/probe.py` does all three (`npm run test:probe` self-tests the harness, including
382
- its own failure branches, because a harness whose failure path has never executed is
383
- the thing it exists to stop). Declare the plant as `Plant(label, path, old, new,
384
- expect=<the guard's own words>)` rather than hand-rolling a fourth copy of the loop.
385
-
386
- **Record the probe.** One line per section, in the change that ships the check.
387
- Otherwise the next reader has to redo it to know whether it was ever done.
388
-
389
- ---
390
-
391
- ## A probe rots, and every way it rots reports green
392
-
393
- The three assertions above prove a probe works **today**. They say nothing about the
394
- day after, and a probe is uniquely exposed: the thing it guards is the thing that moves
395
- it. Four rotted in one release on 2026-08-22, and none of them failed loudly — two
396
- reported `caught`, two reported nothing at all.
397
-
398
- **1. The anchor is a literal the guarded thing moves.** A probe pinned to *"the bundle
399
- is N reference files"* stops landing the day a reference file is added — on the release
400
- that changes the very number it guards. Same for a version, a count word, a phrase a
401
- release rewrites. **Derive the anchor**: read whatever the text currently says and make
402
- *that* wrong. A probe that computes `wrong = stated - 1` never needs maintaining.
403
-
404
- **2. The precondition is inherited from the tree rather than created.** This is the
405
- subtle one, because it is triggered by the system working correctly. A probe that
406
- narrows a declared gap to expose *releases after the newest run stamp* has nothing to
407
- expose the moment a release writes an honest stamp — the newest release is now the
408
- newest stamp. A probe requiring an `## Unreleased` section has none the moment a release
409
- absorbs it. **Both landed. Both proved nothing.** A probe must construct the state it
410
- needs: remove the stamp, write the section, then plant the defect.
411
-
412
- **3. The document quotes the form the probe removes.** Covered as
413
- [`documentation.md`](documentation.md) canon 2's second half, and it belongs here too
414
- because the probe is where it surfaces: prose describing the wrong shape *with real
415
- values in it* is a second instance of the shape. The probe deletes the real one, the
416
- narrative still matches, the guard is silent.
417
-
418
- **How to see it before CI does.** Ask one question per probe: *what does this probe
419
- LOOK FOR, and who is allowed to change it?* Where the answer is "the thing it guards",
420
- it is rotting already. A repository-wide sweep is one grep — a two-digit literal inside
421
- a needle, an `assert` or a `replace` — and the triage is: the number a probe **writes**
422
- is correct, the number it **looks for** is the defect.
423
-
424
349
  ## A ratchet prices the rule, not the exception
425
350
 
426
351
  A floor that counts assertions is a floor that can be lowered by improving the code, if
@@ -477,61 +402,11 @@ the tag-ancestry check, the version-sync check and the run-stamp check have no e
477
402
  opportunity to fire. Locally, run them the way the release does — against the tree you are
478
403
  about to tag, with the suite the release claims.
479
404
 
480
- ## The neighbour probe — plant the evidence outside the subject
481
-
482
- A probe proves a guard rejects **the phrasing its author had in mind**. That is less than
483
- it looks, and the gap has one shape: **a check answered by text that is not its subject.**
484
-
485
- Measured on one project in one session, six times, each found by a reader planting a
486
- defect and watching `PASS`:
487
-
488
- | The check's subject | The text that answered it instead |
489
- |---|---|
490
- | stage 2 names the loop's arming | *"it **arms** the UX track"*, present since an earlier release |
491
- | the section states the authorization floor | the same phrase in a Rationalizations row, and in a section written twenty-nine releases before |
492
- | the run stamps have a cap | the **standing instructions'** `max 10`, in the same table cell — and, once the check was narrowed past it, the same cap moved to the other side of the `·` |
493
- | a section read in full stays capped | rows under **one** heading, while a second heading in the same file held forty more |
494
-
495
- Every one of those guards had a probe. Every probe fired. The probes and the guards were
496
- written in the same hour from the same reading, so they shared the same blind spot.
497
-
498
- **So a guard that reads a scoped span owes a second probe, and it plants in two places at
499
- once:**
500
-
501
- 1. **break the subject** — remove the thing the guard is about;
502
- 2. **plant the guard's own evidence next door** — in a sibling section, an adjacent table
503
- cell, a rationalizations row, the other side of a separator;
504
- 3. require the guard to **still fail**.
505
-
506
- A guard that passes step 3 is reading its subject. A guard that goes green is reading the
507
- neighbourhood, and the ordinary probe cannot tell the difference — which is why this one
508
- is separate rather than a stricter version of it.
509
-
510
- **Step 2 means the literal the predicate matches *today*, read out of the guard.** The
511
- first three neighbour probes written against this section got that wrong on their first
512
- run: one planted the needles of two **retired** predicates, which proves the guards that
513
- used to exist were neighbour-answerable and says nothing about the one that does. A
514
- neighbour probe keyed to a needle the guard no longer reads is the same defect it was
515
- written to catch, one level up.
516
-
517
- **A probe that only deletes is not a neighbour probe.** It may still be a correct probe —
518
- but if it relies on copies that already sit next door, it must **assert they are there**.
519
- Otherwise a later edit removes them and the probe quietly becomes a delete-only test that
520
- still passes, having stopped testing the thing it is named for.
521
-
522
- **Positional narrowing is not scoping.** Three of the six were "fixed" by cutting the
523
- search down to a row, then to everything after a phrase, and fell each time to text that
524
- was still inside the cut. Scope by *what the span is about*: split to the cell, then to
525
- the item, then match on flattened text so an emphasis marker cannot hide the boundary.
526
-
527
- **And state the span in the guard.** One line above the predicate — *what it reads, and
528
- where that ends*. It costs nothing and it is the only part of a check a later reader can
529
- disagree with before the defect arrives.
530
-
531
405
  ## A ratchet's matcher is itself a check, and it needs a near-miss
532
406
 
533
- Reported from another project through `retro.publish`, and it is the neighbour probe's
534
- own class arrived at independently — which is the strongest evidence either has.
407
+ Reported from another project through `retro.publish`, and it is the neighbour
408
+ probe's own class ([`probing.md`](probing.md) → *The neighbour probe*) arrived at
409
+ independently — which is the strongest evidence either has.
535
410
 
536
411
  A run built a ratchet to hold a coverage debt: a list of units with no test, a guard that
537
412
  fails when the list grows, a count printed at the gate. Exactly the shape
@@ -558,49 +433,6 @@ the reason.** In the reporting project the corrected count was *identical* to th
558
433
  and the composition was not: rows credited falsely came back in as rows genuinely paid off
559
434
  went out. A single number with no delta reads as a run where nothing happened.
560
435
 
561
- ## A green probe is evidence only if the mutation is known to have landed
562
-
563
- Also reported from another project, three times in one day, each caught only because the
564
- result was too good:
565
-
566
- 1. a scripted substitution missed on indentation — the file was unchanged and the probe
567
- measured nothing;
568
- 2. an assertion written against a bare identifier kept matching the **import line** after
569
- the field it guarded was deleted;
570
- 3. a file-extension alternation matched the longer extension as though it were the
571
- shorter, reporting nine live files as missing.
572
-
573
- In all three the observable was identical to success. *"See it fail once"* has an unstated
574
- precondition — **that the thing you changed is the thing the check reads** — and a planted
575
- defect that did not land produces the same green as a check that cannot fail.
576
-
577
- **So a probe that mutates an existing file asserts its plant landed, in the same breath as
578
- planting it.** A probe that writes a whole file has no such question: the file exists or
579
- the command failed. This repository measured itself while writing this section and got the
580
- number wrong three times. A hand-rolled classifier said *206 of 206 already carry it*.
581
- The guard written from the rule said **22 did not** — and was itself too narrow, matching
582
- one spelling of the assertion, so six probes that already had it in lower case were
583
- called defective. A sweep then "fixed" those six and **corrupted five**, splitting live
584
- statements. The true figure was **16**, and it took the guard, a compile check and a
585
- restore from git to find it.
586
-
587
- Two things are worth keeping from that. **The check corrected the measurement that
588
- motivated it** — which is the argument for writing checks rather than counting by hand.
589
- And **a check keyed to one spelling of a rule is the class two sections above**: it
590
- reported as defective the probes that obeyed the rule in different words. **Every**
591
- mutating probe carries the assertion now — the figure is deliberately not
592
- written here. A first draft said *201*, which was true of the branch point and false in
593
- the same commit, because the twelve probes added for this release are themselves mutating
594
- probes. The guard computes it; a number in prose beside a check that can count is the
595
- class this bundle calls restating instead of computing. Probes that write a whole file
596
- need none: the file exists or the command
597
- failed.
598
-
599
- **Prefer an assertion that names the construct over one that names a substring of it.** A
600
- guard written against a bare identifier survives the deletion of everything it guarded,
601
- because the identifier still appears in an import. That is case 2 above and it is the same
602
- class as the section before this one, one level down.
603
-
604
436
  ## The false-positive budget
605
437
 
606
438
  Run a new heuristic over the **real corpus** before shipping it and count the false
@@ -772,7 +604,7 @@ of is deleted in the next refactor by someone who assumed it was dead.
772
604
 
773
605
  ## Cross-cutting, at every stage
774
606
 
775
- 5. Cross-cutting, every stage: **when anything is settled — scope, a contract, a
607
+ **When anything is settled — scope, a contract, a
776
608
  name, a policy, a vocabulary — run the Doc Loop
777
609
  (`references/documentation.md`) before the run moves on**: reserve the id,
778
610
  record it, resolve the question it answers, propagate by the matrix, commit
@@ -63,3 +63,16 @@ and no silent downgrade to a cheaper tier.
63
63
  The recommendation is a **reminder, not a block**. If the recommended tier isn't
64
64
  available, keep the current model, state plainly which one is in use, and run. The
65
65
  pipeline never stalls on a model it can't get.
66
+
67
+ ## A harness clause encodes a model's weakness — stress-test it per generation
68
+
69
+ Every fixed control in this pipeline — a cap, a mandatory checkpoint, a fix-loop
70
+ ceiling — encodes an assumption about what the confirmed model cannot yet do
71
+ reliably, and those assumptions expire: Anthropic reports removing an entire
72
+ construct from its own long-running harness when a newer model generation stopped
73
+ needing it (*Harness design for long-running agents*,
74
+ anthropic.com/engineering, read 2026-08-30). So when the confirmed tier moves a
75
+ generation, re-test the clauses that exist to compensate the old one instead of
76
+ carrying them as doctrine — a control the model has outgrown is not free, it is a
77
+ stop nobody can name the reason for. The retro's cold-retirement trigger is the
78
+ same rule applied to standing instructions; this is it applied to the harness.
@@ -63,6 +63,7 @@ a row pointing outside the bundle is the defect this file exists to catch.
63
63
  | The retro: prune, cap, commits, archive | `references/retrospective.md` |
64
64
  | Rules earned by failure | `references/learned.md` |
65
65
  | **The routing default and its boundary** | `templates/routing-rule.md` |
66
+ | **Probing** — plant, run, restore; how a probe rots; the neighbour probe; the landed-mutation rule | `references/probing.md` |
66
67
  | The seeded doc map, registers and gate | `templates/docmap.md`, `templates/decisions.md`, `templates/open-questions.md`, `templates/docgate.sh` |
67
68
  | Which agent-introduced defects are found, and that the agent fixes them rather than the script | `templates/hygiene.sh`, `references/build.md` |
68
69
  | What a stage-3/4 self-review must read back, and that its trace is computed numbers | `references/spec.md`, `references/planning.md`, `references/learned.md` |
@@ -0,0 +1,202 @@
1
+ # Probing — how a check is proven before it is trusted
2
+
3
+ The authoring doctrine for **probes**: planting a defect, watching the guard
4
+ refuse it, and keeping that proof alive after the release that ships it. It
5
+ lived inside [`gates.md`](gates.md) until 2026-08-31 and moved here whole —
6
+ one home, routed from the gate doctrine it proves. The law it serves is
7
+ [`audit.md`](audit.md)'s exit criterion: **a green from a check nobody has
8
+ watched fail is not evidence.**
9
+
10
+ ## Contents
11
+
12
+ - Probing — plant, run, restore
13
+ - A probe rots, and every way it rots reports green
14
+ - The neighbour probe — plant the evidence outside the subject
15
+ - A green probe is evidence only if the mutation is known to have landed
16
+ - Rationalizations
17
+
18
+ ## Probing — plant, run, restore
19
+
20
+ The law is [`audit.md`](audit.md)'s exit criterion. The procedure is this, and it is
21
+ not optional for a check you intend to trust:
22
+
23
+ ```bash
24
+ cp -R . /tmp/probe && cd /tmp/probe
25
+ python3 - <<'PY' # plant IN PYTHON, never sed -i
26
+ p = "docs/DECISIONS.md"
27
+ s = open(p).read().replace("DEC-0001", "DEC-9999", 1)
28
+ open(p, "w").write(s)
29
+ PY
30
+ bash scripts/check-docs.sh; echo "planted -> exit=$?" # MUST be non-zero
31
+ cd - && bash scripts/check-docs.sh; echo "clean -> exit=$?" # MUST be 0
32
+ ```
33
+
34
+ **Assert on `$?`, not on a `FAIL` line in the output.** A line on stdout is a
35
+ decoration; the exit code is what CI reads.
36
+
37
+ **Doubt the probe before you doubt the check.** On the project this comes from,
38
+ **four of five** silent probes were the probe's fault: one added a *definition*
39
+ where the check looks for an unresolved *reference*; one edited a string whose
40
+ whitespace did not match; one hit the first prose mention instead of the table row;
41
+ one flipped a row whose cell was already empty, so nothing was planted. A silent
42
+ check is a claim about two things, and the probe is the one to doubt first — prove
43
+ your edit landed in the text the check actually parses.
44
+
45
+ **Three assertions, and the third is the one hand-rolled probes keep missing.** A
46
+ plant that trips some *other* check has proved that other check works. Three probes
47
+ in one day passed that way: one removed 1 of 3 identical lines and left the shape
48
+ intact; one decremented a number inside an **already-released** section; one deleted
49
+ the shouted spelling of a phrase and left the lowercase one. Each landed somewhere
50
+ real and demonstrated nothing about the guard it was written for. So:
51
+
52
+ 1. **the substitution landed** — `replace` matching nothing returns the string
53
+ unchanged and raises nothing;
54
+ 2. **the exit code is non-zero** — not a `FAIL` line on stdout;
55
+ 3. **the message that fired belongs to the guard under test** — named up front, not
56
+ recognised afterwards.
57
+
58
+ `test/probe.py` does all three (`npm run test:probe` self-tests the harness, including
59
+ its own failure branches, because a harness whose failure path has never executed is
60
+ the thing it exists to stop). Declare the plant as `Plant(label, path, old, new,
61
+ expect=<the guard's own words>)` rather than hand-rolling a fourth copy of the loop.
62
+
63
+ **Record the probe.** One line per section, in the change that ships the check.
64
+ Otherwise the next reader has to redo it to know whether it was ever done.
65
+
66
+ ---
67
+
68
+ ## A probe rots, and every way it rots reports green
69
+
70
+ The three assertions above prove a probe works **today**. They say nothing about the
71
+ day after, and a probe is uniquely exposed: the thing it guards is the thing that moves
72
+ it. Four rotted in one release on 2026-08-22, and none of them failed loudly — two
73
+ reported `caught`, two reported nothing at all.
74
+
75
+ **1. The anchor is a literal the guarded thing moves.** A probe pinned to *"the bundle
76
+ is N reference files"* stops landing the day a reference file is added — on the release
77
+ that changes the very number it guards. Same for a version, a count word, a phrase a
78
+ release rewrites. **Derive the anchor**: read whatever the text currently says and make
79
+ *that* wrong. A probe that computes `wrong = stated - 1` never needs maintaining.
80
+
81
+ **2. The precondition is inherited from the tree rather than created.** This is the
82
+ subtle one, because it is triggered by the system working correctly. A probe that
83
+ narrows a declared gap to expose *releases after the newest run stamp* has nothing to
84
+ expose the moment a release writes an honest stamp — the newest release is now the
85
+ newest stamp. A probe requiring an `## Unreleased` section has none the moment a release
86
+ absorbs it. **Both landed. Both proved nothing.** A probe must construct the state it
87
+ needs: remove the stamp, write the section, then plant the defect.
88
+
89
+ **3. The document quotes the form the probe removes.** Covered as
90
+ [`documentation.md`](documentation.md) canon 2's second half, and it belongs here too
91
+ because the probe is where it surfaces: prose describing the wrong shape *with real
92
+ values in it* is a second instance of the shape. The probe deletes the real one, the
93
+ narrative still matches, the guard is silent.
94
+
95
+ **How to see it before CI does.** Ask one question per probe: *what does this probe
96
+ LOOK FOR, and who is allowed to change it?* Where the answer is "the thing it guards",
97
+ it is rotting already. A repository-wide sweep is one grep — a two-digit literal inside
98
+ a needle, an `assert` or a `replace` — and the triage is: the number a probe **writes**
99
+ is correct, the number it **looks for** is the defect.
100
+
101
+ ## The neighbour probe — plant the evidence outside the subject
102
+
103
+ A probe proves a guard rejects **the phrasing its author had in mind**. That is less than
104
+ it looks, and the gap has one shape: **a check answered by text that is not its subject.**
105
+
106
+ Measured on one project in one session, six times, each found by a reader planting a
107
+ defect and watching `PASS`:
108
+
109
+ | The check's subject | The text that answered it instead |
110
+ |---|---|
111
+ | stage 2 names the loop's arming | *"it **arms** the UX track"*, present since an earlier release |
112
+ | the section states the authorization floor | the same phrase in a Rationalizations row, and in a section written twenty-nine releases before |
113
+ | the run stamps have a cap | the **standing instructions'** `max 10`, in the same table cell — and, once the check was narrowed past it, the same cap moved to the other side of the `·` |
114
+ | a section read in full stays capped | rows under **one** heading, while a second heading in the same file held forty more |
115
+
116
+ Every one of those guards had a probe. Every probe fired. The probes and the guards were
117
+ written in the same hour from the same reading, so they shared the same blind spot.
118
+
119
+ **So a guard that reads a scoped span owes a second probe, and it plants in two places at
120
+ once:**
121
+
122
+ 1. **break the subject** — remove the thing the guard is about;
123
+ 2. **plant the guard's own evidence next door** — in a sibling section, an adjacent table
124
+ cell, a rationalizations row, the other side of a separator;
125
+ 3. require the guard to **still fail**.
126
+
127
+ A guard that passes step 3 is reading its subject. A guard that goes green is reading the
128
+ neighbourhood, and the ordinary probe cannot tell the difference — which is why this one
129
+ is separate rather than a stricter version of it.
130
+
131
+ **Step 2 means the literal the predicate matches *today*, read out of the guard.** The
132
+ first three neighbour probes written against this section got that wrong on their first
133
+ run: one planted the needles of two **retired** predicates, which proves the guards that
134
+ used to exist were neighbour-answerable and says nothing about the one that does. A
135
+ neighbour probe keyed to a needle the guard no longer reads is the same defect it was
136
+ written to catch, one level up.
137
+
138
+ **A probe that only deletes is not a neighbour probe.** It may still be a correct probe —
139
+ but if it relies on copies that already sit next door, it must **assert they are there**.
140
+ Otherwise a later edit removes them and the probe quietly becomes a delete-only test that
141
+ still passes, having stopped testing the thing it is named for.
142
+
143
+ **Positional narrowing is not scoping.** Three of the six were "fixed" by cutting the
144
+ search down to a row, then to everything after a phrase, and fell each time to text that
145
+ was still inside the cut. Scope by *what the span is about*: split to the cell, then to
146
+ the item, then match on flattened text so an emphasis marker cannot hide the boundary.
147
+
148
+ **And state the span in the guard.** One line above the predicate — *what it reads, and
149
+ where that ends*. It costs nothing and it is the only part of a check a later reader can
150
+ disagree with before the defect arrives.
151
+
152
+ ## A green probe is evidence only if the mutation is known to have landed
153
+
154
+ Also reported from another project, three times in one day, each caught only because the
155
+ result was too good:
156
+
157
+ 1. a scripted substitution missed on indentation — the file was unchanged and the probe
158
+ measured nothing;
159
+ 2. an assertion written against a bare identifier kept matching the **import line** after
160
+ the field it guarded was deleted;
161
+ 3. a file-extension alternation matched the longer extension as though it were the
162
+ shorter, reporting nine live files as missing.
163
+
164
+ In all three the observable was identical to success. *"See it fail once"* has an unstated
165
+ precondition — **that the thing you changed is the thing the check reads** — and a planted
166
+ defect that did not land produces the same green as a check that cannot fail.
167
+
168
+ **So a probe that mutates an existing file asserts its plant landed, in the same breath as
169
+ planting it.** A probe that writes a whole file has no such question: the file exists or
170
+ the command failed. This repository measured itself while writing this section and got the
171
+ number wrong three times. A hand-rolled classifier said *206 of 206 already carry it*.
172
+ The guard written from the rule said **22 did not** — and was itself too narrow, matching
173
+ one spelling of the assertion, so six probes that already had it in lower case were
174
+ called defective. A sweep then "fixed" those six and **corrupted five**, splitting live
175
+ statements. The true figure was **16**, and it took the guard, a compile check and a
176
+ restore from git to find it.
177
+
178
+ Two things are worth keeping from that. **The check corrected the measurement that
179
+ motivated it** — which is the argument for writing checks rather than counting by hand.
180
+ And **a check keyed to one spelling of a rule is *The neighbour probe*'s class, above**: it
181
+ reported as defective the probes that obeyed the rule in different words. **Every**
182
+ mutating probe carries the assertion now — the figure is deliberately not
183
+ written here. A first draft said *201*, which was true of the branch point and false in
184
+ the same commit, because the twelve probes added for this release are themselves mutating
185
+ probes. The guard computes it; a number in prose beside a check that can count is the
186
+ class this bundle calls restating instead of computing. Probes that write a whole file
187
+ need none: the file exists or the command
188
+ failed.
189
+
190
+ **Prefer an assertion that names the construct over one that names a substring of it.** A
191
+ guard written against a bare identifier survives the deletion of everything it guarded,
192
+ because the identifier still appears in an import. That is case 2 above and it is the same class as
193
+ [`gates.md`](gates.md) → *A ratchet's matcher is itself a check, and it needs a near-miss*, one level down.
194
+
195
+ ## Rationalizations
196
+
197
+ | Excuse | Reality |
198
+ |---|---|
199
+ | "The check is green, that's evidence" | Only if you have seen it red. An unproven check is a decoration that reports success. |
200
+ | "I probed it when I wrote it" | A probe proves the check works today. The thing it guards is the thing that moves it — re-read the anchor on every release that touches the subject. |
201
+ | "The plant obviously landed, the file is smaller" | `replace` matching nothing returns the string unchanged and raises nothing. Assert the mutation landed, in the same breath as planting it. |
202
+ | "Some guard fired, so the probe passed" | A plant that trips some *other* check has proved that other check works. The message that fired must belong to the guard under test, named up front. |
@@ -63,7 +63,7 @@ changes:
63
63
  task-pipeline v1.34.0 · pipeline-audit · module P1 «the progress print» (1 of 4)
64
64
  0 ✓ 1 ✓ 2 ✓ 3 ▶ 4 · 5 · 6 · 7 · 8 · 9 · 10 ·
65
65
  ███████░░░░░░░░░░░░░░░░░░░ gates 3/11 · now 3 Spec · manual
66
- board B-028 · carry-over 0 rows · exposure 99 never · unlooked 0
66
+ board B-028 · carry-over 0 rows · exposure N never · unlooked 0
67
67
  ```
68
68
 
69
69
  Four lines, and each one answers a question an operator otherwise has to ask:
@@ -142,7 +142,8 @@ its pair is task start and iteration close. The hand-back shares one and adds th
142
142
  end — a rail at task start has nothing to report, and a run that ends without a hand-back
143
143
  is the case this section exists for. A reader resolving *"both boundaries"* against the
144
144
  other section wrote one at task start, where TASK is the only field with content. The run
145
- writes a hand-back with **four sections and two lists**. It is a gate criterion **at stage
145
+ writes a hand-back with **the sections and the two counted lists below** — the
146
+ block is the count, and a number restated here said *four* over a block of six. It is a gate criterion **at stage
146
147
  10**, not a good intention: this file already carried one instruction with no gate behind
147
148
  it (*"copy it, tick it"*), and the v1.37.0 audit found no run had ever obeyed it.
148
149
 
@@ -380,7 +381,7 @@ And why the verb prints
380
381
  `unmeasured` rather than `0` when there are no such lines is the same reason again — the
381
382
  hook being absent and the run reading nothing are **opposite facts** the ledger cannot
382
383
  separate, so it claims neither. A `0` there would be the reassuring answer to a question
383
- nobody asked, over 35 files nobody checked.
384
+ nobody asked, over a directory of reference files nobody checked.
384
385
 
385
386
  It is a disclosure: no floor, no direction, never a target. The moment the number becomes
386
387
  something to raise, a run will open files to raise it.
@@ -416,7 +417,10 @@ written by any run — the detector had no input, and the guard was doctrine wea
416
417
  script's clothes. One file serves both readers: the guard reads the `touch:` lines,
417
418
  this block reads the verdict rows and the iteration counter.
418
419
 
419
- Three kinds of line, appended, never rewritten:
420
+ Appended, never rewritten. The full set of line shapes is declared under
421
+ `## Lines` in [`../templates/run.md`](../templates/run.md) — the list is the count,
422
+ and a count restated here drifted once already. The three these two readers
423
+ consume:
420
424
 
421
425
  ```
422
426
  stage: 3 Spec — gate manual — verdict pass — 2026-08-10T14:02Z