task-pipeline-skill 1.41.0 → 1.42.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,77 @@
1
1
  # Changelog
2
2
 
3
+ ## v1.42.0 — a probe proves the phrasing its author had in mind
4
+
5
+ Six times in one session a guard was defeated by **text that was not its subject**, and
6
+ every one of those guards had a probe, and every probe fired. The probes and the guards
7
+ were written in the same hour from the same reading of the same file, so they shared the
8
+ blind spot rather than covering it.
9
+
10
+ | The check's subject | What answered it instead |
11
+ |---|---|
12
+ | stage 2 names the loop's arming | *"it **arms** the UX track"*, present since v1.7.0 |
13
+ | the section states the authorization floor | the same phrase in a Rationalizations row, and in a section from v1.11.0 |
14
+ | the run stamps have a cap | the **standing instructions'** `max 10` in the same cell — and, once narrowed past it, the same cap on the other side of the `·` |
15
+ | a section read in full stays capped | rows under one heading, while a second heading held forty more |
16
+
17
+ ### The doctrine failed its own first use, and a reader caught that too
18
+
19
+ R-005's fourth consecutive reader read the three new probes against the section that
20
+ defines them. **`nb03` was not a neighbour probe.** It planted the needles of two
21
+ *retired* predicates — proving the guards that used to exist were neighbour-answerable
22
+ and saying nothing about the one that does. The instruction *"plant the guard's own
23
+ evidence"* reads as *"plant something that used to satisfy it"* unless it says **which
24
+ literal**, and now it does: the one the predicate matches today, read out of the guard.
25
+
26
+ Seven more, each planted and watched passing, each now replayed and failing:
27
+
28
+ - **Part 1a's precondition could be replaced by its opposite** and a housekeeping aside in
29
+ a `###` subsection answered for it — the reading `Default off` exists to forbid, shipping
30
+ behind a parenthesis;
31
+ - **the stage-2 guard's declared span was false.** Its comment said *"the GATE bullet and
32
+ nothing else"*; the GATE is the last bullet, so the split handed it everything to the
33
+ section end and a trailing note satisfied it;
34
+ - `- **GATEways to stage 3:**` was taken as the gate — no word boundary;
35
+ - **a decoy row about the archive answered for the live file**, because the retrospective
36
+ check scanned every `|` line rather than the row it names;
37
+ - **the scope depended on a `·`** no doctrine requires: replacing it with `, and` restored
38
+ the v1.41.0 defeat by punctuation;
39
+ - **one disjunct was dead and the live one was a literal** — the source line-wraps, so the
40
+ check was a false positive on any reword and a false negative on inversion;
41
+ - and **rekeying the stage-2 guard silently lost coverage**: deleting the whole queue
42
+ bullet began to pass, while its probe kept firing only because it *also* deleted the
43
+ gate clause. A probe whose stated claim has quietly become another probe's claim.
44
+
45
+ **A probe that only deletes is not a neighbour probe** — it may be correct, but if it
46
+ leans on copies already next door it must assert they are there, or a later edit demotes
47
+ it to a delete-only test that still passes.
48
+
49
+ ### The neighbour probe
50
+
51
+ `gates.md` now asks a guard that reads a scoped span for a **second** probe, planting in
52
+ two places at once: break the subject, plant the guard's own evidence **next door**, and
53
+ require it to still fail. A guard that goes green is reading the neighbourhood, and the
54
+ ordinary probe cannot tell the difference — which is why it is a separate probe rather
55
+ than a stricter one.
56
+
57
+ Three shipped with it, one per guard that fell this way. **The third failed on its first
58
+ run and was right to:** a planted sentence saying *"after-decomposition remains a word in
59
+ the schema"* satisfied the stage-2 guard, proving it neighbour-answerable. The guard is
60
+ now keyed on the stage's **gate** — what it must DO — rather than on prose around it,
61
+ because prose above a gate can say anything.
62
+
63
+ **Positional narrowing is not scoping**, and three of the six were "fixed" that way
64
+ before falling again to text still inside the cut.
65
+
66
+ ### Half of B-057 was already mechanised, and the row overstated it
67
+
68
+ `test/negatives.py` runs `differs_from_repo` on every planted copy, so a probe whose
69
+ needle no longer matches fails **loudly** rather than silently — which is what every
70
+ stale literal in this session actually did. The remaining half is the silent one, and it
71
+ is what the neighbour probe is for.
72
+
73
+ - Guards: 250 → **253**.
74
+
3
75
  ## v1.41.0 — one line per run is a slope, not a bound
4
76
 
5
77
  `retro.md` is read **in full** at stage 0, and its own doctrine called both read sections
package/SKILL-CARD.md CHANGED
@@ -12,7 +12,7 @@ harmless.
12
12
  |---|---|
13
13
  | **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
14
14
  | **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
15
- | **Version** | 1.41.0 |
15
+ | **Version** | 1.42.0 |
16
16
  | **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
17
17
  | **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
18
18
  | **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "task-pipeline-skill",
3
- "version": "1.41.0",
3
+ "version": "1.42.0",
4
4
  "description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
5
5
  "bin": {
6
6
  "task-pipeline": "bin/task-pipeline.js"
@@ -2,7 +2,7 @@
2
2
  "name": "task-pipeline",
3
3
  "displayName": "Task Pipeline",
4
4
  "description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
5
- "version": "1.41.0",
5
+ "version": "1.42.0",
6
6
  "author": {
7
7
  "name": "ssheleg",
8
8
  "url": "https://x.com/sshlg93"
@@ -30,6 +30,7 @@ elsewhere and is not restated here:
30
30
  - Anatomy of a project gate
31
31
  - Writing the check itself
32
32
  - Probing — plant, run, restore
33
+ - The neighbour probe — plant the evidence outside the subject
33
34
  - The false-positive budget
34
35
  - Ratchets
35
36
  - Disclosures — counted like a ratchet, and deliberately not monotone
@@ -295,6 +296,57 @@ Otherwise the next reader has to redo it to know whether it was ever done.
295
296
 
296
297
  ---
297
298
 
299
+ ## The neighbour probe — plant the evidence outside the subject
300
+
301
+ A probe proves a guard rejects **the phrasing its author had in mind**. That is less than
302
+ it looks, and the gap has one shape: **a check answered by text that is not its subject.**
303
+
304
+ Measured on one project in one session, six times, each found by a reader planting a
305
+ defect and watching `PASS`:
306
+
307
+ | The check's subject | The text that answered it instead |
308
+ |---|---|
309
+ | stage 2 names the loop's arming | *"it **arms** the UX track"*, present since an earlier release |
310
+ | the section states the authorization floor | the same phrase in a Rationalizations row, and in a section written twenty-nine releases before |
311
+ | the run stamps have a cap | the **standing instructions'** `max 10`, in the same table cell — and, once the check was narrowed past it, the same cap moved to the other side of the `·` |
312
+ | a section read in full stays capped | rows under **one** heading, while a second heading in the same file held forty more |
313
+
314
+ Every one of those guards had a probe. Every probe fired. The probes and the guards were
315
+ written in the same hour from the same reading, so they shared the same blind spot.
316
+
317
+ **So a guard that reads a scoped span owes a second probe, and it plants in two places at
318
+ once:**
319
+
320
+ 1. **break the subject** — remove the thing the guard is about;
321
+ 2. **plant the guard's own evidence next door** — in a sibling section, an adjacent table
322
+ cell, a rationalizations row, the other side of a separator;
323
+ 3. require the guard to **still fail**.
324
+
325
+ A guard that passes step 3 is reading its subject. A guard that goes green is reading the
326
+ neighbourhood, and the ordinary probe cannot tell the difference — which is why this one
327
+ is separate rather than a stricter version of it.
328
+
329
+ **Step 2 means the literal the predicate matches *today*, read out of the guard.** The
330
+ first three neighbour probes written against this section got that wrong on their first
331
+ run: one planted the needles of two **retired** predicates, which proves the guards that
332
+ used to exist were neighbour-answerable and says nothing about the one that does. A
333
+ neighbour probe keyed to a needle the guard no longer reads is the same defect it was
334
+ written to catch, one level up.
335
+
336
+ **A probe that only deletes is not a neighbour probe.** It may still be a correct probe —
337
+ but if it relies on copies that already sit next door, it must **assert they are there**.
338
+ Otherwise a later edit removes them and the probe quietly becomes a delete-only test that
339
+ still passes, having stopped testing the thing it is named for.
340
+
341
+ **Positional narrowing is not scoping.** Three of the six were "fixed" by cutting the
342
+ search down to a row, then to everything after a phrase, and fell each time to text that
343
+ was still inside the cut. Scope by *what the span is about*: split to the cell, then to
344
+ the item, then match on flattened text so an emphasis marker cannot hide the boundary.
345
+
346
+ **And state the span in the guard.** One line above the predicate — *what it reads, and
347
+ where that ends*. It costs nothing and it is the only part of a check a later reader can
348
+ disagree with before the defect arrives.
349
+
298
350
  ## The false-positive budget
299
351
 
300
352
  Run a new heuristic over the **real corpus** before shipping it and count the false