task-pipeline-skill 1.41.0 → 1.42.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,77 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v1.42.0 — a probe proves the phrasing its author had in mind
|
|
4
|
+
|
|
5
|
+
Six times in one session a guard was defeated by **text that was not its subject**, and
|
|
6
|
+
every one of those guards had a probe, and every probe fired. The probes and the guards
|
|
7
|
+
were written in the same hour from the same reading of the same file, so they shared the
|
|
8
|
+
blind spot rather than covering it.
|
|
9
|
+
|
|
10
|
+
| The check's subject | What answered it instead |
|
|
11
|
+
|---|---|
|
|
12
|
+
| stage 2 names the loop's arming | *"it **arms** the UX track"*, present since v1.7.0 |
|
|
13
|
+
| the section states the authorization floor | the same phrase in a Rationalizations row, and in a section from v1.11.0 |
|
|
14
|
+
| the run stamps have a cap | the **standing instructions'** `max 10` in the same cell — and, once narrowed past it, the same cap on the other side of the `·` |
|
|
15
|
+
| a section read in full stays capped | rows under one heading, while a second heading held forty more |
|
|
16
|
+
|
|
17
|
+
### The doctrine failed its own first use, and a reader caught that too
|
|
18
|
+
|
|
19
|
+
R-005's fourth consecutive reader read the three new probes against the section that
|
|
20
|
+
defines them. **`nb03` was not a neighbour probe.** It planted the needles of two
|
|
21
|
+
*retired* predicates — proving the guards that used to exist were neighbour-answerable
|
|
22
|
+
and saying nothing about the one that does. The instruction *"plant the guard's own
|
|
23
|
+
evidence"* reads as *"plant something that used to satisfy it"* unless it says **which
|
|
24
|
+
literal**, and now it does: the one the predicate matches today, read out of the guard.
|
|
25
|
+
|
|
26
|
+
Seven more, each planted and watched passing, each now replayed and failing:
|
|
27
|
+
|
|
28
|
+
- **Part 1a's precondition could be replaced by its opposite** and a housekeeping aside in
|
|
29
|
+
a `###` subsection answered for it — the reading `Default off` exists to forbid, shipping
|
|
30
|
+
behind a parenthesis;
|
|
31
|
+
- **the stage-2 guard's declared span was false.** Its comment said *"the GATE bullet and
|
|
32
|
+
nothing else"*; the GATE is the last bullet, so the split handed it everything to the
|
|
33
|
+
section end and a trailing note satisfied it;
|
|
34
|
+
- `- **GATEways to stage 3:**` was taken as the gate — no word boundary;
|
|
35
|
+
- **a decoy row about the archive answered for the live file**, because the retrospective
|
|
36
|
+
check scanned every `|` line rather than the row it names;
|
|
37
|
+
- **the scope depended on a `·`** no doctrine requires: replacing it with `, and` restored
|
|
38
|
+
the v1.41.0 defeat by punctuation;
|
|
39
|
+
- **one disjunct was dead and the live one was a literal** — the source line-wraps, so the
|
|
40
|
+
check was a false positive on any reword and a false negative on inversion;
|
|
41
|
+
- and **rekeying the stage-2 guard silently lost coverage**: deleting the whole queue
|
|
42
|
+
bullet began to pass, while its probe kept firing only because it *also* deleted the
|
|
43
|
+
gate clause. A probe whose stated claim has quietly become another probe's claim.
|
|
44
|
+
|
|
45
|
+
**A probe that only deletes is not a neighbour probe** — it may be correct, but if it
|
|
46
|
+
leans on copies already next door it must assert they are there, or a later edit demotes
|
|
47
|
+
it to a delete-only test that still passes.
|
|
48
|
+
|
|
49
|
+
### The neighbour probe
|
|
50
|
+
|
|
51
|
+
`gates.md` now asks a guard that reads a scoped span for a **second** probe, planting in
|
|
52
|
+
two places at once: break the subject, plant the guard's own evidence **next door**, and
|
|
53
|
+
require it to still fail. A guard that goes green is reading the neighbourhood, and the
|
|
54
|
+
ordinary probe cannot tell the difference — which is why it is a separate probe rather
|
|
55
|
+
than a stricter one.
|
|
56
|
+
|
|
57
|
+
Three shipped with it, one per guard that fell this way. **The third failed on its first
|
|
58
|
+
run and was right to:** a planted sentence saying *"after-decomposition remains a word in
|
|
59
|
+
the schema"* satisfied the stage-2 guard, proving it neighbour-answerable. The guard is
|
|
60
|
+
now keyed on the stage's **gate** — what it must DO — rather than on prose around it,
|
|
61
|
+
because prose above a gate can say anything.
|
|
62
|
+
|
|
63
|
+
**Positional narrowing is not scoping**, and three of the six were "fixed" that way
|
|
64
|
+
before falling again to text still inside the cut.
|
|
65
|
+
|
|
66
|
+
### Half of B-057 was already mechanised, and the row overstated it
|
|
67
|
+
|
|
68
|
+
`test/negatives.py` runs `differs_from_repo` on every planted copy, so a probe whose
|
|
69
|
+
needle no longer matches fails **loudly** rather than silently — which is what every
|
|
70
|
+
stale literal in this session actually did. The remaining half is the silent one, and it
|
|
71
|
+
is what the neighbour probe is for.
|
|
72
|
+
|
|
73
|
+
- Guards: 250 → **253**.
|
|
74
|
+
|
|
3
75
|
## v1.41.0 — one line per run is a slope, not a bound
|
|
4
76
|
|
|
5
77
|
`retro.md` is read **in full** at stage 0, and its own doctrine called both read sections
|
package/SKILL-CARD.md
CHANGED
|
@@ -12,7 +12,7 @@ harmless.
|
|
|
12
12
|
|---|---|
|
|
13
13
|
| **Purpose** | Runs a substantial task through ten gated delivery stages — intake grill, docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs+registers, acceptance — refusing to advance until each gate passes |
|
|
14
14
|
| **Owner** | ssheleg ([github.com/ssheleg/task-pipeline](https://github.com/ssheleg/task-pipeline)) |
|
|
15
|
-
| **Version** | 1.
|
|
15
|
+
| **Version** | 1.42.0 |
|
|
16
16
|
| **Surface** | Claude Code (filesystem skill + plugin) and the vercel `skills` CLI. **Not** uploaded to the Skills API; custom Skills do not sync across surfaces |
|
|
17
17
|
| **Dependencies** | None required. Optional: `context7` (MCP), `figma` (MCP), super-ux, agent-sync, graphify, obsidian-wiki. Every stage's doctrine ships in-repo; the one conditional requirement is super-ux for the stage-3 UX track on a user-facing task |
|
|
18
18
|
| **Evaluation status** | Suite authored, 5 categories. One recorded run, **self-observed by the author**; **zero blind runs on zero of three models** — the split, and the numbers, live in [`evals/RESULTS.md`](evals/RESULTS.md) and are computed by `evals/run.py` |
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline-skill",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.42.0",
|
|
4
4
|
"description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"task-pipeline": "bin/task-pipeline.js"
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"name": "task-pipeline",
|
|
3
3
|
"displayName": "Task Pipeline",
|
|
4
4
|
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
|
|
5
|
-
"version": "1.
|
|
5
|
+
"version": "1.42.0",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "ssheleg",
|
|
8
8
|
"url": "https://x.com/sshlg93"
|
|
@@ -30,6 +30,7 @@ elsewhere and is not restated here:
|
|
|
30
30
|
- Anatomy of a project gate
|
|
31
31
|
- Writing the check itself
|
|
32
32
|
- Probing — plant, run, restore
|
|
33
|
+
- The neighbour probe — plant the evidence outside the subject
|
|
33
34
|
- The false-positive budget
|
|
34
35
|
- Ratchets
|
|
35
36
|
- Disclosures — counted like a ratchet, and deliberately not monotone
|
|
@@ -295,6 +296,57 @@ Otherwise the next reader has to redo it to know whether it was ever done.
|
|
|
295
296
|
|
|
296
297
|
---
|
|
297
298
|
|
|
299
|
+
## The neighbour probe — plant the evidence outside the subject
|
|
300
|
+
|
|
301
|
+
A probe proves a guard rejects **the phrasing its author had in mind**. That is less than
|
|
302
|
+
it looks, and the gap has one shape: **a check answered by text that is not its subject.**
|
|
303
|
+
|
|
304
|
+
Measured on one project in one session, six times, each found by a reader planting a
|
|
305
|
+
defect and watching `PASS`:
|
|
306
|
+
|
|
307
|
+
| The check's subject | The text that answered it instead |
|
|
308
|
+
|---|---|
|
|
309
|
+
| stage 2 names the loop's arming | *"it **arms** the UX track"*, present since an earlier release |
|
|
310
|
+
| the section states the authorization floor | the same phrase in a Rationalizations row, and in a section written twenty-nine releases before |
|
|
311
|
+
| the run stamps have a cap | the **standing instructions'** `max 10`, in the same table cell — and, once the check was narrowed past it, the same cap moved to the other side of the `·` |
|
|
312
|
+
| a section read in full stays capped | rows under **one** heading, while a second heading in the same file held forty more |
|
|
313
|
+
|
|
314
|
+
Every one of those guards had a probe. Every probe fired. The probes and the guards were
|
|
315
|
+
written in the same hour from the same reading, so they shared the same blind spot.
|
|
316
|
+
|
|
317
|
+
**So a guard that reads a scoped span owes a second probe, and it plants in two places at
|
|
318
|
+
once:**
|
|
319
|
+
|
|
320
|
+
1. **break the subject** — remove the thing the guard is about;
|
|
321
|
+
2. **plant the guard's own evidence next door** — in a sibling section, an adjacent table
|
|
322
|
+
cell, a rationalizations row, the other side of a separator;
|
|
323
|
+
3. require the guard to **still fail**.
|
|
324
|
+
|
|
325
|
+
A guard that passes step 3 is reading its subject. A guard that goes green is reading the
|
|
326
|
+
neighbourhood, and the ordinary probe cannot tell the difference — which is why this one
|
|
327
|
+
is separate rather than a stricter version of it.
|
|
328
|
+
|
|
329
|
+
**Step 2 means the literal the predicate matches *today*, read out of the guard.** The
|
|
330
|
+
first three neighbour probes written against this section got that wrong on their first
|
|
331
|
+
run: one planted the needles of two **retired** predicates, which proves the guards that
|
|
332
|
+
used to exist were neighbour-answerable and says nothing about the one that does. A
|
|
333
|
+
neighbour probe keyed to a needle the guard no longer reads is the same defect it was
|
|
334
|
+
written to catch, one level up.
|
|
335
|
+
|
|
336
|
+
**A probe that only deletes is not a neighbour probe.** It may still be a correct probe —
|
|
337
|
+
but if it relies on copies that already sit next door, it must **assert they are there**.
|
|
338
|
+
Otherwise a later edit removes them and the probe quietly becomes a delete-only test that
|
|
339
|
+
still passes, having stopped testing the thing it is named for.
|
|
340
|
+
|
|
341
|
+
**Positional narrowing is not scoping.** Three of the six were "fixed" by cutting the
|
|
342
|
+
search down to a row, then to everything after a phrase, and fell each time to text that
|
|
343
|
+
was still inside the cut. Scope by *what the span is about*: split to the cell, then to
|
|
344
|
+
the item, then match on flattened text so an emphasis marker cannot hide the boundary.
|
|
345
|
+
|
|
346
|
+
**And state the span in the guard.** One line above the predicate — *what it reads, and
|
|
347
|
+
where that ends*. It costs nothing and it is the only part of a check a later reader can
|
|
348
|
+
disagree with before the defect arrives.
|
|
349
|
+
|
|
298
350
|
## The false-positive budget
|
|
299
351
|
|
|
300
352
|
Run a new heuristic over the **real corpus** before shipping it and count the false
|