task-pipeline-skill 1.33.0 → 1.38.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +332 -0
- package/CONTRIBUTING.md +66 -0
- package/README.md +23 -0
- package/SKILL-CARD.md +1 -1
- package/cursor/rules/task-pipeline.mdc +34 -9
- package/evals/task-pipeline.evals.json +87 -0
- package/package.json +3 -2
- package/plugins/task-pipeline/.claude-plugin/plugin.json +2 -2
- package/plugins/task-pipeline/commands/task-pipeline.md +167 -56
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +6 -5
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +133 -21
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +12 -5
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +32 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/continuity.md +10 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/exposure.md +8 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/gates.md +18 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/knowledge-sources.md +10 -5
- package/plugins/task-pipeline/skills/task-pipeline/references/learned.md +3 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +52 -2
- package/plugins/task-pipeline/skills/task-pipeline/references/portability.md +1 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/progress.md +190 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/retrospective.md +90 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +56 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +123 -12
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +1 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/run.md +77 -0
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline",
|
|
3
3
|
"displayName": "Task Pipeline",
|
|
4
|
-
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that
|
|
5
|
-
"version": "1.
|
|
4
|
+
"description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that closes with evidence, a work board and a verification ledger that outlive a run, an exposure line naming what shipped unconfirmed, a progress rail computed from the project's own config, a loop guard whose review ceiling measures rather than stops, and stage-3 tracks for what a product does, how it sounds and how it looks. Two modes need no task: `checkup` (what is unverified) and `setup` (audit existing docs). Retro insights can publish upstream as issues, opt-in and redacted.",
|
|
5
|
+
"version": "1.38.0",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "ssheleg",
|
|
8
8
|
"url": "https://x.com/sshlg93"
|
|
@@ -1,62 +1,173 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: Run a task through task-pipeline
|
|
3
|
-
argument-hint: "<one-line task
|
|
2
|
+
description: "Run a task through task-pipeline's gated stages (intake grill → docs → brainstorm → spec → plan → build → tests → deploy → post-deploy → docs/wiki → acceptance). Also: `setup` — audit the docs you already have; `checkup` — what has shipped unverified, with no task running."
|
|
3
|
+
argument-hint: "<one-line task> | setup | checkup"
|
|
4
4
|
---
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
built into the skill** (`references/{knowledge-sources,knowledge-graph,grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance,retrospective,loop-guard}.md`)
|
|
9
|
-
— no companion plugin is required for any of them. **Stage 0 opens with the
|
|
10
|
-
knowledge harvest, before the first question** (`references/knowledge-sources.md`):
|
|
11
|
-
pull what the project already knows about this task from the code, **the code graph**
|
|
12
|
-
if one is built ([graphify](https://github.com/Graphify-Labs/graphify) —
|
|
13
|
-
`references/knowledge-graph.md`; recommended, never required; detect
|
|
14
|
-
`graphify-out/graph.json`; it answers *reach* — what calls this, what breaks if it
|
|
15
|
-
moves — which grep cannot), `CLAUDE.md`,
|
|
16
|
-
`CONTEXT.md`/ADRs, `docs/` + `docs/ux/`, past pipeline briefs, **the retro's standing
|
|
17
|
-
instructions** (`docs/superpowers/retro.md` — read in full, they bind this run;
|
|
18
|
-
`references/retrospective.md`), the **knowledge wiki**
|
|
19
|
-
if one is installed ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki) —
|
|
20
|
-
recommended, never required; detect `~/.obsidian-wiki/config`) and any **other repo
|
|
21
|
-
or hosted doc system the project names as its docs**, then write the **source
|
|
22
|
-
ledger** into the brief. The **intake grill is
|
|
23
|
-
mandatory** (`references/grill.md`): interview the
|
|
24
|
-
operator one question at a time (with a recommended answer each, exploring the
|
|
25
|
-
codebase before asking) until every decision branch is resolved, **validating every
|
|
26
|
-
answer against the harvested sources** — the operator outranks any document, but
|
|
27
|
-
only out loud, and a doc the run proves stale is logged for the stage-9 update —
|
|
28
|
-
applying the
|
|
29
|
-
grill's **domain awareness** (challenge terms against `CONTEXT.md`, sharpen fuzzy
|
|
30
|
-
language, ADRs for hard-to-reverse calls) and covering the **autonomy sweep** (what
|
|
31
|
-
would otherwise stop stages 1→10: docs sources incl. doc repos, the wiki and the code graph, branch/tracker
|
|
32
|
-
policy, test and lint commands, deploy target and authorization, log locations, docs/wiki/graph targets, and for UI tasks whether the design is done visually in Figma or text-only, whether the Figma MCP is connected, and — if it isn't — whether to ship text-only or stop and connect it, since the UX chain degrades on its own and never blocks; **and with Figma on, the design destination: which team/org by name and which file** — the recorded one, a URL the operator gives, or creation in that named team explicitly authorized, written into the project's canonical record before the first frame, because a destination decided at drawing time is how a project ends up with three design files and no way to tell which is real. **Never create while a recorded file resolves; if it doesn't resolve, stop and ask — never create a replacement**) —
|
|
33
|
-
until the brief is locked — including the **REQ table**, the request as an addressable list where every row names how it is verified — so the rest runs autonomously and the final stage can account for all of it. The list is frozen: adding is free, removing needs the operator's agreement. Anything deferred goes into the carry-over ledger the moment it's said, and **the board** (`docs/superpowers/backlog.md`, `references/backlog.md`) — the work-list between runs — is read at stage 0 with its open count quoted in the brief, or seeded when absent; **the verification ledger** (`docs/superpowers/verification.md`, `references/verification.md`) is read at stage 0 for how many rows sit at `never`, written at stage 8 with one row per shipped REQ, and required at stage 10 in both directions. For any user-facing task, recommend/use
|
|
34
|
-
**super-ux**. **If the brief describes a platform rather than a change**, stage 2 also cuts it into modules (`references/decomposition.md`) — module map committed, walking skeleton first, every REQ in exactly one module — and stages 3→10 then run per module, one brick at a time. **If any loop starts undoing an earlier pass** (same file edited twice for the same reason, a closed finding returning, a third entry into one stage), stop and run the loop guard (`references/loop-guard.md`): name both shapes, escalate to the layer that owns the conflict, re-plan the check as an ordered list, then go item by item. **The closing stage opens with the ladder walk** (`references/audit.md`): the REQ table finds what was named and lost, but a comparison needs two sides and an absence has one — so walk each REQ bottom-up through its rungs (decision → spec section → contract *and its failure behavior* → task → change → executed test → surface/docs), check the seam at each step, order findings by seam rather than by file, and turn every absence into a new REQ row **before** the coverage table is written. A green from a check nobody has watched fail against a planted defect is not evidence; a finding class seen twice becomes a script rather than a third ledger row; and every ledger row still `open`, `unresolved` or homed `backlog` leaves stage 10 with a board id and the board's priorities are re-derived (`references/backlog.md`); the carry-over ledger's counts are printed beside every gate verdict, so "green" never reads as "verified". If a searching pass starts finding mostly what the previous pass's own fixes broke, the axis is exhausted — rotate it, don't look harder. **The docs stage closes three artifacts, not two:** module docs, the wiki, **and the code graph** (`/graphify . --update` where `graphify-out/` exists — `references/knowledge-graph.md`), because the graph is what the next run's harvest queries first and a stale one is a false premise carrying the authority of a machine. Then check the graph against the docs: a hub `graphify god-nodes` reports that no document names is an undocumented seam; an edge the docs deny is a leak in the code or a lie in the docs; a doc naming a module the graph no longer has is stale. Doc-side findings are fixed there, absences become REQ rows in the closing stage. **In a project of several repositories, stage 10 closes on the parent too:** a parent records each submodule as a pointer to one commit, and moving the submodule does not move the pointer — so the work can be committed, pushed and green while a clone of the parent still gets the commit before it. Neither repo looks wrong alone, which is why it survives every check that runs inside one. Require `git submodule status` with no line starting `+`, and every repo clean and pushed (`git -C <repo> status --porcelain`, `git -C <repo> log @{u}..HEAD`). The fix is two commands and the second gets forgotten: push the submodule, then `git add <submodule> && git commit`. **The run's last act is the retrospective** (`references/retrospective.md` → `docs/superpowers/retro.md`, one file per project): **stamp the run first** (its commit is what makes the cold-retirement trigger computable), **then prune** — every standing instruction against its three retirement triggers (it became a check; the paths/commands it names are gone; it hasn't fired in five run stamps, or in sixty days — the calendar is the unit that still moves when the stamp counter has stopped), the list held to a hard cap of ten, every deletion logged as one line and never silent — then, only if the run diverged, write the entry: symptom with evidence, the stage it surfaced at, the stage that *owned* it, the root cause, the fix by grade (mechanical check > standing instruction > note that expires in two runs), and the check that catches it next time. Stage 0 reads those standing instructions in full, so the prune is a gate criterion, not a good intention: a rule nobody reads to the end is worse than no rule, because everyone believes it is covered. Honor every stage gate by its type (`auto` = verify yourself;
|
|
35
|
-
`manual` = wait for explicit go). Confirm the **model once at preflight** —
|
|
36
|
-
recommend the most capable one the environment offers, never a hardcoded id — then
|
|
37
|
-
run the whole pipeline on it without re-asking.
|
|
5
|
+
Run the task below through `task-pipeline`'s gated stages. **Every stage's doctrine
|
|
6
|
+
ships inside the skill** — no companion plugin is required for any of it. The
|
|
7
|
+
reference files are indexed in `SKILL.md`; this page is the run order, not the index.
|
|
38
8
|
|
|
39
9
|
Task: $ARGUMENTS
|
|
40
10
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
11
|
+
> **Two other modes.** `setup` audits the documentation this project already has;
|
|
12
|
+
> `checkup` reports what has shipped without a person confirming it, with no task in
|
|
13
|
+
> flight. Both are at the bottom of this page.
|
|
14
|
+
|
|
15
|
+
**Idempotent entry — inspect state first, never restart blindly.** If a pipeline
|
|
16
|
+
TaskList from a previous run exists for this task, **resume** from the first incomplete
|
|
17
|
+
stage. Otherwise begin at stage 0. With no task given, the grill's first question asks
|
|
18
|
+
for it in one line.
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Preflight — settle these once
|
|
23
|
+
|
|
24
|
+
- **Model.** Recommend the most capable one the environment offers, never a hardcoded
|
|
25
|
+
id. Confirmed once, then the whole run uses it without re-asking.
|
|
26
|
+
- **Companions.** Print the detection block once (`references/companion-skills.md`).
|
|
27
|
+
Absent ones state their fallback; none is a gate except the stage-3 UX track on a
|
|
28
|
+
user-facing task.
|
|
29
|
+
|
|
30
|
+
## Stage 0 — the harvest, then the grill
|
|
31
|
+
|
|
32
|
+
**The harvest runs before the first question** (`references/knowledge-sources.md`).
|
|
33
|
+
Pull what the project already knows about *this task*:
|
|
34
|
+
|
|
35
|
+
- the code, and **the code graph** where one is built — [graphify](https://github.com/Graphify-Labs/graphify),
|
|
36
|
+
`references/knowledge-graph.md`; recommended, never required; detect
|
|
37
|
+
`graphify-out/graph.json`. It answers **reach** — what calls this, what breaks if it
|
|
38
|
+
moves — which grep cannot.
|
|
39
|
+
- `CLAUDE.md`, `CONTEXT.md`/ADRs, `docs/` + `docs/ux/`, past briefs and carry-over ledgers.
|
|
40
|
+
- **the retro's standing instructions and run stamps** — `docs/superpowers/retro.md`,
|
|
41
|
+
read in full; both are bounded and they bind this run. Its **Recent log** is
|
|
42
|
+
*queried* by the task's nouns, not read: nothing caps it, and an uncapped section
|
|
43
|
+
inside a binding source is what makes the capped part get skimmed
|
|
44
|
+
(`references/retrospective.md`).
|
|
45
|
+
- the **knowledge wiki** if installed ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki);
|
|
46
|
+
detect `~/.obsidian-wiki/config`), and any other doc system the project names as its docs.
|
|
47
|
+
- **the board** (`docs/superpowers/backlog.md`) — open count quoted in the brief, or
|
|
48
|
+
seeded when absent. **the verification ledger** (`docs/superpowers/verification.md`) —
|
|
49
|
+
how many rows sit at `never`.
|
|
50
|
+
|
|
51
|
+
Write the **source ledger** into the brief: a row per source, or an explicit *none found*.
|
|
52
|
+
|
|
53
|
+
**Then the grill, and it is mandatory** (`references/grill.md`). One question per turn,
|
|
54
|
+
each with a recommended answer, exploring the codebase before asking, until every
|
|
55
|
+
decision branch is resolved.
|
|
56
|
+
|
|
57
|
+
- **Validate every answer against the harvested sources.** The operator outranks any
|
|
58
|
+
document — but only out loud, and a doc the run proves stale is logged for stage 9.
|
|
59
|
+
- **Domain awareness:** challenge terms against `CONTEXT.md`, sharpen fuzzy language,
|
|
60
|
+
write an ADR for a hard-to-reverse call.
|
|
61
|
+
- **The autonomy sweep** pre-resolves what would otherwise stop stages 1→10: doc
|
|
62
|
+
sources, wiki and graph, branch and tracker policy, test and lint commands, deploy
|
|
63
|
+
target and authorization, log locations.
|
|
64
|
+
- **UI tasks add the design surface.** Is the design done visually in Figma or
|
|
65
|
+
text-only? Is the Figma MCP connected? If not — ship text-only, or stop and connect
|
|
66
|
+
it? The UX chain degrades on its own and never blocks, so this choice must be
|
|
67
|
+
recorded rather than discovered.
|
|
68
|
+
- **With Figma on, the destination is named before the first frame:** which team/org,
|
|
69
|
+
which file — the recorded one, a URL the operator gives, or creation in that named
|
|
70
|
+
team explicitly authorized. A destination decided at drawing time is how a project
|
|
71
|
+
ends up with three design files and no way to tell which is real. **Never create
|
|
72
|
+
while a recorded file resolves; if it does not resolve, stop and ask — never create a
|
|
73
|
+
replacement.**
|
|
74
|
+
|
|
75
|
+
**The brief closes on the REQ table** — the request as an addressable list where every
|
|
76
|
+
row names how it is verified. Frozen: adding is free, removing needs the operator.
|
|
77
|
+
Anything deferred enters the carry-over ledger the moment it is said.
|
|
78
|
+
|
|
79
|
+
## Stages 1→10 — the flow
|
|
80
|
+
|
|
81
|
+
| | Stage | The thing that must be true to leave it |
|
|
82
|
+
|---|---|---|
|
|
83
|
+
| 1 | Docs study | every contract the design will lock is grounded on fetched docs, not recall |
|
|
84
|
+
| 2 | Brainstorm + decompose | the design is approved and every REQ is answered by it. **A platform is cut into modules** (`references/decomposition.md`) — map committed, walking skeleton first, every REQ in exactly one module; stages 3→10 then run per module |
|
|
85
|
+
| 3 | Spec | contracts locked; user-facing work runs three tracks — what it **does** (super-ux), how it **sounds** (`copywriting`), how it **looks** (`sheleg-design`); a declined track is recorded, never silent |
|
|
86
|
+
| 4 | Plan | the REQ set-comparison holds: brief REQs == union of `Implements:` |
|
|
87
|
+
| 5 | Build | TDD per task, a review after each, findings fixed or parked with a ruling |
|
|
88
|
+
| 6 | Tests | the **full** suite green; a web surface checked in a browser, not in the diff |
|
|
89
|
+
| 7 | Lint + deploy | outward: the authorization is specific, and the CI verdict is **read** before any tag |
|
|
90
|
+
| 8 | Post-deploy | the verification trio, not one of three; a verification row per shipped REQ |
|
|
91
|
+
| 9 | Docs + wiki | **three artifacts, not two** — module docs, the wiki, **and the code graph** |
|
|
92
|
+
| 10 | Acceptance | the ladder walk first, then the table, then the retrospective |
|
|
93
|
+
|
|
94
|
+
**Honor every gate by its type**: `auto` — verify the check yourself; `manual` — wait
|
|
95
|
+
for an explicit go.
|
|
96
|
+
|
|
97
|
+
## Cross-cutting — the three that fire at any stage
|
|
98
|
+
|
|
99
|
+
**The loop guard** (`references/loop-guard.md`). If a pass starts undoing an earlier one
|
|
100
|
+
— the same file edited twice for the same reason, a closed finding returning, a third
|
|
101
|
+
entry into one stage — stop editing. Name both shapes, escalate to the layer that owns
|
|
102
|
+
the conflict, re-plan the check as an ordered list, then go item by item. The review
|
|
103
|
+
loop has its own ceiling, and at it the run **measures** rather than stops.
|
|
104
|
+
|
|
105
|
+
**The audit's exit** (`references/audit.md`). If a searching pass starts finding mostly
|
|
106
|
+
what the previous pass's own fixes broke, the axis is exhausted — rotate it, do not look
|
|
107
|
+
harder.
|
|
108
|
+
|
|
109
|
+
**Evidence.** A green from a check nobody has watched fail against a planted defect is
|
|
110
|
+
not evidence. A finding class seen twice becomes a script, not a third ledger row.
|
|
111
|
+
|
|
112
|
+
## Stage 9 — the graph is the third artifact
|
|
113
|
+
|
|
114
|
+
Refresh it (`/graphify . --update` where `graphify-out/` exists), then **check it against
|
|
115
|
+
the docs**: a hub `graphify god-nodes` reports that no document names is an undocumented
|
|
116
|
+
seam; an edge the docs deny is a leak in the code or a lie in the docs; a doc naming a
|
|
117
|
+
module the graph no longer has is stale. Doc-side findings are fixed here; absences
|
|
118
|
+
become REQ rows at stage 10.
|
|
119
|
+
|
|
120
|
+
A stale graph is a false premise **carrying the authority of a machine** — a wrong doc
|
|
121
|
+
gets argued with, a wrong graph gets believed.
|
|
122
|
+
|
|
123
|
+
## Stage 10 — the close-out, in order
|
|
124
|
+
|
|
125
|
+
1. **The ladder walk, first.** The REQ table finds what was named and lost; it cannot
|
|
126
|
+
find what was never named, because a comparison needs two sides and an absence has
|
|
127
|
+
one. Walk each REQ bottom-up — decision → spec section → contract *and its failure
|
|
128
|
+
behavior* → task → change → executed test → surface/docs — check the seam at each
|
|
129
|
+
step, and order findings **by seam, not by file**. Every absence becomes a new REQ
|
|
130
|
+
row **before** the coverage table is written.
|
|
131
|
+
2. **The table**, one row per REQ, each with evidence.
|
|
132
|
+
3. **The ledgers close.** Every carry-over row still `open`, `unresolved` or homed
|
|
133
|
+
`backlog` leaves with a board id, and the board's priorities are re-derived. The
|
|
134
|
+
counts print beside every gate verdict, so *green* never reads as *verified*.
|
|
135
|
+
4. **Several repositories? The parent closes too.** A parent records each submodule as a
|
|
136
|
+
pointer to one commit, and moving the submodule does not move the pointer — so work
|
|
137
|
+
can be committed, pushed and green while a clone of the parent still gets the commit
|
|
138
|
+
before it. Neither repo looks wrong alone. Require `git submodule status` with no
|
|
139
|
+
line starting `+`, and every repo clean and pushed. The fix is two commands and the
|
|
140
|
+
second gets forgotten: push the submodule, **then** `git add <submodule> && git commit`.
|
|
141
|
+
5. **The retrospective is the run's last act** (`references/retrospective.md`), in this
|
|
142
|
+
order: **stamp the run first** (its commit makes the cold trigger computable) → **prune**
|
|
143
|
+
every standing instruction against its three retirement triggers, list held to ten,
|
|
144
|
+
every deletion logged → **write an entry only if the run diverged**: symptom with
|
|
145
|
+
evidence, the stage it surfaced at, the stage that *owned* it, the root cause, the
|
|
146
|
+
fix by grade, and the check that catches it next time.
|
|
147
|
+
|
|
148
|
+
Stage 0 reads those standing instructions in full next time, which is why the prune is a
|
|
149
|
+
gate criterion rather than a good intention.
|
|
150
|
+
|
|
151
|
+
---
|
|
152
|
+
|
|
153
|
+
## `/task-pipeline checkup`
|
|
154
|
+
|
|
155
|
+
**Runs with no task in flight, and that is the point.** Accumulated unconfirmed work is
|
|
156
|
+
invisible precisely because nobody is running a pipeline, so a check living only inside a
|
|
157
|
+
run can never say *"stop, fourteen things are unconfirmed."*
|
|
158
|
+
|
|
159
|
+
It takes no brief, opens no grill, and writes nothing on its own. Four sections, each
|
|
160
|
+
read from a file this pipeline already keeps: the **exposure** line with its check-list
|
|
161
|
+
oldest-first, the **board**'s open rows by computed priority, the carry-over ledgers'
|
|
162
|
+
unresolved count, and the code graph's staleness where one exists.
|
|
163
|
+
|
|
54
164
|
Where you ask it to file what it found, it appends board rows whose `Source` names the
|
|
55
|
-
checkup and its date — printing what it would add first, never silently.
|
|
56
|
-
`references/exposure.md`.
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
165
|
+
checkup and its date — printing what it would add first, never silently.
|
|
166
|
+
Doctrine: `references/exposure.md`.
|
|
167
|
+
|
|
168
|
+
## `/task-pipeline setup`
|
|
169
|
+
|
|
170
|
+
**The entry audit instead of a feature.** Seven passes over the documentation this
|
|
171
|
+
project already has, findings reported as `file:line` + the minimal fix ordered by seam,
|
|
172
|
+
and a fix plan the pipeline can run. Offered once at stage 0 when the doc map is absent
|
|
173
|
+
or stale; run it directly any time. Doctrine: `references/setup.md`.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: task-pipeline
|
|
3
|
-
description: "Runs a substantial task through a full delivery pipeline: an intake grill that expands the request into a locked brief, then docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy check, docs/wiki sync and acceptance — gated stages whose doctrine ships inside this skill (no required companions). Use when work changes the repository — a feature, fix, refactor, migration, integration, rewrite, adoption or hardening; фича, фикс, рефактор, миграция, интеграция, доработать, починить, внедрить — or on 'run this through the pipeline' / 'прогони по конвейеру', 'the full cycle' / 'полный цикл', /task-pipeline.
|
|
3
|
+
description: "Runs a substantial task through a full delivery pipeline: an intake grill that expands the request into a locked brief, then docs study, brainstorm, spec, plan, subagent build, tests, lint/deploy, post-deploy check, docs/wiki sync and acceptance — gated stages whose doctrine ships inside this skill (no required companions). Use when work changes the repository — a feature, fix, refactor, migration, integration, rewrite, adoption or hardening; фича, фикс, рефактор, миграция, интеграция, доработать, починить, внедрить — or on 'run this through the pipeline' / 'прогони по конвейеру', 'the full cycle' / 'полный цикл', /task-pipeline. Two modes need no task at all: 'checkup' / 'чекап' reports what has shipped without a person confirming it and what to look at first; 'setup' audits the documentation a project already has. Not for: answering a question, explaining or reading code, a typo or a one-line edit — say 'без пайплайна' / 'quick' to opt out."
|
|
4
4
|
license: MIT
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -67,6 +67,7 @@ gate stops until it is installed.
|
|
|
67
67
|
| **first run · the entry audit** (offered once) | [`references/setup.md`](references/setup.md) |
|
|
68
68
|
| **what travels with the bundle vs stays in a project** | [`references/portability.md`](references/portability.md) |
|
|
69
69
|
| any repeating loop | [`references/loop-guard.md`](references/loop-guard.md) |
|
|
70
|
+
| run-wide · what the run **prints about itself** — the rail, the iteration line | [`references/progress.md`](references/progress.md) |
|
|
70
71
|
| run-wide · how a run keeps going (the loop mode + the context budget) | [`references/continuity.md`](references/continuity.md) |
|
|
71
72
|
| run-wide · the work-list **between** runs, and how its priority is computed | [`references/backlog.md`](references/backlog.md) |
|
|
72
73
|
| run-wide · whether a **human** ever confirmed what shipped, and when | [`references/verification.md`](references/verification.md) |
|
|
@@ -118,10 +119,10 @@ question: pull what the project already knows about this task from the code, the
|
|
|
118
119
|
([`references/knowledge-graph.md`](references/knowledge-graph.md) — graphify;
|
|
119
120
|
recommended, never required),
|
|
120
121
|
`CLAUDE.md`, `CONTEXT.md`/ADRs, **the decision register**, `docs/` + `docs/ux/`,
|
|
121
|
-
past pipeline briefs, **the retro's standing instructions
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
**queried** by the task's nouns
|
|
122
|
+
past pipeline briefs, **the retro's standing instructions and run stamps** —
|
|
123
|
+
`docs/superpowers/retro.md`, read in full because they *bind* this run and are
|
|
124
|
+
bounded by construction (ten rows; one line per run), while its **recent log** and the
|
|
125
|
+
archive under `docs/superpowers/retro/` are **queried** by the task's nouns
|
|
125
126
|
([`references/retrospective.md`](references/retrospective.md)) —
|
|
126
127
|
the **knowledge wiki** if one is installed
|
|
127
128
|
([obsidian-wiki](https://github.com/ar9av/obsidian-wiki) — recommended, never
|
|
@@ -4,18 +4,32 @@
|
|
|
4
4
|
"title": "task-pipeline config",
|
|
5
5
|
"description": "Generic contract for a pipeline config. An ordered list of stages; each stage is run by the host project's own skills/agents and guarded by a typed gate. The framework imposes no specific stages, skills, or gate assignments — those are entirely the host project's config. Copy pipeline.example.json and rewrite it to match your project.",
|
|
6
6
|
"type": "object",
|
|
7
|
-
"required": [
|
|
7
|
+
"required": [
|
|
8
|
+
"stages"
|
|
9
|
+
],
|
|
8
10
|
"additionalProperties": true,
|
|
9
11
|
"properties": {
|
|
10
|
-
"version": {
|
|
12
|
+
"version": {
|
|
13
|
+
"type": "integer",
|
|
14
|
+
"minimum": 1
|
|
15
|
+
},
|
|
11
16
|
"stages": {
|
|
12
17
|
"type": "array",
|
|
13
18
|
"minItems": 1,
|
|
14
19
|
"description": "The pipeline stages, in order. Any number, any names — your project's real stages.",
|
|
15
|
-
"items": {
|
|
20
|
+
"items": {
|
|
21
|
+
"$ref": "#/definitions/stage"
|
|
22
|
+
}
|
|
23
|
+
},
|
|
24
|
+
"release": {
|
|
25
|
+
"$ref": "#/definitions/release"
|
|
26
|
+
},
|
|
27
|
+
"run": {
|
|
28
|
+
"$ref": "#/definitions/run"
|
|
16
29
|
},
|
|
17
|
-
"
|
|
18
|
-
|
|
30
|
+
"retro": {
|
|
31
|
+
"$ref": "#/definitions/retro"
|
|
32
|
+
}
|
|
19
33
|
},
|
|
20
34
|
"definitions": {
|
|
21
35
|
"run": {
|
|
@@ -25,12 +39,17 @@
|
|
|
25
39
|
"properties": {
|
|
26
40
|
"loop": {
|
|
27
41
|
"type": "object",
|
|
28
|
-
"required": [
|
|
42
|
+
"required": [
|
|
43
|
+
"mode"
|
|
44
|
+
],
|
|
29
45
|
"additionalProperties": true,
|
|
30
46
|
"description": "Whether the run advances item by item without a discretionary check-in. It NEVER collapses a manual gate or an outward action — a generic flag is not a specific authorization.",
|
|
31
47
|
"properties": {
|
|
32
48
|
"mode": {
|
|
33
|
-
"enum": [
|
|
49
|
+
"enum": [
|
|
50
|
+
"off",
|
|
51
|
+
"interval"
|
|
52
|
+
],
|
|
34
53
|
"description": "off (the default when absent) = the run pauses between items as it always did. interval = the run is armed with the harness's own loop primitive and advances one item per fire, stopping only at a manual gate, an unresolvable block, a genuine ambiguity, or completion."
|
|
35
54
|
},
|
|
36
55
|
"interval": {
|
|
@@ -43,6 +62,18 @@
|
|
|
43
62
|
"description": "How this harness arms it, e.g. '/loop'. Harness-specific and therefore project-recorded rather than assumed: on a harness with no loop primitive, omit it — the mode then degrades to prose discipline plus the build ledger, and the run says so instead of implying it is armed."
|
|
44
63
|
}
|
|
45
64
|
}
|
|
65
|
+
},
|
|
66
|
+
"review": {
|
|
67
|
+
"type": "object",
|
|
68
|
+
"additionalProperties": true,
|
|
69
|
+
"description": "The review loop's ceiling. Absent, the default is 3 rounds per artifact. It is a DECISION POINT rather than a stop: at the cap the run prints new findings versus findings caused by its own fixes, per round, and either the pair ends the loop or the operator continues it out loud. A flat stop would be wrong — this repository's ten-round runs were still finding real defects on round nine. Doctrine: references/loop-guard.md.",
|
|
70
|
+
"properties": {
|
|
71
|
+
"maxRounds": {
|
|
72
|
+
"type": "integer",
|
|
73
|
+
"minimum": 1,
|
|
74
|
+
"description": "Rounds per artifact before the run stops reviewing and measures. Counted from the run ledger's `touch:` pass numbers, never from memory; a round that finds nothing ends the loop by definition and is not counted."
|
|
75
|
+
}
|
|
76
|
+
}
|
|
46
77
|
}
|
|
47
78
|
}
|
|
48
79
|
},
|
|
@@ -50,48 +81,129 @@
|
|
|
50
81
|
"type": "object",
|
|
51
82
|
"additionalProperties": true,
|
|
52
83
|
"description": "Optional release automation, entirely project-defined and INDIVIDUALLY TOGGLEABLE. Omit the whole object, or set enabled:false, to turn release automation off for this project. The framework ships an example workflow (.github/workflows/release.yml) whose job is gated on a repo variable so each project arms it on its own; this block is the declarative counterpart the orchestrator reads.",
|
|
53
|
-
"required": [
|
|
84
|
+
"required": [
|
|
85
|
+
"enabled"
|
|
86
|
+
],
|
|
54
87
|
"properties": {
|
|
55
|
-
"enabled": {
|
|
56
|
-
|
|
88
|
+
"enabled": {
|
|
89
|
+
"type": "boolean",
|
|
90
|
+
"description": "Master on/off toggle. false (or the object omitted) = no release automation for this project."
|
|
91
|
+
},
|
|
92
|
+
"trigger": {
|
|
93
|
+
"enum": [
|
|
94
|
+
"tag",
|
|
95
|
+
"manual",
|
|
96
|
+
"push",
|
|
97
|
+
"none"
|
|
98
|
+
],
|
|
99
|
+
"description": "What kicks off a release (e.g. a pushed vX.Y.Z tag, manual dispatch)."
|
|
100
|
+
},
|
|
57
101
|
"steps": {
|
|
58
102
|
"type": "array",
|
|
59
|
-
"items": {
|
|
103
|
+
"items": {
|
|
104
|
+
"type": "string",
|
|
105
|
+
"minLength": 1
|
|
106
|
+
},
|
|
60
107
|
"description": "Ordered release actions, in prose — project-specific (e.g. create a GitHub release, npm publish). Human-only steps (2FA publish) are named as such."
|
|
61
108
|
},
|
|
62
109
|
"verify": {
|
|
63
110
|
"type": "array",
|
|
64
|
-
"items": {
|
|
111
|
+
"items": {
|
|
112
|
+
"type": "string",
|
|
113
|
+
"minLength": 1
|
|
114
|
+
},
|
|
65
115
|
"description": "Post-release smoke checks that must pass after a release — this is the release's own post-deploy gate (stage 8 applied to shipping the package itself)."
|
|
66
116
|
}
|
|
67
117
|
}
|
|
68
118
|
},
|
|
69
119
|
"stage": {
|
|
70
120
|
"type": "object",
|
|
71
|
-
"required": [
|
|
121
|
+
"required": [
|
|
122
|
+
"state",
|
|
123
|
+
"skills",
|
|
124
|
+
"gate"
|
|
125
|
+
],
|
|
72
126
|
"additionalProperties": true,
|
|
73
127
|
"properties": {
|
|
74
|
-
"id": {
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
128
|
+
"id": {
|
|
129
|
+
"type": "integer",
|
|
130
|
+
"description": "Optional ordinal."
|
|
131
|
+
},
|
|
132
|
+
"state": {
|
|
133
|
+
"type": "string",
|
|
134
|
+
"minLength": 1,
|
|
135
|
+
"description": "Unique stable key for the stage."
|
|
136
|
+
},
|
|
137
|
+
"name": {
|
|
138
|
+
"type": "string",
|
|
139
|
+
"description": "Optional human label."
|
|
140
|
+
},
|
|
141
|
+
"model": {
|
|
142
|
+
"type": "string",
|
|
143
|
+
"description": "Optional model for the stage. Prefer a provider-agnostic token over a vendor id, which goes stale as generations ship and may not exist on the operator's provider at all: 'default' = the model confirmed for this run (recommended: the most capable reasoning model the environment offers), 'inherit' = whatever the operator is currently on. A literal id is allowed but treated as an example, not a contract."
|
|
144
|
+
},
|
|
78
145
|
"skills": {
|
|
79
146
|
"type": "array",
|
|
80
147
|
"minItems": 1,
|
|
81
148
|
"description": "The skill(s)/agent(s) that execute this stage. Any names your environment resolves — this is where you plug in your OWN skills.",
|
|
82
|
-
"items": {
|
|
149
|
+
"items": {
|
|
150
|
+
"type": "string",
|
|
151
|
+
"minLength": 1
|
|
152
|
+
}
|
|
83
153
|
},
|
|
84
154
|
"gate": {
|
|
85
155
|
"type": "object",
|
|
86
|
-
"required": [
|
|
156
|
+
"required": [
|
|
157
|
+
"type",
|
|
158
|
+
"check"
|
|
159
|
+
],
|
|
87
160
|
"additionalProperties": true,
|
|
88
161
|
"description": "Condition that must pass before advancing to the next stage.",
|
|
89
162
|
"properties": {
|
|
90
163
|
"type": {
|
|
91
|
-
"enum": [
|
|
164
|
+
"enum": [
|
|
165
|
+
"auto",
|
|
166
|
+
"manual"
|
|
167
|
+
],
|
|
92
168
|
"description": "auto = the orchestrator verifies `check` itself (pass/fail); manual = wait for the operator's explicit go."
|
|
93
169
|
},
|
|
94
|
-
"check": {
|
|
170
|
+
"check": {
|
|
171
|
+
"type": "string",
|
|
172
|
+
"minLength": 1,
|
|
173
|
+
"description": "The gate condition, in prose."
|
|
174
|
+
}
|
|
175
|
+
}
|
|
176
|
+
}
|
|
177
|
+
}
|
|
178
|
+
},
|
|
179
|
+
"retro": {
|
|
180
|
+
"type": "object",
|
|
181
|
+
"additionalProperties": true,
|
|
182
|
+
"description": "Optional. Governs what the retrospective does BEYOND writing to the project's own docs/superpowers/retro.md, which always happens. Omit it and nothing leaves the repository: silence arms nothing, exactly as it authorises no deploy.",
|
|
183
|
+
"properties": {
|
|
184
|
+
"publish": {
|
|
185
|
+
"type": "object",
|
|
186
|
+
"required": [
|
|
187
|
+
"repo"
|
|
188
|
+
],
|
|
189
|
+
"additionalProperties": true,
|
|
190
|
+
"description": "Publish skill-level insights as issues on the skill's own repository, so a defect in the pipeline is fixed once instead of rediscovered in every project. OFF unless this object is present — opening an issue in another repository is an outward act, and an outward act taken from a generic flag is one nobody authorized. The body is printed in full before it is sent, and the printed string and the sent string are the same string. Doctrine, including the five redaction rules: references/retrospective.md.",
|
|
191
|
+
"properties": {
|
|
192
|
+
"repo": {
|
|
193
|
+
"type": "string",
|
|
194
|
+
"pattern": "^[^/\\s]+/[^/\\s]+$",
|
|
195
|
+
"description": "owner/name of the SKILL's repository, never the host project's."
|
|
196
|
+
},
|
|
197
|
+
"label": {
|
|
198
|
+
"type": "string",
|
|
199
|
+
"description": "Applied to every issue opened this way, so the operator can read them as a stream rather than find them by accident."
|
|
200
|
+
},
|
|
201
|
+
"redact": {
|
|
202
|
+
"enum": [
|
|
203
|
+
"strict"
|
|
204
|
+
],
|
|
205
|
+
"description": "Only 'strict' exists, and it is not a level among others: the five rules are the contract. A field with one legal value is here so a reader asking 'what redaction applies?' finds the answer in the config rather than assuming none does."
|
|
206
|
+
}
|
|
95
207
|
}
|
|
96
208
|
}
|
|
97
209
|
}
|
|
@@ -59,9 +59,15 @@ at a glance.
|
|
|
59
59
|
> external skill**. A host project may relocate the root via its `CLAUDE.md`; keep
|
|
60
60
|
> the shape, keep the slugs.
|
|
61
61
|
|
|
62
|
-
|
|
63
|
-
stage
|
|
64
|
-
|
|
62
|
+
**Every** run keeps a **git-ignored** run ledger at `.task-pipeline/run.md`, seeded at
|
|
63
|
+
stage 0 from [`../templates/run.md`](../templates/run.md). Three line shapes: a
|
|
64
|
+
`stage:` verdict when a gate returns, an `iter:` line when an iteration closes, and a
|
|
65
|
+
`touch:` line per file per repeating pass. Two readers depend on it —
|
|
66
|
+
[`loop-guard.md`](loop-guard.md) detects churn from the `touch:` lines after a lost
|
|
67
|
+
context, and [`progress.md`](progress.md) derives the stage rail and the iteration
|
|
68
|
+
counter from the other two. It was described as loop-only until 2026-08-10 and, in
|
|
69
|
+
practice, written by no run at all: the guard that calls its own detection *mechanical*
|
|
70
|
+
had no input in any run to date.
|
|
65
71
|
|
|
66
72
|
Stage 5 also creates a **git-ignored** scratch workspace per plan at
|
|
67
73
|
`.task-pipeline/build/<plan-basename>/` — ledger, task briefs, implementer reports,
|
|
@@ -158,12 +164,13 @@ plugins/task-pipeline/
|
|
|
158
164
|
acceptance.md retrospective.md # stage 10 (close-out, then the retro)
|
|
159
165
|
audit.md # cross-cutting: the ladder + seams
|
|
160
166
|
loop-guard.md # cross-cutting: churn detection
|
|
167
|
+
progress.md # cross-cutting: what the run prints about itself
|
|
161
168
|
stages.md model-tiering.md # gates, model policy
|
|
162
169
|
conventions.md artifacts.md # host conventions, this layout
|
|
163
170
|
companion-skills.md # optional companions + preflight
|
|
164
171
|
templates/ # skeletons seeded into a host project
|
|
165
|
-
hygiene.sh
|
|
166
|
-
README.md
|
|
172
|
+
hygiene.sh docgate.sh # -> scripts/check-hygiene.sh, check-docs.sh
|
|
173
|
+
README.md # the index — and it is the list, not a copy of it
|
|
167
174
|
cursor/rules/task-pipeline.mdc # Cursor channel (self-contained rule)
|
|
168
175
|
bin/task-pipeline.js # npx installer (package task-pipeline-skill)
|
|
169
176
|
install.sh # POSIX installer
|
|
@@ -43,7 +43,8 @@ better, plus one that is required only for user-facing work.
|
|
|
43
43
|
|
|
44
44
|
| Skill / tool | Needed for | Required? | Install |
|
|
45
45
|
|---|---|---|---|
|
|
46
|
-
| **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint`) | stage 3 UX track | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
|
|
46
|
+
| **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint` — **and the copy half**: `copywriting`, `brand-voice`, `/brand-init`, `/copy`, `/brand-lint`, plus `/vision`) | stage 3 — the **UX track** *and* the **COPY track**. This row named six surfaces until 2026-08-10 while super-ux shipped eight skills and fifteen commands: the whole brand-and-copy half was invisible to this pipeline, so a run built scenarios and screens and then wrote the interface strings by taste | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
|
|
47
|
+
| **sheleg-design** (`/sheleg-design`) | stage 3 — the **VISUAL track**: tokens and themes, typography and rhythm, motion and how it degrades to rest, the visual language a brand is recognised by. It answers *how it looks*, which no other companion here answers — `super-ux` decides what the interface must do, `copywriting` how it sounds. Before 2026-08-10 this skill appeared once in the whole bundle, as a name in a list | **Recommended** on any task with a visual surface; never a gate. Absent → the run says the visual layer shipped **undesigned**, which is the honest name for picking values at the keyboard | `/plugin marketplace add ssheleg/sheleg-design` → `/plugin install sheleg-design@sheleg-design-skill` |
|
|
47
48
|
| **context7** (MCP — call tools fully qualified: `context7:resolve-library-id`, `context7:query-docs`) | stage 1 docs study | Recommended (web-search fallback) | connect the context7 MCP server |
|
|
48
49
|
| **Figma** (MCP) | stage 3 UX track, when the project designs visually — super-ux mirrors each `SCR-` screen/state into a frame | Optional, **UI + Figma-on only**. Absent → super-ux degrades to text-only *by itself and never blocks*, so shipping a UI feature with no mockups becomes a silent scope call — which is why the stage-0 sweep decides it | connect the Figma MCP server (`/mcp`, or your claude.ai connectors) |
|
|
49
50
|
| **[obsidian-wiki](https://github.com/ar9av/obsidian-wiki)** (`wiki-query`, `wiki-update`) | **stage 0 harvest** (query what's already known) **+ stage 9 sync** | **Recommended** — never a gate; absent → harvest runs on repo docs alone | `pip install obsidian-wiki` → `obsidian-wiki setup --vault /path/to/your/vault` |
|
|
@@ -77,9 +78,16 @@ exchange:
|
|
|
77
78
|
|
|
78
79
|
```
|
|
79
80
|
Pipeline companions (stage doctrine is built in — nothing to install for it):
|
|
80
|
-
✗ super-ux — this task looks user-facing; required for the UX track
|
|
81
|
+
✗ super-ux — this task looks user-facing; required for the UX track,
|
|
82
|
+
and it also owns the COPY track (copywriting, brand-voice):
|
|
81
83
|
/plugin marketplace add ssheleg/super-ux
|
|
82
84
|
/plugin install super-ux@super-ux
|
|
85
|
+
✗ sheleg-design — this task has a visual surface; it owns the VISUAL track:
|
|
86
|
+
tokens, themes, typography, rhythm, motion and its rest state:
|
|
87
|
+
/plugin marketplace add ssheleg/sheleg-design
|
|
88
|
+
/plugin install sheleg-design@sheleg-design-skill
|
|
89
|
+
(running without it — the visual layer ships undesigned,
|
|
90
|
+
and the close-out says so in those words)
|
|
83
91
|
✓ context7 — ready
|
|
84
92
|
✗ Figma MCP — this task is user-facing and the project designs in Figma
|
|
85
93
|
(docs/ux/foundation.md → Design tooling). Without it the
|
|
@@ -114,13 +122,27 @@ Pipeline companions (stage doctrine is built in — nothing to install for it):
|
|
|
114
122
|
🧠 Model for this run: recommended <top tier available>. You're on <current>.
|
|
115
123
|
/model <id> to switch, or "keep current", or name per-stage overrides.
|
|
116
124
|
|
|
125
|
+
⚠ This skill's own behaviour is unverified: evals/RESULTS.md records <N> blind
|
|
126
|
+
run(s) on <M> model(s). Its triggering and its stage compliance are authored,
|
|
127
|
+
not measured. Nothing here is blocked by that — you are told because a skill
|
|
128
|
+
that never says so is one you would assume had been tested.
|
|
129
|
+
|
|
117
130
|
Install the ✗ items you want, answer the model line, then say "continue".
|
|
118
131
|
```
|
|
119
132
|
|
|
120
133
|
Rules:
|
|
121
134
|
|
|
122
135
|
- Only flag **super-ux** when the task implies a UI (the stage-0 grill decides;
|
|
123
|
-
when unsure, flag it — a false positive costs one install).
|
|
136
|
+
when unsure, flag it — a false positive costs one install). It arms **two** tracks,
|
|
137
|
+
not one: the UX chain and the copy layer.
|
|
138
|
+
- **sheleg-design**: flag it when the task has a **visual** surface — a page, a screen,
|
|
139
|
+
a themed component, a landing, a dashboard. Detect via a resolving `/sheleg-design`.
|
|
140
|
+
**A CLI, a library, a backend service or an internal script does not flag it**, and
|
|
141
|
+
neither does a purely structural change to an existing screen: choosing a palette for
|
|
142
|
+
a log parser is how a recommendation is taught to be noise, and this bundle already
|
|
143
|
+
spent a rule learning that about the browser. Absent → the run continues and says the
|
|
144
|
+
visual layer shipped **undesigned**; that is a weaker claim and the close-out records
|
|
145
|
+
it as one, exactly as it does for a surface verified by reading the diff.
|
|
124
146
|
- **obsidian-wiki**: detect via `~/.obsidian-wiki/config` or a resolving
|
|
125
147
|
`wiki-query`/`wiki-update`. Present → say `✓ ready` and use it in the harvest.
|
|
126
148
|
Absent → print the two install lines **once** and continue; never ask twice in a
|
|
@@ -153,6 +175,13 @@ Rules:
|
|
|
153
175
|
reserve ids before minting them. Absent → print the line **once**, continue, and
|
|
154
176
|
**record the run as `ungated`** — never describe the project as protected
|
|
155
177
|
([`documentation.md`](documentation.md) → *Registers are shared state*).
|
|
178
|
+
- **The behaviour line prints whenever `evals/RESULTS.md` records no blind run**, and
|
|
179
|
+
disappears the moment one is recorded — it is a state of the evidence, not a warning
|
|
180
|
+
and not a ratchet. This bundle asks every project it touches for evidence rather than
|
|
181
|
+
assertion, and until 2026-08-10 it made the opposite claim about itself by saying
|
|
182
|
+
nothing: **one self-observed run by the author, zero blind runs on zero models**, with
|
|
183
|
+
a preflight that reported companion availability in detail and its own confidence not
|
|
184
|
+
at all. A skill silent about its own evidence is read as tested.
|
|
156
185
|
- **Never gate any stage on an install** except the stage-3 UX track on a UI task.
|
|
157
186
|
- Optional tools missing → state the fallback, don't block.
|
|
158
187
|
- Re-detect after the operator installs; don't assume.
|