pi-gauntlet 5.9.1 → 5.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.10.0 - 2026-09-18
|
|
4
|
+
|
|
5
|
+
- Council provenance: member findings (except `over-spec`) end with `probed: <source or check> - <result> | none`; the chair tags each suggested edit `grounded:` or `hypothesis:` (`external-ref:` and `over-spec:` clusters stay untagged); `roasting-the-spec` probes `hypothesis` data-shape claims once (bounded, read-only, artifact at hand) before applying, else lands them as one grouped Open Question per artifact; `Applied:` audit lines carry the probe.
|
|
6
|
+
- Standing amend approval: a chat instruction that waives per-diff review ("auto-apply amends, stop only for redraws") lets brainstorming's amendment step render the diff and continue; redraws and the spec gate still stop; the grant is quoted in each amendment commit body. Limitation: back-to-back amendments with no user input between them count as one in telemetry `derived.amendments` (coalesced until the next input event).
|
|
7
|
+
- `skills/brainstorming/SKILL.md` states each rule once; Checklist and all 12 Red Flags link to owning sections, with no gate, step, red flag, or dispatch field removed.
|
|
8
|
+
- CI pins the new persona and skill tokens.
|
|
9
|
+
|
|
3
10
|
## v5.9.1 - 2026-09-17
|
|
4
11
|
|
|
5
12
|
- Process in primary, work by path: `plan_check` roots at the plan's checkout, settings load from the session cwd's checkout toplevel, the branch-switch guard evaluates the command's `-C`/`cd` target, telemetry commits on the spec's checkout, and ship detection accepts `git -C <worktree> push` / `merge --squash`. Stage skills carry the worktree path as a value (dispatch `cwd`, `git -C`, subshell) - a new CI lint keeps them that way; `finishing-a-development-branch` takes `<worktree-path>` as a mandatory argument. Requires git >= 2.31. ([#37](https://github.com/jjuraszek/pi-gauntlet/issues/37))
|
package/README.md
CHANGED
|
@@ -63,7 +63,7 @@ flowchart LR
|
|
|
63
63
|
|
|
64
64
|
<!-- TODO GIF: a real gauntlet run end to end -->
|
|
65
65
|
|
|
66
|
-
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later is a conditional diff-approval stop (brainstorming's `Amending an approved spec`), not a third numbered gate. That's the mechanism. What follows is the machinery behind it.
|
|
66
|
+
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later is a conditional diff-approval stop (brainstorming's `Amending an approved spec`) - or, under a standing grant, a diff render that continues - not a third numbered gate. That's the mechanism. What follows is the machinery behind it.
|
|
67
67
|
|
|
68
68
|
## Architecture
|
|
69
69
|
|
|
@@ -16,7 +16,7 @@ You receive a problem statement and the artifact under review, as defined by you
|
|
|
16
16
|
|
|
17
17
|
You are read-only: you never modify the repository or any input artifact; your only write is your findings file at the dispatched output path.
|
|
18
18
|
|
|
19
|
-
When your dispatching task asks for codebase verification, verify - do not trust assertions about existing files, APIs, or conventions - but bounded: prefer `rg` (it respects `.gitignore`) over recursive `grep`, use `rg`-native bounds (`--max-count`, explicit paths); scope every scan to explicit paths, never a repository root; bound each scan with `timeout` (or `gtimeout`) when available, and do not run it unbounded when neither exists. A scan that times out or cannot be bounded is reported as unverified - never retried broader.
|
|
19
|
+
When your dispatching task asks for codebase verification, verify - do not trust assertions about existing files, APIs, or conventions - but bounded: prefer `rg` (it respects `.gitignore`) over recursive `grep`, use `rg`-native bounds (`--max-count`, explicit paths); scope every scan to explicit paths, never a repository root; bound each scan with `timeout` (or `gtimeout`) when available, and do not run it unbounded when neither exists. A scan that times out or cannot be bounded is reported as unverified - never retried broader. End every finding with `probed:`; `none` is a normal answer.
|
|
20
20
|
|
|
21
21
|
Assess the spec on five axes:
|
|
22
22
|
|
|
@@ -43,12 +43,14 @@ Emit exactly this markdown and nothing else:
|
|
|
43
43
|
verdict: sound | needs-work | unsound
|
|
44
44
|
addresses-problem: yes | partial | no — <why>
|
|
45
45
|
findings:
|
|
46
|
-
- [blocker|major|minor] <kind> @ <section or quote> — <problem> → <suggested edit>
|
|
46
|
+
- [blocker|major|minor] <kind> @ <section or quote> — <problem> → <suggested edit> — probed: <source or check> - <observed result> | none
|
|
47
47
|
lean: nothing to cut | <N> over-spec findings above
|
|
48
48
|
```
|
|
49
49
|
|
|
50
50
|
`<kind>` is one of: gap, oversimplification, ambiguity, scope, not-actionable, external-ref, other, over-spec. `scope` = too little or the wrong problem. `over-spec` = too much. No findings -> keep the `findings:` header, no bullets. `lean:` is always the last line; `<N>` = number of `over-spec` bullets. `lean: nothing to cut` is a normal answer.
|
|
51
51
|
|
|
52
|
+
`probed:` names what you read or ran and what it showed (`probed: rg 'source IN' migrations/ - CHECK lists 3 values`); `probed: none` when the edit rests on the spec text alone. The `over-spec` bullet form below carries no `probed:`.
|
|
53
|
+
|
|
52
54
|
**`over-spec`.** Flag a clause only when all three are true:
|
|
53
55
|
|
|
54
56
|
1. It is clearly outside the stated problem.
|
|
@@ -30,13 +30,15 @@ Emit exactly this markdown and nothing else:
|
|
|
30
30
|
consensus: <one-line overall verdict, e.g. needs-work — 2 of 3 members flagged blockers>
|
|
31
31
|
lean: <k> of <n> members found nothing to cut
|
|
32
32
|
clusters:
|
|
33
|
-
- [blocker|major|minor] <theme> — raised-by: [<model>, <model>] — <consolidated finding> → <suggested edit>
|
|
33
|
+
- [blocker|major|minor] <theme> — raised-by: [<model>, <model>] — <consolidated finding> → grounded|hypothesis: <suggested edit>
|
|
34
34
|
resolved:
|
|
35
35
|
- <contested point> → sided with <position> (<one-clause why>)
|
|
36
36
|
```
|
|
37
37
|
|
|
38
38
|
Every cluster must be pre-resolved — never emit a raw "members disagree" item. Leave `resolved` as a header with no bullets if no members conflicted.
|
|
39
39
|
|
|
40
|
+
Tag every suggested edit: `grounded:` when the raising members' `probed:` results support every factual assertion the edit makes; otherwise `hypothesis:`. A missing `probed:` reads as `none`. Never probe to upgrade a label - the tag records what members checked, not what you could check. `external-ref:` and `over-spec:` clusters carry no tag.
|
|
41
|
+
|
|
40
42
|
When any member raises an `external-ref` finding (load-bearing external context the spec does not inline), surface it as its own cluster with the theme prefixed `external-ref:`, e.g. `- [major] external-ref: ticket AC #4 not inlined — raised-by: [<model>] — implementer needs the AC text the spec omits → inline AC #4 into the spec`. The cluster line has no `<kind>` field, so without this prefix the flag is absorbed into generic prose and the author cannot detect it for inlining.
|
|
41
43
|
|
|
42
44
|
`over-spec` findings get their own cluster, theme prefixed `over-spec:` (the author branches on that prefix, as with `external-ref:`). Rules:
|
package/package.json
CHANGED
|
@@ -15,38 +15,23 @@ Identify the target project → set up an isolated worktree → understand curre
|
|
|
15
15
|
|
|
16
16
|
## HARD CONSTRAINT
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
Do **not** implement until the design is presented and approved, regardless of simplicity. Implementation-heavy requests define spec scope; they do not lift this gate.
|
|
19
19
|
|
|
20
|
-
You **may
|
|
20
|
+
You **may** read code/docs, run the existing system to observe current behavior, write under `doc/specs/`, and `edit` a predecessor spec there to add a [supersession banner](reference/superseding.md).
|
|
21
21
|
|
|
22
|
-
|
|
23
|
-
- Run the existing system to observe its **current** behaviour — this is research and feeds the spec (boot a local service, replay a sample request, capture a baseline classification, etc.)
|
|
24
|
-
- Write to the project's `doc/specs/` directory
|
|
25
|
-
- `edit` a predecessor spec in the project's spec directory to add a supersession banner (see [Marking superseded specs](reference/superseding.md)) — the `edit` prohibition at spec-writing binds the spec being written, not a predecessor file
|
|
22
|
+
You may **not** write outside `doc/specs/`; build, deploy, validate, or exercise the proposed change; run implementation skills; commit the spec on `main`; or start `/skill:writing-plans`.
|
|
26
23
|
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
- Write or edit code outside `doc/specs/`
|
|
30
|
-
- Scaffold a project, or take any action that builds, deploys, or validates the **proposed change** — running new/edited code, exercising behaviour that doesn't exist yet, or "testing the fix" before there is one
|
|
31
|
-
- Run implementation skills (`/skill:test-driven-development`, `/skill:subagent-driven-development`, etc.)
|
|
32
|
-
- Commit the spec directly on `main` in the primary checkout — it must land inside the worktree (see [Worktree First](#worktree-first))
|
|
33
|
-
- Start writing the plan (that's `/skill:writing-plans` — separate phase)
|
|
34
|
-
|
|
35
|
-
The line: exercising the system **as it is today** is research; exercising the **change you're proposing** is implementation and waits for the gate.
|
|
36
|
-
|
|
37
|
-
This skill ends with a **written, user-reviewed spec inside a worktree**. Nothing else.
|
|
24
|
+
The line: current-system observation is research; exercising the proposed change waits for approval.
|
|
38
25
|
|
|
39
26
|
## Foreground dispatch policy
|
|
40
27
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
Foreground does not serialize independent work: preserve existing isolated parallel `tasks` batches and await their terminal results before acceptance or tracker/phase advancement.
|
|
28
|
+
Set top-level `async: false` on every flow-owned dispatch and leave `forceTopLevelAsync` unset or false. If a dispatch still returns an async handle, stop and report; do not poll, relaunch, or advance. A detached child is incomplete work: use the existing coordination path, never accept or duplicate it. Preserve independent parallel `tasks` batches and await all terminal results before acceptance or tracker/phase advancement.
|
|
44
29
|
|
|
45
30
|
## Checklist
|
|
46
31
|
|
|
47
32
|
Work through the items below **in order**. This is your own checklist to follow, not a `plan_tracker` plan — brainstorming is open-ended exploration, and `plan_tracker` is execution-only (the implement phase). The terminal state is the user review gate; after approval the **only** next skill is `/skill:writing-plans`. Do not jump to implementation, and do not silently drop the critique pass.
|
|
48
33
|
|
|
49
|
-
1. **Start brainstorm tracking (fresh epoch)**
|
|
34
|
+
1. **Start brainstorm tracking (fresh epoch)** - as the first action on entry, before reading code, the worktree, or answering the user, reset both trackers and start the phase. A new brainstorm owns a clean slate, so stale phases and tasks from earlier work are cleared; re-entering mid-brainstorm is safe (nothing to lose, same clean slate):
|
|
50
35
|
|
|
51
36
|
```
|
|
52
37
|
phase_tracker({ action: "reset" }) // clears all phases
|
|
@@ -54,149 +39,74 @@ Work through the items below **in order**. This is your own checklist to follow,
|
|
|
54
39
|
phase_tracker({ action: "start", phase: "brainstorm" })
|
|
55
40
|
```
|
|
56
41
|
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
8. **Spec self-review (lint)** — placeholder scan + internal consistency + documentation named, run inline
|
|
70
|
-
9. **Critique pass (auto-dispatched)** — scope + ambiguity; the spec council via `/skill:roasting-the-spec` when `gauntlet_setting` returns verdict `council` (it applies its apply-set, including any external-ref inlining, to the spec before returning — see [Spec Council](#spec-council-optional)), else a fresh `worker` that applies its own fixes in place
|
|
71
|
-
10. **Re-run placeholder scan** — after the critique pass returns, re-scan the **applied** spec for placeholders its edits may have introduced; if a predecessor banner exists, confirm its `<scope>` still matches the applied spec (critique edits can change what is superseded); surface any ambiguity the critique could not safely resolve at the user gate
|
|
72
|
-
11. **Generate spec summary** — dispatch a fresh, spec-only `spec-summarizer` over the **final (post-apply)** spec, writing to an absolute temp-dir path via `outputMode: "file-only"`, then `Read` that file back as the **last content-producing** tool call before composing the gate and render its contents **verbatim** at the top of the gate message — do not paraphrase, condense, re-section, or rewrite it (see [User Review Gate](#user-review-gate)); this is part of the existing gate, not a new one
|
|
73
|
-
12. **User review gate** — user reviews the applied spec's verbatim summary plus the council audit (Applied/Deferred/Rejected), with a revert valve for any applied council edit
|
|
74
|
-
13. **Transition** — only after approval, invoke `/skill:writing-plans`
|
|
42
|
+
2. **Set up the worktree** - see [Worktree First](#worktree-first).
|
|
43
|
+
3. **Gather context** - follow [`gatherer.md`](gatherer.md) without a human stop.
|
|
44
|
+
4. **Understand the idea against the draft** - see [Understand the idea](#3-understand-the-idea).
|
|
45
|
+
5. **Propose 2-3 approaches** - see [Explore approaches](#4-explore-approaches).
|
|
46
|
+
6. **Present the design** - see [Present the design in two rounds](#6-present-the-design-in-two-rounds).
|
|
47
|
+
7. **Write the spec** - follow [Spec Self-Review](#spec-self-review-before-user-review-gate) steps 1-4.
|
|
48
|
+
8. **Spec self-review (lint)** - run the inline checks in [Spec Self-Review](#spec-self-review-before-user-review-gate).
|
|
49
|
+
9. **Critique pass (auto-dispatched)** - use [Spec Council](#spec-council-optional).
|
|
50
|
+
10. **Re-run placeholder scan** - follow [Spec Self-Review](#spec-self-review-before-user-review-gate).
|
|
51
|
+
11. **Generate spec summary** - follow [User Review Gate](#user-review-gate).
|
|
52
|
+
12. **User review gate** - follow [User Review Gate](#user-review-gate).
|
|
53
|
+
13. **Transition** - after approval, invoke `/skill:writing-plans` per [User Review Gate](#user-review-gate).
|
|
75
54
|
|
|
76
55
|
## Project Routing
|
|
77
56
|
|
|
78
|
-
|
|
57
|
+
Identify the target project before brainstorming; this determines the spec directory, instructions, and verification commands. Read top-level and relevant service `AGENTS.md`; follow documented routing, otherwise use repository conventions.
|
|
79
58
|
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
Detection order:
|
|
83
|
-
|
|
84
|
-
1. Issue tracker reference (Linear, Jira, GitHub Issues) → infer from labels, description, or mentioned file paths
|
|
85
|
-
2. cwd inside a service / package directory → use that area
|
|
86
|
-
3. Unclear → ask the developer
|
|
87
|
-
|
|
88
|
-
If the project's `AGENTS.md` calls out additional reading for a specific area (e.g. a TDD workflow doc), read it before brainstorming work in that area.
|
|
59
|
+
Detection order: (1) infer from ticket labels, description, or paths; (2) use the service/package containing cwd; (3) ask if unclear. Read any area-specific material required by `AGENTS.md` before brainstorming.
|
|
89
60
|
|
|
90
61
|
## Worktree First
|
|
91
62
|
|
|
92
|
-
The spec is the
|
|
93
|
-
|
|
94
|
-
1. Invoke `/skill:using-git-worktrees`. Worktrees live under `.worktrees/` at the repo root, or wherever a project-native worktree script places them.
|
|
95
|
-
2. Carry the worktree path from the `using-git-worktrees` Step 4 report (`Worktree ready at <full-path>`) as a value: the spec path is `<full-path>/doc/specs/<filename>.md`, every dispatch below sets `cwd: "<full-path>"`, and git runs as `git -C <full-path> ...`. The process cwd stays where pi was launched.
|
|
96
|
-
3. Write the spec at that path, run self-review, commit with `git -C <full-path>`, hand off.
|
|
63
|
+
The spec is the first commit in a dedicated worktree, never a separate commit on `main`.
|
|
97
64
|
|
|
98
|
-
|
|
65
|
+
1. Invoke `/skill:using-git-worktrees`; use `.worktrees/` or the project-native location.
|
|
66
|
+
2. Carry its `Worktree ready at <full-path>` value: spec path `<full-path>/doc/specs/<filename>.md`; every dispatch uses that `cwd`; git uses `git -C <full-path>`. Keep the process cwd unchanged.
|
|
67
|
+
3. Write, review, and commit there.
|
|
99
68
|
|
|
100
|
-
|
|
69
|
+
Spec, plan, and implementation share this worktree; finishing strips the ephemeral plan before squash. Explicit trivial one-off edits outside this flow need no worktree.
|
|
101
70
|
|
|
102
71
|
## The Process
|
|
103
72
|
|
|
104
73
|
### 1. Check git state
|
|
105
74
|
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
```bash
|
|
109
|
-
git status
|
|
110
|
-
git --no-pager log --oneline -5
|
|
111
|
-
```
|
|
112
|
-
|
|
113
|
-
If on a feature branch with uncommitted or unmerged work, ask:
|
|
114
|
-
|
|
115
|
-
> "You're on `<branch>` with uncommitted changes. Want to finish/merge that first, stash it, or continue here?"
|
|
116
|
-
|
|
117
|
-
Require one of: finish prior work, stash, or explicit "continue here". If the topic is new, set up the worktree per [Worktree First](#worktree-first) before continuing.
|
|
75
|
+
In the primary checkout run `git status` and `git --no-pager log --oneline -5`. On a feature branch with uncommitted or unmerged work, ask whether to finish/merge, stash, or continue; require one choice. For a new topic, create the worktree before continuing.
|
|
118
76
|
|
|
119
77
|
### 2. Scope check
|
|
120
78
|
|
|
121
|
-
One spec is
|
|
79
|
+
One spec is default. Test seemingly independent concerns against `../shape-ticket/reference/split-axes.md` (identity, outcome, closed axis, release timing). For every proposed split render exactly:
|
|
122
80
|
|
|
123
81
|
root cause: <the precipitating failure or missing capability this slice remedies>
|
|
124
82
|
outcome: <what a user observes once it ships>
|
|
125
83
|
axis: <one item from the closed list>
|
|
126
84
|
|
|
127
|
-
|
|
85
|
+
If any test fails, offer no split. Never split by service, package, repo, layer, or team. If all pass, ask whether to brainstorm the independent specs separately or explain their coupling.
|
|
128
86
|
|
|
129
|
-
|
|
87
|
+
### 3. Understand the idea
|
|
130
88
|
|
|
131
|
-
|
|
89
|
+
`Read` the gathered draft unconditionally before question one. Treat it as a helper, not a fence: verify load-bearing claims from primary code, and confirm whether the codebase or ecosystem already solves the problem.
|
|
132
90
|
|
|
133
|
-
|
|
91
|
+
Ask one question per message, preferably multiple choice, about purpose, constraints, success, and affected actors. Before asking, check code, docs, and tracker: look up current-state facts (dispatch a subagent when costly); ask desired-behavior decisions even when a ticket recorded one. Every questionary question, including acceptance of a corrected fact, ends `Recommendation: <answer> - <why>`; other approvals retain their wording.
|
|
134
92
|
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
- `Read` the draft **unconditionally before composing question one**. The on-disk
|
|
139
|
-
copy is canonical — this one rule defeats both a turn-boundary prune after assembly
|
|
140
|
-
and a session restart.
|
|
141
|
-
- The draft is a **helper, not a fence**: judgment still drives exploration. Verify
|
|
142
|
-
load-bearing claims (schemas, contracts, the code being changed) against real code
|
|
143
|
-
via targeted reads (`read_symbol`-grade, not scout's paraphrase) before designing
|
|
144
|
-
against them.
|
|
145
|
-
- **Check if the codebase or ecosystem already solves this** — the draft's recon
|
|
146
|
-
section starts that answer; confirm it before designing from scratch.
|
|
147
|
-
- Ask questions **one at a time** to refine the idea. Prefer multiple-choice; one
|
|
148
|
-
question per message. Focus on: purpose, constraints, success criteria, who/what
|
|
149
|
-
it touches. Before asking, check whether the code, the docs, or the issue tracker
|
|
150
|
-
already answer it: a fact about the current state is looked up, not asked (dispatch
|
|
151
|
-
a subagent when the lookup is costly); a decision about what should happen is asked,
|
|
152
|
-
even when a ticket recorded one earlier. Every questionary question, including the
|
|
153
|
-
ask to accept a corrected fact, ends with the line `Recommendation: <answer> - <why>`
|
|
154
|
-
(the answer, then " - ", then the reason); approvals elsewhere in this skill (git
|
|
155
|
-
state, design rounds, the spec gate) keep their own wording.
|
|
156
|
-
- **Append bar:** append to the draft's `## Appended during questionary` only
|
|
157
|
-
findings the spec will cite — schema shapes, hard constraints, ticket-vs-code
|
|
158
|
-
contradictions, user answers that changed scope. Not a log of every grep.
|
|
159
|
-
(Appending uses `edit`; the `edit` prohibition in the spec-writing step applies
|
|
160
|
-
only there.)
|
|
161
|
-
|
|
162
|
-
Before proposing approaches, state the premise note in chat: what the design depends
|
|
163
|
-
on, which of those claims the sources support and where you saw it (a file and line,
|
|
164
|
-
a doc, a ticket), which they disprove - give the corrected fact and where you found
|
|
165
|
-
it - and which remain unverified, naming the lookup you tried. Write it as a note a
|
|
166
|
-
person can act on: full sentences, no status-keyword lists, no template; when the
|
|
167
|
-
design depends on no claims at all, one sentence saying so is enough. An unverified
|
|
168
|
-
claim is not a stop - it enters the spec as an Open Question or a stated assumption.
|
|
169
|
-
If a claim the design depends on was contradicted, the note is your next message -
|
|
170
|
-
even as questionary question one - and it ends by asking the user to accept the
|
|
171
|
-
corrected fact or explicitly override it; nothing else continues - no other questions,
|
|
172
|
-
no approaches - until they answer, and the outcome is recorded in the draft's
|
|
173
|
-
`## Appended during questionary` so spec-writing carries it into `## Problem` or the
|
|
174
|
-
relevant `## Design` decision.
|
|
93
|
+
Append only citable findings - schemas, hard constraints, contradictions, and scope-changing answers - to `## Appended during questionary` using `edit`.
|
|
94
|
+
|
|
95
|
+
Before approaches, state in chat the supported, disproved, corrected, and unverified premises with sources and attempted lookups. Put unverified claims in Open Questions or assumptions. If a load-bearing claim is contradicted, the premise note is your next message - even as question one - and asks the user to accept the corrected fact or override it; ask nothing else and propose nothing until answered. Record the outcome in the draft for `## Problem` or the relevant design decision.
|
|
175
96
|
|
|
176
97
|
### 4. Explore approaches
|
|
177
98
|
|
|
178
|
-
|
|
179
|
-
- Lead with your recommended option and explain why.
|
|
180
|
-
- Present conversationally — don't dump a comparison table unless the user asks.
|
|
99
|
+
Propose 2-3 approaches with trade-offs; lead with the recommendation and explain it. Use conversational prose unless the user asks for a table.
|
|
181
100
|
|
|
182
101
|
### 5. Design for clarity and isolation
|
|
183
102
|
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
- **Clear boundaries** — clean interfaces between components, easy to test in isolation.
|
|
187
|
-
- **YAGNI ruthlessly** — every component you add is a component you must maintain.
|
|
188
|
-
- **Match existing patterns** — if a service has a convention, follow it. Reference the service's `AGENTS.md` and neighboring code.
|
|
189
|
-
- **Single source of truth** — point at the schema/contract that owns the data (the migration, type definition, or API contract that defines it); don't invent parallel state.
|
|
190
|
-
- **Explicit error and edge cases** — name them. "Out of scope" is a valid answer, but it has to be stated.
|
|
103
|
+
Prefer clear testable boundaries, YAGNI, existing conventions, the owning schema/contract rather than parallel state, and explicit errors and edge cases.
|
|
191
104
|
|
|
192
105
|
### 6. Present the design in two rounds
|
|
193
106
|
|
|
194
|
-
|
|
107
|
+
Use two 300-500-word rounds with one approval each; revisions remain within that approval point. Ask once per round; round-1 approval without correction confirms the predecessor. Round 1 covers architecture, responsibilities, data flow, and `supersedes <path>, <scope>` when applicable. Round 2 covers errors, edges, tests, and `## Documentation impact`.
|
|
195
108
|
|
|
196
|
-
|
|
197
|
-
- Round 2: error handling and edge cases, testing approach, `## Documentation impact`. Ask once.
|
|
198
|
-
|
|
199
|
-
The `## Documentation impact` section is required. Cite the materiality bar in `reference/documentation-impact.md` by relative path rather than restating its categories, and reproduce its template block verbatim:
|
|
109
|
+
`## Documentation impact` is required. Cite `reference/documentation-impact.md` by relative path, do not restate its categories, and reproduce this template verbatim:
|
|
200
110
|
|
|
201
111
|
```markdown
|
|
202
112
|
## Documentation impact
|
|
@@ -205,41 +115,25 @@ The `## Documentation impact` section is required. Cite the materiality bar in `
|
|
|
205
115
|
- Derived / memory docs invalidated: <routers / AGENTS.md sections / topic guides / indexes, or "none">
|
|
206
116
|
```
|
|
207
117
|
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
Be ready to go back and clarify when something doesn't make sense.
|
|
118
|
+
Use doc names, `none`, or `deferred: <trigger>`. Apply `reference/documentation-impact.md`; amend an owner before creating a standalone file. Put project taxonomy in the overrides file's `## documentation` block. Ship doc updates with code and verify them at conformance. Clarify when needed.
|
|
211
119
|
|
|
212
120
|
## Ticket Handling
|
|
213
121
|
|
|
214
|
-
|
|
122
|
+
A ticket is guidance, not sole truth. Fetch it; propose scope, approach, or acceptance changes when code disagrees, and record deviations in the spec.
|
|
215
123
|
|
|
216
124
|
## First-Feature Oversight (Early Project Stages)
|
|
217
125
|
|
|
218
|
-
For the
|
|
126
|
+
For the first two features of a new module, long-lived component, schema area, or repeatable pattern, round 1 explicitly covers structure, naming, shared abstractions, persistence/schema, and proposed AGENTS.md/docs changes. Use no separate confirmation; later features follow established patterns.
|
|
219
127
|
|
|
220
128
|
## Anti-Pattern: "Too simple to need a design"
|
|
221
129
|
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
- Small changes that touch shared schemas, contracts, or invariants need a spec.
|
|
225
|
-
- "Small" often hides assumptions that the spec process would surface.
|
|
226
|
-
- A 5-minute spec saves an hour of rework when the assumption was wrong.
|
|
227
|
-
|
|
228
|
-
If the change truly is mechanical and contained (rename, formatter run, dependency bump), say so explicitly and skip brainstorming. Otherwise: spec first.
|
|
130
|
+
Shared schemas, contracts, and invariants require a spec. If work is mechanical and contained, such as a rename, formatting, or dependency bump, say so explicitly and skip brainstorming; otherwise spec first.
|
|
229
131
|
|
|
230
132
|
## Filename Convention
|
|
231
133
|
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
- With ticket: `YYYY-MM-DD-<ticket-id>-<topic>.md` — `<ticket-id>` is a filename-safe slug of the tracker reference (e.g. `E-12345` for Linear, used verbatim; `gh-123` for GitHub issue `#123`). Plan headers and commit messages use the tracker's native reference form, not the filename slug.
|
|
235
|
-
- Without ticket: `YYYY-MM-DD-<topic>.md`
|
|
236
|
-
|
|
237
|
-
`<topic>` is a short kebab-case slug (3–6 words). Do **not** append `-design` or any other suffix.
|
|
134
|
+
Write under the routed `doc/specs/`: ticketed `YYYY-MM-DD-<ticket-id>-<topic>.md`, otherwise `YYYY-MM-DD-<topic>.md`. Use a filename-safe tracker slug (`E-12345`, `gh-123`), but native references in plan headers and commits. `<topic>` is 3-6 kebab-case words without `-design` or another suffix.
|
|
238
135
|
|
|
239
|
-
|
|
240
|
-
overwrite reuses the path. If the questionary invalidated the slug, rename at
|
|
241
|
-
spec-writing: write the spec at the new path **and delete the old draft file**
|
|
242
|
-
(nothing was committed, so this is free).
|
|
136
|
+
Mint the slug once during gather and reuse it. If questionary invalidates it, write to the new path and delete the uncommitted draft.
|
|
243
137
|
|
|
244
138
|
## Spec Self-Review (Before User Review Gate)
|
|
245
139
|
|
|
@@ -259,7 +153,7 @@ Spec-writing replaces the context draft, in this exact order:
|
|
|
259
153
|
[Marking superseded specs](reference/superseding.md)). This position is fixed:
|
|
260
154
|
the banner is written after any slug rename, so it always cites the final path.
|
|
261
155
|
|
|
262
|
-
After writing the spec to `<project>/doc/specs/<filename>.md` (per [Filename Convention](#filename-convention)) and before showing it to the user, run a self-review pass.
|
|
156
|
+
After writing the spec to `<project>/doc/specs/<filename>.md` (per [Filename Convention](#filename-convention)) and before showing it to the user, run a self-review pass. Read all five bullets first, then act.
|
|
263
157
|
|
|
264
158
|
- **Placeholder scan.** Any `TODO`, `TBD`, `<fill in>`, `[example]`, `xxx`? Either resolve them or convert to explicit "Open Questions" with names.
|
|
265
159
|
- **Internal consistency.** Does Section 4 contradict Section 2? Are component names and field names consistent throughout? If the spec replaces a prior design, confirm the predecessor carries the supersession banner and its href resolves to this spec's final filename.
|
|
@@ -267,35 +161,32 @@ After writing the spec to `<project>/doc/specs/<filename>.md` (per [Filename Con
|
|
|
267
161
|
- **Scope check.** Does every paragraph serve the goal? Cut filler. If something is out of scope, say it's out of scope.
|
|
268
162
|
- **Ambiguity check.** Is every "we should…" backed by a concrete decision? Replace "we could probably" with "we will" or "we won't".
|
|
269
163
|
|
|
270
|
-
The first three
|
|
164
|
+
The first three are the inline **lint**: run them here and fix what they surface. The last two are the **critique pass**, dispatched per [Spec Council](#spec-council-optional). After it returns, re-run the placeholder scan over the applied spec; if a predecessor banner exists, confirm its `<scope>` still matches and reconcile it; carry any ambiguity the critique could not resolve to the [User Review Gate](#user-review-gate).
|
|
271
165
|
|
|
272
|
-
|
|
273
|
-
- **Otherwise** → dispatch one fresh `worker` that applies the scope + ambiguity checks and fixes them in place:
|
|
274
|
-
|
|
275
|
-
```
|
|
276
|
-
subagent({ agent: "worker", context: "fresh", async: false, cwd: "<abs worktree path, from the using-git-worktrees Step 4 report>", task:
|
|
277
|
-
"Problem statement: <the problem the spec addresses + the user's stated intent>.\n" +
|
|
278
|
-
"Read the spec at <abs path to doc/specs/...>. Edit ONLY that file. Apply two checks and\n" +
|
|
279
|
-
"fix what you find in place: (1) Scope — does every paragraph serve the goal? Cut filler;\n" +
|
|
280
|
-
"state out-of-scope explicitly. (2) Ambiguity — is every 'we should' a concrete decision?\n" +
|
|
281
|
-
"Replace 'we could probably' with 'we will'/'we won't'. Also inline any load-bearing\n" +
|
|
282
|
-
"external reference (ticket AC, commit SHA, doc) already given to you in the problem\n" +
|
|
283
|
-
"statement above; if the spec relies on one not provided here, flag it (do NOT fetch) in\n" +
|
|
284
|
-
"your summary. Return a summary of what you changed, and flag any ambiguity you could NOT\n" +
|
|
285
|
-
"safely resolve." })
|
|
286
|
-
```
|
|
166
|
+
## Spec Council (Optional)
|
|
287
167
|
|
|
288
|
-
|
|
168
|
+
After the inline lint and before the user review gate, **brainstorming owns the critique-pass gate**; council **apply mechanics** live in `/skill:roasting-the-spec` (single source of truth - link, don't restate). Resolve the council with `gauntlet_setting({ key: "specCouncil" })` - the tool returns the merged (repo-over-preset) value as `{ verdict, members, chair, malformed, warning, errors }`. **Do not** hand-roll a settings read. When `verdict` is `"council"`, the council *is* the critique pass - invoke `/skill:roasting-the-spec` automatically (no offer, no prompt), passing `members`/`chair`; also pass the verbatim human input (the original prompt, any ticket AC snapshot, and the questionary answers that changed scope) - roasting-the-spec forwards it to members and chair as the `Human input (verbatim; off-limits for over-spec)` block; it applies its apply-set and returns the audit (Applied/Deferred/Rejected). When `verdict` is `"worker"`, dispatch the worker below. If `malformed` is true or `errors` is non-empty, emit the `warning`/error as one line, then branch strictly on `verdict` - `malformed` can accompany *either* verdict (e.g. a bad `chair` with valid `members` still returns `council`), so never infer the worker path from `malformed` alone. If `gauntlet_setting` is unavailable, stop and report - never fall back to a manual bash/JSON settings merge. The already-applied council edits (or the worker's in-place fixes) ride in the same worktree commit. The conceptual precedence rule lives in `verification-before-completion/reference/settings-precedence.md`.
|
|
289
169
|
|
|
290
|
-
|
|
170
|
+
When `verdict` is `"worker"`, dispatch one fresh `worker` that applies the scope + ambiguity checks and fixes them in place:
|
|
291
171
|
|
|
292
|
-
|
|
172
|
+
```
|
|
173
|
+
subagent({ agent: "worker", context: "fresh", async: false, cwd: "<abs worktree path, from the using-git-worktrees Step 4 report>", task:
|
|
174
|
+
"Problem statement: <the problem the spec addresses + the user's stated intent>.\n" +
|
|
175
|
+
"Read the spec at <abs path to doc/specs/...>. Edit ONLY that file. Apply two checks and\n" +
|
|
176
|
+
"fix what you find in place: (1) Scope — does every paragraph serve the goal? Cut filler;\n" +
|
|
177
|
+
"state out-of-scope explicitly. (2) Ambiguity — is every 'we should' a concrete decision?\n" +
|
|
178
|
+
"Replace 'we could probably' with 'we will'/'we won't'. Also inline any load-bearing\n" +
|
|
179
|
+
"external reference (ticket AC, commit SHA, doc) already given to you in the problem\n" +
|
|
180
|
+
"statement above; if the spec relies on one not provided here, flag it (do NOT fetch) in\n" +
|
|
181
|
+
"your summary. Return a summary of what you changed, and flag any ambiguity you could NOT\n" +
|
|
182
|
+
"safely resolve." })
|
|
183
|
+
```
|
|
293
184
|
|
|
294
|
-
|
|
185
|
+
`worker`'s model resolves from `subagents.agentOverrides.worker.model` in `settings.json` (unset → inherits the main loop); the dispatch passes no `model:`.
|
|
295
186
|
|
|
296
187
|
## User Review Gate
|
|
297
188
|
|
|
298
|
-
After self-review and the critique pass (council or worker, both already applied to the spec - see [Spec Council](#spec-council-optional)), dispatch the spec-only summarizer over the **applied** spec, then commit the spec on the worktree branch and stop.
|
|
189
|
+
After self-review and the critique pass (council or worker, both already applied to the spec - see [Spec Council](#spec-council-optional)), dispatch the spec-only summarizer over the **applied** spec, then commit the spec on the worktree branch and stop. The summary is folded into the one human gate, not a new gate.
|
|
299
190
|
|
|
300
191
|
Mint an absolute temp path outside the worktree (so it is never committed), then dispatch the summarizer on a fresh context, reading only the spec, writing to that path via file-only output (no `model:` - it inherits the main loop unless a preset sets `subagents.agentOverrides.spec-summarizer.model`):
|
|
301
192
|
|
|
@@ -320,7 +211,7 @@ Then commit the spec — staging any predecessor spec edited per [Marking supers
|
|
|
320
211
|
|
|
321
212
|
Either way — summary rendered or degraded — then `rm "$SUMMARY_PATH"` (unconditional cleanup; harmless if the file was never created, since it lives outside the worktree under the OS temp dir).
|
|
322
213
|
|
|
323
|
-
|
|
214
|
+
Paste the summary verbatim, unedited in the template below; use adjacent lines for the audit, unresolved ambiguities, and every gap-footer entry:
|
|
324
215
|
|
|
325
216
|
```
|
|
326
217
|
<spec-only summary read back from the temp file — pasted verbatim, unedited>
|
|
@@ -357,7 +248,7 @@ phase_tracker({ action: "complete", phase: "brainstorm" })
|
|
|
357
248
|
Execute this section in place from any later phase. Do not invoke `/skill:brainstorming` (its entry resets both trackers). Worktree, spec commits, and plan survive.
|
|
358
249
|
|
|
359
250
|
1. Edit the spec. Show `git -C <abs worktree path> --no-pager diff -- <spec path>` and one line of impact (affected plan tasks / waves, or "no plan yet").
|
|
360
|
-
2.
|
|
251
|
+
2. Render the diff and impact line. A user instruction in this flow that waives per-diff review for later amends ("auto-apply amends, stop only for redraws", "apply spec fixes without asking") is the approval: quote it in the amendment commit body and continue. Otherwise wait for approval; change request -> revise, re-show. Redraws always wait. A grant never satisfies the spec gate; a grant given with or before spec approval applies to later amends in the same flow; a new brainstorm and a fresh-session resume start with no grant.
|
|
361
252
|
3. No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks; anchor-changed completed tasks are reopened as `in_progress` and re-run the task loop (`update` never sets `pending`). A removed task is deleted from the plan; then re-`init` the tracker with `{ name, status }` elements: preserved tasks keep their order and statuses, reopened tasks are `in_progress` in place, every still-`pending` task (including newly added ones, whatever wave label they carry) trails the non-pending ones, removed tasks are the only deletions (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
362
253
|
|
|
363
254
|
Redraw test: the diff changes the problem statement, adds or removes a component, or moves a component boundary -> redraw. A change inside one component (a persistence mechanism, a worker's HTTP client, dropping a fallback and its task) -> amend. State the call in the same message as the diff; the user overrides either way.
|
|
@@ -366,27 +257,22 @@ Redraw: keep the worktree and the approved spec file. `plan_tracker({ action: "c
|
|
|
366
257
|
|
|
367
258
|
## Key Principles
|
|
368
259
|
|
|
369
|
-
|
|
370
|
-
- **Multiple choice preferred** when possible.
|
|
371
|
-
- **YAGNI ruthlessly.**
|
|
372
|
-
- **Design for testability** — clear boundaries enable TDD.
|
|
373
|
-
- **Explore 2-3 approaches** before settling.
|
|
374
|
-
- **Two design rounds** — one approval per round.
|
|
375
|
-
- **Be flexible** — go back and clarify when something doesn't make sense.
|
|
260
|
+
One question at a time, YAGNI, 2-3 approaches, two design rounds, clarify freely - all owned by [The Process](#the-process).
|
|
376
261
|
|
|
377
262
|
## Red Flags — STOP
|
|
378
263
|
|
|
379
|
-
-
|
|
380
|
-
-
|
|
381
|
-
-
|
|
382
|
-
-
|
|
383
|
-
-
|
|
384
|
-
-
|
|
385
|
-
-
|
|
386
|
-
-
|
|
387
|
-
-
|
|
388
|
-
-
|
|
389
|
-
-
|
|
264
|
+
- Writes outside `doc/specs/` ([owner](#hard-constraint)).
|
|
265
|
+
- Draft overwrite without the same-turn full read, or overwrite via `edit` ([owner](#spec-self-review-before-user-review-gate)).
|
|
266
|
+
- Dispatch while line 1 is the context-draft marker ([owner](#spec-self-review-before-user-review-gate)).
|
|
267
|
+
- Inline scope or ambiguity checks ([owner](#spec-council-optional)).
|
|
268
|
+
- Gate after failed/skipped critique or before placeholder re-scan ([owner](#spec-self-review-before-user-review-gate)).
|
|
269
|
+
- Gate without summary `Read` last, or with a paraphrased summary ([owner](#user-review-gate)).
|
|
270
|
+
- Human stop between gather and question one ([owner](#the-process)).
|
|
271
|
+
- Proposed-change execution before approval ([owner](#hard-constraint)).
|
|
272
|
+
- Plan before approval; brainstorming invocation for an amend ([owner](#user-review-gate)).
|
|
273
|
+
- Missing predecessor banner; invalid multi-spec split ([owner](#spec-self-review-before-user-review-gate); [owner](#2-scope-check)).
|
|
274
|
+
- Approaches while a contradicted premise remains unresolved ([owner](#3-understand-the-idea)).
|
|
275
|
+
- Waiting after an amend grant; auto-applying a redraw ([owner](#amending-an-approved-spec)).
|
|
390
276
|
|
|
391
277
|
## Project overrides
|
|
392
278
|
|
|
@@ -119,6 +119,8 @@ For each cluster in the chair's report, decide one of:
|
|
|
119
119
|
- **defer** — out of scope for this spec; name where it belongs. Do not edit the spec.
|
|
120
120
|
- **reject** — one-line reason. Do not edit the spec.
|
|
121
121
|
|
|
122
|
+
**`hypothesis` clusters that assert data shape, ordering, or semantics** (a parsing rule, a field's meaning, a sort or date order, an identity key) are applied only after a probe. Search once for the artifact: one `rg --max-count` under `timeout` over `<abs worktree path>` and its docs, config, and script directories, for the artifact or for how it is obtained. At hand -> one bounded read-only check: a read, an `rg` over explicit paths, or a project script whose source you read and which only reads local files. Confirmed -> apply as fact. Anything else (not found, inconclusive, timed out, contradicted) -> write the edit into `## Open questions` (create the section if absent) with the probe run, its outcome, the obtain-hint if found, and the settlement path: the user supplies the fact at the gate, or the first plan task that obtains the artifact settles it via brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec). One Open Question per artifact, listing every edit it settles - never one per assertion. Never fetch, build, or run the proposed change to obtain the artifact. A cluster without a tag is read as `hypothesis`; a `hypothesis` cluster that asserts nothing about data applies as any other.
|
|
123
|
+
|
|
122
124
|
Also inline any `external-ref:` cluster you have context for (e.g. a ticket fetched during brainstorming) as part of the apply-set — this is your call, same as any other cluster.
|
|
123
125
|
|
|
124
126
|
An `over-spec:` cluster decided **apply** is executed as deletion or shrink of the quoted clause **and** any acceptance-criteria or testing-approach line that exists only for it. Its audit line reads `Applied: over-spec: <clause> -> cut (was adds: M files / N tests / K ACs)` or `Applied: over-spec: <clause> -> shrunk to <replacement> (was adds: ...)`, so the gate shows what was removed. `defer`/`reject` are unchanged.
|
|
@@ -130,7 +132,7 @@ You are the advocate — decide on scope grounds — and, unlike a dispatched su
|
|
|
130
132
|
Return a structured audit, gate-only (not a committed spec section) — a coverage line plus three labelled lists:
|
|
131
133
|
|
|
132
134
|
- `Coverage:` — `N of M members reported; <slug>: <reason>` — present only when member coverage was partial; omitted at full coverage.
|
|
133
|
-
- `Applied:` — cluster ->
|
|
135
|
+
- `Applied:` — one of `Applied: <cluster> -> <edit> (grounded: <member probe>)`, `Applied: <cluster> -> <edit> (probed: <check> - <result>)` for a confirmed hypothesis, `Applied: <cluster> -> open question (<not found | inconclusive: <check> | contradicted: <result>>)`. The probe rides on the audit line because member files are removed in section 5.
|
|
134
136
|
- `Deferred:` — cluster -> where it belongs.
|
|
135
137
|
- `Rejected:` — cluster -> one-line reason.
|
|
136
138
|
|