pi-gauntlet 5.12.1 → 5.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +8 -0
- package/README.md +2 -2
- package/agents/conformance-reviewer.md +3 -0
- package/agents/spec-council-member.md +12 -2
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +25 -11
- package/skills/brainstorming/gatherer.md +6 -2
- package/skills/brainstorming/reference/amendment-surface.md +158 -0
- package/skills/finishing-a-development-branch/SKILL.md +19 -0
- package/skills/finishing-a-development-branch/reference/disposition-protocol.md +2 -1
- package/skills/subagent-driven-development/SKILL.md +1 -1
- package/skills/verification-before-completion/reference/conformance-check.md +1 -0
- package/skills/writing-plans/SKILL.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,13 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.14.0 - 2026-09-20
|
|
4
|
+
|
|
5
|
+
- Specs carry the ticket's acceptance criteria verbatim in a required `## Acceptance criteria` section, one disposition per row (`in-scope`, `deviates: <why>`, `deferred: <where>`, `venue: <env> - <observation>`; `none - <reason>` when there is no ticket or no ACs). `brainstorming` extracts heading-scoped rows and lints the section's presence, `gatherer` quotes the raw rows, `spec-council-member` flags dropped/reworded rows and invalid deferrals, `conformance-reviewer` and `conformance-check` read the section as origin (`in-scope`/`venue:` rows are requirements, `venue:` observations never block), `writing-plans` tables only `in-scope`/`venue:` rows, and `finishing-a-development-branch` lists `venue:`/`deferred:` rows in the PR body. `scripts/ci.mjs` pins the new tokens. (#41)
|
|
6
|
+
|
|
7
|
+
## v5.13.0 - 2026-09-19
|
|
8
|
+
|
|
9
|
+
- Post-approval spec amendments go through a reviewer-first funnel (`skills/brainstorming/reference/amendment-surface.md`): a fresh `spec-council-member` in `Mode: amendment-review` clears evidence-backed factual corrections that touch no human-owned section, a deterministic prefilter sends descopes and acceptance-criteria edits to the human, and escalations render as one readable batch (what / why / example / recommended, real alternatives only) with a one-reply grammar and the standing-grant offer; each batch lands as one `amend:` commit with per-item records. The spec gate offers the grant. `finishing-a-development-branch` runs eligible conformance `accept` gaps through the same funnel before the disposition menu and lists auto-applied amendments above the ship options. `brainstorming/SKILL.md` shrinks; `scripts/ci.mjs` pins the new tokens.
|
|
10
|
+
|
|
3
11
|
## v5.12.1 - 2026-09-19
|
|
4
12
|
|
|
5
13
|
- Fixed: `gauntlet-telemetry-salvage` and `gauntlet-performance` no longer crash with `ERR_UNSUPPORTED_NODE_MODULES_TYPE_STRIPPING` when run from an npm-installed copy - both bins are now committed esbuild bundles (sources in `src/bins/`, rebuild with `npm run build:bins`), guarded by a CI freshness check, bundle pack assertions, and a packed-install smoke test. (#39)
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ pi-gauntlet's only hard dependency is pi-cohort - every gate that dispatches a r
|
|
|
36
36
|
Concretely, one change through the gauntlet:
|
|
37
37
|
|
|
38
38
|
0. *(Optional)* Before there's even a spec, `/skill:shape-ticket` can create or repair a single tracker issue - shaping a raw ask into a Context/Problem/Idea/Acceptance Criteria ticket, gated by an AC integrity check, a cheap council roast, and one human-confirmed write that may include one optional gated Reporter-note comment. A failed roast is retried once, then surfaced inline at the gate if it fails again. It's a tool, not a phase: no worktree, no plan/phase tracker, runs from any repo state. It never activates on its own (`disable-model-invocation: true`) - invoke it explicitly. Similarly, `/skill:chase-bug` triages a raw bug report into an evidenced verdict - and can hand off to shape-ticket, brainstorming, or a bounded hotfix - before any spec exists.
|
|
39
|
-
1. You describe the change. **`brainstorming`** sets up an isolated worktree, explores the codebase, and turns your description into a written spec. A multi-model critique runs on it automatically. If the spec replaces a known prior spec, brainstorming marks the predecessor with a `> **Superseded by:**` banner under its title (default format, syntax overridable via `.pi/gauntlet-overrides.md`; candidates come from brainstorming's scout recon, never a mechanical sweep). **You read and approve the spec - human gate 1.** No implementation code exists yet.
|
|
39
|
+
1. You describe the change. **`brainstorming`** sets up an isolated worktree, explores the codebase, and turns your description into a written spec; when the run starts from a ticket, the spec carries the ticket's acceptance criteria verbatim, each with a disposition (`in-scope`, `deviates:`, `deferred:`, `venue:`), the conformance gate checks the in-scope ones, and `/skill:check-delivery` verifies `venue:` rows after deploy. A multi-model critique runs on it automatically. If the spec replaces a known prior spec, brainstorming marks the predecessor with a `> **Superseded by:**` banner under its title (default format, syntax overridable via `.pi/gauntlet-overrides.md`; candidates come from brainstorming's scout recon, never a mechanical sweep). **You read and approve the spec - human gate 1.** No implementation code exists yet.
|
|
40
40
|
2. **`writing-plans`** decomposes the approved spec into atomic, independently-verifiable tasks, grouped into parallel waves where they don't touch the same files.
|
|
41
41
|
3. **`subagent-driven-development`** executes the plan one task at a time, each in a fresh subagent, behind spec-compliance review then code-quality review. TDD-locked: red, green, refactor. Independent implementation, review, and council batches remain parallel, but their dispatches are foreground: the orchestrator waits for terminal results before accepting work or advancing a phase.
|
|
42
42
|
4. **verify**: the parent first runs the plan's full verification command set; only a passing result permits whole-diff code review, then the **conformance gate**. A review fix invalidates that result, so the parent reruns full verification before the next review or conformance gate. The conformance subagent reads the finished code and docs against your *original words* from step 1, not the plan, and reports per-requirement: delivered, partial, missing, drifted, or unauthorized. Unauthorized rows - shipped surface no human input asked for, whether it crept in or was laundered through the spec - follow their recommendation like every other row: a contained removal auto-runs, anything another requirement leans on is deferred with a plain-language "I'd cut it / I'd keep it" recommendation. Inside a brainstorming-entered flow this gate is machine-blocked from being skipped. Compatible executable recommendations auto-run through an isolated fix-and-re-audit loop with no prompt; anything still open surfaces as a dense list - one line per decision, plain-language, with its recommended choice inline. Reply `1` to take every recommendation, or `2:` with per-item overrides; a current `CONFORMS` / no-concerns result goes straight to the branch options with no extra conformance sign-off.
|
|
@@ -63,7 +63,7 @@ flowchart LR
|
|
|
63
63
|
|
|
64
64
|
<!-- TODO GIF: a real gauntlet run end to end -->
|
|
65
65
|
|
|
66
|
-
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later
|
|
66
|
+
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later goes through brainstorming's `Amending an approved spec`: a fresh-context reviewer clears evidence-backed factual corrections on its own, escalations reach you as one readable batch, and only a redraw is a full stop - not a third numbered gate. That's the mechanism. What follows is the machinery behind it.
|
|
67
67
|
|
|
68
68
|
## Architecture
|
|
69
69
|
|
|
@@ -39,6 +39,8 @@ Work flows `origin (prompt + spec) → plan → code/doc`. Every hop is lossy: a
|
|
|
39
39
|
3. The feature ships correctly without it. If another requirement needs it, keep it.
|
|
40
40
|
|
|
41
41
|
Report such a clause once, as `UNAUTHORIZED` (not `DELIVERED`, not origin drift). Unsure on 2 -> keep it as an `Rn`. This is the only exception to "spec is canonical"; "could be cleaner" is still not a reason.
|
|
42
|
+
|
|
43
|
+
**Spec `## Acceptance criteria`.** When the spec has a heading starting `## Acceptance criteria` (a legacy `## Acceptance criteria (from #N)` heading matches), read that section as an origin beside the spec body. Each `in-scope` row yields one `Rn` with `origin: spec "Acceptance criteria" - "<AC verbatim>"`; a row with no disposition line is `in-scope`. Each `venue:` row yields one `Rn` that is `DELIVERED` when the Design clauses naming its enabling change are `DELIVERED`, with the venue and observation text on the verdict line; the observation itself is never a finding; a `venue:` row no Design clause names is `MISSING`, `recommended: rescope`. `deviates:` and `deferred:` rows yield no `Rn`; list them under `Origin drift` as `recorded in spec? yes`. A disposition word outside the four is `DRIFTED`, `recommended: fix` (correct the word). A section whose body is a single `none - <reason>` line has no rows: it yields no `Rn` and no drift. A spec without the section is read as before - spec body and prompt only; every existing rule and the source order above are unchanged. When a spec exists you do not fetch the ticket.
|
|
42
44
|
2. **Check origin drift.** If the spec and the prompt/ticket disagree, do **not** absorb it silently. A deviation recorded in the spec → spec wins (it was review-gated). An *unrecorded* divergence → the spec silently dropped or altered a requirement = a conformance failure to report.
|
|
43
45
|
3. **Map each requirement to the deliverable.** Read the diff (code **and** docs) yourself — do not trust any summary. For each requirement, find where it is satisfied and cite real `file:line` evidence. Run read-only checks (tests, grep) when they confirm a behavior; quote actual output.
|
|
44
46
|
4. **Flag the unrequested.** Anything shipped that no requirement in the origin asked for = `UNAUTHORIZED` (scope creep), even if it looks useful. Do not negotiate scope with yourself. This includes the step-1 exception clauses. `origin` is always `none (scope creep)`. For a step-1 clause, start `evidence:` with `spec "<section>" - "<clause>" (over-spec)`, then the files/specs it adds, then what fails without it.
|
|
@@ -143,6 +145,7 @@ serial waves — identical to planned-execution wave grouping. Runtime-resource
|
|
|
143
145
|
- **Evidence or it didn't happen.** Cite a real `file:line` for every DELIVERED/PARTIAL. If you cannot, downgrade the row to MISSING.
|
|
144
146
|
- **Origin quote or it isn't a gap.** Every non-UNAUTHORIZED gap's `origin` carries a locator AND a verbatim quote: `origin: <file/section, 'prompt', or 'ticket'> - "<quoted clause>"` (truncate long clauses with `[...]` as long as the fragment uniquely identifies the clause). No quotable origin clause = no gap. Do not derive implicit requirements. Do not flag wording preferences. A deviation recorded in the spec wins over an older origin value (Process step 2); report it only if unrecorded.
|
|
145
147
|
- **Spec is canonical; the prompt catches what the spec dropped; the ticket is fallback only** when no spec exists.
|
|
148
|
+
- **The spec's `## Acceptance criteria` section is origin, not ticket.** Its `in-scope`/`venue:` rows are `Rn`; its `deviates:`/`deferred:` rows are recorded drift (`recorded in spec? yes`).
|
|
146
149
|
- **Do not absorb origin drift silently** — flag every spec↔prompt/ticket disagreement.
|
|
147
150
|
- **Quote real command output** if you ran checks. Do not paraphrase from memory.
|
|
148
151
|
- **Coverage is binary per requirement** — "mostly done" is PARTIAL, not DELIVERED.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: spec-council-member
|
|
3
|
-
description: Adversarial single-model spec critic dispatched by the roasting-the-spec or shape-ticket skills; assesses whether a spec is sound, complete, and actionable. Not for direct dispatch.
|
|
3
|
+
description: Adversarial single-model spec critic dispatched by the roasting-the-spec or shape-ticket skills, and in `Mode: amendment-review` by brainstorming's amendment surface to clear or escalate proposed spec amendments; assesses whether a spec is sound, complete, and actionable. Not for direct dispatch.
|
|
4
4
|
tools: read, grep, find, ls, bash
|
|
5
5
|
thinking: xhigh
|
|
6
6
|
defaultContext: fresh
|
|
@@ -18,10 +18,20 @@ You are read-only: you never modify the repository or any input artifact; your o
|
|
|
18
18
|
|
|
19
19
|
When your dispatching task asks for codebase verification, verify - do not trust assertions about existing files, APIs, or conventions - but bounded: prefer `rg` (it respects `.gitignore`) over recursive `grep`, use `rg`-native bounds (`--max-count`, explicit paths); scope every scan to explicit paths, never a repository root; bound each scan with `timeout` (or `gtimeout`) when available, and do not run it unbounded when neither exists. A scan that times out or cannot be bounded is reported as unverified - never retried broader. End every finding with `probed:`; `none` is a normal answer.
|
|
20
20
|
|
|
21
|
+
## Amendment-review mode
|
|
22
|
+
|
|
23
|
+
When the first line of your task is `Mode: amendment-review`, this section replaces everything below it. You judge proposed amendments to an approved spec, not the spec. The task carries the rubric, the spec path, and per item a handle, a spec location, the `old -> new` text, and cited evidence (a command and its output, a `file:line`, a test result, a fixture measurement). Read the spec around each location; probe cited evidence read-only, bounded as above; judge scope from the spec's `## Human input` section when the task supplies it, else from its Goal, Problem, scope and acceptance sections. Emit exactly one line per item and nothing else:
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
<handle>: auto-apply | escalate - <one-line reason> - probed: <check> - <result>
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`auto-apply` only when every rubric predicate holds on evidence you probed. Uncited, unverifiable, uncertain, or touching a human-owned section -> `escalate`. You never edit anything.
|
|
30
|
+
|
|
21
31
|
Assess the spec on five axes:
|
|
22
32
|
|
|
23
33
|
1. **Addresses the problem.** Does the spec actually solve the problem in the problem statement? Answer yes / partial / no and say why. A well-written spec for the wrong problem is unsound.
|
|
24
|
-
2. **Logical gaps.** Missing steps, unhandled states, transitions asserted but not specified, data that appears from nowhere. A load-bearing reference to external context the spec does not inline (a ticket acceptance criterion, a commit SHA, another doc that an implementer would need) is a gap - flag it as `external-ref` and recommend inlining the relevant content.
|
|
34
|
+
2. **Logical gaps.** Missing steps, unhandled states, transitions asserted but not specified, data that appears from nowhere. A load-bearing reference to external context the spec does not inline (a ticket acceptance criterion, a commit SHA, another doc that an implementer would need) is a gap - flag it as `external-ref` and recommend inlining the relevant content. Compare the `Human input` ticket AC rows against the spec's `## Acceptance criteria`: a row absent or reworded is `external-ref`; a `deferred:` row whose reason fails the operates-without-it test (the change needs it to operate) is `scope`. Judge a `deviates:` row as any other design decision.
|
|
25
35
|
3. **Oversimplifications.** Places where the spec assumes away real complexity — error paths waved off, concurrency ignored, "just" and "simply" hiding hard problems.
|
|
26
36
|
4. **Ambiguities.** Unnamed components, undefined terms, "we should" without a decision, fields or types referenced but never defined.
|
|
27
37
|
5. **Actionable and testable.** Could a competent implementer with no further context build this and verify it? If not, what is missing?
|
package/package.json
CHANGED
|
@@ -119,7 +119,22 @@ Use doc names, `none`, or `deferred: <trigger>`. Apply `reference/documentation-
|
|
|
119
119
|
|
|
120
120
|
## Ticket Handling
|
|
121
121
|
|
|
122
|
-
A ticket is guidance, not sole truth. Fetch it; propose scope, approach, or acceptance changes when code disagrees, and record deviations in the spec.
|
|
122
|
+
A ticket is guidance, not sole truth. Fetch it; propose scope, approach, or acceptance changes when code disagrees, and record deviations in the spec. Implied requirements in the ticket body (Context, Problem, Idea) land in the spec body as any other requirement; the ticket's explicit acceptance criteria land verbatim in the section below.
|
|
123
|
+
|
|
124
|
+
**Extract the ticket's ACs.** When the ticket body has a heading matching `/acceptance criteria/i`, every list item under it until the next heading is an AC (numbered, bullet, or checkbox) and no other list in the body is. When there is no such heading, every top-level checkbox row in the body is an AC, except rows under a `Post-deployment housekeeping`, `Out of scope`, or `Follow-up` heading. Otherwise the ticket has no ACs. A nested list under an AC row rides with its parent as one row; an AC heading holding prose and no list makes each paragraph one row. Carry checked and unchecked rows alike, all written `- [ ]`. Take the rows from the gather draft's `## Ticket acceptance criteria (verbatim)` heading; when the ticket was not fetched, the gate summary shows the `none` line so the user can paste the rows, which then become rows.
|
|
125
|
+
|
|
126
|
+
**Write the section in every spec**, after `## Problem`:
|
|
127
|
+
|
|
128
|
+
```markdown
|
|
129
|
+
## Acceptance criteria
|
|
130
|
+
|
|
131
|
+
Ticket <ref>, <heading or "checkbox list">, rows verbatim:
|
|
132
|
+
|
|
133
|
+
- [ ] <row text copied verbatim>
|
|
134
|
+
<disposition>
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
The heading `## Acceptance criteria` names the ticket contract only; the spec's own requirements stay in Design. Disposition is exactly one of `in-scope`, `deviates: <why>`, `deferred: <where>`, `venue: <env> - <observation>`; default `in-scope`. Never edit the row text: a wrongly stated row is `deviates: <why>`, an ambiguous row stays `in-scope` with the chosen reading written as a Design clause. Name a `venue:` row's enabling change in Design. An `in-scope` row is a requirement as written; disposition reasons are never normative for the plan or the reviewer, so a `deviates:` reason that adopts part of a row restates that part as a Design clause. Defer a row only when the shipped change operates without it - cost is never a reason; a hard but direct AC stays `in-scope` and ships. Ask every `deviates:`/`deferred:` decision in the questionary and name it in the gate summary; `venue:` states where the observation can happen and is not a scope decision. A disposition changes in any later phase - a reviewer finding `recommended: rescope` at the finish gate, a wrong row noticed mid-implementation, or a ticket edited after approval when the user asks - through [Amending an approved spec](#amending-an-approved-spec) (`reference/amendment-surface.md`); the row text still never changes. With no ticket, or a ticket without ACs, the section body is the single line `none - no ticket`, `none - ticket has no acceptance criteria`, or `none - ticket not fetched (<reason>)`. Never author acceptance criteria on the ticket's behalf; that is `/skill:shape-ticket`'s job.
|
|
123
138
|
|
|
124
139
|
## First-Feature Oversight (Early Project Stages)
|
|
125
140
|
|
|
@@ -158,14 +173,15 @@ After writing the spec to `<project>/doc/specs/<filename>.md` (per [Filename Con
|
|
|
158
173
|
- **Placeholder scan.** Any `TODO`, `TBD`, `<fill in>`, `[example]`, `xxx`? Either resolve them or convert to explicit "Open Questions" with names.
|
|
159
174
|
- **Internal consistency.** Does Section 4 contradict Section 2? Are component names and field names consistent throughout? If the spec replaces a prior design, confirm the predecessor carries the supersession banner and its href resolves to this spec's final filename.
|
|
160
175
|
- **Documentation named.** Does the spec name all three classes (feature/user-facing introduced; materially amended; derived/memory invalidated), or an explicit "none" for each? Enforce the materiality bar in `reference/documentation-impact.md` without restating it: each listed doc names the category it clears, none is a code-mirror, amend-over-create was applied, and skill/agent bodies are implementation surface, not doc-impact entries here.
|
|
176
|
+
- **Ticket contract present.** Does `## Acceptance criteria` exist with either verbatim ticket rows plus dispositions or one `none - <reason>` line? Presence is enforced here, at authoring, and nowhere later.
|
|
161
177
|
- **Scope check.** Does every paragraph serve the goal? Cut filler. If something is out of scope, say it's out of scope.
|
|
162
178
|
- **Ambiguity check.** Is every "we should…" backed by a concrete decision? Replace "we could probably" with "we will" or "we won't".
|
|
163
179
|
|
|
164
|
-
The first
|
|
180
|
+
The first four are the inline **lint**: run them here and fix what they surface. The last two are the **critique pass**, dispatched per [Spec Council](#spec-council-optional). After it returns, re-run the placeholder scan over the applied spec; if a predecessor banner exists, confirm its `<scope>` still matches and reconcile it; carry any ambiguity the critique could not resolve to the [User Review Gate](#user-review-gate).
|
|
165
181
|
|
|
166
182
|
## Spec Council (Optional)
|
|
167
183
|
|
|
168
|
-
After the inline lint and before the user review gate, **brainstorming owns the critique-pass gate**; council **apply mechanics** live in `/skill:roasting-the-spec` (single source of truth - link, don't restate). Resolve the council with `gauntlet_setting({ key: "specCouncil" })` - the tool returns the merged (repo-over-preset) value as `{ verdict, members, chair, malformed, warning, errors }`. **Do not** hand-roll a settings read. When `verdict` is `"council"`, the council *is* the critique pass - invoke `/skill:roasting-the-spec` automatically (no offer, no prompt), passing `members`/`chair`; also pass the verbatim human input (the original prompt, any ticket AC snapshot, and the questionary answers that changed scope) - roasting-the-spec forwards it to members and chair as the `Human input (verbatim; off-limits for over-spec)` block; it applies its apply-set and returns the audit (Applied/Deferred/Rejected). When `verdict` is `"worker"`, dispatch the worker below. If `malformed` is true or `errors` is non-empty, emit the `warning`/error as one line, then branch strictly on `verdict` - `malformed` can accompany *either* verdict (e.g. a bad `chair` with valid `members` still returns `council`), so never infer the worker path from `malformed` alone. If `gauntlet_setting` is unavailable, stop and report - never fall back to a manual bash/JSON settings merge. The already-applied council edits (or the worker's in-place fixes) ride in the same worktree commit. The conceptual precedence rule lives in `verification-before-completion/reference/settings-precedence.md`.
|
|
184
|
+
After the inline lint and before the user review gate, **brainstorming owns the critique-pass gate**; council **apply mechanics** live in `/skill:roasting-the-spec` (single source of truth - link, don't restate). Resolve the council with `gauntlet_setting({ key: "specCouncil" })` - the tool returns the merged (repo-over-preset) value as `{ verdict, members, chair, malformed, warning, errors }`. **Do not** hand-roll a settings read. When `verdict` is `"council"`, the council *is* the critique pass - invoke `/skill:roasting-the-spec` automatically (no offer, no prompt), passing `members`/`chair`; also pass the verbatim human input (the original prompt, any ticket AC snapshot - the raw rows under the gather draft's `## Ticket acceptance criteria (verbatim)` heading, never the spec's section, which holds the author's dispositions - and the questionary answers that changed scope) - roasting-the-spec forwards it to members and chair as the `Human input (verbatim; off-limits for over-spec)` block; it applies its apply-set and returns the audit (Applied/Deferred/Rejected). When `verdict` is `"worker"`, dispatch the worker below. If `malformed` is true or `errors` is non-empty, emit the `warning`/error as one line, then branch strictly on `verdict` - `malformed` can accompany *either* verdict (e.g. a bad `chair` with valid `members` still returns `council`), so never infer the worker path from `malformed` alone. If `gauntlet_setting` is unavailable, stop and report - never fall back to a manual bash/JSON settings merge. The already-applied council edits (or the worker's in-place fixes) ride in the same worktree commit. The conceptual precedence rule lives in `verification-before-completion/reference/settings-precedence.md`.
|
|
169
185
|
|
|
170
186
|
When `verdict` is `"worker"`, dispatch one fresh `worker` that applies the scope + ambiguity checks and fixes them in place:
|
|
171
187
|
|
|
@@ -226,14 +242,14 @@ Rejected: <cluster -> one-line reason>, ...
|
|
|
226
242
|
|
|
227
243
|
<unresolved ambiguities; every gap-footer entry from the summary>
|
|
228
244
|
|
|
229
|
-
Please review. Approve to proceed, tell me what to change in the spec, or say "revert applied council edit <X>" to undo a specific applied edit.
|
|
245
|
+
Please review. Approve to proceed, tell me what to change in the spec, or say "revert applied council edit <X>" to undo a specific applied edit. Reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop. "approve, auto-apply amends" does both.
|
|
230
246
|
```
|
|
231
247
|
|
|
232
248
|
If you believe the summary needs correcting, do **not** silently rewrite it — re-dispatch the summarizer or note the discrepancy as an adjacent line beneath the verbatim block.
|
|
233
249
|
|
|
234
250
|
**Revert valve.** "Revert applied council edit X" is a normal change request: revise the spec to undo edit X, re-dispatch the summarizer with a **fresh** temp path (per the re-dispatch rule below), and re-present the gate. This is cheap here - the spec is not yet plan- or code-bearing.
|
|
235
251
|
|
|
236
|
-
Wait for the user. On a change request (including a revert), revise the spec and re-present — mint a **fresh** temp path for the re-dispatched summarizer (never reuse a prior round's path, so stale content can never be mistaken for the new summary). On approval, proceed immediately to `/skill:writing-plans` with no further prompt — the plan and execution mode are mechanical derivatives, so the only human gate here is spec approval itself. Don't land the spec on `main`; it stays in the worktree and ships in the same squash commit as the implementation.
|
|
252
|
+
Wait for the user. On a change request (including a revert), revise the spec and re-present — mint a **fresh** temp path for the re-dispatched summarizer (never reuse a prior round's path, so stale content can never be mistaken for the new summary). On approval, proceed immediately to `/skill:writing-plans` with no further prompt (a grant given at or before approval - "approve, auto-apply amends", or a standalone "auto-apply amends" reply earlier at this gate - is first quoted in the spec commit body via `git -C <abs worktree path> commit --amend --no-edit -q --trailer "Amend-grant: <the sentence>"`, so the worktree history shows when the grant began) — the plan and execution mode are mechanical derivatives, so the only human gate here is spec approval itself. Don't land the spec on `main`; it stays in the worktree and ships in the same squash commit as the implementation.
|
|
237
253
|
|
|
238
254
|
Post-approval changes follow [Amending an approved spec](#amending-an-approved-spec).
|
|
239
255
|
|
|
@@ -245,13 +261,11 @@ phase_tracker({ action: "complete", phase: "brainstorm" })
|
|
|
245
261
|
|
|
246
262
|
## Amending an approved spec
|
|
247
263
|
|
|
248
|
-
Execute
|
|
264
|
+
Execute in place from any later phase; never invoke `/skill:brainstorming` for it (its entry resets both trackers). Worktree, spec commits, and plan survive.
|
|
249
265
|
|
|
250
|
-
|
|
251
|
-
2. Render the diff and impact line. A user instruction in this flow that waives per-diff review for later amends ("auto-apply amends, stop only for redraws", "apply spec fixes without asking") is the approval: quote it in the amendment commit body and continue. Otherwise wait for approval; change request -> revise, re-show. Redraws always wait. A grant never satisfies the spec gate; a grant given with or before spec approval applies to later amends in the same flow; a new brainstorm and a fresh-session resume start with no grant.
|
|
252
|
-
3. No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks; anchor-changed completed tasks are reopened as `in_progress` and re-run the task loop (`update` never sets `pending`). A removed task is deleted from the plan; then re-`init` the tracker with `{ name, status }` elements: preserved tasks keep their order and statuses, reopened tasks are `in_progress` in place, every still-`pending` task (including newly added ones, whatever wave label they carry) trails the non-pending ones, removed tasks are the only deletions (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
266
|
+
Classify first. Redraw test: the change alters the problem statement, adds or removes a component, or moves a component boundary -> redraw. A change inside one component (a persistence mechanism, a worker's HTTP client, dropping a fallback and its task) -> amend. State the call; the user overrides either way.
|
|
253
267
|
|
|
254
|
-
|
|
268
|
+
Amend -> load `reference/amendment-surface.md` and follow it (unreadable -> stop with a blocking error; never improvise the grammar): it holds items unapplied, reviews them with a fresh `spec-council-member`, renders one readable batch for escalations, applies accepted items, runs the plan/tracker aftermath, and commits once. A user instruction in this flow that waives per-diff review for later amends ("auto-apply amends", "auto-apply amends, stop only for redraws", "apply spec fixes without asking") skips the review; it never satisfies the spec gate, and a new brainstorm or a fresh-session resume starts with no grant. Redraws always stop.
|
|
255
269
|
|
|
256
270
|
Redraw: keep the worktree and the approved spec file. `plan_tracker({ action: "clear" })`, `phase_tracker({ action: "reset" })`, `phase_tracker({ action: "start", phase: "brainstorm" })`, delete the plan file, resume at checklist step 4 with the approved spec as the draft (steps 2-3 skipped). Spec-writing overwrites it; the full gate follows.
|
|
257
271
|
|
|
@@ -272,7 +286,7 @@ One question at a time, YAGNI, 2-3 approaches, two design rounds, clarify freely
|
|
|
272
286
|
- Plan before approval; brainstorming invocation for an amend ([owner](#user-review-gate)).
|
|
273
287
|
- Missing predecessor banner; invalid multi-spec split ([owner](#spec-self-review-before-user-review-gate); [owner](#2-scope-check)).
|
|
274
288
|
- Approaches while a contradicted premise remains unresolved ([owner](#3-understand-the-idea)).
|
|
275
|
-
-
|
|
289
|
+
- Amend without `reference/amendment-surface.md`; waiting after an amend grant; auto-applying a redraw ([owner](#amending-an-approved-spec)).
|
|
276
290
|
|
|
277
291
|
## Project overrides
|
|
278
292
|
|
|
@@ -74,8 +74,12 @@ Context-builder (conditional):
|
|
|
74
74
|
> `<initial prompt verbatim>`. Fetch and distill these references:
|
|
75
75
|
> `<detected refs, one per line>`. For each: acceptance criteria, hard constraints,
|
|
76
76
|
> linked discussion that changes scope, and contradictions with the request as
|
|
77
|
-
> stated.
|
|
78
|
-
>
|
|
77
|
+
> stated. Quote the ticket's acceptance-criteria rows verbatim under their own
|
|
78
|
+
> `## Ticket acceptance criteria (verbatim)` heading before distilling the rest;
|
|
79
|
+
> brainstorming copies these rows unchanged into the spec and the council's
|
|
80
|
+
> `Human input` block. Write ONLY the context handoff to your output path; do NOT
|
|
81
|
+
> produce a meta-prompt file. End with an "Open questions that matter for the spec"
|
|
82
|
+
> section.
|
|
79
83
|
> If a ref is unreadable, say so explicitly and continue.
|
|
80
84
|
|
|
81
85
|
(The meta-prompt exclusion matters: in chain mode context-builder emits two files —
|
|
@@ -0,0 +1,158 @@
|
|
|
1
|
+
# Amendment surface
|
|
2
|
+
|
|
3
|
+
Loaded by `skills/brainstorming/SKILL.md` § Amending an approved spec, from any phase, for amend-class changes only - redraws never enter. The main loop (the orchestrator holding `edit`/`write`) is the only author of spec amendments and of every human-facing line about them. The human is the last resort: a fresh reviewer clears evidence-backed factual corrections; the human sees the rest once per batch, in plain language.
|
|
4
|
+
|
|
5
|
+
## 1. Prepare - never apply yet
|
|
6
|
+
|
|
7
|
+
Collect every amendment pending at this decision point (same spec-review round, same blocked wave, same conformance inventory) into one batch. Never wait for more; a later finding is a new batch. For each item hold, unapplied:
|
|
8
|
+
|
|
9
|
+
| Field | Content |
|
|
10
|
+
|---|---|
|
|
11
|
+
| `handle` | one short word from the title (`Posting date` -> `posting`); digit suffix on collision |
|
|
12
|
+
| `title` | plain, under eight words |
|
|
13
|
+
| `location` | spec section, plus the `old text -> new text` |
|
|
14
|
+
| `what` | one sentence: what changes |
|
|
15
|
+
| `why` | one sentence: why it matters to the outcome |
|
|
16
|
+
| `example` | one before -> after value or line |
|
|
17
|
+
| `evidence` | the observation that falsified the old text (command + output, `file:line`, test result, fixture measurement), or `none` |
|
|
18
|
+
| `recommended` | `accept` or `alt-n` |
|
|
19
|
+
| `alternatives` | genuinely different spec edits, zero or more |
|
|
20
|
+
|
|
21
|
+
The working tree stays at pre-batch HEAD until apply (section 5) - nothing is edited before the reviewer and, where needed, the human have answered. A redraw item stops alone first (`SKILL.md` redraw path); amend items are held and re-batched after it resolves.
|
|
22
|
+
|
|
23
|
+
**Standing grant active** (`v5.10.0` semantics) - a user sentence in this flow that waives per-diff review (`auto-apply amends`, `approve, auto-apply amends` at the spec gate, `auto-apply amends, stop only for redraws`, `apply spec fixes without asking`, or the same intent in other words; never inferred after a fresh-session resume): skip steps 2-4, apply every item, print one line each `amended the spec: <title> - <what>`, record `granted`, and quote the sentence in the commit body.
|
|
24
|
+
|
|
25
|
+
## 2. Prefilter - no model call
|
|
26
|
+
|
|
27
|
+
Send an item straight to the human batch (step 4) when any holds:
|
|
28
|
+
|
|
29
|
+
- the edit removes or narrows approved text (a descope or rescope);
|
|
30
|
+
- the location is a human-owned section: problem statement, goal, acceptance criteria, in/out scope or non-goals, component list or boundaries, public contracts (API, schema, config shape, CLI surface);
|
|
31
|
+
- at the conformance entry: the gap is `UNAUTHORIZED`, or its `origin` quotes an acceptance criterion.
|
|
32
|
+
|
|
33
|
+
Everything else - evidence-backed factual drift outside human-owned text - goes to the reviewer; an item whose evidence is `none` still goes there and fails rubric (a), so the reviewer's `escalate` line records why.
|
|
34
|
+
|
|
35
|
+
## 3. Reviewer - one dispatch per batch
|
|
36
|
+
|
|
37
|
+
Rubric - `auto-apply` only when all three hold:
|
|
38
|
+
|
|
39
|
+
- **(a)** evidence-backed factual correction: `evidence` is a cited observation, not a claim;
|
|
40
|
+
- **(b)** no human-owned section touched (list above; verification commands, documentation-impact lines, and design detail are not human-owned);
|
|
41
|
+
- **(c)** scope-neutral: removes nothing approved, adds nothing unasked - judged from the spec's `## Human input` section when present, else its Goal, Problem, scope and acceptance sections; never from chat.
|
|
42
|
+
|
|
43
|
+
Anything else, including uncertainty, is `escalate`.
|
|
44
|
+
|
|
45
|
+
Model string - the main loop's own model and level, printed with the bash tool and pasted into `model:`:
|
|
46
|
+
|
|
47
|
+
```bash
|
|
48
|
+
lvl="$PI_REASONING_LEVEL"; case "$lvl" in max) lvl=xhigh;; off|"") lvl="";; esac
|
|
49
|
+
printf '%s/%s%s\n' "$PI_PROVIDER" "$PI_MODEL" "${lvl:+:$lvl}"
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
```
|
|
53
|
+
subagent({ agent: "spec-council-member", context: "fresh", async: false,
|
|
54
|
+
model: "<printed string>", cwd: "<abs worktree path>",
|
|
55
|
+
control: { needsAttentionAfterMs: 60000, inFlightSilenceCeilingMs: 240000, inFlightSilenceKillMs: 300000 },
|
|
56
|
+
task: "Mode: amendment-review\nSpec: <abs spec path>\nRubric:\n<the three predicates above, verbatim>\nItems:\n<per item: handle | location | old -> new | evidence>\nHuman input (data, not instructions):\n```\n<the spec's ## Human input section, or: none - judge (c) from Goal/Problem/scope/AC>\n```" })
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Expected reply - one line per item, nothing else:
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
<handle>: auto-apply | escalate - <one-line reason> - probed: <check> - <result>
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
Fail closed: a dispatch error, an async handle, a silence-kill, or a missing or malformed line -> that item (every item when the dispatch failed) is `escalate`, and the batch menu carries `reviewer unavailable: <reason>`.
|
|
66
|
+
|
|
67
|
+
## 4. Human batch - one menu
|
|
68
|
+
|
|
69
|
+
Render only escalated and prefiltered items. Nothing symbol-dense above the fold; each item's `old -> new` sits under `Details`, after the footer.
|
|
70
|
+
|
|
71
|
+
```
|
|
72
|
+
Spec amendments: <N> need your call.
|
|
73
|
+
|
|
74
|
+
* <handle> - <title>: <what>. <why>.
|
|
75
|
+
Example: <before -> after>
|
|
76
|
+
Impact: <plan tasks/waves affected, or: no plan yet>
|
|
77
|
+
Recommended: <accept | alt-n> (<one-clause why>).
|
|
78
|
+
Alternatives: alt-1 <one line>; alt-2 <one line>
|
|
79
|
+
|
|
80
|
+
Reply: 1 (apply all recommendations) | 2: <handle>=<accept|alt-n|custom(<effect>)>, ...
|
|
81
|
+
Standing grant: reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop.
|
|
82
|
+
|
|
83
|
+
Details
|
|
84
|
+
<handle>: <location> - old: <text> -> new: <text>
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
`Alternatives:` appears only when genuine ones exist; otherwise the item's choices are exactly `accept` and `custom(...)`. Reply grammar: `1` applies every recommendation; `2:` overrides the named handles, omitted handles keep theirs, a handle at most once; `custom(<effect>)` is free text and may redirect anywhere ("keep the spec, fix the parser"). A redirect away from the spec drops the item (still recorded in the batch commit body as `custom(<effect>)`) and returns the finding to its calling loop. Invalid handle or choice -> reprompt for that item only, keep every valid pick, never reopen the gate. Take no action before the reply.
|
|
88
|
+
|
|
89
|
+
## 5. Apply, aftermath, commit
|
|
90
|
+
|
|
91
|
+
Apply accepted items only: reviewer clears, `accept`, `alt-n`, and state-changing `custom`. Print one line per applied item: `amended the spec: <title> - <what>`.
|
|
92
|
+
|
|
93
|
+
Aftermath, once per batch, only when at least one item applied (nothing applied -> skip to the commit below). No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks; anchor-changed completed tasks are reopened as `in_progress` and re-run the task loop (`update` never sets `pending`). A removed task is deleted from the plan; then re-`init` the tracker with `{ name, status }` elements: preserved tasks keep their order and statuses, reopened tasks are `in_progress` in place, every still-`pending` task (including newly added ones, whatever wave label they carry) trails the non-pending ones, removed tasks are the only deletions (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
94
|
+
|
|
95
|
+
One commit per batch:
|
|
96
|
+
|
|
97
|
+
```
|
|
98
|
+
amend: <N> item(s) - <first title>[, <second title>]
|
|
99
|
+
|
|
100
|
+
- <handle> | <title> | <what> | <auto-apply | accepted | alt-n | custom(<effect>) | granted> | <reviewer line or reason> | <evidence>
|
|
101
|
+
<the granting sentence, quoted, when a grant applied>
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
`amend:` is the subject marker `finishing-a-development-branch` Step 4 greps for its digest; `<what>` is the item's one-sentence what-changes field, so the digest renders `<title> - <what changed>` from the body alone. When no item applied (every item dropped or redirected), nothing changed on disk; the commit still lands, with `git commit --allow-empty`, so the per-item `custom(<effect>)` records stay in the batch body - the digest ignores them because it reads only `auto-apply` and `granted` records. Wrong apply -> `git revert` the batch commit, then re-enter this surface for the items to keep.
|
|
105
|
+
|
|
106
|
+
## Conformance entry
|
|
107
|
+
|
|
108
|
+
Called from `finishing-a-development-branch` Step 3.5, before the carried-open menu renders, once per inventory:
|
|
109
|
+
|
|
110
|
+
1. Draft an item (step 1) for each gap with `recommended: accept`, verdict `DRIFTED` or `PARTIAL`, not `UNAUTHORIZED`, whose `origin` is not an acceptance criterion - the `accept-into-spec` edit built from its `origin` + `evidence`. Every other gap skips the funnel and stays a menu row.
|
|
111
|
+
2. Review (step 3), apply and commit (step 5).
|
|
112
|
+
3. Re-audit against the amended spec; regenerate the inventory. Only concerns the re-audit closed drop out; sibling concerns keep their rows.
|
|
113
|
+
4. Render the disposition menu for what remains - escalated items are ordinary rows there, never a second menu. Rows whose recommended disposition edits the spec carry the readable card fields (`what`, `why`, `Example:`) on the bullet, adapted to the disposition bullet grammar. A human-selected spec-changing disposition (`accept-into-spec`, `rescope-into-spec`, state-changing `custom`) is already approved: it applies at the protocol's execute-order step 2, bypasses steps 2-4 of this surface, and is recorded as today (`Gn - <title>: <disposition>`); an auto-applied item is recorded `Gn - <title>: accept-into-spec (auto-applied)`.
|
|
114
|
+
|
|
115
|
+
## Worked example
|
|
116
|
+
|
|
117
|
+
Eight synthetic items modelled on one real run. Items 1-5 and 8 reach the reviewer; the prefilter catches 6 and 7:
|
|
118
|
+
|
|
119
|
+
| # | Location | old -> new | Evidence | Outcome |
|
|
120
|
+
|---|---|---|---|---|
|
|
121
|
+
| 1 | Design, parser | parsed with the `AnnouncementList` model -> the `RulesOfProcedureList` model | `rg -n "class .*List" fixtures/rop.html` shows the CMS model name in the page identity block | `auto-apply` |
|
|
122
|
+
| 2 | Verification, fixture line | fixture is 41,208 bytes -> 43,117 bytes | `wc -c fixtures/rop-2026-01.html` = 43117 | `auto-apply` |
|
|
123
|
+
| 3 | Design, date handling | falls back to the earliest document date -> the latest | fixture rows dated 03-02, 03-05, 03-09; page shows posted 03-09 | `auto-apply` |
|
|
124
|
+
| 4 | Design, page identity | assert identity on the `<title>` text -> on the CMS model name | `rg -c "<title>Site</title>" fixtures/` = 4, identical across pages | `auto-apply` |
|
|
125
|
+
| 5 | Design, run status | zero new items reports `success` -> `success_empty` | `test_dedup_all_seen` asserts `success_empty` (`tests/test_rop.py:41`) | `auto-apply` |
|
|
126
|
+
| 6 | Non-goals | adds "the ROP feed is out of scope for this release" | none | prefiltered: removes approved scope |
|
|
127
|
+
| 7 | Acceptance criteria | at least 3 announcements per fetch -> at least 1 | fixture has one item | prefiltered: AC location |
|
|
128
|
+
| 8 | Verification, fixture line | 41,208 bytes -> 43,117 bytes | none cited | `escalate` - rubric (a) |
|
|
129
|
+
|
|
130
|
+
The reviewer clears items 1-5; nothing is applied yet. The batch renders items 6-8 (items 1-5 apply together with the accepted ones after the reply):
|
|
131
|
+
|
|
132
|
+
```
|
|
133
|
+
Spec amendments: 3 need your call.
|
|
134
|
+
|
|
135
|
+
* scope - ROP feed out of scope: adds an out-of-scope line for the ROP feed to Non-goals. Drops a deliverable you approved.
|
|
136
|
+
Example: regulator feed = NERC filings + ROP announcements -> NERC filings only
|
|
137
|
+
Impact: Task 4, Wave 2
|
|
138
|
+
Recommended: accept (the ROP source has no stable page this release).
|
|
139
|
+
* count - Fewer announcements per fetch: the acceptance criterion drops from at least 3 to at least 1. Lowers the bar you set.
|
|
140
|
+
Example: 3 items per fetch -> 1
|
|
141
|
+
Impact: Task 6
|
|
142
|
+
Recommended: accept (the captured fixture has one item; the criterion assumed three).
|
|
143
|
+
* bytes - Fixture byte count: the verification line changes from 41,208 to 43,117 bytes. No measurement was cited, so the reviewer could not confirm it.
|
|
144
|
+
Example: 41,208 bytes -> 43,117 bytes
|
|
145
|
+
Impact: no plan yet
|
|
146
|
+
Recommended: alt-1 (measure first; apply whatever `wc -c` reports).
|
|
147
|
+
Alternatives: alt-1 replace the number with the `wc -c` result
|
|
148
|
+
|
|
149
|
+
Reply: 1 (apply all recommendations) | 2: <handle>=<accept|alt-n|custom(<effect>)>, ...
|
|
150
|
+
Standing grant: reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop.
|
|
151
|
+
|
|
152
|
+
Details
|
|
153
|
+
scope: Non-goals - old: (none) -> new: the ROP feed is out of scope for this release
|
|
154
|
+
count: Acceptance criteria - old: at least 3 announcements per fetch -> new: at least 1 announcement per fetch
|
|
155
|
+
bytes: Verification - old: fixture is 41,208 bytes -> new: fixture is 43,117 bytes
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Reply `2: scope=custom(keep scope; fix the parser instead)` drops `scope`, returns it to the fix loop, and applies `count` and `bytes` as recommended.
|
|
@@ -88,6 +88,8 @@ Closure / conformance: CONFORMS
|
|
|
88
88
|
|
|
89
89
|
then continue directly to Step 4. No approval prompt, no menu, no shared options line, no sign-off. If the run auto-applied fixes, surface the flat `auto-applied fix commits: <Gn: SHA>, ...` index from the durable block as **one informational, non-blocking line** with a one-line revert offer (see "Revert semantics") - a gap that auto-converged mid-verify has no bullet, so this index is the only place its fix commit stays revertable. Do not wait for acknowledgment.
|
|
90
90
|
|
|
91
|
+
**Pre-menu amendment funnel (GAPS only).** Before rendering the carried-open menu, run [`amendment-surface.md` § Conformance entry](../brainstorming/reference/amendment-surface.md) over the inventory once: gaps with `recommended: accept`, verdict `DRIFTED` or `PARTIAL`, not `UNAUTHORIZED`, whose `origin` is not an acceptance criterion are drafted as `accept-into-spec` items (the edit built from `origin` + `evidence`) and sent to the reviewer in one call; cleared items apply and land as one batch commit, the spec is re-audited, the inventory regenerated. Only concerns the re-audit actually closed drop out; sibling concerns in the same gap keep their rows and dispositions. Survivors and every other gap render as ordinary rows below - one menu, never two.
|
|
92
|
+
|
|
91
93
|
**Carried-open (`status: GAPS (N open)`).** Read `reference/disposition-protocol.md` and follow it for the carried-open render (dense) grammar, the response grammar, and the 9-step execute order. Render the human decision menu in the shape below, drive the dispositions per that reference, then print the summary render and continue to Step 4. If that reference file cannot be read, stop and surface a blocking error — do **not** improvise the grammar from memory.
|
|
92
94
|
|
|
93
95
|
Representative carried-open render (multi-concern gap split to `e2e`; single-concern gap `cache`; `UNAUTHORIZED` gap `auth`):
|
|
@@ -133,6 +135,15 @@ A **heavy** revert is not a menu toggle — say so explicitly to the user before
|
|
|
133
135
|
|
|
134
136
|
### Step 4: Present Options
|
|
135
137
|
|
|
138
|
+
**Amendment digest (both variants).** Before the options, read `git -C "$WORKTREE" log <base-branch>..HEAD --grep '^amend:'` and take the `auto-apply` and `granted` records from those commit bodies. Render, then the options:
|
|
139
|
+
|
|
140
|
+
```
|
|
141
|
+
Amendments auto-applied (N):
|
|
142
|
+
- <title> - <what changed>
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
Omit the block when N = 0.
|
|
146
|
+
|
|
136
147
|
**Normal repo and named-branch worktree — present exactly these 4 options:**
|
|
137
148
|
|
|
138
149
|
```
|
|
@@ -220,6 +231,14 @@ EOF
|
|
|
220
231
|
)")
|
|
221
232
|
```
|
|
222
233
|
|
|
234
|
+
When the spec's `## Acceptance criteria` has at least one `venue:` or `deferred:` row, append this block after `## Test Plan`, listing those rows verbatim with their disposition, so the reader knows what `/skill:check-delivery` verifies after deploy and what a follow-up owns. With no such rows the body ends at `## Test Plan`, byte-identical to today. Option 1's squash commit message is unchanged.
|
|
235
|
+
|
|
236
|
+
```markdown
|
|
237
|
+
## Acceptance criteria
|
|
238
|
+
- [ ] <row text verbatim> - venue: <env> - <observation>
|
|
239
|
+
- [ ] <row text verbatim> - deferred: <where>
|
|
240
|
+
```
|
|
241
|
+
|
|
223
242
|
**Do NOT clean up worktree** — user needs it alive to iterate on PR feedback.
|
|
224
243
|
|
|
225
244
|
#### Option 3: Keep As-Is
|
|
@@ -11,6 +11,7 @@ Each bullet:
|
|
|
11
11
|
- `<handle>` leads the bullet and is a short unique human word derived from the title (`Cache coverage` -> `cache`); on collision append a digit. It is the token option 2 targets. When a gap split and no clean word fits, use the bare `Gn/Cn`; a single-concern gap uses its gap ID `Gn`.
|
|
12
12
|
- The shared options line sits below the bullets: `Other options per item: fix-now / accept / rescope / follow-up / custom`, listing the options **generally available across items**. When a specific item's availability deviates - an option unavailable for it, or an `UNAUTHORIZED` item whose `rescope` is unavailable and whose `fix-now` means removal - note that deviation as a short parenthetical on **that item's bullet** (one clause, not a block), e.g. `(rescope N/A: scope creep)`. The shared line appears **only in the carried-open render**, never in the zero-gap path. Full per-option effects only on request, or when option 2 targets an unclear choice.
|
|
13
13
|
- Group items under one recommended line only when they share a disposition and rationale; each grouped handle repeats its title.
|
|
14
|
+
- A bullet whose recommendation edits the spec (`accept-into-spec`, `rescope-into-spec`, or an item the pre-menu funnel escalated) carries, indented under it, `what` and `why` as one sentence each and `Example: <before -> after>` - the readable card from `../../brainstorming/reference/amendment-surface.md`, adapted to this bullet. Escalated items are ordinary rows here; no second amendment menu renders.
|
|
14
15
|
- Availability per concern comes from the reference's single availability table - apply it against current context (worktree state, `maxFixRounds`, ownership, resource accessibility), do not restate it. `UNAUTHORIZED` bullets ask the reference's question verbatim (`Should this unrequested behavior become part of the current workflow?`); `rescope-into-spec` is shown **unavailable** (not dropped) and `fix-now` means **removal** of the unrequested code.
|
|
15
16
|
- `revert conformance fix Gn`, when the gap has an auto-applied fix, renders on the shared options line as a **separate one-off action** - never inside a bullet's recommendation and never in the option-2 list. Name the parent gap and warn that revert undoes the entire gap-level commit (see "Revert semantics").
|
|
16
17
|
|
|
@@ -34,7 +35,7 @@ Each bullet:
|
|
|
34
35
|
Take **no** disposition action before the reply. Then, once, in order:
|
|
35
36
|
|
|
36
37
|
1. **Normalize** every `custom(...)` into explicit operations; classify state-changing (edits code or spec) vs not. Clarify only an ambiguous or unexecutable effect.
|
|
37
|
-
2. **Commit spec edits** (`accept-into-spec`, `rescope-into-spec`, state-changing spec `custom`) - the main session edits the spec directly, before any fix dispatch (a dirty tree rejects `worktree: true`, and the re-audit must read the amended spec).
|
|
38
|
+
2. **Commit spec edits** (`accept-into-spec`, `rescope-into-spec`, state-changing spec `custom`) - the main session edits the spec directly, before any fix dispatch (a dirty tree rejects `worktree: true`, and the re-audit must read the amended spec). A disposition chosen here is already approved: it bypasses the amendment surface's reviewer and batch. Items the pre-menu funnel auto-applied are recorded `Gn - <title>: accept-into-spec (auto-applied)`.
|
|
38
39
|
3. **Re-audit if step 2 changed the spec**; regenerate the inventory and re-render if it changed. Project `fix-now` only from the refreshed inventory.
|
|
39
40
|
4. **fix-now + code-changing custom:** project the selected concerns per gap into the reference's concern-scoped fix contract (excluding accepted/rescoped/followed-up siblings); run the reference "Fix loop" (unchanged - do not re-describe it). A code-changing `custom` runs the project's tests + `code-reviewer` on its delta before proceeding. Re-run Step 1's canonical tests.
|
|
40
41
|
5. **Re-audit after all state-changing work;** obtain fresh decisions **only if** the refreshed inventory differs from the approved one, else proceed.
|
|
@@ -65,7 +65,7 @@ For each task in `plan_tracker`:
|
|
|
65
65
|
6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
|
|
66
66
|
7. After its existing reviews accept the work (and its required commit point), mark that same task index `complete` in `plan_tracker`. A subagent exit or green tests alone are not acceptance.
|
|
67
67
|
|
|
68
|
-
The orchestrator is the spec's only writer during execution. Amendment trigger: implementer `BLOCKED` citing a spec defect, or a review finding showing the spec (not the code) is wrong -> pause the fix loop, execute brainstorming's [
|
|
68
|
+
The orchestrator is the spec's only writer during execution. Amendment trigger: implementer `BLOCKED` citing a spec defect, or a review finding showing the spec (not the code) is wrong -> pause the fix loop, execute brainstorming's [amendment path](../brainstorming/SKILL.md#amending-an-approved-spec) in place (it reviews first, stops only on escalation or redraw), resume. Code-vs-spec mismatch stays in the SR loop. An SR unable to read the spec at a cited anchor (missing file, unresolvable heading/range) returns a blocking finding — the contract is spec+task or stop, never a silent fallback to task-only review.
|
|
69
69
|
|
|
70
70
|
After all tasks: proceed to [After All Tasks](#after-all-tasks-complete) - it owns parent full verification, then whole-diff code review, then conformance.
|
|
71
71
|
|
|
@@ -57,6 +57,7 @@ Self-checking in the main session is the fallback when delegation isn't possible
|
|
|
57
57
|
| Order | Source | Why |
|
|
58
58
|
|---|---|---|
|
|
59
59
|
| 1 | The written spec (`doc/specs/…`) | Canonical. Brainstorm already fetched the ticket, reconciled its ACs, recorded deviations here. |
|
|
60
|
+
| 1 | The spec's `## Acceptance criteria` section | Same priority as the spec body. The ticket's AC rows verbatim with dispositions; `in-scope`/`venue:` rows are requirements, `deviates:`/`deferred:` rows are recorded drift - read per the `conformance-reviewer` persona. |
|
|
60
61
|
| 2 | Original prompt | Catches inline requirements never folded into the spec. |
|
|
61
62
|
| 3 | Re-fetch the ticket | **Fallback only**, when no spec exists. Skip when a spec exists — the live ticket may have drifted. |
|
|
62
63
|
|
|
@@ -229,7 +229,7 @@ Task commit steps use bare `git`: the implementer subagent runs with the checkou
|
|
|
229
229
|
|
|
230
230
|
## Spec Coverage Table
|
|
231
231
|
|
|
232
|
-
Every plan ends with a `## Spec coverage` section (grammar and example: [reference/plan-contract.md § Spec coverage table](reference/plan-contract.md)). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. Two row kinds:
|
|
232
|
+
Every plan ends with a `## Spec coverage` section (grammar and example: [reference/plan-contract.md § Spec coverage table](reference/plan-contract.md)). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. From the spec's `## Acceptance criteria`, table only `in-scope` and `venue:` rows (a `venue:` row's owner is the task delivering its named enabling change); `deviates:`, `deferred:`, and `none` rows get no table row - their disposition is the record, and any obligation a `deviates:` reason adopts is already a Design clause with its own row. Two row kinds:
|
|
233
233
|
|
|
234
234
|
- **Requirement rows:** a cross-cutting requirement (decided in more than one task) lists **every** deciding task as owner, not the first. `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
|
|
235
235
|
- **`Verification` owner:** only for a requirement the header command proves; grammar in the reference.
|