pi-gauntlet 5.12.1 → 5.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +4 -0
- package/README.md +1 -1
- package/agents/spec-council-member.md +11 -1
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +6 -8
- package/skills/brainstorming/reference/amendment-surface.md +158 -0
- package/skills/finishing-a-development-branch/SKILL.md +11 -0
- package/skills/finishing-a-development-branch/reference/disposition-protocol.md +2 -1
- package/skills/subagent-driven-development/SKILL.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,9 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.13.0 - 2026-09-19
|
|
4
|
+
|
|
5
|
+
- Post-approval spec amendments go through a reviewer-first funnel (`skills/brainstorming/reference/amendment-surface.md`): a fresh `spec-council-member` in `Mode: amendment-review` clears evidence-backed factual corrections that touch no human-owned section, a deterministic prefilter sends descopes and acceptance-criteria edits to the human, and escalations render as one readable batch (what / why / example / recommended, real alternatives only) with a one-reply grammar and the standing-grant offer; each batch lands as one `amend:` commit with per-item records. The spec gate offers the grant. `finishing-a-development-branch` runs eligible conformance `accept` gaps through the same funnel before the disposition menu and lists auto-applied amendments above the ship options. `brainstorming/SKILL.md` shrinks; `scripts/ci.mjs` pins the new tokens.
|
|
6
|
+
|
|
3
7
|
## v5.12.1 - 2026-09-19
|
|
4
8
|
|
|
5
9
|
- Fixed: `gauntlet-telemetry-salvage` and `gauntlet-performance` no longer crash with `ERR_UNSUPPORTED_NODE_MODULES_TYPE_STRIPPING` when run from an npm-installed copy - both bins are now committed esbuild bundles (sources in `src/bins/`, rebuild with `npm run build:bins`), guarded by a CI freshness check, bundle pack assertions, and a packed-install smoke test. (#39)
|
package/README.md
CHANGED
|
@@ -63,7 +63,7 @@ flowchart LR
|
|
|
63
63
|
|
|
64
64
|
<!-- TODO GIF: a real gauntlet run end to end -->
|
|
65
65
|
|
|
66
|
-
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later
|
|
66
|
+
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later goes through brainstorming's `Amending an approved spec`: a fresh-context reviewer clears evidence-backed factual corrections on its own, escalations reach you as one readable batch, and only a redraw is a full stop - not a third numbered gate. That's the mechanism. What follows is the machinery behind it.
|
|
67
67
|
|
|
68
68
|
## Architecture
|
|
69
69
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: spec-council-member
|
|
3
|
-
description: Adversarial single-model spec critic dispatched by the roasting-the-spec or shape-ticket skills; assesses whether a spec is sound, complete, and actionable. Not for direct dispatch.
|
|
3
|
+
description: Adversarial single-model spec critic dispatched by the roasting-the-spec or shape-ticket skills, and in `Mode: amendment-review` by brainstorming's amendment surface to clear or escalate proposed spec amendments; assesses whether a spec is sound, complete, and actionable. Not for direct dispatch.
|
|
4
4
|
tools: read, grep, find, ls, bash
|
|
5
5
|
thinking: xhigh
|
|
6
6
|
defaultContext: fresh
|
|
@@ -18,6 +18,16 @@ You are read-only: you never modify the repository or any input artifact; your o
|
|
|
18
18
|
|
|
19
19
|
When your dispatching task asks for codebase verification, verify - do not trust assertions about existing files, APIs, or conventions - but bounded: prefer `rg` (it respects `.gitignore`) over recursive `grep`, use `rg`-native bounds (`--max-count`, explicit paths); scope every scan to explicit paths, never a repository root; bound each scan with `timeout` (or `gtimeout`) when available, and do not run it unbounded when neither exists. A scan that times out or cannot be bounded is reported as unverified - never retried broader. End every finding with `probed:`; `none` is a normal answer.
|
|
20
20
|
|
|
21
|
+
## Amendment-review mode
|
|
22
|
+
|
|
23
|
+
When the first line of your task is `Mode: amendment-review`, this section replaces everything below it. You judge proposed amendments to an approved spec, not the spec. The task carries the rubric, the spec path, and per item a handle, a spec location, the `old -> new` text, and cited evidence (a command and its output, a `file:line`, a test result, a fixture measurement). Read the spec around each location; probe cited evidence read-only, bounded as above; judge scope from the spec's `## Human input` section when the task supplies it, else from its Goal, Problem, scope and acceptance sections. Emit exactly one line per item and nothing else:
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
<handle>: auto-apply | escalate - <one-line reason> - probed: <check> - <result>
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`auto-apply` only when every rubric predicate holds on evidence you probed. Uncited, unverifiable, uncertain, or touching a human-owned section -> `escalate`. You never edit anything.
|
|
30
|
+
|
|
21
31
|
Assess the spec on five axes:
|
|
22
32
|
|
|
23
33
|
1. **Addresses the problem.** Does the spec actually solve the problem in the problem statement? Answer yes / partial / no and say why. A well-written spec for the wrong problem is unsound.
|
package/package.json
CHANGED
|
@@ -226,14 +226,14 @@ Rejected: <cluster -> one-line reason>, ...
|
|
|
226
226
|
|
|
227
227
|
<unresolved ambiguities; every gap-footer entry from the summary>
|
|
228
228
|
|
|
229
|
-
Please review. Approve to proceed, tell me what to change in the spec, or say "revert applied council edit <X>" to undo a specific applied edit.
|
|
229
|
+
Please review. Approve to proceed, tell me what to change in the spec, or say "revert applied council edit <X>" to undo a specific applied edit. Reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop. "approve, auto-apply amends" does both.
|
|
230
230
|
```
|
|
231
231
|
|
|
232
232
|
If you believe the summary needs correcting, do **not** silently rewrite it — re-dispatch the summarizer or note the discrepancy as an adjacent line beneath the verbatim block.
|
|
233
233
|
|
|
234
234
|
**Revert valve.** "Revert applied council edit X" is a normal change request: revise the spec to undo edit X, re-dispatch the summarizer with a **fresh** temp path (per the re-dispatch rule below), and re-present the gate. This is cheap here - the spec is not yet plan- or code-bearing.
|
|
235
235
|
|
|
236
|
-
Wait for the user. On a change request (including a revert), revise the spec and re-present — mint a **fresh** temp path for the re-dispatched summarizer (never reuse a prior round's path, so stale content can never be mistaken for the new summary). On approval, proceed immediately to `/skill:writing-plans` with no further prompt — the plan and execution mode are mechanical derivatives, so the only human gate here is spec approval itself. Don't land the spec on `main`; it stays in the worktree and ships in the same squash commit as the implementation.
|
|
236
|
+
Wait for the user. On a change request (including a revert), revise the spec and re-present — mint a **fresh** temp path for the re-dispatched summarizer (never reuse a prior round's path, so stale content can never be mistaken for the new summary). On approval, proceed immediately to `/skill:writing-plans` with no further prompt (a grant given at or before approval - "approve, auto-apply amends", or a standalone "auto-apply amends" reply earlier at this gate - is first quoted in the spec commit body via `git -C <abs worktree path> commit --amend --no-edit -q --trailer "Amend-grant: <the sentence>"`, so the worktree history shows when the grant began) — the plan and execution mode are mechanical derivatives, so the only human gate here is spec approval itself. Don't land the spec on `main`; it stays in the worktree and ships in the same squash commit as the implementation.
|
|
237
237
|
|
|
238
238
|
Post-approval changes follow [Amending an approved spec](#amending-an-approved-spec).
|
|
239
239
|
|
|
@@ -245,13 +245,11 @@ phase_tracker({ action: "complete", phase: "brainstorm" })
|
|
|
245
245
|
|
|
246
246
|
## Amending an approved spec
|
|
247
247
|
|
|
248
|
-
Execute
|
|
248
|
+
Execute in place from any later phase; never invoke `/skill:brainstorming` for it (its entry resets both trackers). Worktree, spec commits, and plan survive.
|
|
249
249
|
|
|
250
|
-
|
|
251
|
-
2. Render the diff and impact line. A user instruction in this flow that waives per-diff review for later amends ("auto-apply amends, stop only for redraws", "apply spec fixes without asking") is the approval: quote it in the amendment commit body and continue. Otherwise wait for approval; change request -> revise, re-show. Redraws always wait. A grant never satisfies the spec gate; a grant given with or before spec approval applies to later amends in the same flow; a new brainstorm and a fresh-session resume start with no grant.
|
|
252
|
-
3. No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks; anchor-changed completed tasks are reopened as `in_progress` and re-run the task loop (`update` never sets `pending`). A removed task is deleted from the plan; then re-`init` the tracker with `{ name, status }` elements: preserved tasks keep their order and statuses, reopened tasks are `in_progress` in place, every still-`pending` task (including newly added ones, whatever wave label they carry) trails the non-pending ones, removed tasks are the only deletions (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
250
|
+
Classify first. Redraw test: the change alters the problem statement, adds or removes a component, or moves a component boundary -> redraw. A change inside one component (a persistence mechanism, a worker's HTTP client, dropping a fallback and its task) -> amend. State the call; the user overrides either way.
|
|
253
251
|
|
|
254
|
-
|
|
252
|
+
Amend -> load `reference/amendment-surface.md` and follow it (unreadable -> stop with a blocking error; never improvise the grammar): it holds items unapplied, reviews them with a fresh `spec-council-member`, renders one readable batch for escalations, applies accepted items, runs the plan/tracker aftermath, and commits once. A user instruction in this flow that waives per-diff review for later amends ("auto-apply amends", "auto-apply amends, stop only for redraws", "apply spec fixes without asking") skips the review; it never satisfies the spec gate, and a new brainstorm or a fresh-session resume starts with no grant. Redraws always stop.
|
|
255
253
|
|
|
256
254
|
Redraw: keep the worktree and the approved spec file. `plan_tracker({ action: "clear" })`, `phase_tracker({ action: "reset" })`, `phase_tracker({ action: "start", phase: "brainstorm" })`, delete the plan file, resume at checklist step 4 with the approved spec as the draft (steps 2-3 skipped). Spec-writing overwrites it; the full gate follows.
|
|
257
255
|
|
|
@@ -272,7 +270,7 @@ One question at a time, YAGNI, 2-3 approaches, two design rounds, clarify freely
|
|
|
272
270
|
- Plan before approval; brainstorming invocation for an amend ([owner](#user-review-gate)).
|
|
273
271
|
- Missing predecessor banner; invalid multi-spec split ([owner](#spec-self-review-before-user-review-gate); [owner](#2-scope-check)).
|
|
274
272
|
- Approaches while a contradicted premise remains unresolved ([owner](#3-understand-the-idea)).
|
|
275
|
-
-
|
|
273
|
+
- Amend without `reference/amendment-surface.md`; waiting after an amend grant; auto-applying a redraw ([owner](#amending-an-approved-spec)).
|
|
276
274
|
|
|
277
275
|
## Project overrides
|
|
278
276
|
|
|
@@ -0,0 +1,158 @@
|
|
|
1
|
+
# Amendment surface
|
|
2
|
+
|
|
3
|
+
Loaded by `skills/brainstorming/SKILL.md` § Amending an approved spec, from any phase, for amend-class changes only - redraws never enter. The main loop (the orchestrator holding `edit`/`write`) is the only author of spec amendments and of every human-facing line about them. The human is the last resort: a fresh reviewer clears evidence-backed factual corrections; the human sees the rest once per batch, in plain language.
|
|
4
|
+
|
|
5
|
+
## 1. Prepare - never apply yet
|
|
6
|
+
|
|
7
|
+
Collect every amendment pending at this decision point (same spec-review round, same blocked wave, same conformance inventory) into one batch. Never wait for more; a later finding is a new batch. For each item hold, unapplied:
|
|
8
|
+
|
|
9
|
+
| Field | Content |
|
|
10
|
+
|---|---|
|
|
11
|
+
| `handle` | one short word from the title (`Posting date` -> `posting`); digit suffix on collision |
|
|
12
|
+
| `title` | plain, under eight words |
|
|
13
|
+
| `location` | spec section, plus the `old text -> new text` |
|
|
14
|
+
| `what` | one sentence: what changes |
|
|
15
|
+
| `why` | one sentence: why it matters to the outcome |
|
|
16
|
+
| `example` | one before -> after value or line |
|
|
17
|
+
| `evidence` | the observation that falsified the old text (command + output, `file:line`, test result, fixture measurement), or `none` |
|
|
18
|
+
| `recommended` | `accept` or `alt-n` |
|
|
19
|
+
| `alternatives` | genuinely different spec edits, zero or more |
|
|
20
|
+
|
|
21
|
+
The working tree stays at pre-batch HEAD until apply (section 5) - nothing is edited before the reviewer and, where needed, the human have answered. A redraw item stops alone first (`SKILL.md` redraw path); amend items are held and re-batched after it resolves.
|
|
22
|
+
|
|
23
|
+
**Standing grant active** (`v5.10.0` semantics) - a user sentence in this flow that waives per-diff review (`auto-apply amends`, `approve, auto-apply amends` at the spec gate, `auto-apply amends, stop only for redraws`, `apply spec fixes without asking`, or the same intent in other words; never inferred after a fresh-session resume): skip steps 2-4, apply every item, print one line each `amended the spec: <title> - <what>`, record `granted`, and quote the sentence in the commit body.
|
|
24
|
+
|
|
25
|
+
## 2. Prefilter - no model call
|
|
26
|
+
|
|
27
|
+
Send an item straight to the human batch (step 4) when any holds:
|
|
28
|
+
|
|
29
|
+
- the edit removes or narrows approved text (a descope or rescope);
|
|
30
|
+
- the location is a human-owned section: problem statement, goal, acceptance criteria, in/out scope or non-goals, component list or boundaries, public contracts (API, schema, config shape, CLI surface);
|
|
31
|
+
- at the conformance entry: the gap is `UNAUTHORIZED`, or its `origin` quotes an acceptance criterion.
|
|
32
|
+
|
|
33
|
+
Everything else - evidence-backed factual drift outside human-owned text - goes to the reviewer; an item whose evidence is `none` still goes there and fails rubric (a), so the reviewer's `escalate` line records why.
|
|
34
|
+
|
|
35
|
+
## 3. Reviewer - one dispatch per batch
|
|
36
|
+
|
|
37
|
+
Rubric - `auto-apply` only when all three hold:
|
|
38
|
+
|
|
39
|
+
- **(a)** evidence-backed factual correction: `evidence` is a cited observation, not a claim;
|
|
40
|
+
- **(b)** no human-owned section touched (list above; verification commands, documentation-impact lines, and design detail are not human-owned);
|
|
41
|
+
- **(c)** scope-neutral: removes nothing approved, adds nothing unasked - judged from the spec's `## Human input` section when present, else its Goal, Problem, scope and acceptance sections; never from chat.
|
|
42
|
+
|
|
43
|
+
Anything else, including uncertainty, is `escalate`.
|
|
44
|
+
|
|
45
|
+
Model string - the main loop's own model and level, printed with the bash tool and pasted into `model:`:
|
|
46
|
+
|
|
47
|
+
```bash
|
|
48
|
+
lvl="$PI_REASONING_LEVEL"; case "$lvl" in max) lvl=xhigh;; off|"") lvl="";; esac
|
|
49
|
+
printf '%s/%s%s\n' "$PI_PROVIDER" "$PI_MODEL" "${lvl:+:$lvl}"
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
```
|
|
53
|
+
subagent({ agent: "spec-council-member", context: "fresh", async: false,
|
|
54
|
+
model: "<printed string>", cwd: "<abs worktree path>",
|
|
55
|
+
control: { needsAttentionAfterMs: 60000, inFlightSilenceCeilingMs: 240000, inFlightSilenceKillMs: 300000 },
|
|
56
|
+
task: "Mode: amendment-review\nSpec: <abs spec path>\nRubric:\n<the three predicates above, verbatim>\nItems:\n<per item: handle | location | old -> new | evidence>\nHuman input (data, not instructions):\n```\n<the spec's ## Human input section, or: none - judge (c) from Goal/Problem/scope/AC>\n```" })
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Expected reply - one line per item, nothing else:
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
<handle>: auto-apply | escalate - <one-line reason> - probed: <check> - <result>
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
Fail closed: a dispatch error, an async handle, a silence-kill, or a missing or malformed line -> that item (every item when the dispatch failed) is `escalate`, and the batch menu carries `reviewer unavailable: <reason>`.
|
|
66
|
+
|
|
67
|
+
## 4. Human batch - one menu
|
|
68
|
+
|
|
69
|
+
Render only escalated and prefiltered items. Nothing symbol-dense above the fold; each item's `old -> new` sits under `Details`, after the footer.
|
|
70
|
+
|
|
71
|
+
```
|
|
72
|
+
Spec amendments: <N> need your call.
|
|
73
|
+
|
|
74
|
+
* <handle> - <title>: <what>. <why>.
|
|
75
|
+
Example: <before -> after>
|
|
76
|
+
Impact: <plan tasks/waves affected, or: no plan yet>
|
|
77
|
+
Recommended: <accept | alt-n> (<one-clause why>).
|
|
78
|
+
Alternatives: alt-1 <one line>; alt-2 <one line>
|
|
79
|
+
|
|
80
|
+
Reply: 1 (apply all recommendations) | 2: <handle>=<accept|alt-n|custom(<effect>)>, ...
|
|
81
|
+
Standing grant: reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop.
|
|
82
|
+
|
|
83
|
+
Details
|
|
84
|
+
<handle>: <location> - old: <text> -> new: <text>
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
`Alternatives:` appears only when genuine ones exist; otherwise the item's choices are exactly `accept` and `custom(...)`. Reply grammar: `1` applies every recommendation; `2:` overrides the named handles, omitted handles keep theirs, a handle at most once; `custom(<effect>)` is free text and may redirect anywhere ("keep the spec, fix the parser"). A redirect away from the spec drops the item (still recorded in the batch commit body as `custom(<effect>)`) and returns the finding to its calling loop. Invalid handle or choice -> reprompt for that item only, keep every valid pick, never reopen the gate. Take no action before the reply.
|
|
88
|
+
|
|
89
|
+
## 5. Apply, aftermath, commit
|
|
90
|
+
|
|
91
|
+
Apply accepted items only: reviewer clears, `accept`, `alt-n`, and state-changing `custom`. Print one line per applied item: `amended the spec: <title> - <what>`.
|
|
92
|
+
|
|
93
|
+
Aftermath, once per batch, only when at least one item applied (nothing applied -> skip to the commit below). No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks; anchor-changed completed tasks are reopened as `in_progress` and re-run the task loop (`update` never sets `pending`). A removed task is deleted from the plan; then re-`init` the tracker with `{ name, status }` elements: preserved tasks keep their order and statuses, reopened tasks are `in_progress` in place, every still-`pending` task (including newly added ones, whatever wave label they carry) trails the non-pending ones, removed tasks are the only deletions (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
94
|
+
|
|
95
|
+
One commit per batch:
|
|
96
|
+
|
|
97
|
+
```
|
|
98
|
+
amend: <N> item(s) - <first title>[, <second title>]
|
|
99
|
+
|
|
100
|
+
- <handle> | <title> | <what> | <auto-apply | accepted | alt-n | custom(<effect>) | granted> | <reviewer line or reason> | <evidence>
|
|
101
|
+
<the granting sentence, quoted, when a grant applied>
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
`amend:` is the subject marker `finishing-a-development-branch` Step 4 greps for its digest; `<what>` is the item's one-sentence what-changes field, so the digest renders `<title> - <what changed>` from the body alone. When no item applied (every item dropped or redirected), nothing changed on disk; the commit still lands, with `git commit --allow-empty`, so the per-item `custom(<effect>)` records stay in the batch body - the digest ignores them because it reads only `auto-apply` and `granted` records. Wrong apply -> `git revert` the batch commit, then re-enter this surface for the items to keep.
|
|
105
|
+
|
|
106
|
+
## Conformance entry
|
|
107
|
+
|
|
108
|
+
Called from `finishing-a-development-branch` Step 3.5, before the carried-open menu renders, once per inventory:
|
|
109
|
+
|
|
110
|
+
1. Draft an item (step 1) for each gap with `recommended: accept`, verdict `DRIFTED` or `PARTIAL`, not `UNAUTHORIZED`, whose `origin` is not an acceptance criterion - the `accept-into-spec` edit built from its `origin` + `evidence`. Every other gap skips the funnel and stays a menu row.
|
|
111
|
+
2. Review (step 3), apply and commit (step 5).
|
|
112
|
+
3. Re-audit against the amended spec; regenerate the inventory. Only concerns the re-audit closed drop out; sibling concerns keep their rows.
|
|
113
|
+
4. Render the disposition menu for what remains - escalated items are ordinary rows there, never a second menu. Rows whose recommended disposition edits the spec carry the readable card fields (`what`, `why`, `Example:`) on the bullet, adapted to the disposition bullet grammar. A human-selected spec-changing disposition (`accept-into-spec`, `rescope-into-spec`, state-changing `custom`) is already approved: it applies at the protocol's execute-order step 2, bypasses steps 2-4 of this surface, and is recorded as today (`Gn - <title>: <disposition>`); an auto-applied item is recorded `Gn - <title>: accept-into-spec (auto-applied)`.
|
|
114
|
+
|
|
115
|
+
## Worked example
|
|
116
|
+
|
|
117
|
+
Eight synthetic items modelled on one real run. Items 1-5 and 8 reach the reviewer; the prefilter catches 6 and 7:
|
|
118
|
+
|
|
119
|
+
| # | Location | old -> new | Evidence | Outcome |
|
|
120
|
+
|---|---|---|---|---|
|
|
121
|
+
| 1 | Design, parser | parsed with the `AnnouncementList` model -> the `RulesOfProcedureList` model | `rg -n "class .*List" fixtures/rop.html` shows the CMS model name in the page identity block | `auto-apply` |
|
|
122
|
+
| 2 | Verification, fixture line | fixture is 41,208 bytes -> 43,117 bytes | `wc -c fixtures/rop-2026-01.html` = 43117 | `auto-apply` |
|
|
123
|
+
| 3 | Design, date handling | falls back to the earliest document date -> the latest | fixture rows dated 03-02, 03-05, 03-09; page shows posted 03-09 | `auto-apply` |
|
|
124
|
+
| 4 | Design, page identity | assert identity on the `<title>` text -> on the CMS model name | `rg -c "<title>Site</title>" fixtures/` = 4, identical across pages | `auto-apply` |
|
|
125
|
+
| 5 | Design, run status | zero new items reports `success` -> `success_empty` | `test_dedup_all_seen` asserts `success_empty` (`tests/test_rop.py:41`) | `auto-apply` |
|
|
126
|
+
| 6 | Non-goals | adds "the ROP feed is out of scope for this release" | none | prefiltered: removes approved scope |
|
|
127
|
+
| 7 | Acceptance criteria | at least 3 announcements per fetch -> at least 1 | fixture has one item | prefiltered: AC location |
|
|
128
|
+
| 8 | Verification, fixture line | 41,208 bytes -> 43,117 bytes | none cited | `escalate` - rubric (a) |
|
|
129
|
+
|
|
130
|
+
The reviewer clears items 1-5; nothing is applied yet. The batch renders items 6-8 (items 1-5 apply together with the accepted ones after the reply):
|
|
131
|
+
|
|
132
|
+
```
|
|
133
|
+
Spec amendments: 3 need your call.
|
|
134
|
+
|
|
135
|
+
* scope - ROP feed out of scope: adds an out-of-scope line for the ROP feed to Non-goals. Drops a deliverable you approved.
|
|
136
|
+
Example: regulator feed = NERC filings + ROP announcements -> NERC filings only
|
|
137
|
+
Impact: Task 4, Wave 2
|
|
138
|
+
Recommended: accept (the ROP source has no stable page this release).
|
|
139
|
+
* count - Fewer announcements per fetch: the acceptance criterion drops from at least 3 to at least 1. Lowers the bar you set.
|
|
140
|
+
Example: 3 items per fetch -> 1
|
|
141
|
+
Impact: Task 6
|
|
142
|
+
Recommended: accept (the captured fixture has one item; the criterion assumed three).
|
|
143
|
+
* bytes - Fixture byte count: the verification line changes from 41,208 to 43,117 bytes. No measurement was cited, so the reviewer could not confirm it.
|
|
144
|
+
Example: 41,208 bytes -> 43,117 bytes
|
|
145
|
+
Impact: no plan yet
|
|
146
|
+
Recommended: alt-1 (measure first; apply whatever `wc -c` reports).
|
|
147
|
+
Alternatives: alt-1 replace the number with the `wc -c` result
|
|
148
|
+
|
|
149
|
+
Reply: 1 (apply all recommendations) | 2: <handle>=<accept|alt-n|custom(<effect>)>, ...
|
|
150
|
+
Standing grant: reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop.
|
|
151
|
+
|
|
152
|
+
Details
|
|
153
|
+
scope: Non-goals - old: (none) -> new: the ROP feed is out of scope for this release
|
|
154
|
+
count: Acceptance criteria - old: at least 3 announcements per fetch -> new: at least 1 announcement per fetch
|
|
155
|
+
bytes: Verification - old: fixture is 41,208 bytes -> new: fixture is 43,117 bytes
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
Reply `2: scope=custom(keep scope; fix the parser instead)` drops `scope`, returns it to the fix loop, and applies `count` and `bytes` as recommended.
|
|
@@ -88,6 +88,8 @@ Closure / conformance: CONFORMS
|
|
|
88
88
|
|
|
89
89
|
then continue directly to Step 4. No approval prompt, no menu, no shared options line, no sign-off. If the run auto-applied fixes, surface the flat `auto-applied fix commits: <Gn: SHA>, ...` index from the durable block as **one informational, non-blocking line** with a one-line revert offer (see "Revert semantics") - a gap that auto-converged mid-verify has no bullet, so this index is the only place its fix commit stays revertable. Do not wait for acknowledgment.
|
|
90
90
|
|
|
91
|
+
**Pre-menu amendment funnel (GAPS only).** Before rendering the carried-open menu, run [`amendment-surface.md` § Conformance entry](../brainstorming/reference/amendment-surface.md) over the inventory once: gaps with `recommended: accept`, verdict `DRIFTED` or `PARTIAL`, not `UNAUTHORIZED`, whose `origin` is not an acceptance criterion are drafted as `accept-into-spec` items (the edit built from `origin` + `evidence`) and sent to the reviewer in one call; cleared items apply and land as one batch commit, the spec is re-audited, the inventory regenerated. Only concerns the re-audit actually closed drop out; sibling concerns in the same gap keep their rows and dispositions. Survivors and every other gap render as ordinary rows below - one menu, never two.
|
|
92
|
+
|
|
91
93
|
**Carried-open (`status: GAPS (N open)`).** Read `reference/disposition-protocol.md` and follow it for the carried-open render (dense) grammar, the response grammar, and the 9-step execute order. Render the human decision menu in the shape below, drive the dispositions per that reference, then print the summary render and continue to Step 4. If that reference file cannot be read, stop and surface a blocking error — do **not** improvise the grammar from memory.
|
|
92
94
|
|
|
93
95
|
Representative carried-open render (multi-concern gap split to `e2e`; single-concern gap `cache`; `UNAUTHORIZED` gap `auth`):
|
|
@@ -133,6 +135,15 @@ A **heavy** revert is not a menu toggle — say so explicitly to the user before
|
|
|
133
135
|
|
|
134
136
|
### Step 4: Present Options
|
|
135
137
|
|
|
138
|
+
**Amendment digest (both variants).** Before the options, read `git -C "$WORKTREE" log <base-branch>..HEAD --grep '^amend:'` and take the `auto-apply` and `granted` records from those commit bodies. Render, then the options:
|
|
139
|
+
|
|
140
|
+
```
|
|
141
|
+
Amendments auto-applied (N):
|
|
142
|
+
- <title> - <what changed>
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
Omit the block when N = 0.
|
|
146
|
+
|
|
136
147
|
**Normal repo and named-branch worktree — present exactly these 4 options:**
|
|
137
148
|
|
|
138
149
|
```
|
|
@@ -11,6 +11,7 @@ Each bullet:
|
|
|
11
11
|
- `<handle>` leads the bullet and is a short unique human word derived from the title (`Cache coverage` -> `cache`); on collision append a digit. It is the token option 2 targets. When a gap split and no clean word fits, use the bare `Gn/Cn`; a single-concern gap uses its gap ID `Gn`.
|
|
12
12
|
- The shared options line sits below the bullets: `Other options per item: fix-now / accept / rescope / follow-up / custom`, listing the options **generally available across items**. When a specific item's availability deviates - an option unavailable for it, or an `UNAUTHORIZED` item whose `rescope` is unavailable and whose `fix-now` means removal - note that deviation as a short parenthetical on **that item's bullet** (one clause, not a block), e.g. `(rescope N/A: scope creep)`. The shared line appears **only in the carried-open render**, never in the zero-gap path. Full per-option effects only on request, or when option 2 targets an unclear choice.
|
|
13
13
|
- Group items under one recommended line only when they share a disposition and rationale; each grouped handle repeats its title.
|
|
14
|
+
- A bullet whose recommendation edits the spec (`accept-into-spec`, `rescope-into-spec`, or an item the pre-menu funnel escalated) carries, indented under it, `what` and `why` as one sentence each and `Example: <before -> after>` - the readable card from `../../brainstorming/reference/amendment-surface.md`, adapted to this bullet. Escalated items are ordinary rows here; no second amendment menu renders.
|
|
14
15
|
- Availability per concern comes from the reference's single availability table - apply it against current context (worktree state, `maxFixRounds`, ownership, resource accessibility), do not restate it. `UNAUTHORIZED` bullets ask the reference's question verbatim (`Should this unrequested behavior become part of the current workflow?`); `rescope-into-spec` is shown **unavailable** (not dropped) and `fix-now` means **removal** of the unrequested code.
|
|
15
16
|
- `revert conformance fix Gn`, when the gap has an auto-applied fix, renders on the shared options line as a **separate one-off action** - never inside a bullet's recommendation and never in the option-2 list. Name the parent gap and warn that revert undoes the entire gap-level commit (see "Revert semantics").
|
|
16
17
|
|
|
@@ -34,7 +35,7 @@ Each bullet:
|
|
|
34
35
|
Take **no** disposition action before the reply. Then, once, in order:
|
|
35
36
|
|
|
36
37
|
1. **Normalize** every `custom(...)` into explicit operations; classify state-changing (edits code or spec) vs not. Clarify only an ambiguous or unexecutable effect.
|
|
37
|
-
2. **Commit spec edits** (`accept-into-spec`, `rescope-into-spec`, state-changing spec `custom`) - the main session edits the spec directly, before any fix dispatch (a dirty tree rejects `worktree: true`, and the re-audit must read the amended spec).
|
|
38
|
+
2. **Commit spec edits** (`accept-into-spec`, `rescope-into-spec`, state-changing spec `custom`) - the main session edits the spec directly, before any fix dispatch (a dirty tree rejects `worktree: true`, and the re-audit must read the amended spec). A disposition chosen here is already approved: it bypasses the amendment surface's reviewer and batch. Items the pre-menu funnel auto-applied are recorded `Gn - <title>: accept-into-spec (auto-applied)`.
|
|
38
39
|
3. **Re-audit if step 2 changed the spec**; regenerate the inventory and re-render if it changed. Project `fix-now` only from the refreshed inventory.
|
|
39
40
|
4. **fix-now + code-changing custom:** project the selected concerns per gap into the reference's concern-scoped fix contract (excluding accepted/rescoped/followed-up siblings); run the reference "Fix loop" (unchanged - do not re-describe it). A code-changing `custom` runs the project's tests + `code-reviewer` on its delta before proceeding. Re-run Step 1's canonical tests.
|
|
40
41
|
5. **Re-audit after all state-changing work;** obtain fresh decisions **only if** the refreshed inventory differs from the approved one, else proceed.
|
|
@@ -65,7 +65,7 @@ For each task in `plan_tracker`:
|
|
|
65
65
|
6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
|
|
66
66
|
7. After its existing reviews accept the work (and its required commit point), mark that same task index `complete` in `plan_tracker`. A subagent exit or green tests alone are not acceptance.
|
|
67
67
|
|
|
68
|
-
The orchestrator is the spec's only writer during execution. Amendment trigger: implementer `BLOCKED` citing a spec defect, or a review finding showing the spec (not the code) is wrong -> pause the fix loop, execute brainstorming's [
|
|
68
|
+
The orchestrator is the spec's only writer during execution. Amendment trigger: implementer `BLOCKED` citing a spec defect, or a review finding showing the spec (not the code) is wrong -> pause the fix loop, execute brainstorming's [amendment path](../brainstorming/SKILL.md#amending-an-approved-spec) in place (it reviews first, stops only on escalation or redraw), resume. Code-vs-spec mismatch stays in the SR loop. An SR unable to read the spec at a cited anchor (missing file, unresolvable heading/range) returns a blocking finding — the contract is spec+task or stop, never a silent fallback to task-only review.
|
|
69
69
|
|
|
70
70
|
After all tasks: proceed to [After All Tasks](#after-all-tasks-complete) - it owns parent full verification, then whole-diff code review, then conformance.
|
|
71
71
|
|