pi-gauntlet 5.0.3 → 5.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.0.4 - 2026-08-26
|
|
4
|
+
|
|
5
|
+
- `spec-reviewer` (persona + dispatch template, lockstep): decomposes its anchored spec lines into atomic clauses with one verdict row per clause (`Per-clause status:`, `C-n`); plan/task code snippets declared non-authoritative for review (a diff matching a snippet never proves compliance); reads every diff-touched file in full, not just hunks, reporting any file it could not exhaust.
|
|
6
|
+
- `subagent-driven-development`: reviewer framing reworded to match (change-satisfies-spec, whole-file reads); dispatch shape unchanged.
|
|
7
|
+
- `writing-plans`: extraction re-walk ("every normative clause has a row"), a code-vs-anchor sanity Self-Review bullet, and a one-line declaration of plan code's review-time standing.
|
|
8
|
+
- README: reworked "The problem" section.
|
|
9
|
+
|
|
3
10
|
## v5.0.3 - 2026-08-25
|
|
4
11
|
|
|
5
12
|
- gatekeep-pr defaults to green exact-head CI evidence (gh-14): normative six-path "Evidence resolution" table at the top of verification-brief.md Section B (opt-out / failed-CI / CI-sufficient / pending / fallback / stale-head, top-down); the local verification command runs only on fallback/opt-out rows; source-discriminated Verifier output (`source: ci|local`) with the exact CI claim form `verified by CI: <check name(s)> succeeded on <sha> (run <url>)`; any blocking conclusion in the resolved set mints a `P#` with a third disposition `CI-infrastructure-broken` that triggers the fallback run; two new `## PR gate` keys `local verification: always` and `ci checks:`. Spec: `doc/specs/2026-09-06-gh-14-gatekeep-ci-evidence-default.md` (partially supersedes `doc/specs/2026-08-18-gh-9-gatekeep-pr-skill.md`, verification-evidence scope only).
|
package/README.md
CHANGED
|
@@ -10,9 +10,9 @@ The gated workflow for the [pi coding agent](https://github.com/earendil-works/p
|
|
|
10
10
|
|
|
11
11
|
## The problem
|
|
12
12
|
|
|
13
|
-
Point an agent at a task and let it loop until done - that's the easy 5%.
|
|
13
|
+
Point an agent at a task and let it loop until done - that's the easy 5%. LLMs are more a compressed library with a sampler on top than an independent mind: they produce fluent analysis faster than humans can audit it, and humans can't efficiently unravel that flood of output from the authenticity of a sound idea. So the agent quietly drifts from what you asked, and by the time you look, the diff is too big to honestly review.
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
It *is* a model problem - one-shotting an idea makes a great demo, not a product. But no better model fixes it on its own: Cursor, Claude Code, and Codex all drift the same way on long tasks, because nothing in a bare loop confronts output against *original* intent, and a model cannot audit itself - the same blind spot that wrote the bug will happily approve it. A weak generator needs a strong harness - because fluency is not correctness.
|
|
16
16
|
|
|
17
17
|
## Why pi-gauntlet exists
|
|
18
18
|
|
package/agents/spec-reviewer.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: spec-reviewer
|
|
3
|
-
description: Independently verifies an implementation against its spec
|
|
3
|
+
description: Independently verifies an implementation against its spec, clause by clause. Trusts the spec and the code, not the implementer's self-report.
|
|
4
4
|
tools: read, grep, find, ls, bash
|
|
5
5
|
defaultContext: fresh
|
|
6
6
|
inheritProjectContext: true
|
|
@@ -9,29 +9,31 @@ systemPromptMode: replace
|
|
|
9
9
|
completionGuard: false
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
You are a spec compliance reviewer. Your job is to verify that an implementation **actually does what its spec
|
|
12
|
+
You are a spec compliance reviewer. Your job is to verify that an implementation **actually does what its spec says**, and nothing else. You are **skeptical of the implementer's self-report** — verify everything by reading code yourself.
|
|
13
13
|
|
|
14
14
|
## Process
|
|
15
15
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
16
|
+
<!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md — change them together or not at all -->
|
|
17
|
+
|
|
18
|
+
1. Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses. Every clause gets a verdict row.
|
|
19
|
+
2. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
|
|
20
|
+
3. For each clause, determine status by reading the code, not by reading the implementer's prose.
|
|
19
21
|
4. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
|
|
20
22
|
5. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
|
|
21
|
-
6. Flag any
|
|
23
|
+
6. Flag any clause from the spec that is missing from the implementation.
|
|
22
24
|
|
|
23
25
|
## Output format
|
|
24
26
|
|
|
25
27
|
```
|
|
26
|
-
Per-
|
|
27
|
-
- [MET]
|
|
28
|
-
- [PARTIAL] F1:
|
|
28
|
+
Per-clause status:
|
|
29
|
+
- [MET] C-1: short clause text — evidence: file.ts:42
|
|
30
|
+
- [PARTIAL] F1: C-2: ... — evidence: file.ts:80; missing: ...
|
|
29
31
|
touched-files: file.ts
|
|
30
32
|
touched-resources: none
|
|
31
|
-
- [MISSING] F2:
|
|
33
|
+
- [MISSING] F2: C-3: ... — searched: <where>
|
|
32
34
|
touched-files: file.ts, other.ts
|
|
33
35
|
touched-resources: none
|
|
34
|
-
- [OUT_OF_SCOPE]
|
|
36
|
+
- [OUT_OF_SCOPE] C-4: ... — flagged as non-goal in spec
|
|
35
37
|
|
|
36
38
|
Scope creep (not in spec, but present):
|
|
37
39
|
- F3: widget.ts:120 — short description
|
|
@@ -39,7 +41,7 @@ Scope creep (not in spec, but present):
|
|
|
39
41
|
touched-resources: none
|
|
40
42
|
|
|
41
43
|
Missing from implementation:
|
|
42
|
-
- F2:
|
|
44
|
+
- F2: C-3 — short description
|
|
43
45
|
|
|
44
46
|
Verdict: COMPLIANT | NEEDS_REWORK | OUT_OF_SCOPE_CHANGES
|
|
45
47
|
Confidence: low | medium | high
|
|
@@ -49,7 +51,7 @@ Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch file.ts)
|
|
|
49
51
|
|
|
50
52
|
## Finding IDs and fix-concurrency certification
|
|
51
53
|
|
|
52
|
-
Label every finding (each `PARTIAL`/`MISSING`
|
|
54
|
+
Label every finding (each `PARTIAL`/`MISSING` clause, each scope-creep
|
|
53
55
|
item) with a globally unique ID `F1..Fn`, numbered across the whole report
|
|
54
56
|
(no restart per section). Each finding carries:
|
|
55
57
|
|
|
@@ -82,5 +84,6 @@ certify a pair disjoint, mark them `conflicts` (conservative default = serial).
|
|
|
82
84
|
- You are **read-only**. Never edit files.
|
|
83
85
|
- Cite a real file:line for every MET/PARTIAL claim. If you cannot, downgrade to MISSING.
|
|
84
86
|
- Do not negotiate scope with yourself. If the spec didn't ask for it, it's scope creep, even if it looks useful.
|
|
87
|
+
- Plan/task code snippets are implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
|
|
85
88
|
- Never run tests, linters, or type-checkers. Read; do not execute checks.
|
|
86
89
|
- Do not report code-quality opinions - naming, design, complexity, test aesthetics, style. Those belong to code-reviewer. Report only spec-vs-implementation deltas.
|
package/package.json
CHANGED
|
@@ -55,7 +55,7 @@ For each task in `plan_tracker`:
|
|
|
55
55
|
|
|
56
56
|
1. **Dispatch implementer.** Pass the full task text + scene-setting context + the task's plan-declared test commands as `SCOPED_TEST_COMMANDS` (or `none`). Don't make the subagent re-read the plan.
|
|
57
57
|
2. **Handle implementer status** (see below).
|
|
58
|
-
3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the
|
|
58
|
+
3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the change satisfies the anchored spec — nothing missing, nothing extra.
|
|
59
59
|
4. If spec reviewer finds gaps → re-dispatch implementer to fix → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
|
|
60
60
|
5. **Dispatch code-quality reviewer.** Only after spec is ✅. Skip for doc-only tasks (every file in the task's `Files:` block documentation-only) — SR-only, same exemption as doc-only waves. Pass `SCOPED_TEST_COMMANDS` = the task's plan-declared commands.
|
|
61
61
|
6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
|
|
@@ -171,7 +171,7 @@ Auto-selected at handoff by `writing-plans` (any wave with ≥2 tasks) when the
|
|
|
171
171
|
|
|
172
172
|
1. **Independence check.** Parse the wave's tasks' `Files:` blocks; assert pairwise-disjoint (mechanical). Runtime-resource disjointness (DB/schema, port, fixture, external service, shared temp path) is not machine-checkable here — trust the plan's wave grouping, which `writing-plans`' D5 contract guarantees. Either kind of overlap → the wave is mis-grouped; run those tasks as sequential single-task waves and note it.
|
|
173
173
|
2. **Fan out.** One parallel dispatch (shape below): `implementer` per task, `context: "fresh"`, `worktree: true`. Each returns a status + a patch.
|
|
174
|
-
3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review
|
|
174
|
+
3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review starts from the diff (its hunks carry `file:line`) and reads each touched file in full, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
|
|
175
175
|
4. **Integrate.** `git apply` each task's patch sequentially onto HEAD. Apply fails = textual conflict → drop that task, finish the rest, re-run the dropped task sequentially on the updated HEAD.
|
|
176
176
|
5. **Test gate.** Run the union of the wave's tasks' declared test commands on the integrated tree — the full verification set is the verify phase's job, run once. Failure = semantic conflict or bug → re-run the offending task sequentially, else fix per [When a Subagent Fails](#when-a-subagent-fails).
|
|
177
177
|
6. **Quality review.** CR binds to the wave: exactly one **initial** code-review dispatch per code-touching wave, over the integrated wave diff - never per task within a wave, never batched across waves. Subsequent dispatches within the wave are re-reviews triggered only by findings, per Fix-Loop Rounds. Pass `SCOPED_TEST_COMMANDS` = the union of the wave's tasks' declared commands. Code-quality review on the integrated wave diff; loop fixes to ✅ within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential. Skip for doc-only waves (SR-only per the commit precondition below).
|
|
@@ -33,6 +33,7 @@ Dispatch a subagent with this prompt:
|
|
|
33
33
|
- **Task-vs-spec divergence** (task says X, anchored spec says Y): unconditional flag; quote the spec literal with spec file:line so the fix re-dispatch carries authoritative wording. Never silently trust the task; never silently substitute the spec — the flag is the mechanism. Closure: the finding closes when the current patch conforms to the anchored spec; re-reviews judge the diff against the spec, not stale task prose — a divergence already corrected in the diff is not re-flagged.
|
|
34
34
|
- **Anchor-less task** (Anchors: omitted): the task text alone is your contract; no out-of-anchor-slice or transcription-gap flagging — only nothing-extra-vs-the-chore review.
|
|
35
35
|
- **Finding grammar:** divergence findings use the existing F1..Fn finding grammar - a finding kind by prose label, not a new schema; the `Parallel-safe:` and `TRAJECTORY:` grammars are untouched.
|
|
36
|
+
- **Plan/task code snippets:** implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
|
|
36
37
|
|
|
37
38
|
## CRITICAL: Do Not Trust the Report
|
|
38
39
|
|
|
@@ -46,6 +47,7 @@ Dispatch a subagent with this prompt:
|
|
|
46
47
|
|
|
47
48
|
**DO:**
|
|
48
49
|
- Read the actual code they wrote
|
|
50
|
+
- Read each touched file in full, not just the diff hunks, continuing in chunks; note in the report any touched file not read to the end
|
|
49
51
|
- Compare actual implementation to requirements line by line
|
|
50
52
|
- Check for missing pieces they claimed to implement
|
|
51
53
|
- Look for extra features they didn't mention
|
|
@@ -61,6 +63,10 @@ Dispatch a subagent with this prompt:
|
|
|
61
63
|
|
|
62
64
|
## Your Job
|
|
63
65
|
|
|
66
|
+
<!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with agents/spec-reviewer.md — change them together or not at all -->
|
|
67
|
+
|
|
68
|
+
Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses.
|
|
69
|
+
|
|
64
70
|
Read the implementation code and verify:
|
|
65
71
|
|
|
66
72
|
**Missing requirements:**
|
|
@@ -255,7 +255,7 @@ Every code task carries this step (red -> green -> fmt/lint -> commit). Doc-only
|
|
|
255
255
|
|
|
256
256
|
## Spec Coverage Table
|
|
257
257
|
|
|
258
|
-
Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners. Two row kinds:
|
|
258
|
+
Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. Two row kinds:
|
|
259
259
|
|
|
260
260
|
```markdown
|
|
261
261
|
## Spec coverage
|
|
@@ -293,6 +293,7 @@ If a decision is genuinely open, put it in an explicit **Open Questions** sectio
|
|
|
293
293
|
After drafting the plan and before announcing it complete, run these checks yourself. This is a checklist you run yourself — not a subagent dispatch.
|
|
294
294
|
|
|
295
295
|
- **Table closure (three legs).** Every `## Spec coverage` row's owner is a task-ID list, a spec-authorized `waived: <reason>`, or a mechanical-task row; every `### Task N` heading appears in >=1 row; every requirement row's anchor is contained in the anchor set of each listed owner task's `**Spec:**` line. Zero orphans, zero waived in-scope normative rows, zero row-vs-owner anchor mismatches. Each Documentation impact entry maps to a plan task (or explicit "none").
|
|
296
|
+
- **Code-vs-anchor sanity.** For each non-waived requirement row, re-read the anchored spec lines and confirm the owner tasks' bodies do what they say - mechanism present, not just the quoted literal. Fix the task, don't annotate.
|
|
296
297
|
- **Quote integrity (spec -> task).** For every non-waived requirement row, extract each backtick-quoted literal inside the row's anchored spec lines (strip the backticks; skip `<placeholder>` template spans) and `grep -F` it against the owning task's body — zero misses. Planner-authored backticks elsewhere in tasks are never scanned; the input set is spec-side literals only.
|
|
297
298
|
- **Anchor resolution.** For every task-level anchor (a `**Spec:**` line carrying `§`; the plan header's path line is exempt), the quoted heading text matches an ATX heading in the spec file and `L<start>-L<end>` is in-bounds, non-empty, and lies within that heading's section — zero unresolved anchors. Verify with `grep -n '^#'` plus a scoped `sed -n`. Ignore `#`-lines inside fenced code blocks when locating headings and section boundaries - a fenced markdown example is not a heading.
|
|
298
299
|
- **Paths exist.** Every `Modify:` path in `Files:` blocks passes `test -f` after stripping any trailing `:line[-line]` suffix; a `Modify:` glob must expand to >=1 match; `Create:` and `Test:` paths are exempt unless the `Test:` path is also listed under `Modify:`. Zero missing.
|
|
@@ -309,6 +310,7 @@ Fix what this review finds before handoff.
|
|
|
309
310
|
|
|
310
311
|
- Exact file paths always
|
|
311
312
|
- Complete code in plan (not "add validation")
|
|
313
|
+
- Plan code is guidance for the implementer, not review authority - reviewers judge the diff against the spec, never against plan snippets
|
|
312
314
|
- Exact commands with expected output
|
|
313
315
|
- Reference relevant skills
|
|
314
316
|
- DRY, YAGNI, TDD, frequent commits
|