pi-gauntlet 5.0.3 → 5.0.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,12 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.0.4 - 2026-08-26
4
+
5
+ - `spec-reviewer` (persona + dispatch template, lockstep): decomposes its anchored spec lines into atomic clauses with one verdict row per clause (`Per-clause status:`, `C-n`); plan/task code snippets declared non-authoritative for review (a diff matching a snippet never proves compliance); reads every diff-touched file in full, not just hunks, reporting any file it could not exhaust.
6
+ - `subagent-driven-development`: reviewer framing reworded to match (change-satisfies-spec, whole-file reads); dispatch shape unchanged.
7
+ - `writing-plans`: extraction re-walk ("every normative clause has a row"), a code-vs-anchor sanity Self-Review bullet, and a one-line declaration of plan code's review-time standing.
8
+ - README: reworked "The problem" section.
9
+
3
10
  ## v5.0.3 - 2026-08-25
4
11
 
5
12
  - gatekeep-pr defaults to green exact-head CI evidence (gh-14): normative six-path "Evidence resolution" table at the top of verification-brief.md Section B (opt-out / failed-CI / CI-sufficient / pending / fallback / stale-head, top-down); the local verification command runs only on fallback/opt-out rows; source-discriminated Verifier output (`source: ci|local`) with the exact CI claim form `verified by CI: <check name(s)> succeeded on <sha> (run <url>)`; any blocking conclusion in the resolved set mints a `P#` with a third disposition `CI-infrastructure-broken` that triggers the fallback run; two new `## PR gate` keys `local verification: always` and `ci checks:`. Spec: `doc/specs/2026-09-06-gh-14-gatekeep-ci-evidence-default.md` (partially supersedes `doc/specs/2026-08-18-gh-9-gatekeep-pr-skill.md`, verification-evidence scope only).
package/README.md CHANGED
@@ -10,9 +10,9 @@ The gated workflow for the [pi coding agent](https://github.com/earendil-works/p
10
10
 
11
11
  ## The problem
12
12
 
13
- Point an agent at a task and let it loop until done - that's the easy 5%. A bare loop has nothing to aim at, nothing to stop it shipping the wrong thing, and no check that the final output matches what you actually asked for. It holds up on a narrow, well-specified task and drifts on anything open-ended: the agent reinterprets the ask as it goes, nobody catches it until review, and by then the diff is large enough that review is theater too.
13
+ Point an agent at a task and let it loop until done - that's the easy 5%. LLMs are more a compressed library with a sampler on top than an independent mind: they produce fluent analysis faster than humans can audit it, and humans can't efficiently unravel that flood of output from the authenticity of a sound idea. So the agent quietly drifts from what you asked, and by the time you look, the diff is too big to honestly review.
14
14
 
15
- That's not a model problem. Cursor, Claude Code, Codex, Devin all run some version of the same loop, and all of them drift the same way on long tasks - because nothing in the loop confronts the output against the *original* intent.
15
+ It *is* a model problem - one-shotting an idea makes a great demo, not a product. But no better model fixes it on its own: Cursor, Claude Code, and Codex all drift the same way on long tasks, because nothing in a bare loop confronts output against *original* intent, and a model cannot audit itself - the same blind spot that wrote the bug will happily approve it. A weak generator needs a strong harness - because fluency is not correctness.
16
16
 
17
17
  ## Why pi-gauntlet exists
18
18
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: spec-reviewer
3
- description: Independently verifies an implementation against its spec/plan. Trusts the artifacts, not the implementer's self-report.
3
+ description: Independently verifies an implementation against its spec, clause by clause. Trusts the spec and the code, not the implementer's self-report.
4
4
  tools: read, grep, find, ls, bash
5
5
  defaultContext: fresh
6
6
  inheritProjectContext: true
@@ -9,29 +9,31 @@ systemPromptMode: replace
9
9
  completionGuard: false
10
10
  ---
11
11
 
12
- You are a spec compliance reviewer. Your job is to verify that an implementation **actually does what its spec or plan says**, and nothing else. You are **skeptical of the implementer's self-report** — verify everything by reading code yourself.
12
+ You are a spec compliance reviewer. Your job is to verify that an implementation **actually does what its spec says**, and nothing else. You are **skeptical of the implementer's self-report** — verify everything by reading code yourself.
13
13
 
14
14
  ## Process
15
15
 
16
- 1. Read the spec/plan thoroughly. Extract a flat list of every requirement, acceptance criterion, and explicit non-goal.
17
- 2. Read the implementation (diff or relevant files). Do not trust summaries.
18
- 3. For each requirement, determine status by reading the code, not by reading the implementer's prose.
16
+ <!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md change them together or not at all -->
17
+
18
+ 1. Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses. Every clause gets a verdict row.
19
+ 2. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
20
+ 3. For each clause, determine status by reading the code, not by reading the implementer's prose.
19
21
  4. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
20
22
  5. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
21
- 6. Flag any requirement from the spec that is missing from the implementation.
23
+ 6. Flag any clause from the spec that is missing from the implementation.
22
24
 
23
25
  ## Output format
24
26
 
25
27
  ```
26
- Per-requirement status:
27
- - [MET] REQ-1: short requirement text — evidence: file.ts:42
28
- - [PARTIAL] F1: REQ-2: ... — evidence: file.ts:80; missing: ...
28
+ Per-clause status:
29
+ - [MET] C-1: short clause text — evidence: file.ts:42
30
+ - [PARTIAL] F1: C-2: ... — evidence: file.ts:80; missing: ...
29
31
  touched-files: file.ts
30
32
  touched-resources: none
31
- - [MISSING] F2: REQ-3: ... — searched: <where>
33
+ - [MISSING] F2: C-3: ... — searched: <where>
32
34
  touched-files: file.ts, other.ts
33
35
  touched-resources: none
34
- - [OUT_OF_SCOPE] REQ-4: ... — flagged as non-goal in spec
36
+ - [OUT_OF_SCOPE] C-4: ... — flagged as non-goal in spec
35
37
 
36
38
  Scope creep (not in spec, but present):
37
39
  - F3: widget.ts:120 — short description
@@ -39,7 +41,7 @@ Scope creep (not in spec, but present):
39
41
  touched-resources: none
40
42
 
41
43
  Missing from implementation:
42
- - F2: REQ-3 — short description
44
+ - F2: C-3 — short description
43
45
 
44
46
  Verdict: COMPLIANT | NEEDS_REWORK | OUT_OF_SCOPE_CHANGES
45
47
  Confidence: low | medium | high
@@ -49,7 +51,7 @@ Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch file.ts)
49
51
 
50
52
  ## Finding IDs and fix-concurrency certification
51
53
 
52
- Label every finding (each `PARTIAL`/`MISSING` requirement, each scope-creep
54
+ Label every finding (each `PARTIAL`/`MISSING` clause, each scope-creep
53
55
  item) with a globally unique ID `F1..Fn`, numbered across the whole report
54
56
  (no restart per section). Each finding carries:
55
57
 
@@ -82,5 +84,6 @@ certify a pair disjoint, mark them `conflicts` (conservative default = serial).
82
84
  - You are **read-only**. Never edit files.
83
85
  - Cite a real file:line for every MET/PARTIAL claim. If you cannot, downgrade to MISSING.
84
86
  - Do not negotiate scope with yourself. If the spec didn't ask for it, it's scope creep, even if it looks useful.
87
+ - Plan/task code snippets are implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
85
88
  - Never run tests, linters, or type-checkers. Read; do not execute checks.
86
89
  - Do not report code-quality opinions - naming, design, complexity, test aesthetics, style. Those belong to code-reviewer. Report only spec-vs-implementation deltas.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.0.3",
3
+ "version": "5.0.4",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -55,7 +55,7 @@ For each task in `plan_tracker`:
55
55
 
56
56
  1. **Dispatch implementer.** Pass the full task text + scene-setting context + the task's plan-declared test commands as `SCOPED_TEST_COMMANDS` (or `none`). Don't make the subagent re-read the plan.
57
57
  2. **Handle implementer status** (see below).
58
- 3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the diff matches the anchored spec — nothing missing, nothing extra.
58
+ 3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the change satisfies the anchored spec — nothing missing, nothing extra.
59
59
  4. If spec reviewer finds gaps → re-dispatch implementer to fix → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
60
60
  5. **Dispatch code-quality reviewer.** Only after spec is ✅. Skip for doc-only tasks (every file in the task's `Files:` block documentation-only) — SR-only, same exemption as doc-only waves. Pass `SCOPED_TEST_COMMANDS` = the task's plan-declared commands.
61
61
  6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
@@ -171,7 +171,7 @@ Auto-selected at handoff by `writing-plans` (any wave with ≥2 tasks) when the
171
171
 
172
172
  1. **Independence check.** Parse the wave's tasks' `Files:` blocks; assert pairwise-disjoint (mechanical). Runtime-resource disjointness (DB/schema, port, fixture, external service, shared temp path) is not machine-checkable here — trust the plan's wave grouping, which `writing-plans`' D5 contract guarantees. Either kind of overlap → the wave is mis-grouped; run those tasks as sequential single-task waves and note it.
173
173
  2. **Fan out.** One parallel dispatch (shape below): `implementer` per task, `context: "fresh"`, `worktree: true`. Each returns a status + a patch.
174
- 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review is **diff-based**: the diff's hunks carry `file:line`, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
174
+ 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review starts from the diff (its hunks carry `file:line`) and reads each touched file in full, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
175
175
  4. **Integrate.** `git apply` each task's patch sequentially onto HEAD. Apply fails = textual conflict → drop that task, finish the rest, re-run the dropped task sequentially on the updated HEAD.
176
176
  5. **Test gate.** Run the union of the wave's tasks' declared test commands on the integrated tree — the full verification set is the verify phase's job, run once. Failure = semantic conflict or bug → re-run the offending task sequentially, else fix per [When a Subagent Fails](#when-a-subagent-fails).
177
177
  6. **Quality review.** CR binds to the wave: exactly one **initial** code-review dispatch per code-touching wave, over the integrated wave diff - never per task within a wave, never batched across waves. Subsequent dispatches within the wave are re-reviews triggered only by findings, per Fix-Loop Rounds. Pass `SCOPED_TEST_COMMANDS` = the union of the wave's tasks' declared commands. Code-quality review on the integrated wave diff; loop fixes to ✅ within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential. Skip for doc-only waves (SR-only per the commit precondition below).
@@ -33,6 +33,7 @@ Dispatch a subagent with this prompt:
33
33
  - **Task-vs-spec divergence** (task says X, anchored spec says Y): unconditional flag; quote the spec literal with spec file:line so the fix re-dispatch carries authoritative wording. Never silently trust the task; never silently substitute the spec — the flag is the mechanism. Closure: the finding closes when the current patch conforms to the anchored spec; re-reviews judge the diff against the spec, not stale task prose — a divergence already corrected in the diff is not re-flagged.
34
34
  - **Anchor-less task** (Anchors: omitted): the task text alone is your contract; no out-of-anchor-slice or transcription-gap flagging — only nothing-extra-vs-the-chore review.
35
35
  - **Finding grammar:** divergence findings use the existing F1..Fn finding grammar - a finding kind by prose label, not a new schema; the `Parallel-safe:` and `TRAJECTORY:` grammars are untouched.
36
+ - **Plan/task code snippets:** implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
36
37
 
37
38
  ## CRITICAL: Do Not Trust the Report
38
39
 
@@ -46,6 +47,7 @@ Dispatch a subagent with this prompt:
46
47
 
47
48
  **DO:**
48
49
  - Read the actual code they wrote
50
+ - Read each touched file in full, not just the diff hunks, continuing in chunks; note in the report any touched file not read to the end
49
51
  - Compare actual implementation to requirements line by line
50
52
  - Check for missing pieces they claimed to implement
51
53
  - Look for extra features they didn't mention
@@ -61,6 +63,10 @@ Dispatch a subagent with this prompt:
61
63
 
62
64
  ## Your Job
63
65
 
66
+ <!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with agents/spec-reviewer.md — change them together or not at all -->
67
+
68
+ Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses.
69
+
64
70
  Read the implementation code and verify:
65
71
 
66
72
  **Missing requirements:**
@@ -255,7 +255,7 @@ Every code task carries this step (red -> green -> fmt/lint -> commit). Doc-only
255
255
 
256
256
  ## Spec Coverage Table
257
257
 
258
- Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners. Two row kinds:
258
+ Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. Two row kinds:
259
259
 
260
260
  ```markdown
261
261
  ## Spec coverage
@@ -293,6 +293,7 @@ If a decision is genuinely open, put it in an explicit **Open Questions** sectio
293
293
  After drafting the plan and before announcing it complete, run these checks yourself. This is a checklist you run yourself — not a subagent dispatch.
294
294
 
295
295
  - **Table closure (three legs).** Every `## Spec coverage` row's owner is a task-ID list, a spec-authorized `waived: <reason>`, or a mechanical-task row; every `### Task N` heading appears in >=1 row; every requirement row's anchor is contained in the anchor set of each listed owner task's `**Spec:**` line. Zero orphans, zero waived in-scope normative rows, zero row-vs-owner anchor mismatches. Each Documentation impact entry maps to a plan task (or explicit "none").
296
+ - **Code-vs-anchor sanity.** For each non-waived requirement row, re-read the anchored spec lines and confirm the owner tasks' bodies do what they say - mechanism present, not just the quoted literal. Fix the task, don't annotate.
296
297
  - **Quote integrity (spec -> task).** For every non-waived requirement row, extract each backtick-quoted literal inside the row's anchored spec lines (strip the backticks; skip `<placeholder>` template spans) and `grep -F` it against the owning task's body — zero misses. Planner-authored backticks elsewhere in tasks are never scanned; the input set is spec-side literals only.
297
298
  - **Anchor resolution.** For every task-level anchor (a `**Spec:**` line carrying `§`; the plan header's path line is exempt), the quoted heading text matches an ATX heading in the spec file and `L<start>-L<end>` is in-bounds, non-empty, and lies within that heading's section — zero unresolved anchors. Verify with `grep -n '^#'` plus a scoped `sed -n`. Ignore `#`-lines inside fenced code blocks when locating headings and section boundaries - a fenced markdown example is not a heading.
298
299
  - **Paths exist.** Every `Modify:` path in `Files:` blocks passes `test -f` after stripping any trailing `:line[-line]` suffix; a `Modify:` glob must expand to >=1 match; `Create:` and `Test:` paths are exempt unless the `Test:` path is also listed under `Modify:`. Zero missing.
@@ -309,6 +310,7 @@ Fix what this review finds before handoff.
309
310
 
310
311
  - Exact file paths always
311
312
  - Complete code in plan (not "add validation")
313
+ - Plan code is guidance for the implementer, not review authority - reviewers judge the diff against the spec, never against plan snippets
312
314
  - Exact commands with expected output
313
315
  - Reference relevant skills
314
316
  - DRY, YAGNI, TDD, frequent commits