pi-gauntlet 5.3.4 → 5.3.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +10 -0
- package/README.md +1 -1
- package/agents/code-reviewer.md +4 -1
- package/agents/spec-reviewer.md +13 -0
- package/extensions/lib/plan-check.test.ts +29 -0
- package/extensions/lib/plan-check.ts +16 -0
- package/extensions/phase-tracker.ts +1 -1
- package/package.json +1 -1
- package/skills/linear/SKILL.md +3 -1
- package/skills/requesting-code-review/code-reviewer.md +3 -1
- package/skills/subagent-driven-development/SKILL.md +2 -0
- package/skills/subagent-driven-development/code-quality-reviewer-prompt.md +5 -3
- package/skills/subagent-driven-development/spec-reviewer-prompt.md +1 -0
- package/skills/writing-plans/SKILL.md +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,15 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.3.6 - 2026-09-08
|
|
4
|
+
|
|
5
|
+
- `plan_check`: new `waiver-literal` check (9 checks) - a `waived:` coverage row whose requirement names an inline code literal fails; `writing-plans` restates the waiver criterion (out of scope **and** excludes work) and requires cross-cutting requirements to list every deciding task. (#25)
|
|
6
|
+
- `spec-reviewer` (prompt + persona): each value/threshold/trigger clause carries `spec-condition:` / `code-condition:`; a mismatch caps the clause at `PARTIAL` regardless of passing tests. (#25)
|
|
7
|
+
- `code-reviewer` (persona, generic template, SDD prompt): report-level `Behaviour-change: yes | no` sentinel on every report; `subagent-driven-development` routes a `yes` fix round through SR before CR. (#25)
|
|
8
|
+
|
|
9
|
+
## v5.3.5 - 2026-09-08
|
|
10
|
+
|
|
11
|
+
- `linear`: section 3 gains a `Download` row (`linearis files download <url> --output <path>`); section 8 gains a row for the linearis 2026.7.0/2026.8.0 Bearer-prefix bug on personal API keys (a 401 on `files download` while `issues read` works is not an auth problem - do not re-auth; [linearis-oss/linearis#300](https://github.com/linearis-oss/linearis/issues/300)), and the generic 401 row defers to it.
|
|
12
|
+
|
|
3
13
|
## v5.3.4 - 2026-09-07
|
|
4
14
|
|
|
5
15
|
- `spec-council-member`, `spec-council-synthesizer`, `conformance-reviewer`: the `over-spec` rules rewritten as short numbered imperatives so smaller council/closure models follow them; no semantic change (verified by an independent parity review).
|
package/README.md
CHANGED
|
@@ -71,7 +71,7 @@ pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
|
|
|
71
71
|
|
|
72
72
|
- **17 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Four more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`.
|
|
73
73
|
- **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
|
|
74
|
-
- **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers plan_check, a deterministic plan linter (
|
|
74
|
+
- **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers plan_check, a deterministic plan linter (9 mechanical plan-vs-spec checks) whose pass stamp gates implement-start inside a gauntlet flow. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
|
|
75
75
|
|
|
76
76
|
pi-gauntlet is **opinionated**: every non-trivial change is *meant* to ride this one pipeline, entered through `brainstorming`. Enforcement is opt-in by entry, not ambient: once brainstorming starts a flow, the phase-tracker extension mechanically blocks a phase from closing before its gate runs, and warns once if the main loop writes code during implement (subagents own implement-phase edits). A change made *without* entering the flow (a typo, a formatting run, a dependency bump - see "When to use / when NOT to use") is not gated; the discipline of routing real work through the pipeline is a convention the tooling supports, not a trap it springs on every edit.
|
|
77
77
|
|
package/agents/code-reviewer.md
CHANGED
|
@@ -41,6 +41,7 @@ Findings:
|
|
|
41
41
|
|
|
42
42
|
Complexity: net -<N> lines (omit if nothing to cut)
|
|
43
43
|
Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch auth.ts)
|
|
44
|
+
Behaviour-change: yes | no
|
|
44
45
|
```
|
|
45
46
|
|
|
46
47
|
Severity:
|
|
@@ -53,7 +54,7 @@ Label every finding with a globally unique `F1..Fn` ID (no restart per severity)
|
|
|
53
54
|
and a `touched-files:`/`touched-resources:` pair (files/resources a fix would
|
|
54
55
|
touch, or the literal `none`). On any issue-bearing review end the findings
|
|
55
56
|
with one partition line over the `Fn` IDs assigned above; when a task requires
|
|
56
|
-
a trailing `TRAJECTORY:` verdict (re-review), that verdict
|
|
57
|
+
a trailing `TRAJECTORY:` verdict (re-review), that verdict comes after `Behaviour-change:` as the
|
|
57
58
|
true final line:
|
|
58
59
|
|
|
59
60
|
<!-- grammar identical to skills/requesting-code-review/code-reviewer.md — change them together or not at all; writing-plans' plan-time Parallel-safe: line is a deliberately different free-text form, do NOT unify -->
|
|
@@ -72,4 +73,6 @@ forces `conflicts`. Runtime-resource disjointness is estimated over: DB/schema,
|
|
|
72
73
|
port, fixture, external service, shared temp path. When you cannot confidently
|
|
73
74
|
certify a pair disjoint, mark them `conflicts` (conservative default = serial).
|
|
74
75
|
|
|
76
|
+
Footer order: `Parallel-safe:` when present (issue-bearing reviews only), then `Behaviour-change:` on **every** report including clean ones, then `TRAJECTORY:` when a re-review trigger fired - `TRAJECTORY:` stays the true final line. `Behaviour-change: yes` when applying any Critical or Moderate fix would alter observable behaviour - values, control flow, routing, emitted output, persisted state; `no` when every fix is structural or stylistic, and on clean reports.
|
|
77
|
+
|
|
75
78
|
If you ran verification commands, quote them and their output verbatim under a `Verification:` section. If you did not, say so.
|
package/agents/spec-reviewer.md
CHANGED
|
@@ -24,10 +24,23 @@ You are a spec compliance reviewer. Your job is to verify that an implementation
|
|
|
24
24
|
|
|
25
25
|
## Output format
|
|
26
26
|
|
|
27
|
+
**Condition match:** for every anchored clause that fixes a value, threshold, comparison, or trigger ("only when", "unless", "if", a literal), the clause row carries two indented sub-lines, before `touched-files:` where present:
|
|
28
|
+
|
|
29
|
+
```
|
|
30
|
+
spec-condition: <clause fragment quoted from the spec>
|
|
31
|
+
code-condition: <what the code checks, file:line>
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
If the two differ, the clause is `PARTIAL` at most - regardless of passing tests. A plausible condition is not the specified condition.
|
|
35
|
+
|
|
27
36
|
```
|
|
28
37
|
Per-clause status:
|
|
29
38
|
- [MET] C-1: short clause text — evidence: file.ts:42
|
|
39
|
+
spec-condition: "unless the path is absolute"
|
|
40
|
+
code-condition: `!isAbsolute(p)` file.ts:42
|
|
30
41
|
- [PARTIAL] F1: C-2: ... — evidence: file.ts:80; missing: ...
|
|
42
|
+
spec-condition: "only when the path normalizes outside the leaf"
|
|
43
|
+
code-condition: `startsWith("..")` file.ts:80
|
|
31
44
|
touched-files: file.ts
|
|
32
45
|
touched-resources: none
|
|
33
46
|
- [MISSING] F2: C-3: ... — searched: <where>
|
|
@@ -197,6 +197,35 @@ test("check 1 owner-cell grammar: empty waiver reason is malformed, not accepted
|
|
|
197
197
|
);
|
|
198
198
|
});
|
|
199
199
|
|
|
200
|
+
const WAIVED_ROW = '| § "Other" L12-L13 | out of scope thing | waived: out of scope per spec |';
|
|
201
|
+
|
|
202
|
+
test("check 9 waiver-literal: waived row whose requirement names a code literal fails", () => {
|
|
203
|
+
const row = '| § "Other" L12-L13 | `protectedPaths: []` stays empty | waived: out of scope |';
|
|
204
|
+
const mutated = VALID_PLAN.replace(WAIVED_ROW, row);
|
|
205
|
+
const findings = checkPlan(mutated, SPEC_TEXT, alwaysTruePort());
|
|
206
|
+
const wl = findingsFor(findings, "waiver-literal");
|
|
207
|
+
assert.equal(wl.length, 1, `expected exactly one waiver-literal finding, got: ${JSON.stringify(wl)}`);
|
|
208
|
+
assert.equal(wl[0].line, lineOf(mutated, row));
|
|
209
|
+
assert.equal(wl[0].reason, "waived row names a code literal; waive only requirements that exclude work");
|
|
210
|
+
});
|
|
211
|
+
|
|
212
|
+
test("check 9 waiver-literal: same requirement owned by a task passes the whole plan", () => {
|
|
213
|
+
let mutated = VALID_PLAN.replace(WAIVED_ROW, '| § "Other" L12-L13 | `protectedPaths: []` stays empty | Task 1 |');
|
|
214
|
+
mutated = mutated.replace(
|
|
215
|
+
'**Spec:** doc/specs/fixture-spec.md § "Design" L4-L6',
|
|
216
|
+
'**Spec:** doc/specs/fixture-spec.md § "Design" L4-L6, § "Other" L12-L13',
|
|
217
|
+
);
|
|
218
|
+
assert.deepEqual(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), []);
|
|
219
|
+
});
|
|
220
|
+
|
|
221
|
+
test("check 9 waiver-literal: waived prose requirement without backticks passes the whole plan", () => {
|
|
222
|
+
const mutated = VALID_PLAN.replace(
|
|
223
|
+
WAIVED_ROW,
|
|
224
|
+
'| § "Other" L12-L13 | do not add support for protectedPaths | waived: out of scope per spec |',
|
|
225
|
+
);
|
|
226
|
+
assert.deepEqual(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), []);
|
|
227
|
+
});
|
|
228
|
+
|
|
200
229
|
const VERIFICATION_ROW = '| § "Acceptance" L16 | full suite passes | Verification |';
|
|
201
230
|
|
|
202
231
|
function withRow(plan: string, row: string): string {
|
|
@@ -817,6 +817,21 @@ function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
|
|
|
817
817
|
return findings;
|
|
818
818
|
}
|
|
819
819
|
|
|
820
|
+
function checkWaiverLiteral(parsed: ParsedPlan): PlanCheckFinding[] {
|
|
821
|
+
const findings: PlanCheckFinding[] = [];
|
|
822
|
+
for (const row of parsed.coverageRows) {
|
|
823
|
+
if (!row.isWaived) continue;
|
|
824
|
+
if (!/`[^`]+`/.test(row.requirementCell)) continue;
|
|
825
|
+
findings.push({
|
|
826
|
+
check: "waiver-literal",
|
|
827
|
+
line: row.line,
|
|
828
|
+
text: row.text,
|
|
829
|
+
reason: "waived row names a code literal; waive only requirements that exclude work",
|
|
830
|
+
});
|
|
831
|
+
}
|
|
832
|
+
return findings;
|
|
833
|
+
}
|
|
834
|
+
|
|
820
835
|
export function checkPlan(planText: string, specText: string, fs: FsPort): PlanCheckFinding[] {
|
|
821
836
|
try {
|
|
822
837
|
const parsed = parsePlan(planText);
|
|
@@ -851,6 +866,7 @@ export function checkPlan(planText: string, specText: string, fs: FsPort): PlanC
|
|
|
851
866
|
findings.push(...checkWaveFileDisjointness(parsed, fs));
|
|
852
867
|
findings.push(...checkSoloLine(parsed));
|
|
853
868
|
findings.push(...checkHeaderEntrypoint(parsed));
|
|
869
|
+
findings.push(...checkWaiverLiteral(parsed));
|
|
854
870
|
return findings;
|
|
855
871
|
} catch (err) {
|
|
856
872
|
const message = String(err instanceof Error ? err.message : err);
|
|
@@ -694,7 +694,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
694
694
|
name: "plan_check",
|
|
695
695
|
label: "Plan Check",
|
|
696
696
|
description:
|
|
697
|
-
"Deterministically verify an implementation plan against its spec (
|
|
697
|
+
"Deterministically verify an implementation plan against its spec (9 mechanical checks); " +
|
|
698
698
|
"a pass stamps the plan for implement-start.",
|
|
699
699
|
parameters: PlanCheckParams,
|
|
700
700
|
async execute(_toolCallId, params, _signal, _onUpdate, ctx) {
|
package/package.json
CHANGED
package/skills/linear/SKILL.md
CHANGED
|
@@ -108,6 +108,7 @@ treated as absent.
|
|
|
108
108
|
| Labels, teams, users, cycles | `linearis labels list`, `linearis teams list`, `linearis users list`, `linearis cycles list` | Use to resolve names to IDs; see id-cache convention. |
|
|
109
109
|
| Attachments | `linearis attachments create [<issue>] --url <url>` | Positional is optional (`--issue <issue>` alias); link-only, no inline render - see gotcha (d). |
|
|
110
110
|
| Upload | `linearis files upload <file>` | Returns an `assetUrl` for inline embedding - see gotcha (d). |
|
|
111
|
+
| Download | `linearis files download <url> --output <path>` | `<url>` is an attachment/asset URL from `issues read --with-attachments`; asset URLs are short-lived (gotcha d). A 401 here while `issues read` works is not an auth problem - see section 8. |
|
|
111
112
|
|
|
112
113
|
Workspace values above (`<default team>`, `<who>`, etc.) are placeholders bound to
|
|
113
114
|
the override keys in section 2 - never a real urlKey, team prefix, or email.
|
|
@@ -190,7 +191,8 @@ Safety rules, in addition to the write gate above:
|
|
|
190
191
|
|
|
191
192
|
| Symptom | Cause | Fix |
|
|
192
193
|
|---|---|---|
|
|
193
|
-
| 401 | Not authenticated / expired token | `linearis auth status`; re-auth. |
|
|
194
|
+
| 401 | Not authenticated / expired token | `linearis auth status`; re-auth - unless the download row below applies. |
|
|
195
|
+
| 401 on `files download` while `issues read` works | linearis 2026.7.0 and 2026.8.0 prepend `Bearer ` to personal API keys on file downloads ([linearis-oss/linearis#300](https://github.com/linearis-oss/linearis/issues/300)) | Not an auth problem - do not re-auth. Fetch the URL with the bare key, or use a version without the bug once one ships. |
|
|
194
196
|
| Issue not found | Wrong workspace, or issue archived | Confirm workspace; check archived state. |
|
|
195
197
|
| Status not found | Status name doesn't match the team's workflow states | List the team's states before setting one. |
|
|
196
198
|
| Missing `--team` error on create | `--team` is required | Supply `--team <default team>`. |
|
|
@@ -119,7 +119,7 @@ git diff {BASE_SHA}..{HEAD_SHA}
|
|
|
119
119
|
|
|
120
120
|
### Fix-concurrency certification
|
|
121
121
|
|
|
122
|
-
On any issue-bearing review,
|
|
122
|
+
On any issue-bearing review, emit one partition line over the
|
|
123
123
|
`Fn` IDs assigned above:
|
|
124
124
|
|
|
125
125
|
<!-- grammar identical to agents/conformance-reviewer.md (modulo G vs F id prefix) — change them together or not at all; writing-plans' plan-time Parallel-safe: line is a deliberately different free-text form, do NOT unify -->
|
|
@@ -138,6 +138,8 @@ forces `conflicts`. Runtime-resource disjointness is estimated over: DB/schema,
|
|
|
138
138
|
port, fixture, external service, shared temp path. When you cannot confidently
|
|
139
139
|
certify a pair disjoint, mark them `conflicts` (conservative default = serial).
|
|
140
140
|
|
|
141
|
+
Footer order: `Parallel-safe:` when present (issue-bearing reviews only), then `Behaviour-change:` on **every** report including clean ones, then `TRAJECTORY:` when a re-review trigger fired - `TRAJECTORY:` stays the true final line. `Behaviour-change: yes` when applying any Critical or Moderate fix would alter observable behaviour - values, control flow, routing, emitted output, persisted state; `no` when every fix is structural or stylistic, and on clean reports.
|
|
142
|
+
|
|
141
143
|
## Critical Rules
|
|
142
144
|
|
|
143
145
|
**DO:**
|
|
@@ -75,6 +75,8 @@ One rule governs both review loops - spec-compliance and code-quality - in seque
|
|
|
75
75
|
|
|
76
76
|
**Fix fan-out.** When the triggering review's `Parallel-safe:` line certifies a `disjoint` group of ≥ 2 findings, dispatch that fix round per `dispatching-parallel-agents` "Fix fan-out"; the fan-out counts as **one** fix against this budget, its scoped test gate is the consuming task/wave's plan-declared commands, and one re-review of the integrated delta follows.
|
|
77
77
|
|
|
78
|
+
**Behaviour-change reroute.** When the triggering CR report carries `Behaviour-change: yes`, the fix round's re-review is SR first, then CR. The SR dispatch reviews the fix diff against the task's spec anchors as a first review (no `## Previous review report` marker, so no `TRAJECTORY:` line; SR carries no test commands), but its ordinal continues the task's SR-loop count - a task whose SR loop ended at review 2 gets review 3 here, and an issue-bearing rerouted SR after review 4 escalates. Issues follow the normal sequence: fix, then SR re-review pasting this SR's report. CR round numbering is unchanged. `Behaviour-change: no` re-reviews with CR only. A missing or malformed `Behaviour-change:` line is re-asked once like `Parallel-safe:` (see `dispatching-parallel-agents` "Structural probe"); still missing -> route through SR, never default to `no`. In wave mode the SR re-review targets the task(s) whose files the fix touched.
|
|
79
|
+
|
|
78
80
|
Every fix re-dispatch (implementer) and code-review re-review carries the consuming task/wave's `SCOPED_TEST_COMMANDS`; spec-reviewer re-reviews carry none - SR never executes.
|
|
79
81
|
|
|
80
82
|
**The sequence.** Each review that finds issues is a decision point: read the `TRAJECTORY:` line before dispatching anything (review 1 has no line - on issues, dispatch fix 1). Any clean review ends the loop.
|
|
@@ -25,7 +25,9 @@ Dispatch a subagent with the code-reviewer template:
|
|
|
25
25
|
|
|
26
26
|
**Code reviewer returns:** Strengths, Issues (Critical/Moderate/Minor), Assessment
|
|
27
27
|
|
|
28
|
-
Emit finding IDs and the `Parallel-safe:` line per that contract.
|
|
28
|
+
Emit finding IDs and the `Parallel-safe:` line per that contract, then the `Behaviour-change:` line.
|
|
29
|
+
|
|
30
|
+
Footer order: `Parallel-safe:` when present (issue-bearing reviews only), then `Behaviour-change:` on **every** report including clean ones, then `TRAJECTORY:` when a re-review trigger fired - `TRAJECTORY:` stays the true final line. `Behaviour-change: yes` when applying any Critical or Moderate fix would alter observable behaviour - values, control flow, routing, emitted output, persisted state; `no` when every fix is structural or stylistic, and on clean reports.
|
|
29
31
|
|
|
30
32
|
## Re-review: trajectory verdict
|
|
31
33
|
|
|
@@ -34,8 +36,8 @@ the prior review report pasted verbatim under a
|
|
|
34
36
|
`## Previous review report (re-review trigger)` heading:
|
|
35
37
|
|
|
36
38
|
If your task contains a "Previous review report (re-review trigger)" section
|
|
37
|
-
and you found issues, append exactly one more line after `
|
|
38
|
-
line, not `
|
|
39
|
+
and you found issues, append exactly one more line after `Behaviour-change:` — this
|
|
40
|
+
line, not `Behaviour-change:`, is the true final line of the report:
|
|
39
41
|
|
|
40
42
|
TRAJECTORY: CONVERGING (<n_prev> -> <n_now>, max severity <X>)
|
|
41
43
|
TRAJECTORY: DIVERGING
|
|
@@ -33,6 +33,7 @@ Dispatch a subagent with this prompt:
|
|
|
33
33
|
- **Task-vs-spec divergence** (task says X, anchored spec says Y): unconditional flag; quote the spec literal with spec file:line so the fix re-dispatch carries authoritative wording. Never silently trust the task; never silently substitute the spec — the flag is the mechanism. Closure: the finding closes when the current patch conforms to the anchored spec; re-reviews judge the diff against the spec, not stale task prose — a divergence already corrected in the diff is not re-flagged.
|
|
34
34
|
- **Anchor-less task** (Anchors: omitted): the task text alone is your contract; no out-of-anchor-slice or transcription-gap flagging — only nothing-extra-vs-the-chore review.
|
|
35
35
|
- **Finding grammar:** divergence findings use the existing F1..Fn finding grammar - a finding kind by prose label, not a new schema; the `Parallel-safe:` and `TRAJECTORY:` grammars are untouched.
|
|
36
|
+
- **Condition match:** for every anchored clause that fixes a value, threshold, comparison, or trigger ("only when", "unless", "if", a literal), the clause row carries two indented sub-lines, before `touched-files:` where present: `spec-condition: <clause fragment quoted from the spec>` and `code-condition: <what the code checks, file:line>`. If the two differ, the clause is `PARTIAL` at most - regardless of passing tests. A plausible condition is not the specified condition.
|
|
36
37
|
- **Plan/task code snippets:** implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
|
|
37
38
|
|
|
38
39
|
## CRITICAL: Do Not Trust the Report
|
|
@@ -270,7 +270,7 @@ Every plan ends with a `## Spec coverage` section — authored last, placed afte
|
|
|
270
270
|
| - | mechanical: release commit | Task 7 |
|
|
271
271
|
```
|
|
272
272
|
|
|
273
|
-
- **Requirement rows:** anchor + short requirement + owner = task-ID list, or `Verification`, or `waived: <reason
|
|
273
|
+
- **Requirement rows:** anchor + short requirement + owner = task-ID list, or `Verification`, or `waived: <reason>`. A cross-cutting requirement (decided in more than one task) lists **every** deciding task as owner, not the first. `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
|
|
274
274
|
- **`Verification` owner:** use for a requirement the header `**Verification:**` command proves. Write the exact string `Verification`, alone. Quote only literals contained in that header. Anchor the single requirement line. Keep scoped commands task-owned.
|
|
275
275
|
- **Mechanical-task rows:** anchor `-`, requirement `mechanical: <short>`, owner = the task ID. One such row per anchor-less task.
|
|
276
276
|
- The table is plan-authoring-time only — never passed to implementer or reviewer dispatches.
|
|
@@ -295,13 +295,13 @@ If a decision is genuinely open, put it in an explicit **Open Questions** sectio
|
|
|
295
295
|
|
|
296
296
|
After drafting the plan and before announcing it complete, run the deterministic checker, then the judgment checks yourself — not a subagent dispatch.
|
|
297
297
|
|
|
298
|
-
- **Deterministic checker.** Run `plan_check({ planPath })` on the saved plan. Assess and fix every finding yourself (no human involvement), then re-run until it passes — a pass writes the execution stamp that implement-start verifies mechanically. If the same finding survives 3 fix rounds, convert it to an explicit Open Question and stop (the pre-existing Open-Questions halt, resolved by the human in-session — not a new gate). The checker covers table closure, quote integrity, anchor resolution, path existence, placeholder scan, wave file-disjointness, solo-line presence,
|
|
298
|
+
- **Deterministic checker.** Run `plan_check({ planPath })` on the saved plan. Assess and fix every finding yourself (no human involvement), then re-run until it passes — a pass writes the execution stamp that implement-start verifies mechanically. If the same finding survives 3 fix rounds, convert it to an explicit Open Question and stop (the pre-existing Open-Questions halt, resolved by the human in-session — not a new gate). The checker covers table closure, quote integrity, anchor resolution, path existence, placeholder scan, wave file-disjointness, solo-line presence, header-only entrypoint, and waiver-literal.
|
|
299
299
|
- **Code-vs-anchor sanity.** For each task-owned requirement row, re-read the anchored spec lines and confirm the owner tasks' bodies do what they say - mechanism present, not just the quoted literal. For each `Verification` row, confirm the header command exercises the anchored requirement. Fix the task, don't annotate.
|
|
300
300
|
- **Type / API consistency.** Function signatures and field names that appear in multiple tasks must match exactly. The plan is its own contract — internal contradictions surface as bugs during execution.
|
|
301
301
|
- **Scoped-test coverage.** Every code-touching wave declares at least one scoped test command; only doc-only waves may have none.
|
|
302
302
|
- **Runtime-resource disjointness.** For every multi-task wave, confirm no two tasks contend on a shared mutable runtime resource (DB/schema, port, fixture, external service, shared temp path) — `Files:` overlap is checked mechanically, resource contention is not. Contention = mis-grouped wave; split or re-order before handoff.
|
|
303
303
|
- **Solo-reason validity.** Every single-task wave's `Solo:` line (presence is checked mechanically) must name its specific blocker — the blocking task/wave, the contended resource, or `lone remaining task`. Category-only justifications are under-justified; merge or justify before handoff.
|
|
304
|
-
- **Waiver authorization.**
|
|
304
|
+
- **Waiver authorization.** `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
|
|
305
305
|
- **Verification-ownership authorization.** `Verification` on a requirement no header command exercises is a Self-Review failure.
|
|
306
306
|
- **Documentation-impact mapping.** Each Documentation impact entry maps to a plan task (or explicit "none").
|
|
307
307
|
|