pi-gauntlet 5.17.1 → 5.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,9 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.18.0 - 2026-09-23
4
+
5
+ - Projects can declare end-to-end happy-path commands in a `## Happy path` overrides table (`Row | Paths | Command | Timeout`, contract in `README.md`); `writing-plans` selects the row covering the plan's files into an optional `**Happy path:**` header line (`plan_check` validates it and bans the command from tasks and wave prose), `subagent-driven-development` runs it once in the verify phase between code review and the conformance audit under `timeout -k 30s` via `bash -c`, and the `conformance-reviewer` reads the bounded transcript as runtime evidence (`passed` / `failed` / `not run`; a failure in code the change never touched is `rescope`, never `fix`). Fix rounds re-run it only when the round touches the row's paths or an open gap cites the transcript; the closure sentinel and the finish render carry a `happy-path:` line. Optional end to end - no section, no change.
6
+
3
7
  ## v5.17.1 - 2026-09-22
4
8
 
5
9
  - Fix finishing's verification-skip check to inspect the working tree, index, and untracked files outside the configured telemetry directory, rather than comparing commits alone. Require a known clean verified commit; uncertainty or a failed Git check reruns verification. Add executable Git regression tests to CI.
package/README.md CHANGED
@@ -39,7 +39,7 @@ Concretely, one change through the gauntlet:
39
39
  1. You describe the change. **`brainstorming`** sets up an isolated worktree, explores the codebase, and turns your description into a written spec; when the run starts from a ticket, the spec carries the ticket's acceptance criteria verbatim, each with a disposition (`in-scope`, `deviates:`, `deferred:`, `venue:`), the conformance gate checks the in-scope ones, and `/skill:check-delivery` verifies `venue:` rows after deploy. A multi-model critique runs on it automatically. If the spec replaces a known prior spec, brainstorming marks the predecessor with a `> **Superseded by:**` banner under its title (default format, syntax overridable via `.pi/gauntlet-overrides.md`; candidates come from brainstorming's scout recon, never a mechanical sweep). **You read and approve the spec - human gate 1.** No implementation code exists yet.
40
40
  2. **`writing-plans`** decomposes the approved spec into atomic, independently-verifiable tasks, grouped into parallel waves where they don't touch the same files.
41
41
  3. **`subagent-driven-development`** executes the plan one task at a time, each in a fresh subagent, behind spec-compliance review then code-quality review. TDD-locked: red, green, refactor. Independent implementation, review, and council batches remain parallel, but their dispatches are foreground: the orchestrator waits for terminal results before accepting work or advancing a phase.
42
- 4. **verify**: the parent first runs the plan's full verification command set; only a passing result permits whole-diff code review, then the **conformance gate**. A review fix invalidates that result, so the parent reruns full verification before the next review or conformance gate. The conformance subagent reads the finished code and docs against your *original words* from step 1, not the plan, and reports per-requirement: delivered, partial, missing, drifted, or unauthorized. Unauthorized rows - shipped surface no human input asked for, whether it crept in or was laundered through the spec - follow their recommendation like every other row: a contained removal auto-runs, anything another requirement leans on is deferred with a plain-language "I'd cut it / I'd keep it" recommendation. Inside a brainstorming-entered flow this gate is machine-blocked from being skipped. Compatible executable recommendations auto-run through an isolated fix-and-re-audit loop with no prompt; anything still open surfaces as a dense list - one line per decision, plain-language, with its recommended choice inline. Reply `1` to take every recommendation, or `2:` with per-item overrides; a current `CONFORMS` / no-concerns result goes straight to the branch options with no extra conformance sign-off.
42
+ 4. **verify**: the parent first runs the plan's full verification command set; only a passing result permits whole-diff code review, then the **conformance gate**. A review fix invalidates that result, so the parent reruns full verification before the next review or conformance gate. When the overrides file declares a happy path for the touched paths (see [Project-specific overrides](#project-specific-overrides)), the parent runs it once here, between code review and the conformance gate, and the reviewer reads its transcript as runtime evidence - a hang or a failed request shows up in the audit with the transcript quoted - as a gap when an origin clause ties to it, otherwise as a `happy-path: failed - unattributable` line - never as a silent green. The conformance subagent reads the finished code and docs against your *original words* from step 1, not the plan, and reports per-requirement: delivered, partial, missing, drifted, or unauthorized. Unauthorized rows - shipped surface no human input asked for, whether it crept in or was laundered through the spec - follow their recommendation like every other row: a contained removal auto-runs, anything another requirement leans on is deferred with a plain-language "I'd cut it / I'd keep it" recommendation. Inside a brainstorming-entered flow this gate is machine-blocked from being skipped. Compatible executable recommendations auto-run through an isolated fix-and-re-audit loop with no prompt; anything still open surfaces as a dense list - one line per decision, plain-language, with its recommended choice inline. Reply `1` to take every recommendation, or `2:` with per-item overrides; a current `CONFORMS` / no-concerns result goes straight to the branch options with no extra conformance sign-off.
43
43
  5. **`finishing-a-development-branch`**: PR, draft PR, squash, keep, or discard. Once a PR exists, run `/skill:gatekeep-pr <pr>` to verify it against its issue before merging. **Human gate 2** - the only other decision you make.
44
44
  6. *(Optional)* Once the merge lands, `/skill:check-delivery <ref>` can prove delivery - default-branch landing, delivery target, per-AC evidence - before the tracker status advances. Explicit invocation only, no auto-chain: deploys commonly lag merges by minutes to hours, so an auto-run would routinely check too early.
45
45
 
@@ -310,6 +310,25 @@ unknown value.
310
310
  One deploy workflow ships the whole system at once - nothing ships independently.
311
311
  ```
312
312
 
313
+ **`## Happy path` section:** `writing-plans` selects the row whose `Paths` cover the plan's files and copies it into the plan header's `**Happy path:**` line; `subagent-driven-development` runs it once in the verify phase - after full verification and code review, before the conformance audit - and hands a bounded transcript to the conformance reviewer as runtime evidence. Optional: without the section, or with no matching row, nothing downstream changes.
314
+
315
+ ```markdown
316
+ ## Happy path
317
+
318
+ | Row | Paths | Command | Timeout |
319
+ |---|---|---|---|
320
+ | dashboard | `dashboard/` | `script/e2e-dashboard` | 3m |
321
+ | excavation | `excavation/` | `cd excavation && make e2e` | 2m |
322
+ | cross-cutting | `dashboard/`, `excavation/`, `docker-compose.yml`, `Makefile` | `script/e2e-stack` | 5m |
323
+ ```
324
+
325
+ - `Row` is a free label; `cross-cutting` is reserved: it applies when the change matches two or more non-`cross-cutting` rows, or when a path is inside its own `Paths` and inside no other row's `Paths`, and it takes precedence over a single matched row (so root compose files, Makefiles, and lockfiles a stack run depends on select it on their own, while a change inside one project row alone selects that row).
326
+ - `Paths` is a comma-separated list of repo-relative prefixes; a path is inside a row when it starts with one of them. Paths matching no row are ignored, both for selection and for the fix-loop re-run test.
327
+ - `Command` is repo-relative and runs through `bash -c`, so `cd`, `&&`, and env assignments work. It must be self-contained and worktree-safe: boot what it needs, drive one flow, exit, trap `TERM`/`INT` to tear down (containers are not in the process group, so the script owns their teardown), and write only to gitignored paths or outside the worktree. Two worktrees may run it at the same time; pi-gauntlet does not serialize, and knows nothing about brokers, compose files, or ports.
328
+ - `Timeout` is optional, grammar `\d+(s|m|h)`, default `10m`. The parent runs the command under `timeout -k 30s <Timeout>`: `TERM`, then `KILL` after 30s. A malformed cell falls back to the default and the plan header omits the suffix.
329
+ - Exit code contract: `0` = passed; `75` (`EX_TEMPFAIL`) = environment unavailable, reported as `not run`; `126` = not executable, reported as `not run`; any other non-zero = failed. A timeout is failed, not `not run`. Residue left in the worktree is failed regardless of exit code.
330
+ - Host dependency: GNU `timeout` (`timeout` or `gtimeout` on PATH); absent -> `not run - no timeout binary`.
331
+
313
332
  **`## Delivery` section:** `check-delivery` resolves its overrides through the same discovery ladder. Defaults are pessimistic where it matters: an unset `target state` keeps the write comment-only; unset `deploy watch`/`delivery target` skip stage 2 (reported, never silently passed); `browser evidence` defaults to never. `check-delivery` is single-ticket by design - sweep/reconciliation passes over many tickets stay consumer territory, invoking the skill once per ticket. The remaining slots have working defaults shown below:
314
333
 
315
334
  | Slot | Meaning | Default (unset) |
@@ -42,7 +42,7 @@ Work flows `origin (prompt + spec) → plan → code/doc`. Every hop is lossy: a
42
42
 
43
43
  **Spec `## Acceptance criteria`.** When the spec has a heading starting `## Acceptance criteria` (a legacy `## Acceptance criteria (from #N)` heading matches), read that section as an origin beside the spec body. Each `in-scope` row yields one `Rn` with `origin: spec "Acceptance criteria" - "<AC verbatim>"`; a row with no disposition line is `in-scope`. Each `venue:` row yields one `Rn` that is `DELIVERED` when the Design clauses naming its enabling change are `DELIVERED`, with the venue and observation text on the verdict line; the observation itself is never a finding; a `venue:` row no Design clause names is `MISSING`, `recommended: rescope`. `deviates:` and `deferred:` rows yield no `Rn`; list them under `Origin drift` as `recorded in spec? yes`. A disposition word outside the four is `DRIFTED`, `recommended: fix` (correct the word). A section whose body is a single `none - <reason>` line has no rows: it yields no `Rn` and no drift. A spec without the section is read as before - spec body and prompt only; every existing rule and the source order above are unchanged. When a spec exists you do not fetch the ticket.
44
44
  2. **Check origin drift.** If the spec and the prompt/ticket disagree, do **not** absorb it silently. A deviation recorded in the spec → spec wins (it was review-gated). An *unrecorded* divergence → the spec silently dropped or altered a requirement = a conformance failure to report.
45
- 3. **Map each requirement to the deliverable.** Read the diff (code **and** docs) yourself — do not trust any summary. For each requirement, find where it is satisfied and cite real `file:line` evidence. Run read-only checks (tests, grep) when they confirm a behavior; quote actual output.
45
+ 3. **Map each requirement to the deliverable.** Read the diff (code **and** docs) yourself — do not trust any summary. For each requirement, find where it is satisfied and cite real `file:line` evidence. Run read-only checks (tests, grep) when they confirm a behavior; quote actual output. When the dispatch names a happy-path summary (`Happy path: <path> (<outcome>)`), read it as runtime evidence; never run the happy-path command yourself - the summary is the only runtime evidence. Cite transcript evidence as `<abs summary path>:<line>`, a third evidence form beside `file:line` and `absent`. On `passed`, cite transcript lines when they confirm an AC or Design clause. On `failed`, gap only the rows the run demonstrably exercised and failed: verdict `PARTIAL`, `evidence:` quoting the failing transcript lines, `origin:` the row's clause per "Origin quote or it isn't a gap". When the failing component lies outside the audited diff (pre-existing breakage), that gap is `recommended: rescope` with the transcript as evidence, never `fix` - the fix loop must not spend capped rounds on code the change never owned. A failure no origin clause ties to yields no gap card; the `Happy path:` output line carries it as `failed - unattributable`. On `not run`, audit from code as today.
46
46
  4. **Flag the unrequested.** Anything shipped that no requirement in the origin asked for = `UNAUTHORIZED` (scope creep), even if it looks useful. Do not negotiate scope with yourself. This includes the step-1 exception clauses. `origin` is always `none (scope creep)`. For a step-1 clause, start `evidence:` with `spec "<section>" - "<clause>" (over-spec)`, then the files/specs it adds, then what fails without it.
47
47
  5. **Apply the coverage rule.** Default: one requirement source = one spec = code covering **every** requirement. Source and solution must end in sync. Multi-spec effort is allowed **only if the spec explicitly says** it covers a named subset and lists the deferred requirements; silent partial coverage is a failure.
48
48
 
@@ -64,6 +64,9 @@ Origin drift (spec vs prompt/ticket):
64
64
  - <disagreement> — recorded in spec? yes/no — <one-line reconciliation note>
65
65
  (or: none)
66
66
 
67
+ Happy path: passed | failed - attributed to G<n>,... | failed - unattributable | not run - <reason>
68
+ (omitted when the dispatch passed no summary)
69
+
67
70
  Gaps for user decision (only if verdict = GAPS):
68
71
  → emitted as Structured gap blocks (see below), one per non-DELIVERED row — not as free text here.
69
72
  ```
@@ -97,7 +100,7 @@ Fields:
97
100
  | (block label) | The block's label is the stable gap ID (`G1`, `G2`, ...), durable across re-audit rounds - not a field line inside the block |
98
101
  | `verdict` | `PARTIAL` \| `MISSING` \| `DRIFTED` \| `UNAUTHORIZED` (\| `DELIVERED` in re-audit rounds) |
99
102
  | `origin` | Requirement reference; for `UNAUTHORIZED` use the literal `none (scope creep)` |
100
- | `evidence` | `file:line`, or the literal `absent` |
103
+ | `evidence` | `file:line`, `<abs summary path>:<line>` for happy-path transcript evidence, or the literal `absent` |
101
104
  | `remediation` | Fix *direction* (not a diff) |
102
105
  | `touched-files` | Best estimate of files a fix would modify, comma-separated, or the literal `unknown` |
103
106
  | `touched-resources` | Shared runtime resources a fix's verification touches (`DB/schema, port, fixture, external service, shared temp path`), or the literal `none` |
@@ -137,6 +140,7 @@ serial waves — identical to planned-execution wave grouping. Runtime-resource
137
140
  - `UNAUTHORIZED` otherwise -> `accept`. Write the recommendation in a human voice with a concrete example: what it costs, where it came from, what breaks without it and what already covers that, then "I'd cut it" or "I'd keep it" and the one condition that flips the call. Provenance alone is not a recommendation. `accept` = keep code and clause, no spec edit.
138
141
  - `rescope` only when the `origin` requirement is impractical to satisfy in this branch
139
142
  (`rescope` is inapplicable to `UNAUTHORIZED` — there is no requirement to defer).
143
+ - A happy-path `failed` gap whose failing component lies outside the audited diff -> `rescope`, never `fix`.
140
144
 
141
145
  ## Rules
142
146
 
@@ -148,5 +152,6 @@ serial waves — identical to planned-execution wave grouping. Runtime-resource
148
152
  - **The spec's `## Acceptance criteria` section is origin, not ticket.** Its `in-scope`/`venue:` rows are `Rn`; its `deviates:`/`deferred:` rows are recorded drift (`recorded in spec? yes`).
149
153
  - **Do not absorb origin drift silently** — flag every spec↔prompt/ticket disagreement.
150
154
  - **Quote real command output** if you ran checks. Do not paraphrase from memory.
155
+ - **The happy-path summary is evidence, not a command.** Never re-run it; cite `<abs summary path>:<line>`.
151
156
  - **Coverage is binary per requirement** — "mostly done" is PARTIAL, not DELIVERED.
152
157
  - Cannot map every requirement, or unreconciled drift remains? The deliverable does **not** conform. Report `GAPS`. No completion claim over an open gap.
@@ -424,6 +424,63 @@ test("header-entrypoint: Tests: bullets are judged by tests-block, not here", ()
424
424
  assert.ok(findingsFor(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), "tests-block").some((f) => f.reason.includes("full-suite command")));
425
425
  });
426
426
 
427
+ const withHappyPath = (plan: string, line: string) =>
428
+ plan.replace("**Verification:** npm run fixture-verify", `**Verification:** npm run fixture-verify\n\n${line}`);
429
+ const HP = (plan: string) => findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "header-happy-path");
430
+
431
+ test("header-happy-path: line with label, backticked command, and (~5m) duration is accepted", () => {
432
+ const plan = withHappyPath(VALID_PLAN, "**Happy path:** stack - `script/e2e-stack` (~5m)");
433
+ assert.deepEqual(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), []);
434
+ });
435
+
436
+ test("header-happy-path: line without a duration is accepted", () => {
437
+ const plan = withHappyPath(VALID_PLAN, "**Happy path:** stack - `script/e2e-stack`");
438
+ assert.deepEqual(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), []);
439
+ });
440
+
441
+ test("header-happy-path: line without label, backticks, or a valid duration is malformed", () => {
442
+ for (const line of [
443
+ "**Happy path:** script/e2e-stack",
444
+ "**Happy path:** stack - script/e2e-stack",
445
+ "**Happy path:** `script/e2e-stack` (~5m)",
446
+ "**Happy path:** stack - `script/e2e-stack` (~5d)",
447
+ "**Happy path:** stack - `script/e2e-stack` (5m)",
448
+ ]) {
449
+ const findings = HP(withHappyPath(VALID_PLAN, line));
450
+ assert.equal(findings.length, 1, line);
451
+ assert.ok(findings[0].reason.includes("malformed **Happy path:** line"), line);
452
+ }
453
+ });
454
+
455
+ test("header-happy-path: line after the --- separator is ignored", () => {
456
+ const plan = VALID_PLAN.replace("\n---\n", "\n---\n\n**Happy path:** stack - `script/e2e-stack`\n");
457
+ assert.deepEqual(HP(plan), []);
458
+ });
459
+
460
+ test("tests-block: Tests: bullet equal to the happy-path command is a full-suite command", () => {
461
+ const plan = withHappyPath(VALID_PLAN, "**Happy path:** stack - `script/e2e-stack` (~5m)").replace(
462
+ "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`",
463
+ "**Tests:**\n- `script/e2e-stack`",
464
+ );
465
+ const tb = findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "tests-block");
466
+ assert.ok(tb.some((f) => f.reason.includes("full-suite command")), JSON.stringify(tb));
467
+ });
468
+
469
+ test("header-entrypoint: Run: payload and free text repeating the happy-path command fail", () => {
470
+ const base = withHappyPath(VALID_PLAN, "**Happy path:** stack - `script/e2e-stack` (~5m)");
471
+ const run = base.replace("This task handles naming details.", "Run: `script/e2e-stack`");
472
+ assert.ok(HE(run).some((f) => f.text.includes("Run: `script/e2e-stack`")));
473
+ const prose = base.replace("This task handles naming details.", "Then run script/e2e-stack to confirm.");
474
+ assert.ok(HE(prose).some((f) => f.text.includes("script/e2e-stack")));
475
+ });
476
+
477
+ test("table-closure: owner cell 'Happy path' is malformed under the existing reason", () => {
478
+ const row = '| § "Acceptance" L16 | full suite passes | Happy path |';
479
+ const plan = withHappyPath(withRow(VALID_PLAN, row), "**Happy path:** stack - `script/e2e-stack` (~5m)");
480
+ const tc = findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "table-closure");
481
+ assert.ok(tc.some((f) => f.reason === "owner cell is not a 'Task <n>' list, 'Verification', or 'waived: <reason>'" && f.text === row));
482
+ });
483
+
427
484
  test("wave-file-disjointness: Test/Test allowed; Test vs Modify conflict; Modify/Modify conflict", () => {
428
485
  const shared = "- Test: extensions/lib/fixture-shared.test.ts\n";
429
486
  const testTest = VALID_PLAN.replace("- Test: extensions/lib/fixture-task1.test.ts\n", shared).replace("- Test: extensions/lib/fixture-task2.test.ts\n", shared)
@@ -90,6 +90,9 @@ interface Header {
90
90
  specPath: string | undefined;
91
91
  verificationLine: number | undefined;
92
92
  verificationText: string | undefined;
93
+ happyPathLine: number | undefined;
94
+ happyPathText: string | undefined;
95
+ happyPathCommand: string | undefined;
93
96
  separatorLine: number | undefined;
94
97
  }
95
98
 
@@ -123,6 +126,7 @@ function fenceMask(lines: string[]): boolean[] {
123
126
  }
124
127
 
125
128
  const ANCHOR_RE = /\u00a7\s*"([^"]+)"\s*L(\d+)(?:-L(\d+))?/g;
129
+ const HAPPY_PATH_RE = /^(\S.*?) - `([^`]+)`(?: \(~\d+(?:s|m|h)\))?$/;
126
130
 
127
131
  function parseAnchors(text: string): Anchor[] {
128
132
  const anchors: Anchor[] = [];
@@ -151,6 +155,9 @@ function parsePlan(planText: string): ParsedPlan {
151
155
  let specPath: string | undefined;
152
156
  let verificationLine: number | undefined;
153
157
  let verificationText: string | undefined;
158
+ let happyPathLine: number | undefined;
159
+ let happyPathText: string | undefined;
160
+ let happyPathCommand: string | undefined;
154
161
  for (let i = 0; i < headerRangeEnd; i++) {
155
162
  if (mask[i]) continue;
156
163
  const line = lines[i];
@@ -165,6 +172,11 @@ function parsePlan(planText: string): ParsedPlan {
165
172
  verificationLine = i + 1;
166
173
  verificationText = line.replace(/^\*\*Verification:\*\*/, "").trim();
167
174
  }
175
+ if (happyPathLine === undefined && /^\*\*Happy path:\*\*/.test(line)) {
176
+ happyPathLine = i + 1;
177
+ happyPathText = line.replace(/^\*\*Happy path:\*\*/, "").trim();
178
+ happyPathCommand = HAPPY_PATH_RE.exec(happyPathText)?.[2];
179
+ }
168
180
  }
169
181
 
170
182
  const waveRe = /^## Wave (\d+) (\u2014|-) (.+)$/;
@@ -370,7 +382,7 @@ function parsePlan(planText: string): ParsedPlan {
370
382
 
371
383
  return {
372
384
  lines,
373
- header: { specPathLine, specPath, verificationLine, verificationText, separatorLine },
385
+ header: { specPathLine, specPath, verificationLine, verificationText, happyPathLine, happyPathText, happyPathCommand, separatorLine },
374
386
  waves,
375
387
  tasks,
376
388
  coverageTableFound,
@@ -422,10 +434,15 @@ function commandSegments(cmd: string): string[] {
422
434
  }
423
435
 
424
436
  function headerSegments(parsed: ParsedPlan): string[] {
425
- const value = parsed.header.verificationText ?? "";
426
- const spans = backtickSpans(value);
427
- const parts = spans.length > 0 ? spans : [value];
428
- return parts.flatMap((p) => norm(p).split(/\s*(?:&&|\|\||;|,)\s*/)).map((s) => s.trim()).filter(Boolean);
437
+ const values = [parsed.header.verificationText ?? "", parsed.header.happyPathCommand ?? ""].filter(Boolean);
438
+ return values
439
+ .flatMap((value) => {
440
+ const spans = backtickSpans(value);
441
+ const parts = spans.length > 0 ? spans : [value];
442
+ return parts.flatMap((p) => norm(p).split(/\s*(?:&&|\|\||;|,)\s*/));
443
+ })
444
+ .map((s) => s.trim())
445
+ .filter(Boolean);
429
446
  }
430
447
 
431
448
  const RUN_RE = /^\s*(- \[ \] )?Run:\s*(.*)$/;
@@ -491,7 +508,7 @@ function checkTestsBlock(parsed: ParsedPlan, fs: FsPort): PlanCheckFinding[] {
491
508
  for (const t of tokens) {
492
509
  if (!isAnchor(t) && isBroadening(t)) push(task, b.line, b.text, `broadening selector "${t}" in "${seg}"`);
493
510
  }
494
- if (header.includes(seg)) push(task, b.line, b.text, `full-suite command in task: "${seg}" equals a header **Verification:** segment`);
511
+ if (header.includes(seg)) push(task, b.line, b.text, `full-suite command in task: "${seg}" equals a header **Verification:** or **Happy path:** segment`);
495
512
  }
496
513
  }
497
514
  }
@@ -944,6 +961,19 @@ function checkSoloLine(parsed: ParsedPlan): PlanCheckFinding[] {
944
961
  return findings;
945
962
  }
946
963
 
964
+ function checkHappyPathLine(parsed: ParsedPlan): PlanCheckFinding[] {
965
+ const { happyPathLine, happyPathText, happyPathCommand } = parsed.header;
966
+ if (happyPathLine === undefined || happyPathCommand !== undefined) return [];
967
+ return [
968
+ {
969
+ check: "header-happy-path",
970
+ line: happyPathLine,
971
+ text: parsed.lines[happyPathLine - 1],
972
+ reason: `malformed **Happy path:** line "${happyPathText ?? ""}" (expected \`<label> - \\\`<command>\\\`\` with optional \` (~<duration>)\`, duration \\d+(s|m|h))`,
973
+ },
974
+ ];
975
+ }
976
+
947
977
  function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
948
978
  const findings: PlanCheckFinding[] = [];
949
979
  if (parsed.header.verificationText === undefined || parsed.header.separatorLine === undefined) {
@@ -959,6 +989,7 @@ function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
959
989
  if (!entrypoint) return findings;
960
990
 
961
991
  const header = headerSegments(parsed);
992
+ const entrypoints = [entrypoint, parsed.header.happyPathCommand].filter((e): e is string => Boolean(e));
962
993
  const executable = new Set<number>();
963
994
  for (const task of parsed.tasks) {
964
995
  for (const bullet of [...task.tests, ...task.testsVia, ...task.testsNone, ...task.testsMalformed]) {
@@ -993,17 +1024,18 @@ function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
993
1024
  check: "header-entrypoint",
994
1025
  line: ln,
995
1026
  text: line,
996
- reason: `Run: segment "${hit}" equals a header **Verification:** segment (full suite belongs to the verify phase)`,
1027
+ reason: `Run: segment "${hit}" equals a header **Verification:** or **Happy path:** segment (full suite belongs to the verify phase)`,
997
1028
  });
998
1029
  }
999
1030
  continue;
1000
1031
  }
1001
- if (line.includes(entrypoint)) {
1032
+ const hitEntry = entrypoints.find((e) => line.includes(e));
1033
+ if (hitEntry !== undefined) {
1002
1034
  findings.push({
1003
1035
  check: "header-entrypoint",
1004
1036
  line: ln,
1005
1037
  text: line,
1006
- reason: `header entrypoint "${entrypoint}" also appears outside the header (must be header-only)`,
1038
+ reason: `header entrypoint "${hitEntry}" also appears outside the header (must be header-only)`,
1007
1039
  });
1008
1040
  }
1009
1041
  }
@@ -1060,6 +1092,7 @@ export function checkPlan(planText: string, specText: string, fs: FsPort): PlanC
1060
1092
  findings.push(...checkWaveFileDisjointness(parsed, fs));
1061
1093
  findings.push(...checkSoloLine(parsed));
1062
1094
  findings.push(...checkHeaderEntrypoint(parsed));
1095
+ findings.push(...checkHappyPathLine(parsed));
1063
1096
  findings.push(...checkWaiverLiteral(parsed));
1064
1097
  return findings;
1065
1098
  } catch (err) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.17.1",
3
+ "version": "5.18.0",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -78,9 +78,9 @@ This is an **enforced disposition gate**, not a surface-only notice. The user is
78
78
 
79
79
  `verification-before-completion/reference/conformance-check.md` is **canonical** for the durable handoff schema, concern-decomposition rules, the single disposition-availability table, the `UNAUTHORIZED` question text, the `recommended: none` preflight, the freshness rule, and the concern-scoped fix projection. This step owns only **render, response, and execute-order** and consumes the rest by link - it does not restate the availability table, the `UNAUTHORIZED` question, or the preflight prose.
80
80
 
81
- **If no conformance check has run in this flow** (e.g. ad-hoc work that landed without an execution skill): say so, then dispatch a fresh-context `conformance-reviewer` with `cwd: "<worktree-path>"` against the origin (spec + verbatim prompt + full diff vs base) per that reference - it owns the audit-time input rule (stage/commit untracked deliverables before auditing). Closing the loop is cheap relative to shipping unverified intent. Route the raw reviewer verdict through the reference's canonical pipeline (gap/concern partition, auto-fix where eligible, concern decomposition, emission of a durable `## Closure / conformance` block), then consume that block through the branching below exactly as a carried handoff.
81
+ **If no conformance check has run in this flow** (e.g. ad-hoc work that landed without an execution skill): say so, then dispatch a fresh-context `conformance-reviewer` with `cwd: "<worktree-path>"` against the origin (spec + verbatim prompt + full diff vs base) per that reference - it owns the audit-time input rule (stage/commit untracked deliverables before auditing; this path never runs a happy path and passes no happy-path input). Closing the loop is cheap relative to shipping unverified intent. Route the raw reviewer verdict through the reference's canonical pipeline (gap/concern partition, auto-fix where eligible, concern decomposition, emission of a durable `## Closure / conformance` block), then consume that block through the branching below exactly as a carried handoff.
82
82
 
83
- **Freshness precondition - before any verdict branch, including `CONFORMS`.** The durable block opens with a two-line sentinel: `status: CONFORMS (0 open)` or `status: GAPS (N open)`, then `audited-base: <full HEAD SHA at audit time>`. Read the sentinel, then apply the reference's freshness rule (its `## Closure / conformance` block is the single source): compare `audited-base` to the current working tree; any change, doubt, missing/mismatched sentinel, legacy terse row, or malformed structured reviewer block triggers a fresh audit and replacement of the closure block. Never infer `CONFORMS` from the absence of bullets. Only a clean, valid `status: CONFORMS (0 open)` handoff enters the zero-gap fast path.
83
+ **Freshness precondition - before any verdict branch, including `CONFORMS`.** The durable block opens with a sentinel: `status: CONFORMS (0 open)` or `status: GAPS (N open)`, then `audited-base: <full HEAD SHA at audit time>`, then an optional `happy-path: <outcome>` line (present exactly when a happy-path run happened; informational, never a freshness input). Read the sentinel, then apply the reference's freshness rule (its `## Closure / conformance` block is the single source): compare `audited-base` to the current working tree; any change, doubt, missing/mismatched sentinel, legacy terse row, or malformed structured reviewer block triggers a fresh audit and replacement of the closure block. Never infer `CONFORMS` from the absence of bullets. Only a clean, valid `status: CONFORMS (0 open)` handoff enters the zero-gap fast path.
84
84
 
85
85
  **Zero-gap fast path:** print exactly
86
86
 
@@ -90,9 +90,11 @@ Closure / conformance: CONFORMS
90
90
 
91
91
  then continue directly to Step 4. No approval prompt, no menu, no shared options line, no sign-off. If the run auto-applied fixes, surface the flat `auto-applied fix commits: <Gn: SHA>, ...` index from the durable block as **one informational, non-blocking line** with a one-line revert offer (see "Revert semantics") - a gap that auto-converged mid-verify has no bullet, so this index is the only place its fix commit stays revertable. Do not wait for acknowledgment.
92
92
 
93
+ When the sentinel carries a `happy-path:` line, print it directly under `Closure / conformance: CONFORMS` as one informational, non-blocking line in the same shape as the `auto-applied fix commits` line: `happy-path: passed`, `happy-path: failed - unattributable`, or `happy-path: not run - <reason>` (a `failed - attributed to G<n>` value cannot reach this branch: its gaps are open).
94
+
93
95
  **Pre-menu amendment funnel (GAPS only).** Before rendering the carried-open menu, run [`amendment-surface.md` § Conformance entry](../brainstorming/reference/amendment-surface.md) over the inventory once: gaps with `recommended: accept`, verdict `DRIFTED` or `PARTIAL`, not `UNAUTHORIZED`, whose `origin` is not an acceptance criterion are drafted as `accept-into-spec` items (the edit built from `origin` + `evidence`) and sent to the reviewer in one call; cleared items apply and land as one batch commit, the spec is re-audited, the inventory regenerated. Only concerns the re-audit actually closed drop out; sibling concerns in the same gap keep their rows and dispositions. Survivors and every other gap render as ordinary rows below - one menu, never two.
94
96
 
95
- **Carried-open (`status: GAPS (N open)`).** Read `reference/disposition-protocol.md` and follow it for the carried-open render (dense) grammar, the response grammar, and the 9-step execute order. Render the human decision menu in the shape below, drive the dispositions per that reference, then print the summary render and continue to Step 4. If that reference file cannot be read, stop and surface a blocking error — do **not** improvise the grammar from memory.
97
+ **Carried-open (`status: GAPS (N open)`).** Read `reference/disposition-protocol.md` and follow it for the carried-open render (dense) grammar, the response grammar, and the 9-step execute order. Render the human decision menu in the shape below, with the sentinel's `happy-path:` line (when present) printed as one informational line directly under the `Conformance: N decisions needed before shipping.` header - a `failed - attributed to G<n>,...` value names the rows whose evidence is the transcript, a `rescope` recommendation on such a row means the failure lies in code the change never owned - drive the dispositions per that reference, then print the summary render and continue to Step 4. If that reference file cannot be read, stop and surface a blocking error — do **not** improvise the grammar from memory.
96
98
 
97
99
  Representative carried-open render (multi-concern gap split to `e2e`; single-concern gap `cache`; `UNAUTHORIZED` gap `auth`):
98
100
 
@@ -222,8 +222,40 @@ For the fan-out + worktree + patch-integration + conflict mechanics, see `dispat
222
222
  0. Call `phase_tracker({ action: "start", phase: "verify" })`. (The `implement` phase was started at execution start and auto-completes from `plan_tracker` once all tasks are done; this flow runs its own verify gate instead of `/skill:verification-before-completion`, so it must mark verify itself.)
223
223
  1. **Parent full verification.** Run the complete plan-header `**Verification:**` command set foreground in the worktree via `(cd "<worktree>" && <command>)`: tests plus every declared lint, type, format, and build check. A failure must be repaired and the full set rerun successfully before the next step. Before dispatching a verification repair, reopen (`in_progress`) every existing plan-task index whose `Files:` ownership includes its touched files; leave unowned cross-cutting repair work in the verification report. Those indices stay `in_progress` through the successful full rerun **and** step 2's whole-diff review accepting the repair — that acceptance is their completion point, not the passing rerun. Commit any verification-produced tracked changes; use the resulting `HEAD_SHA` in the review task and include the commands/results in its existing `DESCRIPTION`.
224
224
  2. **Whole-diff code review.** Only after passing full verification, dispatch one foreground whole-diff `code-reviewer` per `/skill:requesting-code-review` against that committed HEAD, with `SCOPED_TEST_COMMANDS: none`; the reviewer does not repeat the full suite. Address Critical and Moderate findings. Before dispatching a review repair, reopen (`in_progress`) every existing plan-task index whose `Files:` ownership includes its touched files; leave unowned cross-cutting repair work in the review report. Mark each reopened index `complete` only once the repair is re-verified and the re-review accepts it — this is the same completion point step 1's reopened indices wait for, not an extra gate, and the gate order stays full verification -> whole-diff CR -> conformance. Any repair invalidates prior full verification, so rerun the full set successfully before the next gate.
225
- 3. **Close the loop — conformance check.** The review in step 2 is plan-vs-code (single-step); it inherits any requirement the plan already dropped. Before marking verify complete, dispatch a fresh-context **`conformance-reviewer`** — its **own** dispatch, never fused into the step-2 review — to confront the deliverable (code **and** docs) against the *origin* — the spec **and** the original prompt — per `verification-before-completion/reference/conformance-check.md`. Pass the spec path, the verbatim original prompt, and the full diff. Follow that reference for the partition rule, concern decomposition, and fix-loop mechanics; do not reimplement them here. The fix loop reuses durable `Gn` gap indices as defined in conformance-check.md; it never calls `phase_tracker`. Call `phase_tracker({ action: "complete", phase: "verify" })` only when the reference says the handoff is durably complete: either a current `CONFORMS` result, or a current `## Closure / conformance` inventory whose carried-open concerns all come from valid deferred gaps, including `recommended: fix` gaps carried open because a declared precondition made the fix loop unavailable (`maxFixRounds: 0`, or no eligible named-branch worktree). A started positive-cap fix loop that blocks, fails, or exhausts its rounds with an open `fix` gap is escalation, not completion; on escalation, do not complete verify, stop and report.
226
- 4. Summarize what was implemented (tasks completed, files changed, test counts, code-review verdict). Emit the `## Closure / conformance` block exactly as defined in `verification-before-completion/reference/conformance-check.md`: it must open with the two-line sentinel (`status: CONFORMS (0 open)` or `status: GAPS (N open)`, then `audited-base: <full HEAD SHA>`), then carry the exact durable concern schema by reference with no renamed or reformatted fields. `finishing-a-development-branch` Step 3.5 consumes that block verbatim.
225
+ 3. **Close the loop — conformance check.** The review in step 2 is plan-vs-code (single-step); it inherits any requirement the plan already dropped.
226
+
227
+ **Happy-path run - first action of this step, only when the plan header carries a `**Happy path:**` line or the diff selects a row.** The overrides file's `## Happy path` table (schema in the README overrides contract) names path-prefixed rows; the plan header's line is the plan-time default.
228
+
229
+ 1. Bind `HP_DIR=$(mktemp -d)` first. Re-derive the row from `git -C "<worktree>" diff --name-only <base>..HEAD` (`<base>` = the branch point): two or more non-`cross-cutting` rows, or a path inside the `cross-cutting` row's own `Paths` and inside no other row's -> `cross-cutting`, taking precedence (no `cross-cutting` row -> no run and no outcome line); else paths inside exactly one non-`cross-cutting` row's `Paths` -> that row; none -> no run and no outcome line. Run the diff-derived row; when it differs from the header label, record `row: <diff-derived> (header: <label>)` in `summary.txt`.
230
+ 2. Pre-checks: bind `TO=$(command -v timeout || command -v gtimeout)` -> else `not run - no timeout binary`; the command's first token after any leading `NAME=value` assignments resolves via `(cd "<abs worktree path>" && bash -c 'command -v <token>')` -> else `not run - command not found` (a header present but the script missing from the branch lands here; never a gap). A pre-check failure still writes `$HP_DIR/summary.txt` with only the outcome, `head:`, and `row:` lines (no blank line, no tail - there is no `transcript.log`) and counts as a run for the reviewer input and the closure sentinel.
231
+ 3. Snapshot `git -C "<worktree>" status --porcelain --untracked-files=all`, then run, with `<duration>` = the row's `Timeout` (default `10m`) and `HP_CMD` holding the row's command verbatim:
232
+
233
+ ```bash
234
+ (cd "<abs worktree path>" && "$TO" -k 30s <duration> bash -c "$HP_CMD") >"$HP_DIR/transcript.log" 2>&1
235
+ EXIT=$?
236
+ ```
237
+
238
+ 4. Re-snapshot status; any difference (tracked or untracked residue) is `failed - dirtied worktree: <paths>` regardless of exit code; never stage or commit the listed paths as deliverables - remove untracked residue and restore tracked residue (`git -C "<worktree>" checkout -- <paths>`) before the conformance dispatch, so the tree is clean when the audit-time input rule runs and the closure freshness rule never fires on happy-path artifacts at finish. Then classify the outcome per the table and write the summary:
239
+
240
+ ```bash
241
+ { echo "<outcome line per table>"; echo "head: $(git -C "<abs worktree path>" rev-parse HEAD)"; echo "row: <label>"; echo; tail -n 200 "$HP_DIR/transcript.log"; } >"$HP_DIR/summary.txt"
242
+ ```
243
+
244
+ | Condition | Outcome line |
245
+ |---|---|
246
+ | worktree differs after run | `happy-path: failed - dirtied worktree: <paths>` |
247
+ | exit 0, worktree clean | `happy-path: passed` |
248
+ | exit 75 | `happy-path: not run - environment unavailable: <last line of transcript.log>` |
249
+ | exit 124, or 137 after the `-k` kill | `happy-path: failed - timed out after <duration>` |
250
+ | exit 126 | `happy-path: not run - not executable` |
251
+ | pre-check failed | `happy-path: not run - <reason>` |
252
+ | any other non-zero | `happy-path: failed (exit <n>)` |
253
+ | no header line and no diff-derived row | no run, no outcome line |
254
+
255
+ A timeout is `failed`, not `not run`: a consumer that never receives its message hangs, and the reviewer must see it; `not run` is reserved for a command that never executed. `timeout -k 30s` sends `TERM` then `KILL`; a script that does not trap `TERM` leaves its stack up, and the next run's exit 75 surfaces that. The full `transcript.log` stays in `$HP_DIR` for the human; the reviewer receives `summary.txt` only (outcome, `head:`, `row:`, 200-line tail). A `failed` or `not run` outcome never stops the flow and never becomes a repair item at this step - it is evidence for the audit; `$HP_DIR` and the outcome carry into the fix loop per `conformance-check.md`. Run sub-steps 3-4 (snapshot, run, re-snapshot, classify, summary) in one bash call, or substitute the literal `HP_DIR` and `TO` values into each later command: shell variables do not survive between tool calls.
256
+
257
+ Before marking verify complete, dispatch a fresh-context **`conformance-reviewer`** — its **own** dispatch, never fused into the step-2 review — to confront the deliverable (code **and** docs) against the *origin* — the spec **and** the original prompt — per `verification-before-completion/reference/conformance-check.md`. Pass the spec path, the verbatim original prompt, the full diff, and - when a run happened - `Happy path: <abs path to $HP_DIR/summary.txt> (<outcome>)` (`<outcome>` = the outcome line without its `happy-path: ` prefix). Follow that reference for the partition rule, concern decomposition, and fix-loop mechanics; do not reimplement them here. The fix loop reuses durable `Gn` gap indices as defined in conformance-check.md; it never calls `phase_tracker`. Call `phase_tracker({ action: "complete", phase: "verify" })` only when the reference says the handoff is durably complete: either a current `CONFORMS` result, or a current `## Closure / conformance` inventory whose carried-open concerns all come from valid deferred gaps, including `recommended: fix` gaps carried open because a declared precondition made the fix loop unavailable (`maxFixRounds: 0`, or no eligible named-branch worktree). A started positive-cap fix loop that blocks, fails, or exhausts its rounds with an open `fix` gap is escalation, not completion; on escalation, do not complete verify, stop and report.
258
+ 4. Summarize what was implemented (tasks completed, files changed, test counts, code-review verdict). Emit the `## Closure / conformance` block exactly as defined in `verification-before-completion/reference/conformance-check.md`: it must open with the sentinel (`status: CONFORMS (0 open)` or `status: GAPS (N open)`, then `audited-base: <full HEAD SHA>`, then - exactly when a happy-path run happened - `happy-path: <value of the reviewer's Happy path: line from the final audit, the text after Happy path: >`), then carry the exact durable concern schema by reference with no renamed or reformatted fields. `finishing-a-development-branch` Step 3.5 consumes that block verbatim.
227
259
  5. **Proceed to finishing — no confirmation prompt.** Once verify is complete per step 3's criterion, invoke `/skill:finishing-a-development-branch <abs worktree path>` immediately - the same absolute path every dispatch in this flow carried as `cwd`; finishing stops without it. Its Step 4 menu (PR / draft PR / squash / keep / discard) is the human gate; a separate "ready to finish?" prompt only stacks a second stop in front of it. Carried-open concerns are resolved there per concern via the `## Closure / conformance` block from step 4. Manual testing is a follow-up after the finishing choice (on `<base-branch>` after a squash-merge, or on the PR branch), never a reason to hold this gate.
228
260
 
229
261
  ## Red Flags — STOP
@@ -35,6 +35,7 @@ pass in:
35
35
  - The **spec** (path).
36
36
  - The **original prompt** (verbatim — it holds inline requirements + any ticket ref).
37
37
  - The **diff** to audit (code + docs).
38
+ - The **happy-path summary**, when `subagent-driven-development` step 3 ran one: `Happy path: <abs path to $HP_DIR/summary.txt> (<outcome>)`, `<outcome>` being the outcome line without its `happy-path: ` prefix. Omitted when no run happened; the ad-hoc path in `finishing-a-development-branch` (no plan, no header) never runs one and never passes this input.
38
39
 
39
40
  Dispatch it as its **own** call — do not fold the conformance check into the
40
41
  whole-PR code-quality review. Fusing the two subordinates intent-coverage to a
@@ -164,7 +165,7 @@ Per round:
164
165
  one-task `tasks` wave (never a lone `agent: "implementer"`) and counts as a
165
166
  wave against `maxFixRounds`. A `BLOCKED`/`NEEDS_CONTEXT` return surfaces to the user.
166
167
  4. **Scoped tests** on the integrated tree: run the round's `SCOPED_TEST_COMMANDS` union as `(cd "<conformance-worktree>" && <command>)`. A failure re-enters the failure-handling rules above.
167
- 5. **Re-audit**: foreground re-dispatch `conformance-reviewer` with `async: false`, `cwd` = the conformance worktree, over the fixes **plus** the
168
+ 5. **Re-audit**: Happy-path re-run first, when the prior audit received a happy-path input: if `$HP_DIR` no longer exists (finish-time `fix-now` in a resumed session), re-derive the row per step 3.1 from the overrides table, then re-run the step-3 procedure when `git -C "<worktree>" diff --name-only <audited-base>..HEAD` contains a path inside that row's `Paths` or an open gap's `evidence:` cites the summary path, else pass `not run - no prior transcript`. Otherwise re-run the step-3 procedure into a fresh `$HP_DIR` when either (a) `git -C "<worktree>" diff --name-only <head from current summary.txt>..HEAD` contains a path inside the current row's `Paths` (the summary's `head:` line is the diff base, so changes from skipped rounds accumulate into the next comparison); (b) an open gap's `evidence:` cites the summary path. Otherwise pass the previous summary unchanged. When the re-run's outcome differs from the previous one, extend this re-audit's scope to every row whose `evidence:` cites the summary. Rounds touching only paths outside every row never re-run. Then foreground re-dispatch `conformance-reviewer` with `async: false`, `cwd` = the conformance worktree, over the fixes **plus** the
168
169
  regression guard (any prior-`DELIVERED` requirement whose `evidence` file
169
170
  the fix diff touched). Pass the full prior conformance report (every row,
170
171
  including DELIVERED rows and their `evidence` `file:line`) and the round's
@@ -352,18 +353,20 @@ carried-open `fix` state is valid closure inventory, not escalation.
352
353
  ### Handoff sentinel and freshness anchor - every handoff
353
354
 
354
355
  Every `## Closure / conformance` block - a `CONFORMS` no-card handoff and a
355
- carried-open GAPS handoff alike - **opens with a two-line sentinel** that lets
356
+ carried-open GAPS handoff alike - **opens with a sentinel of two lines (three
357
+ when a happy-path run happened)** that lets
356
358
  the finish gate re-verify freshness after context pruning, with no session
357
359
  history:
358
360
 
359
361
  ```text
360
362
  status: CONFORMS (0 open) # or: status: GAPS (N open)
361
363
  audited-base: <full HEAD SHA at audit time>
364
+ happy-path: passed | failed - attributed to G<n>,... | failed - unattributable | not run - <reason>
362
365
  ```
363
366
 
364
367
  `N` = count of open concerns (decision units), matching the number of emitted
365
368
  concern cards. Record `audited-base` as the full 40-char HEAD SHA at audit time;
366
- never abbreviate. This block is the **single source** for the freshness rule;
369
+ never abbreviate. The `happy-path:` line is present exactly when a happy-path run happened, carrying the value of the reviewer's `Happy path:` output line of the final audit (the text after `Happy path: `; the transcript the final verdict was audited against). It is not a concern card: `N` and the card count ignore it. `finishing-a-development-branch` Step 3.5 renders it as one informational line. This block is the **single source** for the freshness rule;
367
370
  `finishing-a-development-branch` links here rather than restating it.
368
371
 
369
372
  **Freshness rule.** The audit-input rule requires deliverables committed before
@@ -378,7 +381,7 @@ git -C "$ROOT" status --porcelain --untracked-files=all # new/untracked delivera
378
381
  ```
379
382
 
380
383
  Any output from either command, any doubt, a missing/mismatched sentinel, or any
381
- closure block not opening with the two-line sentinel above (e.g. a legacy
384
+ closure block not opening with the sentinel above (e.g. a legacy
382
385
  `Gn: PARTIAL - recommended: ...` row) triggers a fresh audit - never infer
383
386
  `CONFORMS` from the absence of cards. This is a lightweight freshness check (two
384
387
  git commands, no hashing or identity fields).
@@ -59,7 +59,7 @@ subagent({ agent: "scout", context: "fresh", async: false, cwd: "<abs worktree p
59
59
  task: <the fixed template below, with the spec path filled> })
60
60
  ```
61
61
 
62
- > Recon for implementation planning. Read the approved spec at `<abs spec path>` - it is the single source of truth for what is being built. Also read the repo's `AGENTS.md` and, if present, the gauntlet overrides file (checked in order: `.pi/gauntlet-overrides.md`, `gauntlet-overrides.md`, `doc/gauntlet-overrides.md` at the repo root) for conventions. Build an implementation map for the spec: exact file paths to create/modify/delete; existing call sites and tests with line ranges; conventions and patterns the plan must match; the project's test runner and the exact scoped-invocation form for running individual test files (derived from the repo's Makefile/bin/config and the overrides file); the style/lint and auto-format commands in both scoped per-file form and repo-wide form (same sources); separately, the full-suite verification entrypoint and whether it bundles style/format checks. Flag any spec claim that contradicts the code. Read-only recon: do not edit any file except writing your report to your output path. Start your report with the line `# CONTEXT DRAFT - NOT A PLAN - fully replaced at plan-writing` verbatim. End with an "Open questions that matter for the plan" section. Compact handoff, not a dump.
62
+ > Recon for implementation planning. Read the approved spec at `<abs spec path>` - it is the single source of truth for what is being built. Also read the repo's `AGENTS.md` and, if present, the gauntlet overrides file (checked in order: `.pi/gauntlet-overrides.md`, `gauntlet-overrides.md`, `doc/gauntlet-overrides.md` at the repo root) for conventions. Build an implementation map for the spec: exact file paths to create/modify/delete; existing call sites and tests with line ranges; conventions and patterns the plan must match; the project's test runner and the exact scoped-invocation form for running individual test files (derived from the repo's Makefile/bin/config and the overrides file); the style/lint and auto-format commands in both scoped per-file form and repo-wide form (same sources); separately, the full-suite verification entrypoint and whether it bundles style/format checks. When the overrides file has a `## Happy path` section, copy its table verbatim into your report under a `## Happy path` heading. Flag any spec claim that contradicts the code. Read-only recon: do not edit any file except writing your report to your output path. Start your report with the line `# CONTEXT DRAFT - NOT A PLAN - fully replaced at plan-writing` verbatim. End with an "Open questions that matter for the plan" section. Compact handoff, not a dump.
63
63
 
64
64
  Consumption:
65
65
 
@@ -75,6 +75,8 @@ Consumption:
75
75
 
76
76
  List the files this implementation will create, modify, or delete. Group by component. This forces the design decisions out of the task list and into a single review surface.
77
77
 
78
+ **Happy-path row selection.** When the recon report carries a `## Happy path` table, match the union of every task's `Files:` paths against the table's `Paths` prefixes (a path is inside a row when it starts with one of the row's prefixes). Two or more non-`cross-cutting` rows matched, or a path inside the `cross-cutting` row's own `Paths` and inside no other row's -> the `cross-cutting` row (no such row -> no line); this takes precedence. Otherwise exactly one non-`cross-cutting` row matched -> that row. No row matched, or no table -> no line. The selected row becomes the header's `**Happy path:**` line (below); the parent re-derives the row from the real diff at verify time, so this is the plan-time default and the `plan_check` anchor.
79
+
78
80
  ```markdown
79
81
  ## Files
80
82
 
@@ -155,13 +157,14 @@ Each step is **one action, 2-5 minutes**:
155
157
  **Spec:** `<project>/doc/specs/<same-filename-as-this-plan>.md`
156
158
 
157
159
  **Verification:** `<full verification command set — tests + style + format; a single bundling entrypoint, or the listed individual commands; from the recon report / project overrides>` - in a repo with per-service verification commands, list the command of every service the change affects (its own files or code it depends on; a repo-wide shared path such as root config, a lockfile, or a shared library affects every dependent service), taking the per-service commands from the overrides file or AGENTS.md when recon reports a single entrypoint
160
+ **Happy path:** <row label> - `<command from the selected row>` (~<Timeout from the row>)
158
161
 
159
162
  **Ticket:** `<ticket-id>` (omit if none)
160
163
 
161
164
  ---
162
165
  ```
163
166
 
164
- The full verification entrypoint appears only on the `**Verification:**` line — see [reference/plan-contract.md § Header-only entrypoint](reference/plan-contract.md). The verify phase reads it from the plan; execution runs `Tests:` commands only. Execution runs it in the worktree via the subshell form `(cd "<abs worktree path>" && <command>)`; the process cwd stays in the primary checkout, and dispatch `cwd` is the worktree path.
167
+ The full verification entrypoint appears only on the `**Verification:**` line — see [reference/plan-contract.md § Header-only entrypoint](reference/plan-contract.md). The verify phase reads it from the plan; execution runs `Tests:` commands only. Execution runs it in the worktree via the subshell form `(cd "<abs worktree path>" && <command>)`; the process cwd stays in the primary checkout, and dispatch `cwd` is the worktree path. The `**Happy path:**` line is optional: present only when File Structure selected a row, with the row label, the row's command in backticks, and the row's `Timeout` as `(~<duration>)` (omit the suffix when the cell is absent or malformed). Like `Verification`, the happy-path command appears only on this header line - never in a task's `Tests:` or `Run:` block and never as free text in wave scope; `plan_check` enforces both.
165
168
 
166
169
  ## Task Structure
167
170
 
@@ -233,6 +236,7 @@ Every plan ends with a `## Spec coverage` section (grammar and example: [referen
233
236
 
234
237
  - **Requirement rows:** a cross-cutting requirement (decided in more than one task) lists **every** deciding task as owner, not the first. `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
235
238
  - **`Verification` owner:** only for a requirement the header command proves; grammar in the reference.
239
+ - `Happy path` is never an owner. Coverage by the happy-path run is inferred by the conformance reviewer at audit time, never declared in the plan; the checker rejects the cell under the owner-grammar reason.
236
240
  - The table is plan-authoring-time only — never passed to implementer or reviewer dispatches.
237
241
 
238
242
  ## No Placeholders
@@ -27,7 +27,7 @@ Every `### Task N` carries a `**Tests:**` block: the bare line `**Tests:**` dire
27
27
 
28
28
  Absence is never valid. A `- [ ]` step or any non-bullet line ends the block. `via:` and `none:` are unchecked beyond form.
29
29
 
30
- Commands are written repo-relative and run in the worktree via `(cd "<abs worktree path>" && <command>)` - the checker never runs them; the executor does. Each command is split into segments on `&&`, `||`, `;`, `|`; every segment must contain, as a whitespace-delimited token, a `Test:` path of the same task (the path alone, or followed by `::`, `#`, or `:` and a filter). A `Test:` value containing `*`, `?`, `[` or ending in `/` never anchors; any other argument token with those shapes is a broadening selector and fails. `cd `, `sh -c`, `bash -c`, `eval `, `$(` are unsupported. Each `Test:` path must exist or be a `Create:` path of some task. A segment equal to a header `**Verification:**` segment is a full-suite command and fails. Runners with no file-addressable form are out of scope (`go test ./pkg -run X`, `mvn -Dtest=`).
30
+ Commands are written repo-relative and run in the worktree via `(cd "<abs worktree path>" && <command>)` - the checker never runs them; the executor does. Each command is split into segments on `&&`, `||`, `;`, `|`; every segment must contain, as a whitespace-delimited token, a `Test:` path of the same task (the path alone, or followed by `::`, `#`, or `:` and a filter). A `Test:` value containing `*`, `?`, `[` or ending in `/` never anchors; any other argument token with those shapes is a broadening selector and fails. `cd `, `sh -c`, `bash -c`, `eval `, `$(` are unsupported. Each `Test:` path must exist or be a `Create:` path of some task. A segment equal to a header `**Verification:**` or `**Happy path:**` segment is a full-suite command and fails. Runners with no file-addressable form are out of scope (`go test ./pkg -run X`, `mvn -Dtest=`).
31
31
 
32
32
  ## Solo line (`solo-line`)
33
33
 
@@ -39,6 +39,12 @@ The `**Verification:**` line is the **only** place the full verification entrypo
39
39
 
40
40
  The header value's backtick spans (else the raw value) are split on `&&`, `||`, `;`, `,` into segments. No `Run:` step payload segment and no `Tests:` bullet segment may equal a header segment; prose inside waves may not contain the whole header value.
41
41
 
42
+ The optional `**Happy path:**` line's backticked command is a second header entrypoint under the same rule: its segments join the header segment set, and its whole command may not appear in wave prose.
43
+
44
+ ## Happy path line (`header-happy-path`)
45
+
46
+ Optional. When present, the header line reads `**Happy path:** <label> - \`<command>\`` with an optional ` (~<duration>)` suffix, `<duration>` matching `\d+(s|m|h)`. A present line that does not parse is a finding; absence is never a finding. The label names the row of the overrides file's `## Happy path` table the plan selected; the verify phase reads the line from the plan and re-derives the row from the real diff.
47
+
42
48
  ## Spec coverage table (`table-closure`, `waiver-literal`)
43
49
 
44
50
  Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Closure both ways: every `### Task N` appears as an owner in some row; every row's owner task exists.