pi-gauntlet 5.3.7 → 5.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.5.0 - 2026-09-10
4
+
5
+ - `writing-plans`: every task carries a `**Tests:**` block - scoped commands anchored to the task's `Test:` paths, optional `via:` entry point, or `none: <category>`; the plan grammar `plan_check` enforces moves to `skills/writing-plans/reference/plan-contract.md`. (#28)
6
+ - `plan_check`: new `tests-block` check (block present and well-formed, every command segment names a task `Test:` path, no broadening selectors or `cd`/`sh -c`/`eval`/`$(`, no segment equal to a header `**Verification:**` segment); `header-entrypoint` compares `Run:` payloads by segment, so `npm test` no longer slips past `npm test && npm run lint`; `Test:` entries no longer count as file ownership in `wave-file-disjointness`. (#28)
7
+ - `subagent-driven-development`, `test-driven-development`, prompts, `spec-reviewer`: `SCOPED_TEST_COMMANDS` comes from the `Tests:` block; implementer reports `met`/`unmet` per command and per `via:`; the spec reviewer treats the task contract as a supplement the anchored spec overrides. (#28)
8
+
9
+ ## v5.4.0 - 2026-09-10
10
+
11
+ - `subagent-driven-development`: a stalled review fix loop runs one escalated fix round (`implementer`, `context: fresh`, model from new `piGauntlet.escalationLoop.implModel`, default main-loop model + thinking) before stopping; the stop is a one-screen problem note (`stop-note.md`) with concrete fix options, replacing the trajectory-log escalation report. `gauntlet_setting` gains the `escalationLoop` key. (#29)
12
+
3
13
  ## v5.3.7 - 2026-09-10
4
14
 
5
15
  - `brainstorming`: the standalone "does this replace a prior spec" question is gone - the gather scout names candidate predecessor specs and round 1 states them; the design is presented in two rounds (architecture/components/data flow, then errors/testing/docs) with one approval each. (#27)
package/README.md CHANGED
@@ -71,7 +71,7 @@ pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
71
71
 
72
72
  - **17 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Four more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`.
73
73
  - **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
74
- - **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers plan_check, a deterministic plan linter (9 mechanical plan-vs-spec checks) whose pass stamp gates implement-start inside a gauntlet flow. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
74
+ - **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers `plan_check`, which verifies a plan against its spec and against the grammar in [skills/writing-plans/reference/plan-contract.md](./skills/writing-plans/reference/plan-contract.md), including that each task's `Tests:` commands are selective and never the full suite; a pass stamps the plan for implementation. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
75
75
 
76
76
  pi-gauntlet is **opinionated**: every non-trivial change is *meant* to ride this one pipeline, entered through `brainstorming`. Enforcement is opt-in by entry, not ambient: once brainstorming starts a flow, the phase-tracker extension mechanically blocks a phase from closing before its gate runs, and warns once if the main loop writes code during implement (subagents own implement-phase edits). A change made *without* entering the flow (a typo, a formatting run, a dependency bump - see "When to use / when NOT to use") is not gated; the discipline of routing real work through the pipeline is a convention the tooling supports, not a trap it springs on every edit.
77
77
 
@@ -13,14 +13,15 @@ You are a spec compliance reviewer. Your job is to verify that an implementation
13
13
 
14
14
  ## Process
15
15
 
16
- <!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md — change them together or not at all -->
16
+ <!-- clause decomposition / snippet non-authority / whole-file reads / task-contract supplement: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md — change them together or not at all -->
17
17
 
18
18
  1. Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses. Every clause gets a verdict row.
19
- 2. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
20
- 3. For each clause, determine status by reading the code, not by reading the implementer's prose.
21
- 4. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
22
- 5. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
23
- 6. Flag any clause from the spec that is missing from the implementation.
19
+ 2. The task's `**Tests:**` block, `via:`, and its `Files:` paths supplement the anchored spec where it is silent; the anchored spec wins a conflict - report the divergence once, against the plan, never against code corrected to the spec. Findings: a `Create:` path absent from the diff or created elsewhere; a test that does not call the `via:` entry point; a `Tests:` block the diff contradicts. Existing files need no diff touch.
20
+ 3. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
21
+ 4. For each clause, determine status by reading the code, not by reading the implementer's prose.
22
+ 5. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
23
+ 6. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
24
+ 7. Flag any clause from the spec that is missing from the implementation.
24
25
 
25
26
  ## Output format
26
27
 
@@ -4,6 +4,8 @@ import {
4
4
  mergeGauntlet,
5
5
  resolveSpecCouncil,
6
6
  resolveClosureReview,
7
+ resolveEscalationLoop,
8
+ mainLoopModel,
7
9
  resolveFlowGuards,
8
10
  resolveVerifyBeforeShip,
9
11
  settingsErrorWarning,
@@ -24,6 +26,41 @@ test("mergeGauntlet: undefined layers -> {}", () => {
24
26
  assert.deepEqual(mergeGauntlet(undefined, undefined), {});
25
27
  });
26
28
 
29
+ test("mergeGauntlet: repo escalationLoop replaces preset whole-object", () => {
30
+ const preset = { escalationLoop: { implModel: "p/preset:high" }, closureReview: { model: "m" } };
31
+ const repo = { escalationLoop: {} };
32
+ const merged = mergeGauntlet(preset, repo);
33
+ assert.deepEqual(merged.escalationLoop, {});
34
+ assert.deepEqual(merged.closureReview, { model: "m" });
35
+ });
36
+
37
+ test("escalationLoop: absent/empty/null/non-string -> mainLoop", () => {
38
+ const main = "p/main:medium";
39
+ assert.equal(resolveEscalationLoop({}, main).implModel, main);
40
+ assert.equal(resolveEscalationLoop({ escalationLoop: {} }, main).implModel, main);
41
+ assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: "" } }, main).implModel, main);
42
+ assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: null } }, main).implModel, main);
43
+ assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: 42 } }, main).implModel, main);
44
+ });
45
+
46
+ test("escalationLoop: non-empty string wins, trimmed; undefined mainLoop passes through", () => {
47
+ assert.equal(
48
+ resolveEscalationLoop({ escalationLoop: { implModel: " p/x:high " } }, "p/main:medium").implModel,
49
+ "p/x:high",
50
+ );
51
+ assert.equal(resolveEscalationLoop({}, undefined).implModel, undefined);
52
+ });
53
+
54
+ test("mainLoopModel: always suffixed; unset -> off, max -> xhigh, recognised pass through", () => {
55
+ const m = { provider: "p", id: "id" };
56
+ assert.equal(mainLoopModel(m, undefined), "p/id:off");
57
+ assert.equal(mainLoopModel(m, "off"), "p/id:off");
58
+ assert.equal(mainLoopModel(m, "medium"), "p/id:medium");
59
+ assert.equal(mainLoopModel(m, "xhigh"), "p/id:xhigh");
60
+ assert.equal(mainLoopModel(m, "max"), "p/id:xhigh");
61
+ assert.equal(mainLoopModel(undefined, "medium"), undefined);
62
+ });
63
+
27
64
  test("specCouncil: non-empty string array -> council", () => {
28
65
  const r = resolveSpecCouncil({ specCouncil: { members: ["p/m1", " p/m2 "], chair: "p/c" } });
29
66
  assert.equal(r.verdict, "council");
@@ -8,6 +8,7 @@ export interface PiGauntlet {
8
8
  closureReview?: { enforce?: unknown; model?: unknown; maxFixRounds?: unknown };
9
9
  flowGuards?: { enforce?: unknown; specDirs?: unknown };
10
10
  verifyBeforeShip?: { testCommands?: unknown; warningReference?: unknown };
11
+ escalationLoop?: { implModel?: unknown };
11
12
  }
12
13
 
13
14
  // Whole-object second-level merge: each piGauntlet key present in the repo layer
@@ -87,6 +88,29 @@ export function resolveClosureReview(g: PiGauntlet): ClosureReviewResolved {
87
88
  return { model, enforce, maxFixRounds };
88
89
  }
89
90
 
91
+ export interface EscalationLoopResolved {
92
+ implModel: string | undefined;
93
+ }
94
+
95
+ export function resolveEscalationLoop(g: PiGauntlet, mainLoop: string | undefined): EscalationLoopResolved {
96
+ const raw = g.escalationLoop?.implModel;
97
+ return { implModel: nonEmptyString(raw) ? raw.trim() : mainLoop };
98
+ }
99
+
100
+ const THINKING_SUFFIXES = new Set(["off", "minimal", "low", "medium", "high", "xhigh"]);
101
+
102
+ // Always emit a suffix so pi-cohort's applyThinkingSuffix never falls back to the
103
+ // implementer's configured thinking; pi's "max" has no pi-cohort equivalent -> xhigh.
104
+ export function mainLoopModel(
105
+ model: { provider: string; id: string } | undefined,
106
+ thinkingLevel: string | undefined,
107
+ ): string | undefined {
108
+ if (!model) return undefined;
109
+ const level =
110
+ thinkingLevel === "max" ? "xhigh" : THINKING_SUFFIXES.has(thinkingLevel ?? "") ? thinkingLevel : "off";
111
+ return `${model.provider}/${model.id}:${level}`;
112
+ }
113
+
90
114
  export interface FlowGuardsResolved {
91
115
  enforce: boolean;
92
116
  specDirs: string[];
@@ -38,6 +38,11 @@ const VALID_PLAN = `# Fixture Plan
38
38
  **Files:**
39
39
  - Create: extensions/lib/fixture-task1.ts
40
40
  - Modify: extensions/lib/fixture-shared.ts
41
+ - Test: extensions/lib/fixture-task1.test.ts
42
+
43
+ **Tests:**
44
+ - \`node --test extensions/lib/fixture-task1.test.ts\`
45
+ - via: \`helperFn()\`
41
46
 
42
47
  This task implements helperFn() for parsing.
43
48
 
@@ -48,6 +53,10 @@ This task implements helperFn() for parsing.
48
53
  **Files:**
49
54
  - Create: extensions/lib/fixture-task2.ts
50
55
  - Modify: extensions/lib/fixture-other.ts
56
+ - Test: extensions/lib/fixture-task2.test.ts
57
+
58
+ **Tests:**
59
+ - \`node --test extensions/lib/fixture-task2.test.ts\`
51
60
 
52
61
  This task handles naming details.
53
62
 
@@ -61,6 +70,10 @@ Solo: lone remaining task
61
70
 
62
71
  **Files:**
63
72
  - Modify: extensions/lib/fixture-task3.ts
73
+ - Test: extensions/lib/fixture-task3.test.ts
74
+
75
+ **Tests:**
76
+ - \`node --test extensions/lib/fixture-task3.test.ts\`
64
77
 
65
78
  The literal TODO is intentionally documented here per spec quote-integrity requirement.
66
79
 
@@ -136,8 +149,8 @@ test("check 1 fail-closed: missing '## Spec coverage' table entirely", () => {
136
149
 
137
150
  test("check 2 quote-integrity: required verbatim literal missing from owner task body", () => {
138
151
  const mutated = VALID_PLAN.replace(
139
- "This task implements helperFn() for parsing.",
140
- "This task implements the helper for parsing.",
152
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
153
+ "- via: `parserSeam()`\n\nThis task implements the helper for parsing.",
141
154
  );
142
155
  const findings = checkPlan(mutated, SPEC_TEXT, alwaysTruePort());
143
156
  const qi = findingsFor(findings, "quote-integrity");
@@ -374,8 +387,61 @@ test("Verification quote-integrity: task body containing only a sub-command of a
374
387
  assert.deepEqual(findingsFor(findings, "quote-integrity"), []);
375
388
  });
376
389
 
390
+ const HE = (plan: string) => findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "header-entrypoint");
391
+ const WD = (plan: string) => findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "wave-file-disjointness");
392
+
393
+ test("header-entrypoint: Run: with backticked payload npm test under header npm test && npm run lint fails (regression for the hole)", () => {
394
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test && npm run lint")
395
+ .replace("This task handles naming details.", "- [ ] **Step 1: verify**\n\n Run: `npm test`\n Expected: PASS");
396
+ assert.ok(HE(mutated).some((f) => f.text.includes("Run: `npm test`")));
397
+ });
398
+
399
+ test("header-entrypoint: Run: npm test && echo ok fails", () => {
400
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test")
401
+ .replace("This task handles naming details.", "Run: npm test && echo ok");
402
+ assert.ok(HE(mutated).some((f) => f.text.includes("Run: npm test && echo ok")));
403
+ });
404
+
405
+ test("header-entrypoint: Run: npm test -- x.test.ts passes under bare and backticked header npm test", () => {
406
+ for (const header of ["**Verification:** npm test", "**Verification:** `npm test`"]) {
407
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header)
408
+ .replace("This task handles naming details.", "Run: npm test -- x.test.ts");
409
+ assert.deepEqual(HE(mutated), [], header);
410
+ }
411
+ });
412
+
413
+ test("header-entrypoint: Run: line with two backtick spans - both are payload", () => {
414
+ const mutated = VALID_PLAN.replace("This task handles naming details.", "Run: `echo a` then `npm run fixture-verify`");
415
+ assert.equal(HE(mutated).length, 1);
416
+ });
417
+
418
+ test("header-entrypoint: Tests: bullets are judged by tests-block, not here", () => {
419
+ const mutated = VALID_PLAN.replace(
420
+ "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`",
421
+ "**Tests:**\n- `npm run fixture-verify`",
422
+ );
423
+ assert.deepEqual(HE(mutated), []);
424
+ assert.ok(findingsFor(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), "tests-block").some((f) => f.reason.includes("full-suite command")));
425
+ });
426
+
427
+ test("wave-file-disjointness: Test/Test allowed; Test vs Modify conflict; Modify/Modify conflict", () => {
428
+ const shared = "- Test: extensions/lib/fixture-shared.test.ts\n";
429
+ const testTest = VALID_PLAN.replace("- Test: extensions/lib/fixture-task1.test.ts\n", shared).replace("- Test: extensions/lib/fixture-task2.test.ts\n", shared)
430
+ .replace("- `node --test extensions/lib/fixture-task1.test.ts`", "- `node --test extensions/lib/fixture-shared.test.ts`")
431
+ .replace("- `node --test extensions/lib/fixture-task2.test.ts`", "- `node --test extensions/lib/fixture-shared.test.ts`");
432
+ assert.deepEqual(WD(testTest), []);
433
+ const testModify = VALID_PLAN.replace("- Test: extensions/lib/fixture-task2.test.ts\n", "- Test: extensions/lib/fixture-shared.ts\n")
434
+ .replace("- `node --test extensions/lib/fixture-task2.test.ts`", "- `node --test extensions/lib/fixture-shared.ts`");
435
+ assert.ok(WD(testModify).some((f) => f.reason.includes("fixture-shared.ts")));
436
+ const modifyModify = VALID_PLAN.replace("- Modify: extensions/lib/fixture-other.ts", "- Modify: extensions/lib/fixture-shared.ts");
437
+ assert.ok(WD(modifyModify).some((f) => f.reason.includes("fixture-shared.ts")));
438
+ });
439
+
377
440
  test("Verification quote-integrity: task-owned literal check unchanged", () => {
378
- const mutated = VALID_PLAN.replace("This task implements helperFn() for parsing.", "This task implements the helper.");
441
+ const mutated = VALID_PLAN.replace(
442
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
443
+ "- via: `parserSeam()`\n\nThis task implements the helper.",
444
+ );
379
445
  const qi = findingsFor(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), "quote-integrity");
380
446
  assert.ok(qi.some((f) => f.reason.includes("Task 1 body does not contain the required verbatim literal `helperFn()`")));
381
447
  });
@@ -405,6 +471,10 @@ Solo: lone remaining task
405
471
 
406
472
  **Files:**
407
473
  - Modify: extensions/lib/fixture-task1.ts
474
+ - Test: extensions/lib/plan-check.test.ts
475
+
476
+ **Tests:**
477
+ - \`node --test extensions/lib/plan-check.test.ts\`
408
478
 
409
479
  Run node --test extensions/lib/plan-check.test.ts and confirm green.
410
480
 
@@ -423,8 +493,8 @@ test("quote-integrity: header entrypoint literal on an anchored line is satisfie
423
493
 
424
494
  test("quote-integrity: scoped command on the same anchored line is still required in the task body", () => {
425
495
  const mutated = ENTRYPOINT_PLAN.replace(
426
- "Run node --test extensions/lib/plan-check.test.ts and confirm green.",
427
- "Run the scoped test and confirm green.",
496
+ "- Test: extensions/lib/plan-check.test.ts\n\n**Tests:**\n- `node --test extensions/lib/plan-check.test.ts`\n\nRun node --test extensions/lib/plan-check.test.ts and confirm green.",
497
+ "- Test: extensions/lib/other.test.ts\n\n**Tests:**\n- `node --test extensions/lib/other.test.ts`\n\nRun the scoped test and confirm green.",
428
498
  );
429
499
  const qi = findingsFor(checkPlan(mutated, ENTRYPOINT_SPEC, alwaysTruePort()), "quote-integrity");
430
500
  assert.equal(qi.length, 1);
@@ -522,7 +592,7 @@ test("check 4 fail-closed: invalid glob (port throws)", () => {
522
592
 
523
593
  test("check 4 fail-closed: task missing **Files:** block", () => {
524
594
  const mutated = VALID_PLAN.replace(
525
- "**Files:**\n- Modify: extensions/lib/fixture-task3.ts\n\n",
595
+ "**Files:**\n- Modify: extensions/lib/fixture-task3.ts\n- Test: extensions/lib/fixture-task3.test.ts\n\n**Tests:**\n- `node --test extensions/lib/fixture-task3.test.ts`\n\n",
526
596
  "",
527
597
  );
528
598
  const findings = checkPlan(mutated, SPEC_TEXT, alwaysTruePort());
@@ -627,6 +697,10 @@ Solo: only task in this wave
627
697
 
628
698
  **Files:**
629
699
  - Create: extensions/lib/exemption-task1.ts
700
+ - Test: extensions/lib/exemption-task1.test.ts
701
+
702
+ **Tests:**
703
+ - \`node --test extensions/lib/exemption-task1.test.ts\`
630
704
 
631
705
  The literal TODO is intentionally documented here per spec quote-integrity requirement.
632
706
 
@@ -672,8 +746,8 @@ test("aggregate: independent mutations across three checks are all reported toge
672
746
  let mutated = VALID_PLAN;
673
747
  mutated = mutated.replace("Solo: lone remaining task\n\n", "");
674
748
  mutated = mutated.replace(
675
- "This task implements helperFn() for parsing.",
676
- "This task implements the helper for parsing.",
749
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
750
+ "- via: `parserSeam()`\n\nThis task implements the helper for parsing.",
677
751
  );
678
752
  mutated = mutated.replace(
679
753
  "- Modify: extensions/lib/fixture-other.ts",
@@ -727,3 +801,202 @@ test("sha256 returns lowercase hex of the expected length", () => {
727
801
  assert.equal(digest.length, 64);
728
802
  assert.match(digest, /^[0-9a-f]+$/);
729
803
  });
804
+
805
+ const T2_TESTS = "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`\n";
806
+ const T2_FILES_TEST = "- Test: extensions/lib/fixture-task2.test.ts\n";
807
+
808
+ function tb(plan: string, spec = SPEC_TEXT, fs: FsPort = alwaysTruePort()): PlanCheckFinding[] {
809
+ return findingsFor(checkPlan(plan, spec, fs), "tests-block");
810
+ }
811
+
812
+ test("tests-block: block missing when a task has no **Tests:**", () => {
813
+ const mutated = VALID_PLAN.replace(T2_TESTS, "");
814
+ assert.ok(tb(mutated).some((f) => f.reason.includes("block missing") && f.reason.includes("Task 2")));
815
+ });
816
+
817
+ test("tests-block: block missing when **Files:** is absent", () => {
818
+ const mutated = VALID_PLAN.replace(
819
+ "**Files:**\n- Create: extensions/lib/fixture-task2.ts\n- Modify: extensions/lib/fixture-other.ts\n" + T2_FILES_TEST + "\n" + T2_TESTS,
820
+ "",
821
+ );
822
+ const f = tb(mutated);
823
+ assert.ok(f.some((x) => x.reason.includes("block missing") && x.reason.includes("Task 2")));
824
+ });
825
+
826
+ test("tests-block: misplaced block (before Files:) fires without block missing", () => {
827
+ const mutated = VALID_PLAN.replace(
828
+ "**Files:**\n- Create: extensions/lib/fixture-task2.ts",
829
+ T2_TESTS + "\n**Files:**\n- Create: extensions/lib/fixture-task2.ts",
830
+ ).replace("\n" + T2_TESTS + "\nThis task handles naming details.", "\nThis task handles naming details.");
831
+ const f = tb(mutated);
832
+ assert.ok(f.some((x) => x.reason.includes("misplaced")));
833
+ assert.ok(!f.some((x) => x.reason.includes("block missing")));
834
+ });
835
+
836
+ test("tests-block: duplicated **Tests:** heading is misplaced", () => {
837
+ const mutated = VALID_PLAN.replace("This task handles naming details.", T2_TESTS + "\nThis task handles naming details.");
838
+ assert.ok(tb(mutated).some((x) => x.reason.includes("misplaced")));
839
+ });
840
+
841
+ test("tests-block: text after the heading is misplaced, not missing", () => {
842
+ const mutated = VALID_PLAN.replace("**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`", "**Tests:** see below\n- `node --test extensions/lib/fixture-task2.test.ts`");
843
+ const f = tb(mutated);
844
+ assert.ok(f.some((x) => x.reason.includes("misplaced")));
845
+ assert.ok(!f.some((x) => x.reason.includes("block missing")));
846
+ });
847
+
848
+ test("tests-block: a Delete: bullet before **Tests:** is legal", () => {
849
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, T2_FILES_TEST + "- Delete: extensions/lib/fixture-legacy.ts\n");
850
+ assert.deepEqual(tb(mutated), []);
851
+ });
852
+
853
+ test("tests-block: block empty (via: alone)", () => {
854
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- via: `naming()`\n");
855
+ assert.ok(tb(mutated).some((x) => x.reason.includes("block empty")));
856
+ });
857
+
858
+ test("tests-block: malformed bullet after a valid one; block continues; a - [ ] step terminates", () => {
859
+ const mutated = VALID_PLAN.replace(
860
+ T2_TESTS,
861
+ T2_TESTS + "- node --test x\n- via: `naming()`\n- [ ] **Step 1: nothing**\n",
862
+ );
863
+ const f = tb(mutated);
864
+ assert.equal(f.filter((x) => x.reason.includes("malformed")).length, 1);
865
+ assert.ok(!f.some((x) => x.reason.includes("block empty")));
866
+ });
867
+
868
+ test("tests-block: none: with a command, none: with via:, two none: -> contradictory", () => {
869
+ for (const block of [
870
+ "**Tests:**\n- none: docs\n- `node --test extensions/lib/fixture-task2.test.ts`\n",
871
+ "**Tests:**\n- none: docs\n- via: `naming()`\n",
872
+ "**Tests:**\n- none: docs\n- none: config\n",
873
+ ]) {
874
+ const mutated = VALID_PLAN.replace(T2_TESTS, block);
875
+ assert.ok(tb(mutated).some((x) => x.reason.includes("contradictory")), block);
876
+ }
877
+ });
878
+
879
+ test("tests-block: none: while a Test: path is declared -> unused Test: path", () => {
880
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- none: docs\n");
881
+ assert.ok(tb(mutated).some((x) => x.reason.includes("unused `Test:` path")));
882
+ });
883
+
884
+ test("tests-block: none: docs, config with no Test: path is valid", () => {
885
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, "").replace(T2_TESTS, "**Tests:**\n- none: docs, config\n");
886
+ assert.deepEqual(tb(mutated), []);
887
+ });
888
+
889
+ test("tests-block: via: with commands passes", () => {
890
+ const mutated = VALID_PLAN.replace(T2_TESTS, T2_TESTS + "- via: `naming()` - the seam\n");
891
+ assert.deepEqual(tb(mutated), []);
892
+ });
893
+
894
+ test("tests-block: unknown Test: path unless it exists or another task Create:s it", () => {
895
+ const fs: FsPort = { exists: () => false, glob: () => [] };
896
+ assert.ok(tb(VALID_PLAN, SPEC_TEXT, fs).some((x) => x.reason.includes("unknown `Test:` path")));
897
+ const created = VALID_PLAN.replace(
898
+ "- Create: extensions/lib/fixture-task1.ts",
899
+ "- Create: extensions/lib/fixture-task1.ts\n- Create: extensions/lib/fixture-task1.test.ts\n- Create: extensions/lib/fixture-task2.test.ts\n- Create: extensions/lib/fixture-task3.test.ts",
900
+ );
901
+ assert.deepEqual(tb(created, SPEC_TEXT, fs), []);
902
+ });
903
+
904
+ test("tests-block: segment without a Test: token is not anchored", () => {
905
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts && echo done`\n");
906
+ assert.ok(tb(mutated).some((x) => x.reason.includes("not anchored")));
907
+ });
908
+
909
+ test("tests-block: pipe segment equal to a header segment; tee log has no anchor", () => {
910
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test")
911
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
912
+ .replace(T2_TESTS, "**Tests:**\n- `npm test | tee log && node --test x.test.ts`\n");
913
+ const f = tb(mutated);
914
+ assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.reason.includes("npm test")));
915
+ assert.ok(f.some((x) => x.reason.includes("not anchored") && x.reason.includes("tee log")));
916
+ });
917
+
918
+ test("tests-block: broadening selectors tests/ and tests/*.py fail; Test: value tests/ never anchors", () => {
919
+ const dir = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts tests/`\n");
920
+ assert.ok(tb(dir).some((x) => x.reason.includes("broadening")));
921
+ const glob = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts tests/*.py`\n");
922
+ assert.ok(tb(glob).some((x) => x.reason.includes("broadening")));
923
+ const dirAnchor = VALID_PLAN.replace(T2_FILES_TEST, "- Test: tests/\n").replace(
924
+ T2_TESTS,
925
+ "**Tests:**\n- `node --test tests/`\n",
926
+ );
927
+ assert.ok(tb(dirAnchor).some((x) => x.reason.includes("not anchored")));
928
+ });
929
+
930
+ test("tests-block: unsupported shell (cd, sh -c, bash -c, eval, $( )", () => {
931
+ for (const cmd of [
932
+ "cd pkg && pytest tests/a.py",
933
+ "sh -c 'node --test extensions/lib/fixture-task2.test.ts'",
934
+ "bash -c 'node --test extensions/lib/fixture-task2.test.ts'",
935
+ "eval node --test extensions/lib/fixture-task2.test.ts",
936
+ "node --test $(echo extensions/lib/fixture-task2.test.ts)",
937
+ ]) {
938
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `" + cmd + "`\n");
939
+ assert.ok(tb(mutated).some((x) => x.reason.includes("unsupported shell")), cmd);
940
+ }
941
+ });
942
+
943
+ test("tests-block: segment equal to a header segment (bare, &&, comma-listed, trailing prose, a && b)", () => {
944
+ for (const header of [
945
+ "**Verification:** npm test",
946
+ "**Verification:** `npm test`",
947
+ "**Verification:** npm test && npm run lint",
948
+ "**Verification:** `npm test`, `npm run lint`",
949
+ "**Verification:** `npm test` (bundles lint)",
950
+ ]) {
951
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header).replace(
952
+ T2_TESTS,
953
+ "**Tests:**\n- `npm run lint`\n- `npm test`\n",
954
+ );
955
+ const f = tb(mutated);
956
+ assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.text === "- `npm test`"), header);
957
+ if (header.includes("npm run lint")) assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.text.includes("lint")), header);
958
+ }
959
+ const ab = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** a && b").replace(T2_TESTS, "**Tests:**\n- `b`\n");
960
+ assert.ok(tb(ab).some((x) => x.reason.includes("full-suite command")));
961
+ });
962
+
963
+ test("tests-block: npm test -- x.test.ts passes under header npm test (bare and backticked); the header segment npm test alone fails", () => {
964
+ for (const header of ["**Verification:** npm test", "**Verification:** `npm test`"]) {
965
+ const ok = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header)
966
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
967
+ .replace(T2_TESTS, "**Tests:**\n- `npm test -- x.test.ts`\n");
968
+ assert.deepEqual(tb(ok, SPEC_TEXT, alwaysTruePort()), [], header);
969
+ }
970
+ });
971
+
972
+ test("tests-block: header segment that is itself scoped is still rejected as a bullet", () => {
973
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** node --test x.test.ts && npm run lint")
974
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
975
+ .replace(T2_TESTS, "**Tests:**\n- `node --test x.test.ts`\n");
976
+ assert.ok(tb(mutated).some((x) => x.reason.includes("full-suite command")));
977
+ });
978
+
979
+ test("tests-block: passing forms - multi-path, ::filter -v, -k name, -- passthrough, line-suffixed Test:", () => {
980
+ const multi = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts extensions/lib/fixture-task1.test.ts`\n")
981
+ .replace(T2_FILES_TEST, T2_FILES_TEST + "- Test: extensions/lib/fixture-task1.test.ts\n");
982
+ assert.deepEqual(tb(multi), []);
983
+ const filt = VALID_PLAN.replace(T2_FILES_TEST, "- Test: tests/a.py\n").replace(T2_TESTS, "**Tests:**\n- `pytest tests/a.py::test_x -v -k name --filter x`\n");
984
+ assert.deepEqual(tb(filt), []);
985
+ const passthrough = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `npm run fixture-verify -- extensions/lib/fixture-task2.test.ts`\n");
986
+ assert.deepEqual(tb(passthrough), []);
987
+ const suffixed = VALID_PLAN.replace(T2_FILES_TEST, "- Test: extensions/lib/fixture-task2.test.ts:10-20\n");
988
+ assert.deepEqual(tb(suffixed), []);
989
+ });
990
+
991
+ test("tests-block: no path normalization - ./x and x differ, as in Files:", () => {
992
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, "- Test: ./extensions/lib/fixture-task2.test.ts\n");
993
+ assert.ok(tb(mutated).some((x) => x.reason.includes("not anchored")));
994
+ });
995
+
996
+ test("tests-block: fenced **Tests:** lines are ignored", () => {
997
+ const mutated = VALID_PLAN.replace(
998
+ "This task handles naming details.",
999
+ "This task handles naming details.\n\n```markdown\n**Tests:**\n- none: docs\n```",
1000
+ );
1001
+ assert.deepEqual(tb(mutated), []);
1002
+ });