pi-gauntlet 5.4.0 → 5.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.5.1 - 2026-09-10
4
+
5
+ - `linear`: copyable, version-scoped recovery for attachment-download 401s resolves the decrypted credential through linearis instead of reading encrypted token storage. Restricts credential delivery to HTTPS Linear uploads, rejects redirects, and checks downloaded bytes; an offline regression executes the documented example.
6
+
7
+ ## v5.5.0 - 2026-09-10
8
+
9
+ - `writing-plans`: every task carries a `**Tests:**` block - scoped commands anchored to the task's `Test:` paths, optional `via:` entry point, or `none: <category>`; the plan grammar `plan_check` enforces moves to `skills/writing-plans/reference/plan-contract.md`. (#28)
10
+ - `plan_check`: new `tests-block` check (block present and well-formed, every command segment names a task `Test:` path, no broadening selectors or `cd`/`sh -c`/`eval`/`$(`, no segment equal to a header `**Verification:**` segment); `header-entrypoint` compares `Run:` payloads by segment, so `npm test` no longer slips past `npm test && npm run lint`; `Test:` entries no longer count as file ownership in `wave-file-disjointness`. (#28)
11
+ - `subagent-driven-development`, `test-driven-development`, prompts, `spec-reviewer`: `SCOPED_TEST_COMMANDS` comes from the `Tests:` block; implementer reports `met`/`unmet` per command and per `via:`; the spec reviewer treats the task contract as a supplement the anchored spec overrides. (#28)
12
+
3
13
  ## v5.4.0 - 2026-09-10
4
14
 
5
15
  - `subagent-driven-development`: a stalled review fix loop runs one escalated fix round (`implementer`, `context: fresh`, model from new `piGauntlet.escalationLoop.implModel`, default main-loop model + thinking) before stopping; the stop is a one-screen problem note (`stop-note.md`) with concrete fix options, replacing the trajectory-log escalation report. `gauntlet_setting` gains the `escalationLoop` key. (#29)
package/README.md CHANGED
@@ -71,7 +71,7 @@ pi-gauntlet ships three kinds of pieces, layered on top of pi-cohort's dispatch:
71
71
 
72
72
  - **17 skills** - the workflow logic. Thirteen activate automatically when pi sees the matching kind of task, and each one gates the next: `brainstorming`, `writing-plans`, `roasting-the-spec`, `test-driven-development`, `subagent-driven-development`, `dispatching-parallel-agents`, `verification-before-completion`, `requesting-code-review`, `receiving-code-review`, `using-git-worktrees`, `finishing-a-development-branch`, `writing-skills`, `linear` (reads/searches/comments on/manages Linear tickets via the `linearis` CLI; owns all linearis mechanics and the `## Issue tracker` overrides schema; tracker-facing skills route to it). Four more are explicit-invocation-only (`disable-model-invocation: true`): `shape-ticket` creates or repairs one tracker issue per run against a Context/Problem/Idea/Acceptance-Criteria template, gated by an AC integrity check, a cheap council roast, and a single human-confirmed write - run it with `/skill:shape-ticket`. `gatekeep-pr` is consent-gated pre-merge verification of a PR against its issue - read-only gathering, verification evidence resolved CI-first (green checks on the exact assessed head count as evidence; the project's verification command runs only as fallback), a rubric-based review, then a deterministic authorship-aware menu with stable finding IDs (P#/L#/C#/F#) and numbered pre-composed courses (fixes execute as a single parallel-safe wave: one gate run, one re-review, one push); nothing mutates (fixes, pushes, reviews, merges) until you pick a row - run it with `/skill:gatekeep-pr <pr>`. `check-delivery` is a post-merge detective control: proves an issue actually shipped (default-branch landing, delivery target, per-AC evidence) before its tracker status advances; it never writes a terminal status - run it with `/skill:check-delivery <ref>`. `chase-bug` is human-only bug triage: read-only root-cause discovery to an evidenced verdict menu (real bug -> ticket/brainstorm/hotfix/respond; five negative verdicts), then a gated response to the reporter for addressable origins (GitHub issue / tracker ticket) and a rendered verdict summary otherwise - it never fixes during triage; the hotfix row hands off to `skills/chase-bug/hotfix.md` after the menu - run it with `/skill:chase-bug`.
73
73
  - **7 subagent personas** - the specialized child agents the skills dispatch via pi-cohort: `implementer`, `code-reviewer`, `spec-reviewer`, `conformance-reviewer`, `spec-summarizer`, `spec-council-member`, `spec-council-synthesizer`. See [doc/personas.md](./doc/personas.md) for what each one does and why its permissions are scoped the way they are.
74
- - **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers plan_check, a deterministic plan linter (9 mechanical plan-vs-spec checks) whose pass stamp gates implement-start inside a gauntlet flow. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
74
+ - **3 runtime extensions** - the enforcement layer. `plan-tracker` and `phase-tracker` are tools skills call to track progress (with a TUI widget); `verify-before-ship` is a hook that warns if you push or open a PR without a passing test run since your last edit; a phase-tracker flow guard reminds on implement-phase commits missing spec/code review. In a brainstorming-entered flow, phase-tracker rejects `implement` or `verify` completion while tracker tasks remain pending or in progress; see [its configuration reference](./doc/configuration.md#phase-tracker). phase-tracker also registers `plan_check`, which verifies a plan against its spec and against the grammar in [skills/writing-plans/reference/plan-contract.md](./skills/writing-plans/reference/plan-contract.md), including that each task's `Tests:` commands are selective and never the full suite; a pass stamps the plan for implementation. See [doc/configuration.md](./doc/configuration.md) for the settings each one reads.
75
75
 
76
76
  pi-gauntlet is **opinionated**: every non-trivial change is *meant* to ride this one pipeline, entered through `brainstorming`. Enforcement is opt-in by entry, not ambient: once brainstorming starts a flow, the phase-tracker extension mechanically blocks a phase from closing before its gate runs, and warns once if the main loop writes code during implement (subagents own implement-phase edits). A change made *without* entering the flow (a typo, a formatting run, a dependency bump - see "When to use / when NOT to use") is not gated; the discipline of routing real work through the pipeline is a convention the tooling supports, not a trap it springs on every edit.
77
77
 
@@ -13,14 +13,15 @@ You are a spec compliance reviewer. Your job is to verify that an implementation
13
13
 
14
14
  ## Process
15
15
 
16
- <!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md — change them together or not at all -->
16
+ <!-- clause decomposition / snippet non-authority / whole-file reads / task-contract supplement: keep in lockstep with skills/subagent-driven-development/spec-reviewer-prompt.md — change them together or not at all -->
17
17
 
18
18
  1. Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses. Every clause gets a verdict row.
19
- 2. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
20
- 3. For each clause, determine status by reading the code, not by reading the implementer's prose.
21
- 4. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
22
- 5. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
23
- 6. Flag any clause from the spec that is missing from the implementation.
19
+ 2. The task's `**Tests:**` block, `via:`, and its `Files:` paths supplement the anchored spec where it is silent; the anchored spec wins a conflict - report the divergence once, against the plan, never against code corrected to the spec. Findings: a `Create:` path absent from the diff or created elsewhere; a test that does not call the `via:` entry point; a `Tests:` block the diff contradicts. Existing files need no diff touch.
20
+ 3. Read the implementation. Do not trust summaries. Read every diff-touched file in full, not just the hunks - continue in chunks until the file is exhausted; if you cannot exhaust it, say so in the report instead of treating the file as covered. A statement elsewhere in a touched file that the change now contradicts is in scope.
21
+ 4. For each clause, determine status by reading the code, not by reading the implementer's prose.
22
+ 5. Never run tests, linters, or type-checkers. Your evidence is the diff and the files you read. Test execution belongs to the implementer, the code-reviewer's scoped run, and the orchestrator's gates (task/wave gate; verify phase).
23
+ 6. Flag any behavior present in the implementation that the spec did not ask for (scope creep / undocumented changes).
24
+ 7. Flag any clause from the spec that is missing from the implementation.
24
25
 
25
26
  ## Output format
26
27
 
@@ -38,6 +38,11 @@ const VALID_PLAN = `# Fixture Plan
38
38
  **Files:**
39
39
  - Create: extensions/lib/fixture-task1.ts
40
40
  - Modify: extensions/lib/fixture-shared.ts
41
+ - Test: extensions/lib/fixture-task1.test.ts
42
+
43
+ **Tests:**
44
+ - \`node --test extensions/lib/fixture-task1.test.ts\`
45
+ - via: \`helperFn()\`
41
46
 
42
47
  This task implements helperFn() for parsing.
43
48
 
@@ -48,6 +53,10 @@ This task implements helperFn() for parsing.
48
53
  **Files:**
49
54
  - Create: extensions/lib/fixture-task2.ts
50
55
  - Modify: extensions/lib/fixture-other.ts
56
+ - Test: extensions/lib/fixture-task2.test.ts
57
+
58
+ **Tests:**
59
+ - \`node --test extensions/lib/fixture-task2.test.ts\`
51
60
 
52
61
  This task handles naming details.
53
62
 
@@ -61,6 +70,10 @@ Solo: lone remaining task
61
70
 
62
71
  **Files:**
63
72
  - Modify: extensions/lib/fixture-task3.ts
73
+ - Test: extensions/lib/fixture-task3.test.ts
74
+
75
+ **Tests:**
76
+ - \`node --test extensions/lib/fixture-task3.test.ts\`
64
77
 
65
78
  The literal TODO is intentionally documented here per spec quote-integrity requirement.
66
79
 
@@ -136,8 +149,8 @@ test("check 1 fail-closed: missing '## Spec coverage' table entirely", () => {
136
149
 
137
150
  test("check 2 quote-integrity: required verbatim literal missing from owner task body", () => {
138
151
  const mutated = VALID_PLAN.replace(
139
- "This task implements helperFn() for parsing.",
140
- "This task implements the helper for parsing.",
152
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
153
+ "- via: `parserSeam()`\n\nThis task implements the helper for parsing.",
141
154
  );
142
155
  const findings = checkPlan(mutated, SPEC_TEXT, alwaysTruePort());
143
156
  const qi = findingsFor(findings, "quote-integrity");
@@ -374,8 +387,61 @@ test("Verification quote-integrity: task body containing only a sub-command of a
374
387
  assert.deepEqual(findingsFor(findings, "quote-integrity"), []);
375
388
  });
376
389
 
390
+ const HE = (plan: string) => findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "header-entrypoint");
391
+ const WD = (plan: string) => findingsFor(checkPlan(plan, SPEC_TEXT, alwaysTruePort()), "wave-file-disjointness");
392
+
393
+ test("header-entrypoint: Run: with backticked payload npm test under header npm test && npm run lint fails (regression for the hole)", () => {
394
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test && npm run lint")
395
+ .replace("This task handles naming details.", "- [ ] **Step 1: verify**\n\n Run: `npm test`\n Expected: PASS");
396
+ assert.ok(HE(mutated).some((f) => f.text.includes("Run: `npm test`")));
397
+ });
398
+
399
+ test("header-entrypoint: Run: npm test && echo ok fails", () => {
400
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test")
401
+ .replace("This task handles naming details.", "Run: npm test && echo ok");
402
+ assert.ok(HE(mutated).some((f) => f.text.includes("Run: npm test && echo ok")));
403
+ });
404
+
405
+ test("header-entrypoint: Run: npm test -- x.test.ts passes under bare and backticked header npm test", () => {
406
+ for (const header of ["**Verification:** npm test", "**Verification:** `npm test`"]) {
407
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header)
408
+ .replace("This task handles naming details.", "Run: npm test -- x.test.ts");
409
+ assert.deepEqual(HE(mutated), [], header);
410
+ }
411
+ });
412
+
413
+ test("header-entrypoint: Run: line with two backtick spans - both are payload", () => {
414
+ const mutated = VALID_PLAN.replace("This task handles naming details.", "Run: `echo a` then `npm run fixture-verify`");
415
+ assert.equal(HE(mutated).length, 1);
416
+ });
417
+
418
+ test("header-entrypoint: Tests: bullets are judged by tests-block, not here", () => {
419
+ const mutated = VALID_PLAN.replace(
420
+ "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`",
421
+ "**Tests:**\n- `npm run fixture-verify`",
422
+ );
423
+ assert.deepEqual(HE(mutated), []);
424
+ assert.ok(findingsFor(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), "tests-block").some((f) => f.reason.includes("full-suite command")));
425
+ });
426
+
427
+ test("wave-file-disjointness: Test/Test allowed; Test vs Modify conflict; Modify/Modify conflict", () => {
428
+ const shared = "- Test: extensions/lib/fixture-shared.test.ts\n";
429
+ const testTest = VALID_PLAN.replace("- Test: extensions/lib/fixture-task1.test.ts\n", shared).replace("- Test: extensions/lib/fixture-task2.test.ts\n", shared)
430
+ .replace("- `node --test extensions/lib/fixture-task1.test.ts`", "- `node --test extensions/lib/fixture-shared.test.ts`")
431
+ .replace("- `node --test extensions/lib/fixture-task2.test.ts`", "- `node --test extensions/lib/fixture-shared.test.ts`");
432
+ assert.deepEqual(WD(testTest), []);
433
+ const testModify = VALID_PLAN.replace("- Test: extensions/lib/fixture-task2.test.ts\n", "- Test: extensions/lib/fixture-shared.ts\n")
434
+ .replace("- `node --test extensions/lib/fixture-task2.test.ts`", "- `node --test extensions/lib/fixture-shared.ts`");
435
+ assert.ok(WD(testModify).some((f) => f.reason.includes("fixture-shared.ts")));
436
+ const modifyModify = VALID_PLAN.replace("- Modify: extensions/lib/fixture-other.ts", "- Modify: extensions/lib/fixture-shared.ts");
437
+ assert.ok(WD(modifyModify).some((f) => f.reason.includes("fixture-shared.ts")));
438
+ });
439
+
377
440
  test("Verification quote-integrity: task-owned literal check unchanged", () => {
378
- const mutated = VALID_PLAN.replace("This task implements helperFn() for parsing.", "This task implements the helper.");
441
+ const mutated = VALID_PLAN.replace(
442
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
443
+ "- via: `parserSeam()`\n\nThis task implements the helper.",
444
+ );
379
445
  const qi = findingsFor(checkPlan(mutated, SPEC_TEXT, alwaysTruePort()), "quote-integrity");
380
446
  assert.ok(qi.some((f) => f.reason.includes("Task 1 body does not contain the required verbatim literal `helperFn()`")));
381
447
  });
@@ -405,6 +471,10 @@ Solo: lone remaining task
405
471
 
406
472
  **Files:**
407
473
  - Modify: extensions/lib/fixture-task1.ts
474
+ - Test: extensions/lib/plan-check.test.ts
475
+
476
+ **Tests:**
477
+ - \`node --test extensions/lib/plan-check.test.ts\`
408
478
 
409
479
  Run node --test extensions/lib/plan-check.test.ts and confirm green.
410
480
 
@@ -423,8 +493,8 @@ test("quote-integrity: header entrypoint literal on an anchored line is satisfie
423
493
 
424
494
  test("quote-integrity: scoped command on the same anchored line is still required in the task body", () => {
425
495
  const mutated = ENTRYPOINT_PLAN.replace(
426
- "Run node --test extensions/lib/plan-check.test.ts and confirm green.",
427
- "Run the scoped test and confirm green.",
496
+ "- Test: extensions/lib/plan-check.test.ts\n\n**Tests:**\n- `node --test extensions/lib/plan-check.test.ts`\n\nRun node --test extensions/lib/plan-check.test.ts and confirm green.",
497
+ "- Test: extensions/lib/other.test.ts\n\n**Tests:**\n- `node --test extensions/lib/other.test.ts`\n\nRun the scoped test and confirm green.",
428
498
  );
429
499
  const qi = findingsFor(checkPlan(mutated, ENTRYPOINT_SPEC, alwaysTruePort()), "quote-integrity");
430
500
  assert.equal(qi.length, 1);
@@ -522,7 +592,7 @@ test("check 4 fail-closed: invalid glob (port throws)", () => {
522
592
 
523
593
  test("check 4 fail-closed: task missing **Files:** block", () => {
524
594
  const mutated = VALID_PLAN.replace(
525
- "**Files:**\n- Modify: extensions/lib/fixture-task3.ts\n\n",
595
+ "**Files:**\n- Modify: extensions/lib/fixture-task3.ts\n- Test: extensions/lib/fixture-task3.test.ts\n\n**Tests:**\n- `node --test extensions/lib/fixture-task3.test.ts`\n\n",
526
596
  "",
527
597
  );
528
598
  const findings = checkPlan(mutated, SPEC_TEXT, alwaysTruePort());
@@ -627,6 +697,10 @@ Solo: only task in this wave
627
697
 
628
698
  **Files:**
629
699
  - Create: extensions/lib/exemption-task1.ts
700
+ - Test: extensions/lib/exemption-task1.test.ts
701
+
702
+ **Tests:**
703
+ - \`node --test extensions/lib/exemption-task1.test.ts\`
630
704
 
631
705
  The literal TODO is intentionally documented here per spec quote-integrity requirement.
632
706
 
@@ -672,8 +746,8 @@ test("aggregate: independent mutations across three checks are all reported toge
672
746
  let mutated = VALID_PLAN;
673
747
  mutated = mutated.replace("Solo: lone remaining task\n\n", "");
674
748
  mutated = mutated.replace(
675
- "This task implements helperFn() for parsing.",
676
- "This task implements the helper for parsing.",
749
+ "- via: `helperFn()`\n\nThis task implements helperFn() for parsing.",
750
+ "- via: `parserSeam()`\n\nThis task implements the helper for parsing.",
677
751
  );
678
752
  mutated = mutated.replace(
679
753
  "- Modify: extensions/lib/fixture-other.ts",
@@ -727,3 +801,202 @@ test("sha256 returns lowercase hex of the expected length", () => {
727
801
  assert.equal(digest.length, 64);
728
802
  assert.match(digest, /^[0-9a-f]+$/);
729
803
  });
804
+
805
+ const T2_TESTS = "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`\n";
806
+ const T2_FILES_TEST = "- Test: extensions/lib/fixture-task2.test.ts\n";
807
+
808
+ function tb(plan: string, spec = SPEC_TEXT, fs: FsPort = alwaysTruePort()): PlanCheckFinding[] {
809
+ return findingsFor(checkPlan(plan, spec, fs), "tests-block");
810
+ }
811
+
812
+ test("tests-block: block missing when a task has no **Tests:**", () => {
813
+ const mutated = VALID_PLAN.replace(T2_TESTS, "");
814
+ assert.ok(tb(mutated).some((f) => f.reason.includes("block missing") && f.reason.includes("Task 2")));
815
+ });
816
+
817
+ test("tests-block: block missing when **Files:** is absent", () => {
818
+ const mutated = VALID_PLAN.replace(
819
+ "**Files:**\n- Create: extensions/lib/fixture-task2.ts\n- Modify: extensions/lib/fixture-other.ts\n" + T2_FILES_TEST + "\n" + T2_TESTS,
820
+ "",
821
+ );
822
+ const f = tb(mutated);
823
+ assert.ok(f.some((x) => x.reason.includes("block missing") && x.reason.includes("Task 2")));
824
+ });
825
+
826
+ test("tests-block: misplaced block (before Files:) fires without block missing", () => {
827
+ const mutated = VALID_PLAN.replace(
828
+ "**Files:**\n- Create: extensions/lib/fixture-task2.ts",
829
+ T2_TESTS + "\n**Files:**\n- Create: extensions/lib/fixture-task2.ts",
830
+ ).replace("\n" + T2_TESTS + "\nThis task handles naming details.", "\nThis task handles naming details.");
831
+ const f = tb(mutated);
832
+ assert.ok(f.some((x) => x.reason.includes("misplaced")));
833
+ assert.ok(!f.some((x) => x.reason.includes("block missing")));
834
+ });
835
+
836
+ test("tests-block: duplicated **Tests:** heading is misplaced", () => {
837
+ const mutated = VALID_PLAN.replace("This task handles naming details.", T2_TESTS + "\nThis task handles naming details.");
838
+ assert.ok(tb(mutated).some((x) => x.reason.includes("misplaced")));
839
+ });
840
+
841
+ test("tests-block: text after the heading is misplaced, not missing", () => {
842
+ const mutated = VALID_PLAN.replace("**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts`", "**Tests:** see below\n- `node --test extensions/lib/fixture-task2.test.ts`");
843
+ const f = tb(mutated);
844
+ assert.ok(f.some((x) => x.reason.includes("misplaced")));
845
+ assert.ok(!f.some((x) => x.reason.includes("block missing")));
846
+ });
847
+
848
+ test("tests-block: a Delete: bullet before **Tests:** is legal", () => {
849
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, T2_FILES_TEST + "- Delete: extensions/lib/fixture-legacy.ts\n");
850
+ assert.deepEqual(tb(mutated), []);
851
+ });
852
+
853
+ test("tests-block: block empty (via: alone)", () => {
854
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- via: `naming()`\n");
855
+ assert.ok(tb(mutated).some((x) => x.reason.includes("block empty")));
856
+ });
857
+
858
+ test("tests-block: malformed bullet after a valid one; block continues; a - [ ] step terminates", () => {
859
+ const mutated = VALID_PLAN.replace(
860
+ T2_TESTS,
861
+ T2_TESTS + "- node --test x\n- via: `naming()`\n- [ ] **Step 1: nothing**\n",
862
+ );
863
+ const f = tb(mutated);
864
+ assert.equal(f.filter((x) => x.reason.includes("malformed")).length, 1);
865
+ assert.ok(!f.some((x) => x.reason.includes("block empty")));
866
+ });
867
+
868
+ test("tests-block: none: with a command, none: with via:, two none: -> contradictory", () => {
869
+ for (const block of [
870
+ "**Tests:**\n- none: docs\n- `node --test extensions/lib/fixture-task2.test.ts`\n",
871
+ "**Tests:**\n- none: docs\n- via: `naming()`\n",
872
+ "**Tests:**\n- none: docs\n- none: config\n",
873
+ ]) {
874
+ const mutated = VALID_PLAN.replace(T2_TESTS, block);
875
+ assert.ok(tb(mutated).some((x) => x.reason.includes("contradictory")), block);
876
+ }
877
+ });
878
+
879
+ test("tests-block: none: while a Test: path is declared -> unused Test: path", () => {
880
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- none: docs\n");
881
+ assert.ok(tb(mutated).some((x) => x.reason.includes("unused `Test:` path")));
882
+ });
883
+
884
+ test("tests-block: none: docs, config with no Test: path is valid", () => {
885
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, "").replace(T2_TESTS, "**Tests:**\n- none: docs, config\n");
886
+ assert.deepEqual(tb(mutated), []);
887
+ });
888
+
889
+ test("tests-block: via: with commands passes", () => {
890
+ const mutated = VALID_PLAN.replace(T2_TESTS, T2_TESTS + "- via: `naming()` - the seam\n");
891
+ assert.deepEqual(tb(mutated), []);
892
+ });
893
+
894
+ test("tests-block: unknown Test: path unless it exists or another task Create:s it", () => {
895
+ const fs: FsPort = { exists: () => false, glob: () => [] };
896
+ assert.ok(tb(VALID_PLAN, SPEC_TEXT, fs).some((x) => x.reason.includes("unknown `Test:` path")));
897
+ const created = VALID_PLAN.replace(
898
+ "- Create: extensions/lib/fixture-task1.ts",
899
+ "- Create: extensions/lib/fixture-task1.ts\n- Create: extensions/lib/fixture-task1.test.ts\n- Create: extensions/lib/fixture-task2.test.ts\n- Create: extensions/lib/fixture-task3.test.ts",
900
+ );
901
+ assert.deepEqual(tb(created, SPEC_TEXT, fs), []);
902
+ });
903
+
904
+ test("tests-block: segment without a Test: token is not anchored", () => {
905
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts && echo done`\n");
906
+ assert.ok(tb(mutated).some((x) => x.reason.includes("not anchored")));
907
+ });
908
+
909
+ test("tests-block: pipe segment equal to a header segment; tee log has no anchor", () => {
910
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** npm test")
911
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
912
+ .replace(T2_TESTS, "**Tests:**\n- `npm test | tee log && node --test x.test.ts`\n");
913
+ const f = tb(mutated);
914
+ assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.reason.includes("npm test")));
915
+ assert.ok(f.some((x) => x.reason.includes("not anchored") && x.reason.includes("tee log")));
916
+ });
917
+
918
+ test("tests-block: broadening selectors tests/ and tests/*.py fail; Test: value tests/ never anchors", () => {
919
+ const dir = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts tests/`\n");
920
+ assert.ok(tb(dir).some((x) => x.reason.includes("broadening")));
921
+ const glob = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts tests/*.py`\n");
922
+ assert.ok(tb(glob).some((x) => x.reason.includes("broadening")));
923
+ const dirAnchor = VALID_PLAN.replace(T2_FILES_TEST, "- Test: tests/\n").replace(
924
+ T2_TESTS,
925
+ "**Tests:**\n- `node --test tests/`\n",
926
+ );
927
+ assert.ok(tb(dirAnchor).some((x) => x.reason.includes("not anchored")));
928
+ });
929
+
930
+ test("tests-block: unsupported shell (cd, sh -c, bash -c, eval, $( )", () => {
931
+ for (const cmd of [
932
+ "cd pkg && pytest tests/a.py",
933
+ "sh -c 'node --test extensions/lib/fixture-task2.test.ts'",
934
+ "bash -c 'node --test extensions/lib/fixture-task2.test.ts'",
935
+ "eval node --test extensions/lib/fixture-task2.test.ts",
936
+ "node --test $(echo extensions/lib/fixture-task2.test.ts)",
937
+ ]) {
938
+ const mutated = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `" + cmd + "`\n");
939
+ assert.ok(tb(mutated).some((x) => x.reason.includes("unsupported shell")), cmd);
940
+ }
941
+ });
942
+
943
+ test("tests-block: segment equal to a header segment (bare, &&, comma-listed, trailing prose, a && b)", () => {
944
+ for (const header of [
945
+ "**Verification:** npm test",
946
+ "**Verification:** `npm test`",
947
+ "**Verification:** npm test && npm run lint",
948
+ "**Verification:** `npm test`, `npm run lint`",
949
+ "**Verification:** `npm test` (bundles lint)",
950
+ ]) {
951
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header).replace(
952
+ T2_TESTS,
953
+ "**Tests:**\n- `npm run lint`\n- `npm test`\n",
954
+ );
955
+ const f = tb(mutated);
956
+ assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.text === "- `npm test`"), header);
957
+ if (header.includes("npm run lint")) assert.ok(f.some((x) => x.reason.includes("full-suite command") && x.text.includes("lint")), header);
958
+ }
959
+ const ab = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** a && b").replace(T2_TESTS, "**Tests:**\n- `b`\n");
960
+ assert.ok(tb(ab).some((x) => x.reason.includes("full-suite command")));
961
+ });
962
+
963
+ test("tests-block: npm test -- x.test.ts passes under header npm test (bare and backticked); the header segment npm test alone fails", () => {
964
+ for (const header of ["**Verification:** npm test", "**Verification:** `npm test`"]) {
965
+ const ok = VALID_PLAN.replace("**Verification:** npm run fixture-verify", header)
966
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
967
+ .replace(T2_TESTS, "**Tests:**\n- `npm test -- x.test.ts`\n");
968
+ assert.deepEqual(tb(ok, SPEC_TEXT, alwaysTruePort()), [], header);
969
+ }
970
+ });
971
+
972
+ test("tests-block: header segment that is itself scoped is still rejected as a bullet", () => {
973
+ const mutated = VALID_PLAN.replace("**Verification:** npm run fixture-verify", "**Verification:** node --test x.test.ts && npm run lint")
974
+ .replace(T2_FILES_TEST, "- Test: x.test.ts\n")
975
+ .replace(T2_TESTS, "**Tests:**\n- `node --test x.test.ts`\n");
976
+ assert.ok(tb(mutated).some((x) => x.reason.includes("full-suite command")));
977
+ });
978
+
979
+ test("tests-block: passing forms - multi-path, ::filter -v, -k name, -- passthrough, line-suffixed Test:", () => {
980
+ const multi = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `node --test extensions/lib/fixture-task2.test.ts extensions/lib/fixture-task1.test.ts`\n")
981
+ .replace(T2_FILES_TEST, T2_FILES_TEST + "- Test: extensions/lib/fixture-task1.test.ts\n");
982
+ assert.deepEqual(tb(multi), []);
983
+ const filt = VALID_PLAN.replace(T2_FILES_TEST, "- Test: tests/a.py\n").replace(T2_TESTS, "**Tests:**\n- `pytest tests/a.py::test_x -v -k name --filter x`\n");
984
+ assert.deepEqual(tb(filt), []);
985
+ const passthrough = VALID_PLAN.replace(T2_TESTS, "**Tests:**\n- `npm run fixture-verify -- extensions/lib/fixture-task2.test.ts`\n");
986
+ assert.deepEqual(tb(passthrough), []);
987
+ const suffixed = VALID_PLAN.replace(T2_FILES_TEST, "- Test: extensions/lib/fixture-task2.test.ts:10-20\n");
988
+ assert.deepEqual(tb(suffixed), []);
989
+ });
990
+
991
+ test("tests-block: no path normalization - ./x and x differ, as in Files:", () => {
992
+ const mutated = VALID_PLAN.replace(T2_FILES_TEST, "- Test: ./extensions/lib/fixture-task2.test.ts\n");
993
+ assert.ok(tb(mutated).some((x) => x.reason.includes("not anchored")));
994
+ });
995
+
996
+ test("tests-block: fenced **Tests:** lines are ignored", () => {
997
+ const mutated = VALID_PLAN.replace(
998
+ "This task handles naming details.",
999
+ "This task handles naming details.\n\n```markdown\n**Tests:**\n- none: docs\n```",
1000
+ );
1001
+ assert.deepEqual(tb(mutated), []);
1002
+ });
@@ -36,6 +36,12 @@ interface FileEntry {
36
36
  path: string;
37
37
  }
38
38
 
39
+ interface TestsBullet {
40
+ line: number;
41
+ text: string;
42
+ value: string;
43
+ }
44
+
39
45
  interface Task {
40
46
  number: number;
41
47
  line: number;
@@ -48,6 +54,12 @@ interface Task {
48
54
  specAnchorLine: number | undefined;
49
55
  anchors: Anchor[];
50
56
  anchorParseError: boolean;
57
+ testsHeadingLine: number | undefined;
58
+ testsMisplacedLines: number[];
59
+ tests: TestsBullet[];
60
+ testsVia: TestsBullet[];
61
+ testsNone: TestsBullet[];
62
+ testsMalformed: TestsBullet[];
51
63
  }
52
64
 
53
65
  interface Wave {
@@ -235,6 +247,50 @@ function parsePlan(planText: string): ParsedPlan {
235
247
  }
236
248
  }
237
249
 
250
+ let testsHeadingLine: number | undefined;
251
+ const testsMisplacedLines: number[] = [];
252
+ const tests: TestsBullet[] = [];
253
+ const testsVia: TestsBullet[] = [];
254
+ const testsNone: TestsBullet[] = [];
255
+ const testsMalformed: TestsBullet[] = [];
256
+
257
+ let headingIdx = -1;
258
+ if (filesLine !== undefined) {
259
+ let lastEntry = filesLine - 1;
260
+ let p = filesLine;
261
+ while (p <= bodyEndIdx) {
262
+ const l = lines[p];
263
+ if (l.trim() === "") { p++; continue; }
264
+ if (/^- \w+: /.test(l)) { lastEntry = p; p++; continue; }
265
+ break;
266
+ }
267
+ let q = lastEntry + 1;
268
+ while (q <= bodyEndIdx && lines[q].trim() === "") q++;
269
+ if (q <= bodyEndIdx && !mask[q] && lines[q] === "**Tests:**") {
270
+ headingIdx = q;
271
+ testsHeadingLine = q + 1;
272
+ }
273
+ }
274
+ for (let k = i; k <= bodyEndIdx; k++) {
275
+ if (mask[k] || k === headingIdx) continue;
276
+ if (/^\*\*Tests:\*\*/.test(lines[k])) testsMisplacedLines.push(k + 1);
277
+ }
278
+ if (headingIdx !== -1) {
279
+ for (let p = headingIdx + 1; p <= bodyEndIdx; p++) {
280
+ const l = lines[p];
281
+ if (l.trim() === "") continue;
282
+ if (mask[p] || !l.startsWith("- ") || l.startsWith("- [ ]")) break;
283
+ const cmd = /^- `([^`]+)`$/.exec(l);
284
+ const via = /^- via: (\S.*)$/.exec(l);
285
+ const none = /^- none: (\S.*)$/.exec(l);
286
+ const bullet = { line: p + 1, text: l, value: (cmd ?? via ?? none)?.[1] ?? "" };
287
+ if (cmd) tests.push(bullet);
288
+ else if (via) testsVia.push(bullet);
289
+ else if (none) testsNone.push(bullet);
290
+ else testsMalformed.push(bullet);
291
+ }
292
+ }
293
+
238
294
  tasks.push({
239
295
  number: Number(m[1]),
240
296
  line: i + 1,
@@ -247,6 +303,12 @@ function parsePlan(planText: string): ParsedPlan {
247
303
  specAnchorLine,
248
304
  anchors,
249
305
  anchorParseError,
306
+ testsHeadingLine,
307
+ testsMisplacedLines,
308
+ tests,
309
+ testsVia,
310
+ testsNone,
311
+ testsMalformed,
250
312
  });
251
313
  }
252
314
 
@@ -343,6 +405,99 @@ function isGlob(p: string): boolean {
343
405
  return /[*?{[\]]/.test(p);
344
406
  }
345
407
 
408
+ function norm(s: string): string {
409
+ return s.replaceAll("`", "").replace(/\s+/g, " ").trim();
410
+ }
411
+
412
+ function backtickSpans(s: string): string[] {
413
+ const out: string[] = [];
414
+ const re = /`([^`]+)`/g;
415
+ let m: RegExpExecArray | null;
416
+ while ((m = re.exec(s))) out.push(m[1]);
417
+ return out;
418
+ }
419
+
420
+ function commandSegments(cmd: string): string[] {
421
+ return norm(cmd).split(/\s*(?:&&|\|\||;|\|)\s*/).map((s) => s.trim()).filter(Boolean);
422
+ }
423
+
424
+ function headerSegments(parsed: ParsedPlan): string[] {
425
+ const value = parsed.header.verificationText ?? "";
426
+ const spans = backtickSpans(value);
427
+ const parts = spans.length > 0 ? spans : [value];
428
+ return parts.flatMap((p) => norm(p).split(/\s*(?:&&|\|\||;|,)\s*/)).map((s) => s.trim()).filter(Boolean);
429
+ }
430
+
431
+ const RUN_RE = /^\s*(- \[ \] )?Run:\s*(.*)$/;
432
+
433
+ function runPayloadSegments(line: string): string[] | undefined {
434
+ const m = RUN_RE.exec(line);
435
+ if (!m) return undefined;
436
+ const spans = backtickSpans(m[2]);
437
+ return (spans.length > 0 ? spans : [m[2]]).flatMap(commandSegments);
438
+ }
439
+
440
+ const UNSUPPORTED_SHELL = ["cd ", "sh -c", "bash -c", "eval ", "$("];
441
+
442
+ function isBroadening(token: string): boolean {
443
+ return /[*?[]/.test(token) || token.endsWith("/");
444
+ }
445
+
446
+ function checkTestsBlock(parsed: ParsedPlan, fs: FsPort): PlanCheckFinding[] {
447
+ const findings: PlanCheckFinding[] = [];
448
+ const push = (task: Task, line: number, text: string, reason: string) =>
449
+ findings.push({ check: "tests-block", line, text, reason: `Task ${task.number}: ${reason}` });
450
+ const createPaths = new Set(
451
+ parsed.tasks.flatMap((t) => t.files.filter((f) => f.kind === "create").map((f) => stripLineSuffix(f.path))),
452
+ );
453
+ const header = headerSegments(parsed);
454
+
455
+ for (const task of parsed.tasks) {
456
+ const testEntries = task.files.filter((f) => f.kind === "test");
457
+ const testPaths = testEntries.map((f) => stripLineSuffix(f.path));
458
+
459
+ for (const f of testEntries) {
460
+ const p = stripLineSuffix(f.path);
461
+ if (!createPaths.has(p) && !fs.exists(p)) push(task, f.line, f.text, `unknown \`Test:\` path "${p}" (neither exists nor is a Create: path of any task)`);
462
+ }
463
+
464
+ if (task.testsMisplacedLines.length > 0) {
465
+ for (const ln of task.testsMisplacedLines) push(task, ln, parsed.lines[ln - 1], "misplaced block: `**Tests:**` must be the bare line directly after the Files: entries");
466
+ } else if (task.testsHeadingLine === undefined) {
467
+ push(task, task.line, task.text, "block missing: no `**Tests:**` directly after the Files: entries");
468
+ }
469
+ if (task.testsHeadingLine === undefined) continue;
470
+
471
+ const headingText = parsed.lines[task.testsHeadingLine - 1];
472
+ if (task.tests.length === 0 && task.testsNone.length === 0) push(task, task.testsHeadingLine, headingText, "block empty: no command bullet and no `none:`");
473
+ for (const b of task.testsMalformed) push(task, b.line, b.text, "malformed bullet: expected `- \\`command\\``, `- via: <seam>`, or `- none: <category>`");
474
+ if (task.testsNone.length > 1 || (task.testsNone.length > 0 && (task.tests.length > 0 || task.testsVia.length > 0))) {
475
+ push(task, task.testsNone[0].line, task.testsNone[0].text, "contradictory block: `none:` with a command or `via:`, or more than one `none:`");
476
+ }
477
+ if (task.testsNone.length > 0) {
478
+ for (const f of testEntries) push(task, f.line, f.text, "unused `Test:` path: task declares `none:`");
479
+ }
480
+
481
+ const anchors = testPaths.filter((p) => !isBroadening(p));
482
+ for (const b of task.tests) {
483
+ if (UNSUPPORTED_SHELL.some((s) => b.value.includes(s))) {
484
+ push(task, b.line, b.text, "unsupported shell: `cd `, `sh -c`, `bash -c`, `eval `, `$(` are not allowed");
485
+ continue;
486
+ }
487
+ for (const seg of commandSegments(b.value)) {
488
+ const tokens = seg.split(" ");
489
+ const isAnchor = (t: string) => anchors.some((p) => t === p || t.startsWith(p + "::") || t.startsWith(p + "#") || t.startsWith(p + ":"));
490
+ if (!tokens.some(isAnchor)) push(task, b.line, b.text, `segment not anchored: "${seg}" names no Test: path of this task`);
491
+ for (const t of tokens) {
492
+ if (!isAnchor(t) && isBroadening(t)) push(task, b.line, b.text, `broadening selector "${t}" in "${seg}"`);
493
+ }
494
+ if (header.includes(seg)) push(task, b.line, b.text, `full-suite command in task: "${seg}" equals a header **Verification:** segment`);
495
+ }
496
+ }
497
+ }
498
+ return findings;
499
+ }
500
+
346
501
  function taskBodyText(task: Task, lines: string[]): string {
347
502
  return lines.slice(task.bodyStartLine - 1, task.bodyEndLine).join("\n");
348
503
  }
@@ -687,10 +842,14 @@ function checkPlaceholderScan(parsed: ParsedPlan, requiredLiterals: Map<number,
687
842
  return findings;
688
843
  }
689
844
 
690
- function fileEntries(task: Task): { path: string; kind: "literal" | "glob" }[] {
845
+ function fileEntries(task: Task): { path: string; kind: "literal" | "glob"; role: "test" | "write" }[] {
691
846
  return task.files.map((f) => {
692
847
  const p = stripLineSuffix(f.path);
693
- return { path: p, kind: (isGlob(p) ? "glob" : "literal") as "literal" | "glob" };
848
+ return {
849
+ path: p,
850
+ kind: (isGlob(p) ? "glob" : "literal") as "literal" | "glob",
851
+ role: f.kind === "test" ? "test" : "write",
852
+ };
694
853
  });
695
854
  }
696
855
 
@@ -719,6 +878,7 @@ function checkWaveFileDisjointness(parsed: ParsedPlan, fs: FsPort): PlanCheckFin
719
878
  const b = fileEntries(tasks[j]);
720
879
  for (const ea of a) {
721
880
  for (const eb of b) {
881
+ if (ea.role === "test" && eb.role === "test") continue;
722
882
  let overlap = false;
723
883
  let errFinding: PlanCheckFinding | undefined;
724
884
  if (ea.kind === "literal" && eb.kind === "literal") {
@@ -798,6 +958,15 @@ function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
798
958
  const entrypoint = parsed.header.verificationText.trim();
799
959
  if (!entrypoint) return findings;
800
960
 
961
+ const header = headerSegments(parsed);
962
+ const executable = new Set<number>();
963
+ for (const task of parsed.tasks) {
964
+ for (const bullet of [...task.tests, ...task.testsVia, ...task.testsNone, ...task.testsMalformed]) {
965
+ executable.add(bullet.line);
966
+ }
967
+ }
968
+ const mask = fenceMask(parsed.lines);
969
+
801
970
  const waveBoundaryRe = /^##\s/;
802
971
  const inScope = new Set<number>();
803
972
  for (const wave of parsed.waves) {
@@ -814,6 +983,21 @@ function checkHeaderEntrypoint(parsed: ParsedPlan): PlanCheckFinding[] {
814
983
 
815
984
  for (const ln of [...inScope].sort((a, b) => a - b)) {
816
985
  const line = parsed.lines[ln - 1];
986
+ if (executable.has(ln)) continue;
987
+ const segments = runPayloadSegments(line);
988
+ if (segments) {
989
+ if (mask[ln - 1]) continue;
990
+ const hit = segments.find((segment) => header.includes(segment));
991
+ if (hit !== undefined) {
992
+ findings.push({
993
+ check: "header-entrypoint",
994
+ line: ln,
995
+ text: line,
996
+ reason: `Run: segment "${hit}" equals a header **Verification:** segment (full suite belongs to the verify phase)`,
997
+ });
998
+ }
999
+ continue;
1000
+ }
817
1001
  if (line.includes(entrypoint)) {
818
1002
  findings.push({
819
1003
  check: "header-entrypoint",
@@ -870,6 +1054,7 @@ export function checkPlan(planText: string, specText: string, fs: FsPort): PlanC
870
1054
  findings.push(...checkQuoteIntegrity(parsed, specLines));
871
1055
  findings.push(...checkAnchorResolution(parsed, specLines));
872
1056
  findings.push(...checkPathsExist(parsed, fs));
1057
+ findings.push(...checkTestsBlock(parsed, fs));
873
1058
  const requiredLiterals = computeRequiredLiteralsPerTask(parsed, specLines);
874
1059
  findings.push(...checkPlaceholderScan(parsed, requiredLiterals));
875
1060
  findings.push(...checkWaveFileDisjointness(parsed, fs));
@@ -707,7 +707,12 @@ const FIXTURE_PLAN = `# Fixture Plan
707
707
 
708
708
  **Files:**
709
709
  - Create: lib/task1.ts
710
+ - Create: lib/task1.test.ts
710
711
  - Modify: file-a.ts
712
+ - Test: lib/task1.test.ts
713
+
714
+ **Tests:**
715
+ - \`node --test lib/task1.test.ts\`
711
716
 
712
717
  This task implements helperFn() for parsing.
713
718
 
@@ -717,7 +722,12 @@ This task implements helperFn() for parsing.
717
722
 
718
723
  **Files:**
719
724
  - Create: lib/task2.ts
725
+ - Create: lib/task2.test.ts
720
726
  - Modify: file-b.ts
727
+ - Test: lib/task2.test.ts
728
+
729
+ **Tests:**
730
+ - \`node --test lib/task2.test.ts\`
721
731
 
722
732
  This task handles naming details.
723
733
 
@@ -705,7 +705,7 @@ export default function (pi: ExtensionAPI) {
705
705
  name: "plan_check",
706
706
  label: "Plan Check",
707
707
  description:
708
- "Deterministically verify an implementation plan against its spec (9 mechanical checks); " +
708
+ "Deterministically verify an implementation plan against its spec (mechanical checks); " +
709
709
  "a pass stamps the plan for implement-start.",
710
710
  parameters: PlanCheckParams,
711
711
  async execute(_toolCallId, params, _signal, _onUpdate, ctx) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.4.0",
3
+ "version": "5.5.1",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -24,7 +24,9 @@ this skill alters re-gates.
24
24
  Preferred: `linearis` on PATH and authenticated (`linearis auth status`). Token
25
25
  resolution order: `--api-token`, `LINEAR_API_TOKEN`, `~/.linearis/token`. This is a
26
26
  preference, not a precondition - a missing or unauthenticated CLI degrades Linear
27
- functionality and is reported, never blocks the run.
27
+ functionality and is reported, never blocks the run. The token file is encrypted
28
+ storage: never use its contents as an HTTP credential or infer the credential type
29
+ from its `v1:` storage-format prefix. Resolve it through linearis instead.
28
30
 
29
31
  > **No `linearis` installed?** If `command -v linearis` fails, fall back to a
30
32
  > **Linear MCP server** when the harness has one configured - its tools cover the
@@ -192,7 +194,7 @@ Safety rules, in addition to the write gate above:
192
194
  | Symptom | Cause | Fix |
193
195
  |---|---|---|
194
196
  | 401 | Not authenticated / expired token | `linearis auth status`; re-auth - unless the download row below applies. |
195
- | 401 on `files download` while `issues read` works | linearis 2026.7.0 and 2026.8.0 prepend `Bearer ` to personal API keys on file downloads ([linearis-oss/linearis#300](https://github.com/linearis-oss/linearis/issues/300)) | Not an auth problem - do not re-auth. Fetch the URL with the bare key, or use a version without the bug once one ships. |
197
+ | 401 on `files download` while `issues read` works | linearis 2026.7.0 and 2026.8.0 prepend `Bearer ` to personal API keys on file downloads ([linearis-oss/linearis#300](https://github.com/linearis-oss/linearis/issues/300)) | Not an auth problem - do not re-auth. Use the recovery below, or use a version without the bug once one ships. |
196
198
  | Issue not found | Wrong workspace, or issue archived | Confirm workspace; check archived state. |
197
199
  | Status not found | Status name doesn't match the team's workflow states | List the team's states before setting one. |
198
200
  | Missing `--team` error on create | `--team` is required | Supply `--team <default team>`. |
@@ -203,6 +205,45 @@ Safety rules, in addition to the write gate above:
203
205
  | Read is slow | Big ticket with many comments/attachments | Drop `--with-*` flags not needed. |
204
206
  | Parser-shape failure on a documented invocation: unknown command/option, unexpected argument | Section 3's snapshot may have drifted from the installed CLI | Re-read that subcommand's `--help`; report the row stale **only if** help actually contradicts it, then follow help |
205
207
 
208
+ For that download-only 401, this 2026.7.0/2026.8.0 workaround calls linearis's
209
+ version-specific internal `getApiToken` API. Supply the fresh `uploads.linear.app` URL
210
+ from `issues read --with-attachments` and an output path. It rejects other hosts and
211
+ redirects, sends the resolved key without `Bearer`, writes only a non-empty response,
212
+ and never prints or stores the key separately:
213
+
214
+ <!-- linear-download-recovery:start -->
215
+ ```bash
216
+ download_linear_asset() {
217
+ LINEARIS_BIN="${LINEARIS_BIN:-$(command -v linearis)}" node --input-type=module - "$1" "$2" <<'NODE'
218
+ import { realpathSync, writeFileSync } from "node:fs";
219
+ import { dirname, join } from "node:path";
220
+ import { pathToFileURL } from "node:url";
221
+
222
+ const [urlText, output] = process.argv.slice(2);
223
+ const url = new URL(urlText);
224
+ if (url.protocol !== "https:" || url.hostname !== "uploads.linear.app")
225
+ throw new Error("refusing to send a credential outside https://uploads.linear.app");
226
+ const packageRoot = dirname(dirname(realpathSync(process.env.LINEARIS_BIN)));
227
+ const { getApiToken } = await import(pathToFileURL(join(packageRoot, "dist/common/auth.js")));
228
+ const response = await fetch(url, {
229
+ headers: { Authorization: getApiToken({}) },
230
+ redirect: "error",
231
+ });
232
+ if (!response.ok) throw new Error(`download failed: HTTP ${response.status}`);
233
+ const bytes = new Uint8Array(await response.arrayBuffer());
234
+ if (bytes.byteLength === 0) throw new Error("download failed: empty response");
235
+ writeFileSync(output, bytes);
236
+ console.log(`downloaded ${bytes.byteLength} bytes to ${output}`);
237
+ NODE
238
+ }
239
+ download_linear_asset 'https://uploads.linear.app/...' '/tmp/attachment'
240
+ ```
241
+ <!-- linear-download-recovery:end -->
242
+
243
+ Do not declare the attachment inaccessible until this recovery used the resolved
244
+ credential; reading `~/.linearis/token` directly does not count. Afterward, confirm the
245
+ reported byte count and inspect the file type before consuming or extracting it.
246
+
206
247
  The last row's trigger is deliberately narrow. Data, auth, status-name, and root-thread
207
248
  validation errors have their own rows above and are **not** drift - routing them to "the
208
249
  skill is stale" would misdiagnose ordinary failures. This row is the reactive path for
@@ -55,11 +55,13 @@ Before the first task, enter the implement phase: `phase_tracker({ action: "star
55
55
 
56
56
  For each task in `plan_tracker`:
57
57
 
58
- 1. **Start, then dispatch implementer.** Mark the task's existing `plan_tracker` index `in_progress` before dispatch. Pass the full task text + scene-setting context + the task's plan-declared test commands as `SCOPED_TEST_COMMANDS` (or `none`). Don't make the subagent re-read the plan.
58
+ **`SCOPED_TEST_COMMANDS`** = the task's `**Tests:**` command bullets, backticks stripped, one per line; `- none:` -> `none`; a wave = the union of its tasks' commands. `via:` bullets and `- Test:` paths are contract, not commands: they ride with the task text, never in `SCOPED_TEST_COMMANDS`.
59
+
60
+ 1. **Start, then dispatch implementer.** Mark the task's existing `plan_tracker` index `in_progress` before dispatch. Pass the full task text + scene-setting context + its `SCOPED_TEST_COMMANDS` (definition above) and its `**Tests:**` block verbatim. Don't make the subagent re-read the plan.
59
61
  2. **Handle implementer status** (see below).
60
- 3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the change satisfies the anchored spec — nothing missing, nothing extra.
62
+ 3. **Dispatch spec reviewer.** Pass the task text, the patch diff, the absolute spec path, and the task's `**Spec:**` anchors, plus the task's `**Tests:**` block and `- Test:` paths as contract (never as commands - SR never executes) — SR reads the anchored ranges from the spec file itself; never inline spec excerpts. The spec wins every dispute; the authority hierarchy and finding labels live in `./spec-reviewer-prompt.md`. Anchor-less tasks (no `**Spec:**` line): task text alone is the contract. Verify the change satisfies the anchored spec — nothing missing, nothing extra.
61
63
  4. If spec reviewer finds gaps → re-dispatch implementer to fix → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
62
- 5. **Dispatch code-quality reviewer.** Only after spec is ✅. Skip for doc-only tasks (every file in the task's `Files:` block documentation-only) — SR-only, same exemption as doc-only waves. Pass `SCOPED_TEST_COMMANDS` = the task's plan-declared commands.
64
+ 5. **Dispatch code-quality reviewer.** Only after spec is ✅. Skip for doc-only tasks (every file in the task's `Files:` block documentation-only) — SR-only, same exemption as doc-only waves. Pass the task's `SCOPED_TEST_COMMANDS`.
63
65
  6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
64
66
  7. After its existing reviews accept the work (and its required commit point), mark that same task index `complete` in `plan_tracker`. A subagent exit or green tests alone are not acceptance.
65
67
 
@@ -73,11 +75,11 @@ One rule governs both review loops - spec-compliance and code-quality - in seque
73
75
 
74
76
  **Re-review dispatch rule:** every re-review task includes the complete prior review report verbatim under the marker `## Previous review report (re-review trigger)`, plus the trajectory block from the reviewer's prompt template. The marker's presence is what obligates the reviewer to emit the `TRAJECTORY:` line. You never select, summarize, or diff findings yourself - pattern-match the sentinel line only.
75
77
 
76
- **Fix fan-out.** When the triggering review's `Parallel-safe:` line certifies a `disjoint` group of ≥ 2 findings, dispatch that fix round per `dispatching-parallel-agents` "Fix fan-out"; the fan-out counts as **one** fix against this budget, its scoped test gate is the consuming task/wave's plan-declared commands, and one re-review of the integrated delta follows.
78
+ **Fix fan-out.** When the triggering review's `Parallel-safe:` line certifies a `disjoint` group of ≥ 2 findings, dispatch that fix round per `dispatching-parallel-agents` "Fix fan-out"; the fan-out counts as **one** fix against this budget, its scoped test gate is the consuming task/wave's `SCOPED_TEST_COMMANDS`, and one re-review of the integrated delta follows.
77
79
 
78
- **Behaviour-change reroute.** When the triggering CR report carries `Behaviour-change: yes`, the fix round's re-review is SR first, then CR. The SR dispatch reviews the fix diff against the task's spec anchors as a first review (no `## Previous review report` marker, so no `TRAJECTORY:` line; SR carries no test commands), but its ordinal continues the task's SR-loop count - a task whose SR loop ended at review 2 gets review 3 here, and an issue-bearing rerouted SR after review 4 escalates. Issues follow the normal sequence: fix, then SR re-review pasting this SR's report. CR round numbering is unchanged. `Behaviour-change: no` re-reviews with CR only. A missing or malformed `Behaviour-change:` line is re-asked once like `Parallel-safe:` (see `dispatching-parallel-agents` "Structural probe"); still missing -> route through SR, never default to `no`. In wave mode the SR re-review targets the task(s) whose files the fix touched.
80
+ **Behaviour-change reroute.** When the triggering CR report carries `Behaviour-change: yes`, the fix round's re-review is SR first, then CR. The SR dispatch reviews the fix diff against the task's spec anchors as a first review (no `## Previous review report` marker, so no `TRAJECTORY:` line; SR carries no test commands but does carry the task's `**Tests:**` block and `- Test:` paths), but its ordinal continues the task's SR-loop count - a task whose SR loop ended at review 2 gets review 3 here, and an issue-bearing rerouted SR after review 4 escalates. Issues follow the normal sequence: fix, then SR re-review pasting this SR's report. CR round numbering is unchanged. `Behaviour-change: no` re-reviews with CR only. A missing or malformed `Behaviour-change:` line is re-asked once like `Parallel-safe:` (see `dispatching-parallel-agents` "Structural probe"); still missing -> route through SR, never default to `no`. In wave mode the SR re-review targets the task(s) whose files the fix touched.
79
81
 
80
- Every fix re-dispatch (implementer) and code-review re-review carries the consuming task/wave's `SCOPED_TEST_COMMANDS`; spec-reviewer re-reviews carry none - SR never executes.
82
+ Every fix re-dispatch (implementer) and code-review re-review carries the consuming task/wave's `SCOPED_TEST_COMMANDS`; spec-reviewer re-reviews carry the `**Tests:**` block and `- Test:` paths, no commands - SR never executes.
81
83
 
82
84
  **The sequence.** Each review that finds issues is a decision point: read the `TRAJECTORY:` line before dispatching anything (review 1 has no line - on issues, dispatch fix 1). Any clean review ends the loop.
83
85
 
@@ -143,7 +145,7 @@ subagent({ agent: "implementer", async: false, task: "<task text + context + SCO
143
145
  subagent({ agent: "implementer", model: "<implModel>", context: "fresh", async: false, task: "<the just-dispatched fix payload + prior review report verbatim>" })
144
146
 
145
147
  // spec compliance
146
- subagent({ agent: "spec-reviewer", async: false, task: "<task text + patch diff + absolute spec path + task's Spec: anchors>" })
148
+ subagent({ agent: "spec-reviewer", async: false, task: "<task text + patch diff + absolute spec path + task's Spec: anchors + Tests: block + Test: paths>" })
147
149
 
148
150
  // code quality
149
151
  subagent({ agent: "code-reviewer", async: false, task: "<diff range + SCOPED_TEST_COMMANDS (task commands; wave: union; whole-diff: none) + ask: production-ready?>" })
@@ -177,12 +179,12 @@ Auto-selected at handoff by `writing-plans` (any wave with ≥2 tasks) when the
177
179
 
178
180
  **Per-wave loop:**
179
181
 
180
- 1. **Independence check.** Parse the wave's tasks' `Files:` blocks; assert pairwise-disjoint (mechanical). Runtime-resource disjointness (DB/schema, port, fixture, external service, shared temp path) is not machine-checkable here — trust the plan's wave grouping, which `writing-plans`' D5 contract guarantees. Either kind of overlap → the wave is mis-grouped; run those tasks as sequential single-task waves and note it.
182
+ 1. **Independence check.** Parse the wave's tasks' `Files:` blocks; assert pairwise-disjoint (mechanical, mirroring `wave-file-disjointness`: `Test`/`Test` on one path is not overlap, `Test` vs another task's `Create`/`Modify` is, `Modify`/`Modify` is). Runtime-resource disjointness (DB/schema, port, fixture, external service, shared temp path) is not machine-checkable here — trust the plan's wave grouping, which `writing-plans`' D5 contract guarantees. Either kind of overlap → the wave is mis-grouped; run those tasks as sequential single-task waves and note it.
181
183
  2. **Start, then fan out.** Mark every wave index `in_progress` before one parallel foreground dispatch (shape below): `implementer` per task, `context: "fresh"`, `worktree: true`. Each returns a status + a patch.
182
- 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review starts from the diff (its hunks carry `file:line`) and reads each touched file in full, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's declared test commands). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
184
+ 3. **Status + spec review per task.** Parse each `DONE`/`BLOCKED`/etc. (see [Implementer Status](#implementer-status)) **first**. Then **dispatch a `spec-reviewer` per accepted patch** (`DONE`, or a `DONE_WITH_CONCERNS` you proceeded with) in one parallel fan-out — `context: "fresh"`, `cwd: <this worktree>`, **no `worktree` flag** (read-only) — each passed its task text, the returned **patch diff**, the absolute spec path, and the task's `**Spec:**` anchors, plus its `**Tests:**` block and `- Test:` paths as contract — SR reads the anchored ranges itself (never inline excerpts; authority hierarchy in `./spec-reviewer-prompt.md`; anchor-less tasks are task-text-only). Review starts from the diff (its hunks carry `file:line`) and reads each touched file in full, and test execution is never the reviewer's job - in either mode (persona rule; the wave test gate in step 5 runs the wave's `SCOPED_TEST_COMMANDS`). Inline verdicts are fine at normal wave sizes; large waves use `output:` + `outputMode: "file-only"` to keep verdicts out of your context. **Re-dispatch by cause:** `BLOCKED`/`NEEDS_CONTEXT` per the [Implementer Status](#implementer-status) matrix; a **spec gap** re-dispatches the implementer (fresh, `worktree: true`) carrying the prior patch + the reviewer's findings, the new patch superseding the old at step 4. Loop until accepted + spec ✅, within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential.
183
185
  4. **Integrate.** `git apply` each task's patch sequentially onto HEAD. Apply fails = textual conflict → drop that task, finish the rest, re-run the dropped task sequentially on the updated HEAD.
184
- 5. **Test gate.** Run the union of the wave's tasks' declared test commands on the integrated tree — the full verification set is the verify phase's job, run once. Failure = semantic conflict or bug → re-run the offending task sequentially, else fix per [When a Subagent Fails](#when-a-subagent-fails).
185
- 6. **Quality review.** CR binds to the wave: exactly one **initial** code-review dispatch per code-touching wave, over the integrated wave diff - never per task within a wave, never batched across waves. Subsequent dispatches within the wave are re-reviews triggered only by findings, per Fix-Loop Rounds. Pass `SCOPED_TEST_COMMANDS` = the union of the wave's tasks' declared commands. Code-quality review on the integrated wave diff; loop fixes to ✅ within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential. Skip for doc-only waves (SR-only per the commit precondition below).
186
+ 5. **Test gate.** Run the wave's `SCOPED_TEST_COMMANDS` (union of its tasks' `Tests:` commands) on the integrated tree — the full verification set is the verify phase's job, run once. Failure = semantic conflict or bug → re-run the offending task sequentially, else fix per [When a Subagent Fails](#when-a-subagent-fails).
187
+ 6. **Quality review.** CR binds to the wave: exactly one **initial** code-review dispatch per code-touching wave, over the integrated wave diff - never per task within a wave, never batched across waves. Subsequent dispatches within the wave are re-reviews triggered only by findings, per Fix-Loop Rounds. Pass the wave's `SCOPED_TEST_COMMANDS`. Code-quality review on the integrated wave diff; loop fixes to ✅ within [Fix-Loop Rounds](#fix-loop-rounds), same as sequential. Skip for doc-only waves (SR-only per the commit precondition below).
186
188
  7. **Commit and complete the wave.** After the gate passes and the wave commits, mark all of its existing indices `complete`. Leaves a clean tree; the next wave's children branch from this commit and so see the integrated work.
187
189
 
188
190
  **Two-stage review is preserved:** spec review per task (pre-integration, dispatched `spec-reviewer` — not inline), quality review per wave (post-integration). A wave commit requires one spec-review verdict per accepted task, plus one code-review verdict on the integrated diff for waves that touch code. A doc-only wave (every task's `Files:` block documentation-only, per `writing-plans`' Wave Grouping) is SR-only — the CR gate does not apply.
@@ -243,11 +245,11 @@ For the fan-out + worktree + patch-integration + conflict mechanics, see `dispat
243
245
  - Writing code yourself instead of dispatching.
244
246
  - Pausing between tasks for anything other than `NEEDS_CONTEXT`, `BLOCKED`, a fix-loop escalation, a workflow warning, or a spec amendment.
245
247
  - Dispatching parallel implementers on overlapping files, on a shared mutable runtime resource, or without `worktree: true`.
246
- - Making a subagent read the plan, inlining spec excerpts to the spec reviewer, or dispatching without a `SCOPED_TEST_COMMANDS` value.
248
+ - Making a subagent read the plan, inlining spec excerpts to the spec reviewer, or dispatching with a `SCOPED_TEST_COMMANDS` value missing or not copied from the plan's `Tests:` bullets.
247
249
  - Dispatching `code-reviewer` before every in-scope spec-review verdict is ✅, or per task inside a wave.
248
250
  - Moving to the next task with either review still showing issues, or skipping the `Implementer Status` parse.
249
251
  - Dispatching fix 3 without a reviewer-emitted `CONVERGING` verdict, or continuing past `STAGNANT` instead of escalating.
250
- - Running the full verification entrypoint during the implement phase.
252
+ - Running the full verification entrypoint, or any test command outside `SCOPED_TEST_COMMANDS`, during the implement phase.
251
253
  - Dispatching a repair before reopening the plan-task indices that own its files, whole-diff CR before parent verification passes, or conformance before the CR result is accepted.
252
254
  - Polling, joining, or relaunching an unexpectedly asynchronous dispatch, or starting on main without explicit user consent.
253
255
 
@@ -14,7 +14,7 @@ Dispatch a subagent with the code-reviewer template:
14
14
  PLAN_OR_REQUIREMENTS: Task N from [plan-file]
15
15
  BASE_SHA: [commit before task]
16
16
  HEAD_SHA: [current commit]
17
- SCOPED_TEST_COMMANDS: [the consuming task's plan-declared commands; wave reviews: the union of the wave's tasks' declared commands; `none` for the whole-diff verify-phase review]
17
+ SCOPED_TEST_COMMANDS: [the consuming task's Tests: commands; wave reviews: the union of the wave's tasks' Tests: commands; `none` for the whole-diff verify-phase review]
18
18
  ```
19
19
 
20
20
  **In addition to standard code quality concerns, the reviewer should check:**
@@ -38,12 +38,17 @@ Dispatch a subagent with this prompt:
38
38
 
39
39
  Work from: [directory]
40
40
 
41
- SCOPED_TEST_COMMANDS: [the task's plan-declared test commands, verbatim | none]
41
+ SCOPED_TEST_COMMANDS: [the task's Tests: bullets, backticks stripped, verbatim | none]
42
42
 
43
43
  Run ONLY these commands for verification. Never run a repo-wide suite,
44
44
  linter, or type-checker. If the value is `none`, run nothing and say so
45
45
  in your report.
46
46
 
47
+ TEST_CONTRACT: [the task's **Tests:** block verbatim (commands, via:, none:) and its Files: Create:/Test: paths]
48
+
49
+ Tests call the via: seam directly - not a wrapper, not the internals behind it.
50
+ Create every Create: path at exactly that path.
51
+
47
52
  **While you work:** If you encounter something unexpected or unclear, **ask questions**.
48
53
  It's always OK to pause and clarify. Don't guess or make assumptions.
49
54
 
@@ -109,6 +114,9 @@ Dispatch a subagent with this prompt:
109
114
  - **Status:** `DONE` | `DONE_WITH_CONCERNS` | `BLOCKED` | `NEEDS_CONTEXT`
110
115
  - What you implemented (or what you attempted, if blocked)
111
116
  - What you tested and test results
117
+ - Test contract: one line per SCOPED_TEST_COMMANDS command - `met` (exit 0; quote the last output line) or `unmet <reason>`;
118
+ one line per via: - `met <test file:line calling it>` or `unmet <reason>`; `none` when the block is `none:`.
119
+ Any `unmet` -> `DONE_WITH_CONCERNS`.
112
120
  - Files changed
113
121
  - Self-review findings (if any)
114
122
  - Any issues or concerns
@@ -24,6 +24,7 @@ Dispatch a subagent with this prompt:
24
24
 
25
25
  Spec: [absolute spec path]
26
26
  Anchors: [the task's **Spec:** anchor list, e.g. § "Design" L34-L37 — or "omitted: anchor-less mechanical task"]
27
+ Task contract: [the task's **Tests:** block verbatim + its `Files:` paths]
27
28
 
28
29
  The spec is the sole authority — human-approved; the task never wins a dispute. Read the anchored ranges from the spec file yourself. Requirements in scope are ONLY the cited anchor ranges; do not extract, review, or flag the rest of the spec file.
29
30
 
@@ -35,6 +36,7 @@ Dispatch a subagent with this prompt:
35
36
  - **Finding grammar:** divergence findings use the existing F1..Fn finding grammar - a finding kind by prose label, not a new schema; the `Parallel-safe:` and `TRAJECTORY:` grammars are untouched.
36
37
  - **Condition match:** for every anchored clause that fixes a value, threshold, comparison, or trigger ("only when", "unless", "if", a literal), the clause row carries two indented sub-lines, before `touched-files:` where present: `spec-condition: <clause fragment quoted from the spec>` and `code-condition: <what the code checks, file:line>`. If the two differ, the clause is `PARTIAL` at most - regardless of passing tests. A plausible condition is not the specified condition.
37
38
  - **Plan/task code snippets:** implementation guidance, not review authority; a diff matching a snippet never proves compliance. For anchor-less tasks the task text's prose requirements remain your contract.
39
+ - **Task contract (Tests:/via:/Files:):** The task's `**Tests:**` block, `via:`, and its `Files:` paths supplement the anchored spec where it is silent; the anchored spec wins a conflict - report the divergence once, against the plan, never against code corrected to the spec. Findings: a `Create:` path absent from the diff or created elsewhere; a test that does not call the `via:` entry point; a `Tests:` block the diff contradicts. Existing files need no diff touch. You never run the commands.
38
40
 
39
41
  ## CRITICAL: Do Not Trust the Report
40
42
 
@@ -64,7 +66,7 @@ Dispatch a subagent with this prompt:
64
66
 
65
67
  ## Your Job
66
68
 
67
- <!-- clause decomposition / snippet non-authority / whole-file reads: keep in lockstep with agents/spec-reviewer.md — change them together or not at all -->
69
+ <!-- clause decomposition / snippet non-authority / whole-file reads / task-contract supplement: keep in lockstep with agents/spec-reviewer.md — change them together or not at all -->
68
70
 
69
71
  Decompose the binding contract - the anchored spec lines, or the task text when anchors are omitted - into atomic clauses, covering every requirement, acceptance criterion, and explicit non-goal. Each independently checkable statement is one clause; a sentence listing three requirements yields three clauses.
70
72
 
@@ -117,11 +117,11 @@ Don't add features, refactor other code, or "improve" beyond what the test requi
117
117
 
118
118
  Run the test. Confirm:
119
119
  - New test passes
120
- - The task's scoped commands pass (full-suite verification belongs to the verify phase)
120
+ - The task's `Tests:` commands pass (full-suite verification belongs to the verify phase)
121
121
  - Output is pristine (no errors, no warnings)
122
122
 
123
123
  **Test fails?** Fix code, not test.
124
- **Scoped commands fail?** Fix now — don't move on with broken tests.
124
+ **`Tests:` commands fail?** Fix now — don't move on with broken tests.
125
125
 
126
126
  ### REFACTOR — Clean Up
127
127
 
@@ -176,7 +176,7 @@ Before marking work complete:
176
176
  - [ ] Watched each test fail before implementing
177
177
  - [ ] Each test failed for expected reason (feature missing, not typo)
178
178
  - [ ] Wrote minimal code to pass each test
179
- - [ ] The task's scoped commands pass (full suite belongs to the verify phase)
179
+ - [ ] The task's `Tests:` commands pass (full suite belongs to the verify phase)
180
180
  - [ ] Output pristine (no errors, warnings)
181
181
  - [ ] Tests use real code (mocks only if unavoidable)
182
182
  - [ ] Edge cases and errors covered
@@ -126,14 +126,13 @@ If you can't list the files, the spec isn't ready: amend or redraw per brainstor
126
126
 
127
127
  Group tasks into **waves** so the executor can parallelize independent work (see `subagent-driven-development` Parallel-Wave Mode). A wave is a maximal set of tasks that (a) have no ordering dependency on each other, (b) own **pairwise-disjoint files**, and (c) contend on **no shared mutable runtime resource** (same DB/schema, port, fixture file, external service, shared temp path).
128
128
 
129
- - Tasks nest under `## Wave N — <label>` headers; `### Task N` headers sit inside a wave.
130
129
  - Group independent tasks into the same wave by default. A wave with one task is legal **only with a named-blocker justification**: a body line directly under the `## Wave N — <label>` header, `Solo: <reason>`, where the reason names the blocking task/wave, the contended runtime resource, or `lone remaining task` (reserved for the genuinely final unmatched task; doc-only trailing waves qualify). Category-only justifications ("dependency" with no named task) do not satisfy the rule.
131
130
  - A pure dependency chain yields one task per wave — no parallelism, which is correct; each such wave carries its `Solo:` line naming the prior-wave dependency.
132
131
  - Each wave after the first states its dependency on prior waves.
133
132
 
134
- **File-ownership contract.** The per-task `**Files:**` block *is* the ownership declaration — no new syntax. Rule: **within a wave, the union of every task's declared paths must be pairwise disjoint.** Globs are allowed for `Modify` when exact paths are unknown, but must not overlap another same-wave task's paths. A task that must touch another's file belongs in a later wave.
133
+ **File-ownership contract.** See [reference/plan-contract.md § Files](reference/plan-contract.md).
135
134
 
136
- **Test-command contract.** Every code-touching wave declares at least one scoped test command across its tasks' steps. A wave with zero test commands is legal only when every task's `Files:` block is documentation-only (the trailing doc-only wave below).
135
+ **Test contract.** Every task that creates or modifies code declares a `Test:` path and an executable `**Tests:**` command anchored to it (grammar: [reference/plan-contract.md § Tests](reference/plan-contract.md)); `- none: <category>` only when no tests apply. Anchoring is presence, not coverage. When the anchored spec names the thing under test, the task carries `- via:` naming it; a fixture path the spec names goes under `Create:`.
137
136
 
138
137
  **Runtime-resource disjointness.** File-disjoint is necessary but not sufficient: two tasks with disjoint files that both mutate the same DB, bind the same port, or share a fixture are **not** parallel-safe and must land in different waves. The executor auto-selects parallel for *every* multi-task wave, so this grouping is the sole parallel-safety guarantee — there is no selection-time judgment downstream. No new mandatory per-task syntax; when a shared runtime resource is the reason two file-disjoint tasks sit in different waves, record it in an inline note on the later wave.
139
138
 
@@ -192,7 +191,7 @@ Each step is **one action, 2-5 minutes**:
192
191
  ---
193
192
  ```
194
193
 
195
- The `**Verification:**` line is the **only** place the full verification entrypoint may appear — never in any task or wave step. The verify phase reads it from the plan instead of re-deriving it; execution runs scoped commands only.
194
+ The full verification entrypoint appears only on the `**Verification:**` line — see [reference/plan-contract.md § Header-only entrypoint](reference/plan-contract.md). The verify phase reads it from the plan; execution runs `Tests:` commands only.
196
195
 
197
196
  ## Task Structure
198
197
 
@@ -210,6 +209,10 @@ Each task uses `- [ ]` checkbox steps so execution tools (and humans) can track
210
209
  - Modify: `exact/path/to/existing.py:123-145`
211
210
  - Test: `tests/exact/path/to/test.py`
212
211
 
212
+ **Tests:**
213
+ - `uv run pytest tests/exact/path/to/test.py`
214
+ - via: `function()`
215
+
213
216
  - [ ] **Step 1: Write the failing test**
214
217
 
215
218
  ```python
@@ -250,28 +253,14 @@ Each task uses `- [ ]` checkbox steps so execution tools (and humans) can track
250
253
 
251
254
  Every code task carries this step (red -> green -> fmt/lint -> commit). Doc-only tasks omit it unless the project formats Markdown.
252
255
 
253
- **Anchor rules.** The task's `**Spec:**` line cites the plan header's spec path; multiple anchors sit comma-separated on one line (`§ "A" L10-L18, § "C" L40-L44`). Checks key on the `§` marker, so the header's path-only `**Spec:**` line is never matched. Anchors are captured against the gated spec at plan-writing time; a change to the approved spec follows brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec), executed in place. A task with no anchorable requirement (pure-mechanics chore) omits the `**Spec:**` line entirely (never `**Spec:** none`) and carries a mechanical-task row in `## Spec coverage` — silence is never valid.
256
+ **Anchor rules.** See [reference/plan-contract.md § Spec anchors](reference/plan-contract.md). A task with no anchorable requirement omits the `**Spec:**` line and carries a mechanical-task row in `## Spec coverage` — silence is never valid.
254
257
 
255
258
  ## Spec Coverage Table
256
259
 
257
- Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. Two row kinds:
258
-
259
- ```markdown
260
- ## Spec coverage
261
-
262
- | anchor | requirement (short) | owner |
263
- |---|---|---|
264
- | § "Design" L34-L37 | anchor line in task template | Task 2 |
265
- | § "Edge cases" L120 | stale anchor = blocking SR finding | Task 4, Task 5 |
266
- | § "Testing" L84 | checker fixtures: `node --test extensions/lib/plan-check.test.ts` | Task 3 |
267
- | § "Acceptance" L88 | full suite passes: `npm test` | Verification |
268
- | § "Out of scope" L131 | fix-round anchoring | waived: out of scope per spec |
269
- | - | mechanical: release commit | Task 7 |
270
- ```
260
+ Every plan ends with a `## Spec coverage` section (grammar and example: [reference/plan-contract.md § Spec coverage table](reference/plan-contract.md)). Build it extraction-first: walk the spec top to bottom and write one row per normative requirement **before** assigning owners — every Design imperative (Add/Remove/Keep/Replace-style directives, not any fixed lexical form), every Edge-cases rule, every Acceptance criterion, every Out-of-scope entry, and every non-none Documentation-impact entry. Then assign owners, then re-walk the spec once: every normative clause has a row. Two row kinds:
271
261
 
272
- - **Requirement rows:** anchor + short requirement + owner = task-ID list, or `Verification`, or `waived: <reason>`. A cross-cutting requirement (decided in more than one task) lists **every** deciding task as owner, not the first. `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
273
- - **`Verification` owner:** use for a requirement the header `**Verification:**` command proves. Write the exact string `Verification`, alone. Quote only literals contained in that header. Anchor the single requirement line. Keep scoped commands task-owned.
274
- - **Mechanical-task rows:** anchor `-`, requirement `mechanical: <short>`, owner = the task ID. One such row per anchor-less task.
262
+ - **Requirement rows:** a cross-cutting requirement (decided in more than one task) lists **every** deciding task as owner, not the first. `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
263
+ - **`Verification` owner:** only for a requirement the header command proves; grammar in the reference.
275
264
  - The table is plan-authoring-time only — never passed to implementer or reviewer dispatches.
276
265
 
277
266
  ## No Placeholders
@@ -282,11 +271,10 @@ Every plan failure mode:
282
271
  - ❌ `# Implement the rest of the function` — incomplete code is invalid code.
283
272
  - ❌ "Add tests for edge cases" — name the edge cases.
284
273
  - ❌ "Wire it up to the existing system" — give file paths and call sites.
285
- - ❌ "timeout/gtimeout ladder" when the spec fixes the literal `timeout 30` — never paraphrase an exact-string requirement (setting keys, error messages, banner/format strings, command names and invocations, API shapes); transcribe it as a backtick-quoted spec literal: `timeout 30`. Spec-side backtick spans containing `<placeholder>` segments are templates the plan instantiates, not exact-string requirements — exempt from quote integrity.
286
274
  - ❌ "Similar to Task N" — repeat the code. Implementers (and subagents with fresh context) may read tasks out of order; pointing at a sibling task is not a substitute for showing the code.
287
275
  - ❌ References to types, functions, methods, or fields not defined in any task in this plan. If it shows up in Task 5, it must be introduced by Task 1–4 or already exist in the codebase (with a file:line citation).
288
- - ❌ `[fill in]`, `<example>`, `xxx` markers anywhere in the doc.
289
276
  - ❌ "Probably also need to update the docs" — either yes (which doc) or no. Docs are named plan tasks, sourced from the spec's Documentation impact section (materiality bar in `brainstorming/reference/documentation-impact.md`).
277
+ - Quote integrity and the banned-token list: [reference/plan-contract.md § Placeholders and quote integrity](reference/plan-contract.md).
290
278
 
291
279
  If a decision is genuinely open, put it in an explicit **Open Questions** section at the top and resolve before execution starts.
292
280
 
@@ -294,10 +282,10 @@ If a decision is genuinely open, put it in an explicit **Open Questions** sectio
294
282
 
295
283
  After drafting the plan and before announcing it complete, run the deterministic checker, then the judgment checks yourself — not a subagent dispatch.
296
284
 
297
- - **Deterministic checker.** Run `plan_check({ planPath })` on the saved plan. Assess and fix every finding yourself (no human involvement), then re-run until it passes — a pass writes the execution stamp that implement-start verifies mechanically. If the same finding survives 3 fix rounds, convert it to an explicit Open Question and stop (the pre-existing Open-Questions halt, resolved by the human in-session — not a new gate). The checker covers table closure, quote integrity, anchor resolution, path existence, placeholder scan, wave file-disjointness, solo-line presence, header-only entrypoint, and waiver-literal.
285
+ - **Deterministic checker.** Run `plan_check({ planPath })` on the saved plan. Assess and fix every finding yourself (no human involvement), then re-run until it passes — a pass writes the execution stamp that implement-start verifies mechanically. If the same finding survives 3 fix rounds, convert it to an explicit Open Question and stop (the pre-existing Open-Questions halt, resolved by the human in-session — not a new gate). Findings are defined in [reference/plan-contract.md](reference/plan-contract.md).
298
286
  - **Code-vs-anchor sanity.** For each task-owned requirement row, re-read the anchored spec lines and confirm the owner tasks' bodies do what they say - mechanism present, not just the quoted literal. For each `Verification` row, confirm the header command exercises the anchored requirement. Fix the task, don't annotate.
299
287
  - **Type / API consistency.** Function signatures and field names that appear in multiple tasks must match exactly. The plan is its own contract — internal contradictions surface as bugs during execution.
300
- - **Scoped-test coverage.** Every code-touching wave declares at least one scoped test command; only doc-only waves may have none.
288
+ - **Test contract.** Every code task's `Tests:` commands are anchored to its `Test:` path(s); `none:` only where no tests apply; a spec-named seam appears as `via:`, a spec-named fixture path as `Create:`.
301
289
  - **Runtime-resource disjointness.** For every multi-task wave, confirm no two tasks contend on a shared mutable runtime resource (DB/schema, port, fixture, external service, shared temp path) — `Files:` overlap is checked mechanically, resource contention is not. Contention = mis-grouped wave; split or re-order before handoff.
302
290
  - **Solo-reason validity.** Every single-task wave's `Solo:` line (presence is checked mechanically) must name its specific blocker — the blocking task/wave, the contended resource, or `lone remaining task`. Category-only justifications are under-justified; merge or justify before handoff.
303
291
  - **Waiver authorization.** `waived: <reason>` is only for requirements the spec marks out of scope **and** that exclude work from the change. A requirement whose text carries an inline code span (`` `literal` ``) names concrete behaviour and is never waivable - it maps to a task or `Verification`. A waiver on an in-scope normative requirement is a Self-Review failure — there is no human plan-review gate to catch it downstream.
@@ -0,0 +1,70 @@
1
+ # Plan contract
2
+
3
+ The grammar `plan_check` enforces. Each section names its check(s). Findings resolve in `writing-plans` Self-Review: fix, re-run.
4
+
5
+ ## Waves and tasks
6
+
7
+ A wave is a maximal set of tasks that have no ordering dependency on each other, own pairwise-disjoint files, and contend on no shared mutable runtime resource (same DB/schema, port, fixture file, external service, shared temp path).
8
+
9
+ - Tasks nest under `## Wave N — <label>` headers; `### Task N` headers sit inside a wave.
10
+
11
+ ## Files (`wave-file-disjointness`, `paths-exist`)
12
+
13
+ **File-ownership contract.** The per-task `**Files:**` block *is* the ownership declaration — no new syntax. Rule: **within a wave, the union of every task's declared paths must be pairwise disjoint.** Globs are allowed for `Modify` when exact paths are unknown, but must not overlap another same-wave task's paths. A task that must touch another's file belongs in a later wave.
14
+
15
+ `Test:` entries are run anchors, not ownership: `Test`/`Test` on the same path across same-wave tasks is allowed; `Test` vs another task's `Create`/`Modify` is a conflict; writer/writer stays a conflict. `Modify:` paths must exist.
16
+
17
+ ## Spec anchors (`anchor-resolution`)
18
+
19
+ **Anchor rules.** The task's `**Spec:**` line cites the plan header's spec path; multiple anchors sit comma-separated on one line (`§ "A" L10-L18, § "C" L40-L44`). Checks key on the `§` marker, so the header's path-only `**Spec:**` line is never matched. Anchors are captured against the gated spec at plan-writing time; a change to the approved spec follows brainstorming's [Amending an approved spec](../../brainstorming/SKILL.md#amending-an-approved-spec), executed in place. A task with no anchorable requirement (pure-mechanics chore) omits the `**Spec:**` line entirely (never `**Spec:** none`) and carries a mechanical-task row in `## Spec coverage` — silence is never valid.
20
+
21
+ ## Tests (`tests-block`)
22
+
23
+ Every `### Task N` carries a `**Tests:**` block: the bare line `**Tests:**` directly after the last `Files:` entry (blank lines allowed). Bullets, in either form:
24
+
25
+ - `- ` + backtick + command + backtick - one scoped test command; `- via: <entry point>` - the seam the tests call directly (function, route, CLI, module - e.g. `each_finding`), zero or more
26
+ - `- none: <category>` - the task runs no tests; `<category>` is free text derived from the project (docs, config, fixtures, generated assets, ...); exactly one, and no `- Test:` path in `Files:`
27
+
28
+ Absence is never valid. A `- [ ]` step or any non-bullet line ends the block. `via:` and `none:` are unchecked beyond form.
29
+
30
+ Commands run from the repo root. Each command is split into segments on `&&`, `||`, `;`, `|`; every segment must contain, as a whitespace-delimited token, a `Test:` path of the same task (the path alone, or followed by `::`, `#`, or `:` and a filter). A `Test:` value containing `*`, `?`, `[` or ending in `/` never anchors; any other argument token with those shapes is a broadening selector and fails. `cd `, `sh -c`, `bash -c`, `eval `, `$(` are unsupported. Each `Test:` path must exist or be a `Create:` path of some task. A segment equal to a header `**Verification:**` segment is a full-suite command and fails. Runners with no file-addressable form are out of scope (`go test ./pkg -run X`, `mvn -Dtest=`).
31
+
32
+ ## Solo line (`solo-line`)
33
+
34
+ A wave with one task carries, directly under its `## Wave N — <label>` header, the line `Solo: <reason>`. The reason names the blocking task/wave, the contended runtime resource, or `lone remaining task`. Presence is checked here; validity is `writing-plans` Self-Review.
35
+
36
+ ## Header-only entrypoint (`header-entrypoint`)
37
+
38
+ The `**Verification:**` line is the **only** place the full verification entrypoint may appear — never in any task or wave step. The verify phase reads it from the plan instead of re-deriving it; execution runs scoped commands only.
39
+
40
+ The header value's backtick spans (else the raw value) are split on `&&`, `||`, `;`, `,` into segments. No `Run:` step payload segment and no `Tests:` bullet segment may equal a header segment; prose inside waves may not contain the whole header value.
41
+
42
+ ## Spec coverage table (`table-closure`, `waiver-literal`)
43
+
44
+ Every plan ends with a `## Spec coverage` section — authored last, placed after all Task sections (owner IDs do not exist earlier). Closure both ways: every `### Task N` appears as an owner in some row; every row's owner task exists.
45
+
46
+ ```markdown
47
+ ## Spec coverage
48
+
49
+ | anchor | requirement (short) | owner |
50
+ |---|---|---|
51
+ | § "Design" L34-L37 | anchor line in task template | Task 2 |
52
+ | § "Edge cases" L120 | stale anchor = blocking SR finding | Task 4, Task 5 |
53
+ | § "Testing" L84 | checker fixtures: `node --test extensions/lib/plan-check.test.ts` | Task 3 |
54
+ | § "Acceptance" L88 | full suite passes: `npm test` | Verification |
55
+ | § "Out of scope" L131 | fix-round anchoring | waived: out of scope per spec |
56
+ | - | mechanical: release commit | Task 7 |
57
+ ```
58
+
59
+ **Requirement rows:** anchor + short requirement + owner = task-ID list, or `Verification`, or `waived: <reason>`. A `waived:` row whose requirement cell contains an inline code span fails `waiver-literal`.
60
+
61
+ **`Verification` owner:** use for a requirement the header `**Verification:**` command proves. Write the exact string `Verification`, alone. Quote only literals contained in that header. Anchor the single requirement line. Keep scoped commands task-owned.
62
+
63
+ **Mechanical-task rows:** anchor `-`, requirement `mechanical: <short>`, owner = the task ID. One such row per anchor-less task.
64
+
65
+ ## Placeholders and quote integrity (`placeholder-scan`, `quote-integrity`)
66
+
67
+ - ❌ "timeout/gtimeout ladder" when the spec fixes the literal `timeout 30` — never paraphrase an exact-string requirement (setting keys, error messages, banner/format strings, command names and invocations, API shapes); transcribe it as a backtick-quoted spec literal: `timeout 30`. Spec-side backtick spans containing `<placeholder>` segments are templates the plan instantiates, not exact-string requirements — exempt from quote integrity.
68
+ - ❌ `[fill in]`, `<example>`, `xxx` markers anywhere in the doc.
69
+
70
+ The banned token list is `BANNED_TOKENS` in `extensions/lib/plan-check.ts`; a token inside a literal the anchored spec requires is exempt.