@gobing-ai/spur 0.3.95 → 0.3.96

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (155) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/config/plugin-scripts.json +0 -54
  3. package/config/rules/boundary/sp-script-placement.yaml +19 -0
  4. package/config/rules/strict/runtime-boundaries.yaml +5 -0
  5. package/config/rules/structure/test-location.yaml +2 -0
  6. package/config/rules/typescript/no-syscall-emulation-in-boundary-mock.yaml +1 -1
  7. package/config/rules/typescript/output-boundaries.yaml +15 -2
  8. package/config/script-placement-baseline.json +51 -0
  9. package/config/workflows/feature-verification.yaml +1 -1
  10. package/config/workflows/history-anatomy.yaml +23 -11
  11. package/config/workflows/idea-pipeline.yaml +80 -85
  12. package/config/workflows/pr-review.yaml +18 -9
  13. package/config/workflows/task-pipeline.yaml +42 -95
  14. package/config/workflows/wrapup-pipeline.yaml +89 -37
  15. package/package.json +1 -1
  16. package/plugins/sp/README.md +8 -10
  17. package/plugins/sp/agents/super-reviewer.md +43 -18
  18. package/plugins/sp/commands/dev-fixgha.md +83 -0
  19. package/plugins/sp/commands/dev-gitmsg.md +8 -6
  20. package/plugins/sp/commands/dev-gtd.md +2 -2
  21. package/plugins/sp/commands/dev-idea.md +22 -12
  22. package/plugins/sp/commands/dev-plan.md +8 -9
  23. package/plugins/sp/commands/dev-review.md +22 -13
  24. package/plugins/sp/commands/dev-verifyall.md +1 -1
  25. package/plugins/sp/commands/spur-init.md +2 -2
  26. package/plugins/sp/lib/history-anatomy.generated.d.mts +112 -0
  27. package/plugins/sp/lib/history-anatomy.generated.mjs +686 -0
  28. package/plugins/sp/lib/idea-handoff.generated.mjs +5 -4
  29. package/plugins/sp/lib/inline-run.generated.d.mts +11 -0
  30. package/plugins/sp/lib/inline-run.generated.mjs +25 -8
  31. package/plugins/sp/lib/quality-gate.generated.d.mts +104 -0
  32. package/plugins/sp/lib/quality-gate.generated.mjs +438 -0
  33. package/plugins/sp/lib/residual-scan.generated.d.mts +62 -0
  34. package/plugins/sp/lib/residual-scan.generated.mjs +210 -0
  35. package/plugins/sp/lib/spur-bin.ts +36 -0
  36. package/plugins/sp/lib/step-profile.generated.d.mts +71 -0
  37. package/plugins/sp/lib/step-profile.generated.mjs +174 -0
  38. package/plugins/sp/plugin.json +1 -1
  39. package/plugins/sp/references/roles.md +1 -1
  40. package/plugins/sp/scripts/history-anatomy-cache.mjs +20 -19
  41. package/plugins/sp/scripts/history-anatomy-cache.ts +23 -928
  42. package/plugins/sp/scripts/inline-run-setup.mjs +85 -320
  43. package/plugins/sp/scripts/inline-run-setup.ts +104 -667
  44. package/plugins/sp/scripts/quality-gate.mjs +34 -19
  45. package/plugins/sp/scripts/quality-gate.ts +12 -658
  46. package/plugins/sp/scripts/residual-scan.mjs +135 -156
  47. package/plugins/sp/scripts/residual-scan.ts +102 -499
  48. package/plugins/sp/scripts/script-root.mjs +5 -1
  49. package/plugins/sp/scripts/script-root.ts +5 -1
  50. package/plugins/sp/scripts/workflow-step-profile.mjs +25 -17
  51. package/plugins/sp/scripts/workflow-step-profile.ts +21 -315
  52. package/plugins/sp/scripts/wrapup-drift-probe.mjs +8 -2
  53. package/plugins/sp/scripts/wrapup-drift-probe.ts +3 -2
  54. package/plugins/sp/scripts/wrapup-steps.mjs +10 -32
  55. package/plugins/sp/scripts/wrapup-steps.ts +8 -36
  56. package/plugins/sp/skills/code-improvement/SKILL.md +5 -4
  57. package/plugins/sp/skills/code-verification/SKILL.md +34 -9
  58. package/plugins/sp/skills/code-verification/references/verdict-schema.md +3 -3
  59. package/plugins/sp/skills/functional-review/SKILL.md +7 -4
  60. package/plugins/sp/skills/functional-review/references/verdict-schema.md +1 -1
  61. package/plugins/sp/skills/history-anatomy/references/modes.md +2 -1
  62. package/plugins/sp/skills/next-feature/references/handoff-routing.md +1 -1
  63. package/plugins/sp/skills/next-router/references/routing-table.md +1 -1
  64. package/plugins/sp/skills/spur-cli/SKILL.md +3 -3
  65. package/plugins/sp/skills/spur-cli/references/agent.md +10 -10
  66. package/plugins/sp/skills/spur-cli/references/features.md +17 -6
  67. package/plugins/sp/skills/spur-cli/references/init.md +17 -16
  68. package/plugins/sp/skills/spur-cli/references/self.md +3 -2
  69. package/plugins/sp/skills/spur-cli/references/serve.md +10 -10
  70. package/plugins/sp/skills/spur-cli/references/tasks/verbs.md +21 -6
  71. package/plugins/sp/skills/spur-cli/references/tasks.md +14 -8
  72. package/plugins/sp/skills/spur-dev/references/ac-style-guide.md +4 -3
  73. package/plugins/sp/skills/spur-dev/references/cross-cutting.md +11 -9
  74. package/plugins/sp/skills/spur-dev/references/decision-brief.md +1 -1
  75. package/plugins/sp/skills/spur-dev/references/dev-operations.md +86 -70
  76. package/plugins/sp/skills/spur-dev/references/done-housekeeping.md +3 -3
  77. package/plugins/sp/skills/spur-dev/references/execution-batch.md +77 -34
  78. package/plugins/sp/skills/spur-dev/references/execution-workflow.md +5 -9
  79. package/plugins/sp/skills/spur-dev/references/feature-link-helper.md +3 -3
  80. package/plugins/sp/skills/spur-dev/references/flag-glossary.md +31 -21
  81. package/plugins/sp/skills/spur-dev/references/gate-checklists.md +19 -19
  82. package/plugins/sp/skills/spur-dev/references/idea-evaluation.md +4 -3
  83. package/plugins/sp/skills/spur-dev/references/inline-pipeline-driver.md +29 -3
  84. package/plugins/sp/skills/spur-doctor/SKILL.md +1 -1
  85. package/plugins/sp/skills/sys-architecture/SKILL.md +3 -2
  86. package/spur.js +1495 -369
  87. package/web/_astro/{BoardApp.Cpxntzad.js → BoardApp.BIjMatT1.js} +1 -1
  88. package/web/_astro/{BoardApp.BxJuwD7I.js → BoardApp.GvjIe9Z6.js} +42 -42
  89. package/web/_astro/{TaskDetail.C5bW4WGV.js → TaskDetail.BZM3EAFx.js} +1 -1
  90. package/web/_astro/{arc.CyjRvNMY.js → arc.HkRiZnoI.js} +1 -1
  91. package/web/_astro/{architectureDiagram-3BPJPVTR.B4lWRzJA.js → architectureDiagram-3BPJPVTR.BDdK-tgZ.js} +1 -1
  92. package/web/_astro/{blockDiagram-GPEHLZMM.Be38USQ2.js → blockDiagram-GPEHLZMM.CZ9kOBxx.js} +1 -1
  93. package/web/_astro/{c4Diagram-AAUBKEIU.DT9Fj5Qx.js → c4Diagram-AAUBKEIU.DXK9qeRD.js} +1 -1
  94. package/web/_astro/channel.DlhX2MEt.js +1 -0
  95. package/web/_astro/{chunk-2J33WTMH.SeSBWLg5.js → chunk-2J33WTMH.BTD2WeX5.js} +1 -1
  96. package/web/_astro/{chunk-4BX2VUAB.DoO4VE14.js → chunk-4BX2VUAB.Ci6Qbcvk.js} +1 -1
  97. package/web/_astro/{chunk-55IACEB6.DrmTFA53.js → chunk-55IACEB6.yl3zsj7p.js} +1 -1
  98. package/web/_astro/{chunk-727SXJPM.Bi-TUb_V.js → chunk-727SXJPM.ClXZsfyR.js} +1 -1
  99. package/web/_astro/{chunk-AQP2D5EJ.B-xp8Uhi.js → chunk-AQP2D5EJ.BmKWZBcP.js} +1 -1
  100. package/web/_astro/{chunk-FMBD7UC4.DGWrDnUU.js → chunk-FMBD7UC4.By30fcb7.js} +1 -1
  101. package/web/_astro/{chunk-ND2GUHAM._BagvPDy.js → chunk-ND2GUHAM.24NmrD-k.js} +1 -1
  102. package/web/_astro/{chunk-QZHKN3VN.VOboYWQ0.js → chunk-QZHKN3VN.BN4sKdcS.js} +1 -1
  103. package/web/_astro/{classDiagram-4FO5ZUOK.D_apfHS7.js → classDiagram-4FO5ZUOK.CJpoMPb5.js} +1 -1
  104. package/web/_astro/{classDiagram-v2-Q7XG4LA2.D_apfHS7.js → classDiagram-v2-Q7XG4LA2.CJpoMPb5.js} +1 -1
  105. package/web/_astro/{cose-bilkent-S5V4N54A.DNLU_L-x.js → cose-bilkent-S5V4N54A.BraxQ2Nt.js} +1 -1
  106. package/web/_astro/{cynefin-OW5HDTMX.D8borhV-.js → cynefin-OW5HDTMX.gzoU73oL.js} +1 -1
  107. package/web/_astro/{dagre-BM42HDAG.Uc0lvjcV.js → dagre-BM42HDAG.wyDfWySU.js} +1 -1
  108. package/web/_astro/{diagram-2AECGRRQ.DDIxtDBo.js → diagram-2AECGRRQ.h2_EbMxV.js} +1 -1
  109. package/web/_astro/{diagram-5GNKFQAL.DVoDnD3H.js → diagram-5GNKFQAL.Db6dcs89.js} +1 -1
  110. package/web/_astro/{diagram-KO2AKTUF.CRubTJ-0.js → diagram-KO2AKTUF.NVwzxu_U.js} +1 -1
  111. package/web/_astro/{diagram-LMA3HP47.m9SeYcXQ.js → diagram-LMA3HP47.B-eNgDQk.js} +1 -1
  112. package/web/_astro/{diagram-OG6HWLK6.C6s2ZjV5.js → diagram-OG6HWLK6.Bsa8ZRbJ.js} +1 -1
  113. package/web/_astro/{erDiagram-TEJ5UH35.nSzr5KTj.js → erDiagram-TEJ5UH35.BospRE5Q.js} +1 -1
  114. package/web/_astro/{flowDiagram-I6XJVG4X.DrZcys16.js → flowDiagram-I6XJVG4X.DMzvHwAA.js} +1 -1
  115. package/web/_astro/{ganttDiagram-6RSMTGT7.Do4M1b-K.js → ganttDiagram-6RSMTGT7.Dmdxgd-M.js} +1 -1
  116. package/web/_astro/{gitGraphDiagram-PVQCEYII.BWc9oJ4l.js → gitGraphDiagram-PVQCEYII.CaqdO0rL.js} +1 -1
  117. package/web/_astro/index.EoZlzLK-.css +1 -0
  118. package/web/_astro/{infoDiagram-5YYISTIA.BIdOctFk.js → infoDiagram-5YYISTIA.DDVD9Wn7.js} +1 -1
  119. package/web/_astro/{ishikawaDiagram-YF4QCWOH.CV5SuG2t.js → ishikawaDiagram-YF4QCWOH.EKyCUFsL.js} +1 -1
  120. package/web/_astro/{journeyDiagram-JHISSGLW.CxE-rEJ-.js → journeyDiagram-JHISSGLW.By2-RkYB.js} +1 -1
  121. package/web/_astro/{kanban-definition-UN3LZRKU.BPRdKWs_.js → kanban-definition-UN3LZRKU.ar-WmlyR.js} +1 -1
  122. package/web/_astro/{linear.zRsuuDTE.js → linear.BnzMBgo_.js} +1 -1
  123. package/web/_astro/{mermaid.core.jAJTcMKc.js → mermaid.core.DSqeFu2Y.js} +4 -4
  124. package/web/_astro/{mindmap-definition-RKZ34NQL.BffJxfCr.js → mindmap-definition-RKZ34NQL.1XcF8Wh2.js} +1 -1
  125. package/web/_astro/{pieDiagram-4H26LBE5.CacwWaE3.js → pieDiagram-4H26LBE5.DnItHOdK.js} +1 -1
  126. package/web/_astro/{quadrantDiagram-W4KKPZXB.hXELD06-.js → quadrantDiagram-W4KKPZXB.BQQyrnBt.js} +1 -1
  127. package/web/_astro/{requirementDiagram-4Y6WPE33.DdEbPcqT.js → requirementDiagram-4Y6WPE33.Dmzxun82.js} +1 -1
  128. package/web/_astro/{sankeyDiagram-5OEKKPKP.DgaDocKv.js → sankeyDiagram-5OEKKPKP.BBIeRKYB.js} +1 -1
  129. package/web/_astro/{sequenceDiagram-3UESZ5HK.B0KScnAu.js → sequenceDiagram-3UESZ5HK.CUNVdBBf.js} +1 -1
  130. package/web/_astro/{stateDiagram-AJRCARHV.CV9M_WNc.js → stateDiagram-AJRCARHV.DlFf2sJm.js} +1 -1
  131. package/web/_astro/{stateDiagram-v2-BHNVJYJU.BOAo74Et.js → stateDiagram-v2-BHNVJYJU.DXR7I1H5.js} +1 -1
  132. package/web/_astro/{timeline-definition-PNZ67QCA.CqqepUMe.js → timeline-definition-PNZ67QCA.BcO-GyX2.js} +1 -1
  133. package/web/_astro/{vennDiagram-CIIHVFJN.DDqe4dOG.js → vennDiagram-CIIHVFJN.ChCECmmW.js} +1 -1
  134. package/web/_astro/{wardleyDiagram-YWT4CUSO.B3R9OPdO.js → wardleyDiagram-YWT4CUSO.CRaT12mD.js} +1 -1
  135. package/web/_astro/{xychartDiagram-2RQKCTM6.B3SvgXpQ.js → xychartDiagram-2RQKCTM6.B3lsHVpj.js} +1 -1
  136. package/web/index.html +2 -2
  137. package/plugins/sp/lib/artifact-digest.generated.d.mts +0 -7
  138. package/plugins/sp/lib/artifact-digest.generated.mjs +0 -48
  139. package/plugins/sp/scripts/feature-sync-bounded.mjs +0 -301
  140. package/plugins/sp/scripts/feature-sync-bounded.ts +0 -481
  141. package/plugins/sp/scripts/idea-coverage-check.ts +0 -168
  142. package/plugins/sp/scripts/inline-pipeline-parity-check.ts +0 -298
  143. package/plugins/sp/scripts/record-feature-sync.mjs +0 -63
  144. package/plugins/sp/scripts/record-feature-sync.ts +0 -84
  145. package/plugins/sp/scripts/script-contract-check.ts +0 -506
  146. package/plugins/sp/scripts/stage-registry-adapter.ts +0 -1533
  147. package/plugins/sp/scripts/surface-drift-inventory.ts +0 -989
  148. package/plugins/sp/scripts/task-evidence-precheck.ts +0 -189
  149. package/plugins/sp/scripts/task-size-precheck.ts +0 -212
  150. package/plugins/sp/scripts/transition-shim-check.ts +0 -238
  151. package/plugins/sp/scripts/validate-commands.ts +0 -689
  152. package/plugins/sp/scripts/validate-flag-contracts.ts +0 -890
  153. package/plugins/sp/scripts/verify-answer-lint.ts +0 -549
  154. package/web/_astro/channel.CI6N_tCg.js +0 -1
  155. package/web/_astro/index.CENnIEqT.css +0 -1
@@ -64,6 +64,10 @@ function normalizeArgs(raw: Args): Args {
64
64
 
65
65
  - If `--feature FOO` is present and `--tasks` is absent, treat the effective selector as `feature:FOO`.
66
66
  - If both are present, `--tasks` wins (with a one-line note in the batch report).
67
+ - **Per-command admission filters (I33 1023).** Step 1 is the shared baseline; commands may layer
68
+ stricter grammar on top — e.g. `/sp:dev-review` rejects `ready`/status pseudo-lists and the mixed
69
+ `--tasks` + `--feature` combination (exit 2), and accepts multi-id `--feature <id>,<id>` as
70
+ caller-level sugar expanded by the command layer before the resolver.
67
71
 
68
72
  **Feature-derived strict preflight (R2, task 0510).** After normalization, if the **effective
69
73
  selector** is `feature:<id>` (whether via `--tasks feature:<id>` or the `--feature <id>` sugar),
@@ -291,8 +295,8 @@ Each pipeline run ends in one of two terminal states:
291
295
  `.spur/run/<wbs>-verify-answer.txt` AC table is exactly four columns:
292
296
  `| AC | Status | Evidence Type | Evidence |`. The evidence-type token
293
297
  (`test`, `command`, `static-ref`, `manual-review`, `llm-judge`, `n/a`, or a `+`
294
- compound) is isolated in cell 3. A token merged into the evidence cell fails
295
- `verify-answer-lint`.
298
+ compound) is isolated in cell 3. A token merged into the evidence cell fails the
299
+ `spur task verdict` answer lint.
296
300
 
297
301
  **Driver acceptance (0930 R3).** The trace row and `.spur/run/<wbs>-verdict.json` are accepted as
298
302
  terminal evidence only if BOTH hold:
@@ -328,40 +332,35 @@ node "$(superskill script path sp batch-preflight.mjs)" --wbs <wbs> --status <st
328
332
  Helper: `recoveryHint(status, wbs)` in `plugins/sp/scripts/batch-preflight.ts`. Tables remain SSOT
329
333
  in next-router; this only maps status → primary TABLE A hop for recovery.
330
334
 
331
- ### 3.3c Bounded feature-sync retry suppression (task 0411)
335
+ ### 3.3c Feature-sync retry suppression (task 0411; 1004 R3 moved it into the sync service)
332
336
 
333
337
  During a batch, the per-task `record` step and the wrap-up `feature-transition` step each invoke
334
338
  feature status sync. When a feature is L4-gate-blocked (e.g. not all linked tasks are `done`), the
335
339
  identical blocked proposal repeats on every call with no intervening input change — in the H9
336
- dogfood, 4 redundant sync calls produced the same blocked result. The orchestration seam fixes
337
- this, not the engine.
340
+ dogfood, 4 redundant sync calls produced the same blocked result. The service fixes this, not the
341
+ engine.
338
342
 
339
343
  Both `task-pipeline.yaml` (`record` step) and `wrapup-pipeline.yaml` (`feature-transition` step)
340
- invoke the bounded wrapper instead of raw `feature sync`:
341
-
342
- ```bash
343
- node "$(superskill script path sp feature-sync-bounded.mjs)" <feature-id> --spur-bin "<spurBin>" --json
344
- ```
345
-
346
- The wrapper:
347
-
348
- 1. Reads an input fingerprint (feature file content hash, linked task statuses, verdict artifact
349
- mtimes) **before** invoking `feature sync`.
350
- 2. Classifies the structured result — `gateBlocked` checked first (a partial hop can have
351
- `applied: true` while still gate-blocked), then `applied`, then `no-op`.
352
- 3. On a **blocked** result, persists `.spur/run/feature-sync-blocked-<id>.json` and, on the next
353
- call with an **identical fingerprint**, suppresses the redundant sync and replays the prior
354
- blocked result.
355
- 4. On **applied** or **no-op** results, passes through unchanged (no suppression).
356
- 5. When the fingerprint **changes** (a task completed, a verdict file updated), suppression is
344
+ invoke `spur feature sync <feature-id> --json` directly — retry suppression lives inside the
345
+ `FeatureService.syncFeature` implementation:
346
+
347
+ 1. On a **blocked** result (`gateBlocked` checked first — a partial hop can have `applied: true`
348
+ while still gate-blocked — then an unapplied from≠to deferral), the service persists
349
+ `.spur/run/feature-sync-blocked-<id>.json` keyed by an input fingerprint (feature file content
350
+ hash, linked task statuses, verdict artifact mtimes).
351
+ 2. On the next call with an **identical fingerprint**, the service suppresses the redundant sync
352
+ and replays the prior blocked result (`suppressed: true`) without re-deriving hops.
353
+ 3. On **applied** or **no-op** results, the state file is cleared (no suppression).
354
+ 4. When the fingerprint **changes** (a task completed, a verdict file updated), suppression is
357
355
  invalidated and a fresh sync runs.
356
+ 5. `--force` (and an explicit confirm re-attempt) bypass the replay and re-derive live; dry-run
357
+ never reads or writes the state.
358
358
 
359
- **Batch driver contract:** the orchestrator does **nothing extra** — the wrapper lives inside the
360
- pipeline's `record` step and the wrap-up's `feature-transition` step. The driver still launches
361
- `task-pipeline.yaml` verbatim (R4.1). Suppression is transparent: the wrapper emits the same
362
- `FeatureSyncResult` JSON shape as `feature sync --json`, so downstream report logic is unchanged.
363
- The only observable difference is fewer redundant `feature sync` invocations and a one-line
364
- `feature-sync-bounded:` annotation on stderr when a duplicate is suppressed.
359
+ **Batch driver contract:** the orchestrator does **nothing extra** — the suppression lives inside
360
+ the pipeline's `record` step and the wrap-up's `feature-transition` step. The driver still
361
+ launches `task-pipeline.yaml` verbatim (R4.1). Suppression is transparent: the sync emits the same
362
+ `FeatureSyncResult` JSON shape (`suppressed: true` added on replay), so downstream report logic is
363
+ unchanged. The only observable difference is fewer redundant `feature sync` derivations.
365
364
 
366
365
  ### 3.4 Metadata-only host controller (R5, task 0510)
367
366
 
@@ -513,6 +512,19 @@ obligation; neither do root-qualified paths (`knowledge-kit/.spur/run/…`, `/ab
513
512
  which cite another project's evidence — cite foreign run artifacts that way, never bare. A citation missing in BOTH trees, a divergent cited file (never overwritten — reconcile
514
513
  by hand), an unreadable task file, or more than 64 distinct cited files fails the pass → WT-5.
515
514
 
515
+ **Owned evidence rides it too (1012).** With at least one `--task-file`, persist-out also treats as
516
+ copy obligations the worktree's `.spur/run/` direct children named `<wbs>-…` (the WBS is each
517
+ forwarded task file's leading four digits before `_`) or `<runId>-…` (every run row in the worktree
518
+ DB, whichever task it ran) — `<wbs>-verdict.json`, check receipts, route reasons — whether or not
519
+ the task file cites them. They join the cited set: same copy / byte-identical no-op /
520
+ divergent-refuse handling. The 64-file cap bounds citations alone; owned names are bounded per
521
+ owner (each `<wbs>-` / `<runId>-` prefix gets its own 64-file budget, task 1034), so the bound
522
+ scales with the batch and one runaway owner refuses by name before any write. `<runId>.md` /
523
+ `<runId>.state.json` stay with the record copy (a conflict there is a reported skip). Files
524
+ matching neither a citation nor an ownership prefix are left behind. An absent worktree `.spur/run/`
525
+ means nothing is owned; any other listing failure (not a directory, permission denied) fails the
526
+ pass before the invoking tree is written → WT-5. Without `--task-file` nothing is enumerated.
527
+
516
528
  The shapes are pinned (task 0975 R1; `record-missing` and citation behavior per 0984): idempotent on re-persist;
517
529
  success exits 0 printing
518
530
  `{"ok":true,"persisted":<n>,"skipped":[{"id":<run-id>,"reason":"id-exists"|"external-key-conflict"|"record-conflict:<file>"|"record-missing:<file>"|"cited-directory:<name>"|"cited-symlink:<name>"|"cited-non-file:<name>"}]}`
@@ -544,6 +556,15 @@ The wrap receives only what it would accept:
544
556
  3. When the done subset is **empty**, skip the wrap entirely with the reason (e.g. `batch wrap
545
557
  skipped: no done tasks`) instead of invoking wrapup-pipeline on an empty set.
546
558
 
559
+ **Repo-wide tripwire (1037).** After the doc-sync exits converge, the pipeline's `doc-tripwire` hop
560
+ runs the TRUSTED CONFIG ONLY `docTripwireCmd` over the still-uncommitted wrap diff before
561
+ metrics-record. The default probes `package.json` for a `test-repo-wide` script and runs
562
+ `bun run test-repo-wide` only when it is declared (a no-op in other projects), so the batch driver
563
+ passes no extra vars. Batch callers override it like any wrap var (`docTripwireCmd` in `--vars`);
564
+ an empty string disables the check while still recording PASS. A FAIL routes the wrap to `failed`
565
+ with already-written learnings/docs preserved — fix the flagged working-diff violation and re-run
566
+ the wrap.
567
+
547
568
  Filtering lives here, in the batch driver — no change to wrapup-pipeline.yaml or wrapup-steps.ts;
548
569
  the wrap's refusal of non-done tasks remains the hard invariant.
549
570
 
@@ -582,11 +603,14 @@ non-PASS verify verdict, or a HITL pause that ends the run take the WT-5 retenti
582
603
  full pipeline is eligible — `--worktree --mode implement` is rejected (WT-7), because that mode is
583
604
  the pipeline's implement stage and already runs in the driver's tree.
584
605
 
585
- **Review triage `dev-review` (run of one).** `/sp:dev-review <target> --triage --worktree [<name>]`
606
+ **Review triage `dev-review` (run of one).** `/sp:dev-review [--tasks <selector> | --feature <id>[,<id>] | --scope <path>[,<path>]] --triage --worktree [<name>]`
586
607
  runs this lifecycle around one review-plus-triage pass: WT-1…WT-6 apply unchanged, the marker's
587
- `command` is `dev-review` and its `selector` is the review target, and the slug is the WBS or the
588
- path's basename (`sp/review-<slug>-<short-id>`). It skips `quickReadiness` (there is no task set;
589
- admission is "the target resolves"). WT-4 success reads as "every direct fix passed its check and the
608
+ `command` is `dev-review` and its `selector` records the full normalized target list, and the slug
609
+ is `sp/review-<first>-and-<N>-<short-id>` for a multi-target run (N = target count) or
610
+ `sp/review-<slug>-<short-id>` for a single target (the WBS or the path's basename). It skips
611
+ `quickReadiness` (there is no task set; admission is "every target resolves" — each WBS/path must
612
+ resolve before the tree is cut). Under `--triage` the findings are bucketed across all targets once
613
+ (identical `file:line` findings deduped). WT-4 success reads as "every direct fix passed its check and the
590
614
  project gate is green"; anything else takes WT-5. Contract: [dev-operations.md § 2. review](dev-operations.md#2-review).
591
615
 
592
616
  One flag, two modes (see the glossary entry for the ownership rule). Bare `--worktree` is **create
@@ -979,6 +1003,25 @@ Resume, merge, or discard:
979
1003
  discard: git worktree remove <worktree-path> && git branch -D <branch> && spur projects remove <worktree-path>
980
1004
  ```
981
1005
 
1006
+ When the halt cause is `non-FF base ref`, the report replaces the one-line `merge:` hint with this
1007
+ ordered divergence recipe, run by the operator — the driver never merges, rebases, or resolves
1008
+ conflicts itself. Other halt causes (task failure, HITL pause) keep the hint as printed:
1009
+
1010
+ ```
1011
+ # 1. integrate as a merge commit — never a rebase; task evidence cites the branch's commit SHAs
1012
+ git checkout <base-ref> && git merge --no-ff --no-commit <branch>
1013
+ # 2. resolve source conflicts by hand; generated files are then regenerated with the project's
1014
+ # generator, never hand-merged
1015
+ # (this repo: bun run build:plugin-lib && bun run --filter @gobing-ai/spur build:bundle)
1016
+ # 3. stage every resolved path and regenerated bundle (git add …) — an unmerged or
1017
+ # unstaged path makes step 5 abort
1018
+ # 4. run qualityGateCmd once, after ALL conflicts are resolved
1019
+ # 5. commit the merge with the prepared message file
1020
+ git commit -F <message-file>
1021
+ # 6. persist evidence out (WT-4a), then WT-4b/4c cleanup, and set the marker to merged
1022
+ inline-run-setup --persist-out --from <worktree> --task-file …
1023
+ ```
1024
+
982
1025
  The report reuses the [`--next` chain contract](flag-glossary.md#--next-chain-contract) halt-report
983
1026
  shape (halt cause + where + why), not new vocabulary. Retention is the right default: these batches
984
1027
  are long and already resumable via `--continue`; auto-deleting is data loss, auto-merging is a
@@ -1156,11 +1199,11 @@ per-task with `/sp:dev-run <wbs> --worktree <branch>`.
1156
1199
  ### Generated regions — defer the sync, regenerate once (R5)
1157
1200
 
1158
1201
  The only per-task writer of feature files is the `record` step's post-record feature sync
1159
- (`task-pipeline.yaml`, the `feature-sync-bounded` wrapper). Parallel launches set
1202
+ (`task-pipeline.yaml`). Parallel launches set
1160
1203
  the pipeline var `deferFeatureSync: "true"` (default `"false"`): the record step appends
1161
1204
  `feature sync deferred to batch integration` to the task report and skips the sync, so task
1162
1205
  branches never touch feature files or `docs/features/INDEX.md`. After the last integration, on the
1163
- base ref, the orchestrator runs the same bounded wrapper plus `spur feature refresh --feature <f>`
1206
+ base ref, the orchestrator runs `spur feature sync <f> --json` (service-level suppression) plus `spur feature refresh --feature <f>`
1164
1207
  once per touched feature and commits the result as one `chore(corpus)` commit. Sequential and
1165
1208
  inline runs keep the default `"false"` and are unchanged. Any rebase conflict — on a generated
1166
1209
  path or any other — is an R4 `integration-conflict`; there is no path-based exception.
@@ -47,7 +47,7 @@ one thing and yields, so the **pipeline (not the agent) owns the loop**.
47
47
  | ------- | ----------- | ------------ |
48
48
  | `implement` | `/sp:dev-run --mode implement <wbs>` — write the code that satisfies the task; author `## Solution`. | [dev-operations.md §4 run](dev-operations.md) → `sp:code-implementation` |
49
49
  | `test` → (`test-fix` ↔ `test-recheck`) → `review` \| `failed` | **Project quality gate** (not `/sp:dev-unit`). Soft shell probe of `${vars.qualityGateCmd}` (default `bun run spur-check`) — green path pays **one** full gate run. On FAIL: bounded `/sp:dev-fixall` loop (`qualityGateMaxFixAttempts`, default 2) with soft recheck; exhausted attempts route to pipeline `failed`. `/sp:dev-unit` remains **coverage gap-fill** (router C3/C5 / standalone). | [dev-operations.md §10 fixall](dev-operations.md); unit op still §1 |
50
- | `review` | `/sp:dev-review <wbs>` — SECUA-framework review of the diff. | [dev-operations.md §2 review](dev-operations.md) |
50
+ | `review` | `/sp:dev-review --tasks <wbs>` — SECUA-framework review of the diff. | [dev-operations.md §2 review](dev-operations.md) |
51
51
  | `verify` | `sp:code-verification` — requirements traceability + verdict. | [dev-operations.md §3 verify](dev-operations.md) |
52
52
 
53
53
  Interactive omit/`inline` executes these model stages through the
@@ -96,10 +96,7 @@ cache-conservation discipline (`plugins/sp/skills/dogfood-testing/references/mon
96
96
 
97
97
  ## Step 2: Pipeline run
98
98
 
99
- > **Pre-launch size-gate pre-check (R1 / 0478).** Before launching `spur workflow run task-pipeline.yaml`, probe the task's `## Plan` checklist item count (`spur task show <wbs> --json`). The default cap is 8 items (`maxImplementPlanItems: 8`). If the plan item count exceeds 8:
100
- >
101
- > - Without `--auto`: warn the operator before calling `spur workflow run` and prompt for confirmation or a plan-item override via `--vars '{"maxImplementPlanItems":"<count>"}'`.
102
- > - With `--auto`: automatically append `"maxImplementPlanItems": "<count>"` to `--vars` and log a single-line notice (e.g. `Notice: task <wbs> has N plan items (>8 default cap); injecting maxImplementPlanItems override`).
99
+ > **Pre-launch size-gate pre-check (R1 / 0478; task 1002).** Before launching `spur workflow run task-pipeline.yaml`, run `spur task check <wbs> --precheck --json` — the same command the pipeline's precheck guard runs. It fails with a `precheck-size` finding above 10 requirements or 16 Plan items. The limits are fixed (no `--vars` override): on a failure, split the task instead of launching.
103
100
 
104
101
  **`--worktree [<name>]` wraps Step 2, on either surface.** When `/sp:dev-run --mode full` carries
105
102
  [`--worktree`](flag-glossary.md#flag-worktree), create or adopt the worktree *before* launching the
@@ -336,10 +333,9 @@ task.** A `cheap`/`standard`-tier model handed a task that big does not fail fas
336
333
  entire `implementTimeoutMs` and exits 3 with a partial tree (run `ca130182` — 7 reqs / 9 plan
337
334
  items / 12+ files → 30 minutes, 6 of 12 files, no tests, no docs, no `## Solution`).
338
335
 
339
- The precheck size gate is count-only: it writes FAIL above 10 requirements or 16 Plan items and
340
- never consults the executor's capability tier. Clear a FAIL deliberately — split the task, or raise
341
- the cap with `--vars '{"maxImplementReqs":<n>}'` — but the caps only accept a big task, they do not
342
- make a flash model able to finish one.
336
+ The precheck size gate (`spur task check <wbs> --precheck`) is count-only: it fails above 10
337
+ requirements or 16 Plan items and never consults the executor's capability tier. The limits are
338
+ fixed — clear a failure by splitting the task.
343
339
 
344
340
  The empty-implement guard (`requireDiff` on the task-pipeline `implement` step, R3) fails the
345
341
  run fast when an implement exits 0 with zero non-corpus changes — a no-op never drifts into
@@ -11,13 +11,13 @@ see_also:
11
11
  **Scope:** opt-in, strictness-triggered — never gate-time, never automatic.
12
12
 
13
13
  This helper resolves a deferred `feature_id` edge when the operator explicitly invokes or intends
14
- `--strict` rigor, or asks to "link this task to a feature." It is **NOT** part of the `--strict-core`
14
+ `--strict` rigor, or asks to "link this task to a feature." It is **NOT** part of the `--as done`
15
15
  done-gate, NOT in any `--next` chain, and NOT triggered automatically. Invoking it is always an
16
16
  explicit operator choice.
17
17
 
18
18
  **Design boundaries (enforced):**
19
19
 
20
- - `feature_id: null` is a valid, supported state under the default done-gate (`--strict-core`). Deferral is legitimate.
20
+ - `feature_id: null` is a valid, supported state under the default done-gate (`--as done`). Deferral is legitimate.
21
21
  - This helper fires only when the operator opts in — it does NOT change the L4 warning severity.
22
22
  - It NEVER creates a new feature without operator confirmation.
23
23
  - It ALWAYS prefers matching an **existing** feature before proposing creation.
@@ -30,7 +30,7 @@ explicit operator choice.
30
30
  - A deliberate traceability audit: `spur task check --strict` across the corpus reveals N orphan tasks.
31
31
 
32
32
  **Do NOT invoke from:**
33
- - The `--strict-core` done-gate (it must stay feature_id-agnostic)
33
+ - The `--as done` done-gate (it must stay feature_id-agnostic)
34
34
  - Any `--next` chain or automated pipeline step
35
35
  - Any context where the operator has not explicitly requested strict rigor or linking
36
36
 
@@ -43,7 +43,7 @@ as a follow-up, not quietly left inconsistent.
43
43
  `implementAgent` override, objective triggers, and surface-derivation logic — lives in
44
44
  [cross-cutting.md](cross-cutting.md#inline-default-execution-surface).
45
45
  The value table below is the C3a cross-file parity surface (kept in lockstep with the SSOT by
46
- `validate-flag-contracts.ts`), not an independent restatement.
46
+ `scripts/commands/validate-flag-contracts.ts`), not an independent restatement.
47
47
 
48
48
  | Value | Who does the work | Derived surface |
49
49
  | ------------------------------- | --------------------------------------------------------------------------- | --------------------------------------------------------------------------- |
@@ -126,6 +126,10 @@ Skip objective HITL confirmations inside this command (feature-check, batch-crea
126
126
  gate). Taste gates and irreversible HITL gates (e.g. `--merge`) still pause even under `--auto`.
127
127
  Only declared where the command already has at least one HITL gate the flag can skip.
128
128
 
129
+ **Planning exception** (`dev-idea`, `dev-plan`): `--auto` also accepts the recommendation at the
130
+ taste gates (idea-eval, design-approval) — every gate there is reversible corpus writes. A gate
131
+ without an actionable recommendation (missing eval recommendation, FAIL design check) still pauses.
132
+
129
133
  ### `--keep-going` — batch failure policy: skip dependents, continue independents
130
134
 
131
135
  **Anchor:** `#flag-keep-going`.
@@ -177,7 +181,8 @@ forced. Never bypasses lifecycle status transitions or irreversible HITL gates.
177
181
 
178
182
  Scope the operation to all tasks under a feature id (`^[A-Z][1-9]*$`). On feature-advancing
179
183
  commands (`dev-wrapall`) it also advances the feature through legal lifecycle edges with guards
180
- honored.
184
+ honored. On `dev-review` it takes a comma list (`--feature <id>[,<id>]`) — sugar for the union of
185
+ the `feature:<id>` sets, resolved once and frozen.
181
186
 
182
187
  ### `--check <cmd>` — validation command for iterate-and-check loops
183
188
 
@@ -190,11 +195,13 @@ establishes a baseline with it before the first change and re-runs it after each
190
195
 
191
196
  **Anchor:** `#flag-focus`.
192
197
 
193
- Constrain the operation to a named subset of dimensions — review dimensions on `dev-review`/
194
- `dev-verify`/`dev-verifyall` (`all|stack|dependencies|data|flows|api|security|quality|performance`),
195
- a refactor lens set on `dev-refactor` (`api|architect|tests|ui|auto`), a refine focus mode on
198
+ Constrain the operation to a named subset of dimensions — review dimensions on `dev-review`
199
+ (vocabulary SSOT: [code-verification/SKILL.md](../../code-verification/SKILL.md) review mode —
200
+ `dev-verify`/`dev-verifyall` keep their SECUA-only lens set), a refactor lens set on
201
+ `dev-refactor` (`api|architect|tests|ui|auto`), a refine focus mode on
196
202
  `dev-refine`/`dev-refineall` (`all|requirements|background|constraints|acceptance|quick` —
197
- [dev-operations.md](dev-operations.md) § refine), or a reconstruction lens on `dev-reverse`.
203
+ [dev-operations.md](dev-operations.md) § refine), or a reconstruction lens on `dev-reverse`
204
+ (`all|stack|dependencies|data|flows|api|security|quality|performance`).
198
205
  Narrowing reduces token cost; omitting runs
199
206
  all dimensions.
200
207
 
@@ -203,7 +210,9 @@ all dimensions.
203
210
  **Anchor:** `#flag-scope`.
204
211
 
205
212
  Limit the operation to a file or directory path (`dev-arch`, `dev-debug`, `dev-fixall`,
206
- `dev-gitmsg`, `dev-gtd`, `dev-refactor`, `dev-simplify`) to bound the working set.
213
+ `dev-gitmsg`, `dev-gtd`, `dev-refactor`, `dev-review`, `dev-simplify`) to bound the working set.
214
+ On `dev-review` it takes a comma list of paths — normalized, nested and duplicate paths merged,
215
+ one advisory sub-review per surviving path.
207
216
 
208
217
  ### `--all` — widen the operation to everything in its domain
209
218
 
@@ -226,11 +235,13 @@ is the contract; divergence between `--dry-run` and the real run is a bug.
226
235
 
227
236
  **Anchor:** `#flag-tasks`.
228
237
 
229
- Batch operation only (`dev-parallel`, `dev-refineall`, `dev-runall`, `dev-verifyall`). An explicit selector — WBS
230
- list, status pseudo-list (`todo`, `wip`), `feature:<id>`, or `ready` — resolving to the set the
231
- batch runs over. Required on `dev-parallel`, `dev-runall`, and `dev-verifyall`, where `--feature` is an optional
232
- restrictor. On `dev-refineall` it is instead one of a required pair — supply exactly one of
233
- `--feature` or `--tasks`.
238
+ Batch operation (`dev-parallel`, `dev-refineall`, `dev-review`, `dev-runall`, `dev-verifyall`). An
239
+ explicit selector — WBS list, status pseudo-list (`todo`, `wip`), `feature:<id>`, or `ready` —
240
+ resolving to the set the batch runs over. On `dev-review` the selector is restricted to the
241
+ review-safe forms: comma WBS list and `feature:<id>` — status pseudo-lists and `ready` are
242
+ rejected (they select work to do, not work to review). Required on `dev-parallel`, `dev-runall`,
243
+ and `dev-verifyall`, where `--feature` is an optional restrictor. On `dev-refineall` it is instead
244
+ one of a required pair — supply exactly one of `--feature` or `--tasks`.
234
245
 
235
246
  ### `--mode <kind>` — select an execution mode
236
247
 
@@ -352,11 +363,18 @@ pauses, even under `--auto`.
352
363
 
353
364
  **Anchor:** `#flag-max-retry`.
354
365
 
355
- Bound the retry loop on fix-family commands (`dev-dogfood`, `dev-fixall`, `dev-gtd`). After `n` consecutive
366
+ Bound the retry loop on fix-family commands (`dev-dogfood`, `dev-fixall`, `dev-fixgha`, `dev-gtd`). After `n` consecutive
356
367
  failed fix attempts, stop and ask the operator rather than looping indefinitely. On `dev-dogfood`
357
368
  the default is `2` (fix mode) and `--max-retry 0` selects observe-only — matching the backing
358
369
  `sp:dogfood-testing` skill; the command table and the skill must not drift on this default.
359
370
 
371
+ ### `--no-push` — commit locally, stop before push
372
+
373
+ **Anchor:** `#flag-no-push`.
374
+
375
+ Commit the fixes locally but do not `git push` or run the post-push `gh` verification
376
+ (`dev-fixgha`, `dev-gtd`). The report lists the unpushed commits.
377
+
360
378
  ### `--full` — rewrite a `--next` run as full pipeline
361
379
 
362
380
  **Anchor:** `#flag-full`.
@@ -421,14 +439,6 @@ remainder through `spur task` — never fix straight from the raw findings list.
421
439
  features ([dev-operations.md § 2. review](dev-operations.md#2-review)); `dev-review-session` keeps
422
440
  the stricter direct-fix bar (pure docs / one-to-two-line fixes).
423
441
 
424
- ### `--approve-taste` — pre-clear all taste gates this run
425
-
426
- **Anchor:** `#flag-approve-taste`.
427
-
428
- Planning commands (`dev-idea`, `dev-plan`): with `--auto`, skip all remaining taste pauses this
429
- run (idea-eval + design-approval). Sets `idea_approved=true` and `design_approved=true`. One CLI
430
- flag sets both.
431
-
432
442
  ### `--worktree [<name>]` — run the batch in an isolated git worktree (create or reuse)
433
443
 
434
444
  **Anchor:** `#flag-worktree`.
@@ -72,16 +72,15 @@ Entered before `task-pipeline.yaml` `precheck` state runs `spur task check <wbs>
72
72
  - [ ] The `## Plan` section is an ordered checklist (not prose).
73
73
  - [ ] The `## Design` section, if present, does not contradict the parent feature's design.
74
74
  - [ ] No `TODO`, `TBD`, or `???` placeholders in Requirements, AC, Design, or Plan.
75
- - [ ] The evidence-channel precheck (0726 R2) status file is consulted by the
76
- pipeline guard: `plugins/sp/scripts/task-evidence-precheck.ts` parses the task
77
- content for an exact `evidence-channel: history_tool_call.args_raw[pi]`
75
+ - [ ] The evidence-channel precheck (0726 R2, folded into the pipeline guard by 1002) is
76
+ enforced by `spur task check <wbs> --precheck` — the guard command itself. It parses the
77
+ task content for an exact `evidence-channel: history_tool_call.args_raw[pi]`
78
78
  declaration and, when present, counts live pi rows with `args_raw` on
79
- `.spur/spur.db` via bun:sqlite. Tasks without a declaration pass without opening
80
- SQLite; unknown declarations, a missing database/table, and a zero count write
81
- FAIL. Both precheck guard conjuncts (`precheck-size.status` and
82
- `precheck-evidence.status`) must read PASS — tasks declaring a live-data
83
- evidence channel must import real history (safe importer, non-dry-run) before
84
- implementation begins.
79
+ `.spur/spur.db` via the domain read (fail-closed: unknown declarations, a missing
80
+ database/table, and a zero count error). Tasks without a declaration pass without
81
+ opening the database; size limits (10 R-items / 16 plan items) are checked in the
82
+ same call. Tasks declaring a live-data evidence channel must import real history
83
+ (safe importer, non-dry-run) before implementation begins.
85
84
 
86
85
  ## review gate
87
86
 
@@ -102,12 +101,13 @@ Entered before `task-pipeline.yaml` `review` state dispatches `sp:code-verificat
102
101
  Entered before `task-pipeline.yaml` `verify` state produces a task verdict.
103
102
 
104
103
  - [ ] The verify answer file (`.spur/run/<wbs>-verify-answer.txt`) is lint-clean before
105
- verdict derivation: `plugins/sp/scripts/verify-answer-lint.ts <wbs>` (0726 R3)
106
- rejects missing/duplicate/unknown R IDs, AC identities that are not an exact task
107
- checklist label or linked-feature scenario title, invalid status/evidence-type
108
- values, and empty evidence on any row. A lint failure fails the verify
109
- step (fail-closed) before `spur task verdict` runs.
110
- - [ ] `spur task check <wbs> --strict-core --json` returns PASS.
104
+ verdict derivation: `spur task verdict <wbs> --from-answer …` (0726 R3) lints the
105
+ file first and rejects missing/duplicate/unknown R IDs, AC identities that are not an
106
+ exact task checklist label or linked-feature scenario title, invalid status/evidence-type
107
+ values, and empty evidence on any row. A lint failure fails the verdict step
108
+ (fail-closed) before verdict derivation; an unresolvable task file fails open so
109
+ pre-lint historical runs (8001–8005) stay reproducible.
110
+ - [ ] `spur task check <wbs> --as done --json` returns PASS.
111
111
  - [ ] Every AC scenario has a corresponding verify command that exited 0.
112
112
  - [ ] The `## Solution` section is filled (not the placeholder comment).
113
113
  - [ ] The `## Testing` evidence (commands run + outcomes) is present in the verdict artifact —
@@ -143,19 +143,19 @@ bounded retry does not.
143
143
  ### The three `testing → done` gate layers
144
144
 
145
145
  The CLI verdict-artifact check runs first. The lifecycle adapter then checks provenance, Review L3,
146
- and finally the workflow's strict-core shell guard. The table groups the two complementary
147
- strict-core/verdict checks as one defense-in-depth layer even though they bracket the adapter checks.
146
+ and finally the workflow's `task check --as done` shell guard. The table groups the two complementary
147
+ done-row/verdict checks as one defense-in-depth layer even though they bracket the adapter checks.
148
148
  The first denial wins; each denial names its own remediation. In verify-0293, the artifact check
149
149
  passed, so provenance denied first and Review L3 denied on the retry.
150
150
 
151
151
  | # | Gate layer | Triggers denial when | Remediation |
152
152
  |---|------------|----------------------|-------------|
153
- | 1 | **Strict-core + verdict artifact** (`spur task check <wbs> --strict-core` + `done-transition-guard.ts`) | The strict-core check fails, or `.spur/run/<wbs>-verdict.json` is **missing** or has a non-PASS aggregate. **Missing artifact is a deny** (not a silent allow — closes the 0349 "done without verdict" class). The aggregate is recomputed from requirement/AC rows; the harsher of stored and computed wins. | Re-run `/sp:dev-verify <wbs>` until PASS (writes the artifact), or explicitly override with `spur task update <wbs> done --force-done --reason "<why>"`. Docs-only procedures meet the same layer: read-only measured verification
153
+ | 1 | **Done-row check + verdict artifact** (`spur task check <wbs> --as done` + `done-transition-guard.ts`) | The done-row check fails, or `.spur/run/<wbs>-verdict.json` is **missing** or has a non-PASS aggregate. **Missing artifact is a deny** (not a silent allow — closes the 0349 "done without verdict" class). The aggregate is recomputed from requirement/AC rows; the harsher of stored and computed wins. | Re-run `/sp:dev-verify <wbs>` until PASS (writes the artifact), or explicitly override with `spur task update <wbs> done --force-done --reason "<why>"`. Docs-only procedures meet the same layer: read-only measured verification
154
154
  (answer file + `spur task verdict`) writes the standard `.spur/run/<wbs>-verdict.json` artifact
155
155
  under proof-input digest bracketing; missing or non-PASS evidence is a refusal, never a synthetic
156
156
  PASS stub. |
157
157
  | 2 | **Provenance guard** (`lifecycle-adapter.ts`) | No pipeline-kind run link exists for `<wbs>`. | Run `/sp:dev-run <wbs>` through the full pipeline, use `/sp:dev-run <wbs> --mode implement --auto --next` for the explicit step chain, or record the audited bypass with `--provenance-bypass` on `spur task update`. |
158
- | 3 | **Review L3** (`task-check.ts`) | `### Review` is empty, placeholder-only, or lacks a populated P1–P4 findings table. | Run `/sp:dev-review <wbs>`; verify cannot write Review because of the Step 10 prohibition above. |
158
+ | 3 | **Review L3** (`task-check.ts`) | `### Review` is empty, placeholder-only, or lacks a populated P1–P4 findings table. | Run `/sp:dev-review --tasks <wbs>`; verify cannot write Review because of the Step 10 prohibition above. |
159
159
 
160
160
  When the verdict is **PARTIAL/FAIL**, or any gate layer fails: stop as review-pending — surface
161
161
  the verdict (or the gate's blocking finding), leave the task at its current status, do NOT
@@ -33,7 +33,7 @@ before any processing); the `## Requirement inventory` items trace back to it.
33
33
  <one-paragraph refined statement of what the idea actually requires — the "real requirement" after discovery sharpens the vague input>
34
34
 
35
35
  ## Requirement inventory
36
- <mandatory — the coverage gate (idea-coverage-check) parses this section, so keep the exact `- I<n> — ` item form>
36
+ <mandatory — the coverage gate (`feature check --inventory`) parses this section, so keep the exact `- I<n> — ` item form>
37
37
  - I1 — <requirement stated as an ask, quoting or paraphrasing the source line from the run's idea-input artifact> (source: "<quoted fragment from the operator's idea>")
38
38
  - I2 — <next requirement>
39
39
  - I<n> — <optional: a requirement explicitly out of scope> [deferred: <reason>]
@@ -68,6 +68,7 @@ Score guide:
68
68
 
69
69
  ## Recommendation
70
70
  <proceed | reshape | drop> — <one-line rationale linking scores, premises, and pros/cons>
71
+ <!-- the first line under this heading MUST start with exactly one of proceed / reshape / drop — `--auto` routes on it -->
71
72
 
72
73
  Stakes: <plain-English cost of proceeding vs not; reversibility; blast radius>
73
74
 
@@ -83,8 +84,8 @@ Stakes: <plain-English cost of proceeding vs not; reversibility; blast radius>
83
84
  |------|--------|
84
85
  | Filled instance path | `.spur/run/idea-eval-report.md` |
85
86
  | Template home | this file |
86
- | Requirement inventory | mandatory `## Requirement inventory` section (0887 R3); consumed by `idea-coverage-check` (R4) |
87
+ | Requirement inventory | mandatory `## Requirement inventory` section (0887 R3); consumed by the `feature check --inventory` coverage gate (R4) |
87
88
  | HITL state | `idea-eval` in `idea-pipeline.yaml` |
88
89
  | Approve | continue → `feature-create` |
89
90
  | Reject / cancel | → `cancelled` (no feature) |
90
- | `--auto` | still pauses unless taste pre-cleared (`--approve-taste` → `idea_approved=true`) |
91
+ | `--auto` (`idea_approved=true`) | `proceed`/`reshape` → `feature-create`; `drop` → `cancelled`; missing/unparseable recommendation → pauses |
@@ -2,7 +2,7 @@
2
2
  name: inline-pipeline-driver
3
3
  description: "Interactive host-session interpreter for Spur state-machine pipelines: execute the existing FSM without a workflow agent subprocess while preserving actions, guards, artifacts, and provenance."
4
4
  owner: spur-dev-maintainers
5
- retirement-criterion: "The per-task interpreter retires once the engine covers per-task execution for /sp:dev-runall with real terminal runs and the parity check (plugins/sp/scripts/inline-pipeline-parity-check.ts) is green (D8 decision D7). Batch orchestration wrapper may remain."
5
+ retirement-criterion: "The per-task interpreter retires once the engine covers per-task execution for /sp:dev-runall with real terminal runs and the parity check (scripts/commands/inline-pipeline-parity-check.ts) is green (D8 decision D7). Batch orchestration wrapper may remain."
6
6
  see_also:
7
7
  - spur-dev
8
8
  - execution-workflow
@@ -18,7 +18,7 @@ see_also:
18
18
  ## Supported action and guard set (0755 R2 parity contract)
19
19
 
20
20
  The action and guard kinds this driver implements. The parity check
21
- (`plugins/sp/scripts/inline-pipeline-parity-check.ts`) compares this set against
21
+ (`scripts/commands/inline-pipeline-parity-check.ts`) compares this set against
22
22
  the resolved actions and guards of every `.spur/workflows/*.yaml`; any element present
23
23
  in one and absent in the other fails the check. Add a new kind here when the driver
24
24
  implements it; remove the entry when the corresponding kind is dropped from the YAML.
@@ -201,7 +201,7 @@ Expected artifacts per stage (all run-scoped under `.spur/run/<run-id>-*`):
201
201
  | start | `-idea-input.md` (verbatim idea), `-idea-precheck-doctor.status` |
202
202
  | discovery | `-idea-eval-report.md` (with `## Requirement inventory`), `-idea-needs-design.json` |
203
203
  | feature-create | `-idea-feature-id.txt`, `-idea-goal.md`, `-idea-scope.md` |
204
- | ac-generate | `-idea-ac-content.md`, `-idea-ac-check.status`, `-idea-coverage.status` |
204
+ | ac-generate | `-idea-ac-content.md`, `-idea-ac-check.status` |
205
205
  | system-design | `-idea-design-review.md`, `-idea-design-check.status` |
206
206
  | decompose | `-idea-task-batch.json`, `-idea-task-order.json` |
207
207
  | batch-create-run | `-idea-batch-create-result.json`, `-idea-batch-create.done`/`.failed` |
@@ -247,6 +247,13 @@ Action semantics come from the YAML and the workflow action contract:
247
247
  actions/guards.
248
248
  - `hitl.confirm` — under `profile=auto`, follow the YAML's auto-skip transition. Otherwise pause,
249
249
  surface the prompt, and resume from the same state with the operator's answer.
250
+ **Host-session rendering:** ask the gate as ONE `AskUserQuestion`
251
+ [decision brief](decision-brief.md), never a typed yes/no. Read the state's evidence (the
252
+ artifacts its prompt names and the recorded `.status` files) and derive the recommendation;
253
+ map options to the `yes` / `no` / `cancel` answers the guards route on, recommended option
254
+ first, each with a one-line reason from that evidence. When a `no` needs operator feedback
255
+ (design-approval's `## Operator feedback`), take it from the operator's notes/"Other" text
256
+ and write it into the named artifact yourself before resuming. Never ask them to edit the file.
250
257
  - `hitl.input` — pause, surface the declared prompt (the agent's operator question, 0933), and
251
258
  resume from the same state with the operator's answer written into the declared var (default
252
259
  `__hitlInput`); the subsequent guards route on answer presence exactly as the engine does.
@@ -500,6 +507,25 @@ The driver reaches it through the existing run delegate (`$SETUP_SCRIPT`,
500
507
  (0887 R8), so `completed_at − started_at == duration_ms` exactly; a back-date failure is
501
508
  recorded (`action.backdate`) and never affects the run.
502
509
 
510
+ - **A state with several actions (1007 R5)** — emit the whole state's boundaries in one call
511
+ instead of one `--action` invocation per action. Write a JSON array
512
+ (`[{node,kind,status,ok,durationMs}, …]` — same fields the `--action` flags carry) to a temp
513
+ file and pass it with `--actions-file`:
514
+
515
+ ```bash
516
+ bun "$SETUP_SCRIPT" --actions-file <actions.json> --run-id "$RUN_ID"
517
+ ```
518
+
519
+ Every row is recorded through the same writer as `--action` (one `action_runs` row per entry);
520
+ the batch is validated in full before the first write, so an invalid row or unreadable file
521
+ exits `1` with `{"ok":false}` and leaves **no** partial rows — fix the batch and re-emit. On
522
+ success it prints `{"ok":true,"runId":…,"recorded":<n>}` and exits `0`. Row emission stays
523
+ best-effort exactly like `--action`: if a row's write fails mid-batch, the failure is recorded
524
+ to the run record and the call reports `{"ok":false,…,"error":…}` but still exits `0` — the run
525
+ continues; never retry the batch or backfill by hand. `--actions-file` is exclusive with the
526
+ other mode flags (`--action`, `--decide`, `--close`, …): mixing them is a usage error (exit 2),
527
+ and a mixed call must be corrected, not silently split.
528
+
503
529
  - **A `decide` action (0941)** — the driver never executes the DecisionMaker itself; it delegates
504
530
  to the same app runner the engine registers, which writes the resultFile row (schemaVersion 1)
505
531
  and returns the decision, then the delegate records the `action_runs` row (`kind=decide`)
@@ -105,7 +105,7 @@ rows stay per file. Flagged rows wait for the operator's answer.
105
105
 
106
106
  ## Workflow step profile and cache-window flags
107
107
 
108
- The step profile (`plugins/sp/scripts/workflow-step-profile`, ADR-065 plugin entrypoint) reads
108
+ The step profile (`plugins/sp/scripts/workflow-step-profile`, ADR-065 plugin entrypoint whose aggregation core lives in `packages/app/src/workflow/step-profile.ts`) reads
109
109
  `spur workflow trace` for a workflow's last N completed, non-dry runs. Per node and action kind it
110
110
  reports run count, executions, p50 and max `durationMs`, p50 idle gap before the step, session mode
111
111
  (`fresh`, `resumed` or `mixed`) and `cacheHit` p50 with its coverage — satellite §10 step evidence.
@@ -73,8 +73,9 @@ codebase (or a named module tree) for **shallow modules and deepening opportunit
73
73
  them as candidates for the planning half. This is a *generator*, not a fixer — it never refactors;
74
74
  it produces a ranked candidate report an operator can turn into a task.
75
75
 
76
- **Not `/sp:dev-review`.** `/sp:dev-review` is a per-task DIFF review (a WBS, forward, findings written
77
- to the task's `## Review`, backed by `sp:code-verification`). The survey has no WBS and no diff — it
76
+ **Not `/sp:dev-review`.** `/sp:dev-review` is a DIFF review of a task set (`--tasks`/`--feature`) or an advisory `--scope` path review
77
+ (task targets: forward, findings written to each task's `## Review`; paths: advisory report, no task
78
+ mutation; backed by `sp:code-verification`). The survey has no task set and no diff — it
78
79
  audits the standing codebase and feeds the planning half. Folding it into `dev-review` would overload
79
80
  that verb and pollute `code-verification` with a codebase scanner; it earns its own operation here.
80
81