@gobing-ai/spur 0.3.90 → 0.3.92

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (130) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/config/pipeline-budgets.json +9 -2
  3. package/config/plugin-scripts.json +17 -1
  4. package/config/workflows/decision-routing-example.yaml +134 -0
  5. package/config/workflows/feature-lifecycle.yaml +14 -5
  6. package/config/workflows/feature-verification.yaml +35 -27
  7. package/config/workflows/history-anatomy.yaml +14 -25
  8. package/config/workflows/idea-pipeline.yaml +17 -5
  9. package/config/workflows/task-pipeline.yaml +187 -4
  10. package/config/workflows/wayfinder-resolution.yaml +2 -0
  11. package/config/workflows/wrapup-pipeline.yaml +49 -5
  12. package/package.json +9 -9
  13. package/plugins/sp/README.md +1 -0
  14. package/plugins/sp/agents/expert-spur.md +2 -2
  15. package/plugins/sp/agents/super-planner.md +14 -5
  16. package/plugins/sp/commands/dev-dogfood.md +4 -4
  17. package/plugins/sp/commands/dev-fixall.md +8 -5
  18. package/plugins/sp/commands/dev-run.md +6 -0
  19. package/plugins/sp/commands/dev-runall.md +12 -6
  20. package/plugins/sp/commands/dev-verify.md +9 -0
  21. package/plugins/sp/commands/dev-verifyall.md +5 -0
  22. package/plugins/sp/lib/idea-handoff.generated.mjs +301 -300
  23. package/plugins/sp/lib/inline-run.generated.d.mts +17 -0
  24. package/plugins/sp/lib/inline-run.generated.mjs +1460 -0
  25. package/plugins/sp/plugin.json +1 -1
  26. package/plugins/sp/references/environment-lens.md +1 -1
  27. package/plugins/sp/scripts/dogfood-testing/validate-report.mjs +136 -0
  28. package/plugins/sp/scripts/dogfood-testing/validate-report.ts +196 -2
  29. package/plugins/sp/scripts/feature-verification-steps.mjs +174 -0
  30. package/plugins/sp/scripts/feature-verification-steps.ts +275 -0
  31. package/plugins/sp/scripts/history-anatomy-cache.mjs +104 -4
  32. package/plugins/sp/scripts/history-anatomy-cache.ts +137 -13
  33. package/plugins/sp/scripts/inline-pipeline-parity-check.ts +2 -0
  34. package/plugins/sp/scripts/inline-run-setup.mjs +349 -0
  35. package/plugins/sp/scripts/inline-run-setup.ts +192 -75
  36. package/plugins/sp/scripts/record-feature-sync.mjs +63 -0
  37. package/plugins/sp/scripts/record-feature-sync.ts +84 -0
  38. package/plugins/sp/scripts/residual-scan.mjs +476 -0
  39. package/plugins/sp/scripts/residual-scan.ts +614 -0
  40. package/plugins/sp/scripts/surface-drift-inventory.ts +71 -6
  41. package/plugins/sp/scripts/task-evidence-precheck.ts +8 -3
  42. package/plugins/sp/scripts/task-size-precheck.ts +8 -3
  43. package/plugins/sp/scripts/validate-flag-contracts.ts +3 -3
  44. package/plugins/sp/skills/branch-workflow/SKILL.md +1 -0
  45. package/plugins/sp/skills/branch-workflow/references/worktree-patterns.md +2 -0
  46. package/plugins/sp/skills/code-implementation/SKILL.md +17 -0
  47. package/plugins/sp/skills/code-verification/SKILL.md +21 -0
  48. package/plugins/sp/skills/code-verification/references/verdict-schema.md +1 -0
  49. package/plugins/sp/skills/dogfood-testing/SKILL.md +5 -3
  50. package/plugins/sp/skills/dogfood-testing/references/monitor-ledger.md +63 -26
  51. package/plugins/sp/skills/dogfood-testing/references/report-template.md +33 -10
  52. package/plugins/sp/skills/history-anatomy/references/modes.md +5 -3
  53. package/plugins/sp/skills/next-feature/references/ranking-rubric.md +1 -1
  54. package/plugins/sp/skills/next-router/references/routing-table.md +7 -0
  55. package/plugins/sp/skills/parallel-execution/references/dispatch-surface.md +1 -1
  56. package/plugins/sp/skills/session-review/SKILL.md +12 -2
  57. package/plugins/sp/skills/spur-cli/references/agent.md +11 -2
  58. package/plugins/sp/skills/spur-cli/references/features.md +1 -1
  59. package/plugins/sp/skills/spur-cli/references/projects.md +3 -1
  60. package/plugins/sp/skills/spur-cli/references/self.md +5 -1
  61. package/plugins/sp/skills/spur-cli/references/workflows.md +42 -20
  62. package/plugins/sp/skills/spur-dev/SKILL.md +11 -4
  63. package/plugins/sp/skills/spur-dev/references/cross-cutting.md +25 -3
  64. package/plugins/sp/skills/spur-dev/references/dev-operations.md +2 -2
  65. package/plugins/sp/skills/spur-dev/references/execution-batch.md +285 -57
  66. package/plugins/sp/skills/spur-dev/references/execution-workflow.md +14 -5
  67. package/plugins/sp/skills/spur-dev/references/flag-glossary.md +19 -4
  68. package/plugins/sp/skills/spur-dev/references/glossary.md +9 -1
  69. package/plugins/sp/skills/spur-dev/references/inline-pipeline-driver.md +51 -16
  70. package/plugins/sp/skills/spur-dev/references/planning-workflow.md +5 -0
  71. package/plugins/sp/skills/wayfinder/SKILL.md +1 -1
  72. package/spur.js +23422 -19364
  73. package/web/_astro/BoardApp.CerSBgis.js +192 -0
  74. package/web/_astro/BoardApp.eoTz0pZs.js +1 -0
  75. package/web/_astro/{TaskDetail.DwTmbQp5.js → TaskDetail.DCqiC-OZ.js} +1 -1
  76. package/web/_astro/arc.CPwg6Rw0.js +1 -0
  77. package/web/_astro/{architectureDiagram-3BPJPVTR.jvdDahWM.js → architectureDiagram-3BPJPVTR.DM_vp_hO.js} +1 -1
  78. package/web/_astro/{blockDiagram-GPEHLZMM.zSg4AmFD.js → blockDiagram-GPEHLZMM.DXVIiv0p.js} +1 -1
  79. package/web/_astro/{c4Diagram-AAUBKEIU.BkUIUQWH.js → c4Diagram-AAUBKEIU.BbF_zCxW.js} +1 -1
  80. package/web/_astro/channel.MYZLKNwy.js +1 -0
  81. package/web/_astro/{chunk-2J33WTMH.DvfQ_f50.js → chunk-2J33WTMH.CAgQHpPC.js} +1 -1
  82. package/web/_astro/{chunk-4BX2VUAB.DuI4gQqX.js → chunk-4BX2VUAB.BN-5tpw4.js} +1 -1
  83. package/web/_astro/{chunk-55IACEB6.D3BWBOpF.js → chunk-55IACEB6.CnPkEEr0.js} +1 -1
  84. package/web/_astro/{chunk-727SXJPM.3QSi0a9M.js → chunk-727SXJPM.BQzQeMVm.js} +4 -4
  85. package/web/_astro/{chunk-AQP2D5EJ.xazCQrAF.js → chunk-AQP2D5EJ.B6xNyDnL.js} +1 -1
  86. package/web/_astro/{chunk-FMBD7UC4.B2g6u4rA.js → chunk-FMBD7UC4.C7f9Ih78.js} +1 -1
  87. package/web/_astro/{chunk-ND2GUHAM.wWwWs99t.js → chunk-ND2GUHAM.CNV1dFXT.js} +1 -1
  88. package/web/_astro/{chunk-QZHKN3VN.BD5g3qa9.js → chunk-QZHKN3VN.Cudn2TkJ.js} +1 -1
  89. package/web/_astro/{classDiagram-4FO5ZUOK.C7CzCdsX.js → classDiagram-4FO5ZUOK.D1NwP50q.js} +1 -1
  90. package/web/_astro/{classDiagram-v2-Q7XG4LA2.C7CzCdsX.js → classDiagram-v2-Q7XG4LA2.D1NwP50q.js} +1 -1
  91. package/web/_astro/{cose-bilkent-S5V4N54A.Xyiau0gw.js → cose-bilkent-S5V4N54A.B1wSL-Xb.js} +1 -1
  92. package/web/_astro/{cynefin-OW5HDTMX.BeC5MWas.js → cynefin-OW5HDTMX.BmK52w8G.js} +1 -1
  93. package/web/_astro/{dagre-BM42HDAG.yZbMN9vc.js → dagre-BM42HDAG.Bfy5CTDT.js} +2 -2
  94. package/web/_astro/diagram-2AECGRRQ.DhNnvUvX.js +43 -0
  95. package/web/_astro/diagram-5GNKFQAL.lTX5KwnS.js +10 -0
  96. package/web/_astro/{diagram-KO2AKTUF.CR6k3Y3G.js → diagram-KO2AKTUF.CW_vMJ4z.js} +3 -3
  97. package/web/_astro/{diagram-LMA3HP47.x7mwu8jz.js → diagram-LMA3HP47.B_8ZGF67.js} +1 -1
  98. package/web/_astro/{diagram-OG6HWLK6.D8aTTvUr.js → diagram-OG6HWLK6.BppnHsdS.js} +1 -1
  99. package/web/_astro/{erDiagram-TEJ5UH35.BoBqcKXQ.js → erDiagram-TEJ5UH35.BEuHXcjJ.js} +5 -5
  100. package/web/_astro/{flowDiagram-I6XJVG4X.D3mTQdrU.js → flowDiagram-I6XJVG4X.CH-UlnGr.js} +4 -4
  101. package/web/_astro/{ganttDiagram-6RSMTGT7.H-cqgIh-.js → ganttDiagram-6RSMTGT7.BO81S85v.js} +1 -1
  102. package/web/_astro/{gitGraphDiagram-PVQCEYII.B6s9zbfC.js → gitGraphDiagram-PVQCEYII.XnPxPPZN.js} +1 -1
  103. package/web/_astro/index.Hjbr15fG.css +1 -0
  104. package/web/_astro/{infoDiagram-5YYISTIA.BzgCoV6P.js → infoDiagram-5YYISTIA.JyjYRu_T.js} +1 -1
  105. package/web/_astro/{ishikawaDiagram-YF4QCWOH.BZzVhy1-.js → ishikawaDiagram-YF4QCWOH.BBRBF-Fo.js} +5 -5
  106. package/web/_astro/{journeyDiagram-JHISSGLW.BV3195Py.js → journeyDiagram-JHISSGLW.C_iymSyp.js} +1 -1
  107. package/web/_astro/{kanban-definition-UN3LZRKU.BjRd2DWz.js → kanban-definition-UN3LZRKU.DdfW-Oqt.js} +7 -7
  108. package/web/_astro/{linear.BILTgS5N.js → linear.C2_IkbZT.js} +1 -1
  109. package/web/_astro/mermaid.core.GAOYeSR0.js +303 -0
  110. package/web/_astro/{mindmap-definition-RKZ34NQL.BiEjaI4-.js → mindmap-definition-RKZ34NQL.DAZIxQSK.js} +2 -2
  111. package/web/_astro/{pieDiagram-4H26LBE5.i_8V5pIn.js → pieDiagram-4H26LBE5.CN8sIhKM.js} +3 -3
  112. package/web/_astro/{quadrantDiagram-W4KKPZXB.BWaW3MHn.js → quadrantDiagram-W4KKPZXB.3dGcX5GP.js} +1 -1
  113. package/web/_astro/{requirementDiagram-4Y6WPE33.CzddBbtg.js → requirementDiagram-4Y6WPE33.BV2y4dd6.js} +3 -3
  114. package/web/_astro/{sankeyDiagram-5OEKKPKP.X2ww0e-D.js → sankeyDiagram-5OEKKPKP.Cqo15Tvo.js} +4 -4
  115. package/web/_astro/{sequenceDiagram-3UESZ5HK.DSA4kTcc.js → sequenceDiagram-3UESZ5HK.CROCPMJB.js} +1 -1
  116. package/web/_astro/{stateDiagram-AJRCARHV.D0DtFSpR.js → stateDiagram-AJRCARHV.RfXZrkFE.js} +1 -1
  117. package/web/_astro/{stateDiagram-v2-BHNVJYJU.BfQq0zQv.js → stateDiagram-v2-BHNVJYJU.CPXmbBs9.js} +1 -1
  118. package/web/_astro/{timeline-definition-PNZ67QCA.Dmlrgi1m.js → timeline-definition-PNZ67QCA.DdgKTiO8.js} +3 -3
  119. package/web/_astro/{vennDiagram-CIIHVFJN.D5mpl00Z.js → vennDiagram-CIIHVFJN.CPNVSHF1.js} +5 -5
  120. package/web/_astro/{wardleyDiagram-YWT4CUSO.Df4BdzO4.js → wardleyDiagram-YWT4CUSO.CQhA0Jyr.js} +3 -3
  121. package/web/_astro/{xychartDiagram-2RQKCTM6.DiTRreKN.js → xychartDiagram-2RQKCTM6.n61BWyy4.js} +1 -1
  122. package/web/index.html +2 -2
  123. package/web/_astro/BoardApp.CDUcHlTJ.js +0 -188
  124. package/web/_astro/BoardApp.CaCGU_uX.js +0 -1
  125. package/web/_astro/arc.BzF71EFI.js +0 -1
  126. package/web/_astro/channel.SSVY0JPQ.js +0 -1
  127. package/web/_astro/diagram-2AECGRRQ.Cmo2zQM-.js +0 -43
  128. package/web/_astro/diagram-5GNKFQAL.D033eSVi.js +0 -10
  129. package/web/_astro/index.CcU5weKX.css +0 -1
  130. package/web/_astro/mermaid.core.DBy_WKeW.js +0 -301
@@ -346,6 +346,67 @@ export function checkNounVerbFlags(
346
346
  }
347
347
  }
348
348
 
349
+ // ─── Semantic operand layer (I7 / task 0906 R2) ─────────────────────────────
350
+
351
+ /**
352
+ * Metadata-only operands the I6 sweep classified. This is the audit boundary, not a
353
+ * CLI section matrix: a key that is ALSO a genuine body heading for the noun (feature
354
+ * `Scope`) is accepted per-noun in checkSectionOperand; anything outside this set is
355
+ * not statically classifiable and stays unverified.
356
+ */
357
+ const SECTION_METADATA_KEYS = new Set([
358
+ 'tags',
359
+ 'priority',
360
+ 'status',
361
+ 'phase',
362
+ 'id',
363
+ 'parent',
364
+ 'name',
365
+ 'owner',
366
+ 'scope',
367
+ ]);
368
+
369
+ /** Body headings that overlap the swept metadata keys, per noun (case-insensitive). */
370
+ const NOUN_BODY_SECTIONS: Record<string, ReadonlySet<string>> = {
371
+ feature: new Set(['scope']),
372
+ };
373
+
374
+ /**
375
+ * Semantic operand check (I7 / task 0906 R2): `--section` is a real flag on
376
+ * task/feature update, so existence parity passes while the operand names a
377
+ * metadata-only key instead of a body section. Classify literal operands only —
378
+ * dynamic operands (placeholders) stay unverified under the existing scanner
379
+ * convention, and quoted/equal forms are handled alongside the spaced form.
380
+ * Comparison is case-normalized; the original operand text stays in the row evidence.
381
+ */
382
+ export function checkSectionOperand(spanRaw: string, occ: { file: string; line: number }): void {
383
+ const parsed = parseInvocation(spanRaw);
384
+ if (!parsed) return;
385
+ const noun = parsed.nouns[0];
386
+ if ((noun !== 'task' && noun !== 'feature') || !parsed.verbs.includes('update')) return;
387
+ for (const m of spanRaw.matchAll(/--section(?:=|\s+)("([^"]*)"|'([^']*)'|[^\s]+)/g)) {
388
+ const operand = (m[2] ?? m[3] ?? m[1] ?? '').replace(/[.,;:]+$/, '');
389
+ if (!operand || PLACEHOLDER.test(operand)) continue; // dynamic — unverified by convention
390
+ const key = operand.toLowerCase();
391
+ if ((NOUN_BODY_SECTIONS[noun] ?? new Set()).has(key)) continue; // genuine body heading
392
+ if (!SECTION_METADATA_KEYS.has(key)) continue; // outside the audit boundary — unverified
393
+ // P3 (0906 review): feature update owns the generic --field/--value pair; task update
394
+ // exposes metadata only via dedicated flags (no --field/--value, no --tags), so the hint
395
+ // must not suggest a pair that would fail with VALIDATION_FAILED on task rows.
396
+ const remediation =
397
+ noun === 'feature'
398
+ ? `use --field ${operand} --value <value>`
399
+ : `task update exposes metadata only via dedicated flags (e.g. --priority <value>), not --section`;
400
+ record(
401
+ `spur ${noun} update --section ${operand}`,
402
+ 'semantic-operand(I6 metadata keys vs noun body sections)',
403
+ 'mismatch',
404
+ `--section writes a body section; "${operand}" is a metadata-only ${noun} operand — ${remediation}`,
405
+ occ,
406
+ );
407
+ }
408
+ }
409
+
349
410
  export function sweepPluginTrees(root: string = PLUGIN_ROOT): void {
350
411
  const files = [
351
412
  ...walk(join(root, 'commands'), ['.md']),
@@ -398,6 +459,7 @@ export function sweepPluginTrees(root: string = PLUGIN_ROOT): void {
398
459
  continue;
399
460
  }
400
461
  checkNounVerbFlags(parsed.nouns, parsed.verbs, parsed.flags, occ);
462
+ checkSectionOperand(span, occ);
401
463
  }
402
464
  if (file.endsWith('.ts')) return;
403
465
  if (inVerbTable && refNoun && /^\|/.test(line)) {
@@ -430,7 +492,10 @@ export function sweepPluginTrees(root: string = PLUGIN_ROOT): void {
430
492
  } else {
431
493
  for (const span of lineInvocationSpans(line)) {
432
494
  const parsed = parseInvocation(span);
433
- if (parsed) checkNounVerbFlags(parsed.nouns, parsed.verbs, parsed.flags, occ);
495
+ if (parsed) {
496
+ checkNounVerbFlags(parsed.nouns, parsed.verbs, parsed.flags, occ);
497
+ checkSectionOperand(span, occ);
498
+ }
434
499
  }
435
500
  }
436
501
  });
@@ -764,11 +829,11 @@ export function sweepWorkflows(opts: { run?: CliRunner; wfDir?: string; link?: s
764
829
  .forEach((line: string, i: number) => {
765
830
  for (const span of lineInvocationSpans(line)) {
766
831
  const parsed = parseInvocation(span);
767
- if (parsed)
768
- checkNounVerbFlags(parsed.nouns, parsed.verbs, parsed.flags, {
769
- file: path,
770
- line: i + 1,
771
- });
832
+ if (parsed) {
833
+ const occ = { file: path, line: i + 1 };
834
+ checkNounVerbFlags(parsed.nouns, parsed.verbs, parsed.flags, occ);
835
+ checkSectionOperand(span, occ);
836
+ }
772
837
  }
773
838
  });
774
839
  }
@@ -74,10 +74,15 @@ function parseArgs(argv: string[]): { wbs: string; spurBin: string } {
74
74
  if (arg === '--spur-bin') {
75
75
  spurBin = argv[i + 1] ?? defaultSpurBin();
76
76
  i += 2;
77
- } else if (!arg.startsWith('--')) {
78
- wbs = arg;
79
- i++;
77
+ } else if (arg.startsWith('-')) {
78
+ // 0948 R4: an unknown flag is a mis-invocation, not something to swallow.
79
+ // The old `else { i++; }` let `script 0926 --task-file x.md` run against x.md.
80
+ console.error(`task-evidence-precheck: unknown flag: ${arg}`);
81
+ usage();
80
82
  } else {
83
+ // 0948 R4: first positional wins — a later positional must never overwrite
84
+ // `wbs` (it used to write a garbage-named status file and mask the exit code).
85
+ if (wbs === '') wbs = arg;
81
86
  i++;
82
87
  }
83
88
  }
@@ -124,10 +124,15 @@ function parseArgs(argv: string[]): {
124
124
  } else if (arg === '--max-plan-items') {
125
125
  maxPlanItems = Number(argv[i + 1]) || 16;
126
126
  i += 2;
127
- } else if (!arg.startsWith('--')) {
128
- wbs = arg;
129
- i++;
127
+ } else if (arg.startsWith('-')) {
128
+ // 0948 R4: an unknown flag is a mis-invocation, not something to swallow.
129
+ // The old `else { i++; }` let `script 0926 --task-file x.md` run against x.md.
130
+ console.error(`task-size-precheck: unknown flag: ${arg}`);
131
+ usage();
130
132
  } else {
133
+ // 0948 R4: first positional wins — a later positional must never overwrite
134
+ // `wbs` (it used to write a garbage-named status file and mask the exit code).
135
+ if (wbs === '') wbs = arg;
131
136
  i++;
132
137
  }
133
138
  }
@@ -406,9 +406,9 @@ export function extractTriggerTable(crossCuttingRaw: string): string[] | null {
406
406
  function adrAgentClaims(adrRaw: string): Map<string, SurfaceBehavior> | null {
407
407
  if (adrRaw.includes('## ADR-047')) {
408
408
  const out = new Map<string, SurfaceBehavior>();
409
- // G5 amendment (feature G5 / task 0565): explicit inline is host-session-only — headless
410
- // surfaces reject it with the stable special error; 0508 native-subagent eligibility
411
- // applies to omitted --agent only, never explicit inline.
409
+ // ADR-087 (task 0687) retired the G5 frozen rejection: explicit `inline` on a headless
410
+ // surface resolves via role/tier substitution with one warning. This claims map mirrors
411
+ // the ADR-047 amendment text as written; ADR-087-aware parsing is a follow-up.
412
412
  out.set('inline', {
413
413
  surfaces: new Set(['inline']),
414
414
  conditional: false,
@@ -81,6 +81,7 @@ git push
81
81
  ```bash
82
82
  git branch -d feature/<slug> # Delete merged branch
83
83
  git worktree remove ../<project>-feature-<slug> # Remove worktree if used
84
+ spur projects remove ../<project>-feature-<slug> 2>/dev/null || true # Deregister from projects.json if registered
84
85
  git worktree prune # Clean up stale worktree references
85
86
  ```
86
87
 
@@ -46,6 +46,7 @@ Shows all worktrees with their branches and paths.
46
46
 
47
47
  ```bash
48
48
  git worktree remove ../project-hotfix
49
+ spur projects remove ../project-hotfix 2>/dev/null || true
49
50
  ```
50
51
 
51
52
  ### Prune (clean up stale references)
@@ -83,6 +84,7 @@ After removing a worktree, run `git worktree prune` and `git gc` to reclaim spac
83
84
 
84
85
  ```bash
85
86
  git worktree remove ../old-worktree
87
+ spur projects remove ../old-worktree 2>/dev/null || true
86
88
  git worktree prune
87
89
  git gc --aggressive
88
90
  ```
@@ -52,6 +52,23 @@ When this skill is entered via `/sp:dev-run --mode implement <wbs>` (the form
52
52
  The structural guard is the slash form itself (`--mode implement`). Prose in the workflow YAML
53
53
  `agent.run` `input` is the wrong place for this rule; it belongs here and in `dev-run.md`.
54
54
 
55
+ ## Escalation contract (task 0933, implement step only)
56
+
57
+ When the pipeline runs implement headless, you have a bounded channel to surface a question
58
+ instead of guessing. The rules, in decision order:
59
+
60
+ - **Decide without asking** when the answer follows from the task's frozen Design /
61
+ Requirements / Q&A, the project and global instructions, or the codebase itself. Ambiguity
62
+ you can resolve from evidence is not a question.
63
+ - **Escalate only when a requirement or design ambiguity would change scope, correctness, or
64
+ authorization.** Style preferences, implementation details, and curiosity are not escalations.
65
+ - **To escalate:** write ONE concise question — with options and a recommendation — to the
66
+ escalation file, then exit 0 **without further edits**. The pipeline pauses the run and an
67
+ operator answers; the transcript of prior Q/A pairs lives at the `--escalation-file` path
68
+ (the ask loop is bounded, default 2 — a third pause fails the task).
69
+ - **Read `--escalation-file` if it exists and treat its answers as binding.** They are the
70
+ operator's recorded decisions for this pass, not suggestions.
71
+
55
72
  ## One WBS per implement pass (task 0487 R1)
56
73
 
57
74
  The target WBS is the **only** task you implement. Sibling tasks in the corpus are context you do
@@ -259,13 +259,34 @@ the deterministic Testing writer `spur task record` (section authorship never ha
259
259
 
260
260
  ```bash
261
261
  # write .spur/run/<wbs>-verdict.json (shape in references/verdict-schema.md), then:
262
+ # F96 residual sweep (observe-only): scan + fold BEFORE record, under every --fix mode.
263
+ RESIDUAL=$(superskill script path sp residual-scan.mjs)
264
+ node "$RESIDUAL" scan --wbs <wbs> --base .spur/run/<wbs>-base.sha
265
+ node "$RESIDUAL" fold --wbs <wbs> --verdict .spur/run/<wbs>-verdict.json
262
266
  spur task record <wbs> --verdict-file .spur/run/<wbs>-verdict.json # renders ## Testing
263
267
  ```
264
268
 
269
+ > **Residual fold (F96).** `scan` reads `.spur/run/` artifacts and writes `residuals.json` +
270
+ > `residual-report.md`; `fold` rewrites the **just-written** verdict artifact's `residual-sweep`
271
+ > check in place (PASS → PARTIAL when a blocking residual exists), so `record` transcribes the
272
+ > downgraded verdict — a manual verify cannot certify what the pipeline would reject. Resolve the
273
+ > script via `superskill script path sp residual-scan.mjs`; shipped surfaces never reference
274
+ > `plugins/sp/scripts/` directly (script-contract-check rule 4). When the task reaches `done`
275
+ > through `--next`, run `residual-scan settle` (links follow-up tasks; best-effort).
276
+
265
277
  > **Corrections: the answer file is the source of truth.** `spur task record` re-transcribes
266
278
  > `## Testing` from the verdict artifact — direct `--section Testing` writes are futile. Fix
267
279
  > `.spur/run/<wbs>-verify-answer.txt` → `spur task verdict <wbs> --from-answer <file>` → re-record.
268
280
 
281
+ > **Scenario-key carry-forward (standalone `--force` re-verifies).** When you author a fresh
282
+ > verdict artifact for a task whose `## Testing` already carries feature scenario-title rows
283
+ > (`Scenario: <title>` / `AC-N` keys with MET status), **copy those rows into the new artifact**
284
+ > keyed the same way — record re-transcribes `## Testing` wholesale, so a fresh artifact keyed by
285
+ > bare `R1`-style ids drops the scenario keys `spur feature check` needs for satisfaction
286
+ > (`L4.scenario-unverified` regresses; 0921/D63). Since 0936, `record` warns on stderr (exit 0)
287
+ > for each dropped MET-matched scenario key and when the new rows match no feature scenario —
288
+ > treat those warnings as a re-key instruction, not noise.
289
+
269
290
  > **Do not write `## Review` directly, ever.** The `## Review` section is owned by the
270
291
  > `review` coordinator (`/sp:dev-review` → `sp:super-reviewer`), which merges
271
292
  > `functional-review` + `code-verification` review mode + `code-improvement` fragments. The
@@ -136,6 +136,7 @@ Wave C verification can emit the following additive `checks[]` rows:
136
136
  | `evidence-rule-pass` | All behavior-bearing AC rows had executable evidence or were explicitly non-behavioral. |
137
137
  | `evidence-rule-failed` | One or more MET behavior-bearing AC rows lacked `test` / `command` evidence and were downgraded to PARTIAL. |
138
138
  | `cli-golden-path-present` | CLI-surface tasks supplied, or failed to supply, one golden-path command evidence row. |
139
+ | `residual-sweep` | Post-verdict residual scan (F96): `fail` when blocking leftovers exist (P1–P3 findings, added diff markers, unchecked boxes); evidence lists blocking/deferrable/advisory/housekeeping counts plus blocking and deferrable item ids. A `fail` downgrades an otherwise-PASS verdict to PARTIAL via the fold step. |
139
140
 
140
141
  ## How the gate reads it
141
142
 
@@ -194,8 +194,9 @@ On **every** step resolve:
194
194
 
195
195
  The final report MUST include a `### 3. Monitor Ledger` section containing those rows (cardinality:
196
196
  row count == the declared executed steps). Cardinality, full methodology, column contract,
197
- token/cache estimation, multi-source Cost honesty, the cache-health finding rule, and the
198
- **cache-conservation discipline** live in
197
+ token/context-reuse estimation, multi-source Cost honesty, the evidence-based reuse and
198
+ pipeline-provenance observation rules (supported 0912 baseline findings, owner handoffs,
199
+ INSUFFICIENT_EVIDENCE limits — task 0913), and the **cache-conservation discipline** live in
199
200
  **[monitor-ledger.md](references/monitor-ledger.md)** — apply conservation while monitoring; low
200
201
  cache% is usually the driver re-fetching data it already holds.
201
202
 
@@ -532,7 +533,8 @@ Do NOT:
532
533
  - [references/report-template.md](references/report-template.md) — report section contract,
533
534
  mandatory footer, task-sink L3 rule.
534
535
  - [references/monitor-ledger.md](references/monitor-ledger.md) — live-ledger column contract,
535
- token/cache estimation, cache-health finding rule.
536
+ token/context-reuse estimation, evidence-based reuse observation rule, pipeline-run provenance
537
+ observation rules.
536
538
 
537
539
  ## Platform Notes
538
540
 
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: monitor-ledger
3
- description: "The dogfood monitor methodology + on-disk live ledger column contract + dual-write + token/cache estimation heuristic + the cache-health finding rule. The on-disk ledger is the single source of truth the report is assembled from — recorded live, per step, never reconstructed."
3
+ description: "The dogfood monitor methodology + on-disk live ledger column contract + dual-write + token/cache estimation heuristic + evidence-based context-reuse observation rules. The on-disk ledger is the single source of truth the report is assembled from — recorded live, per step, never reconstructed."
4
4
  see_also:
5
5
  - dogfood-testing
6
6
  - report-template
@@ -99,14 +99,16 @@ in the report's §6 Findings (no exemption applies).
99
99
  | `Finding` | One-line finding surfaced at this step, or `—`. A finding does **not** change `Outcome`. |
100
100
  | `Fresh Tokens` | Estimated fresh context for the step. Prefix with `~`. |
101
101
  | `Cached Tokens` | Estimated reused context for the step. Prefix with `~`. |
102
- | `Cache %` | `Cached Tokens / (Fresh Tokens + Cached Tokens)`, rounded to the nearest whole percent. An `~unknown` row carries `—`, never `0%` — unknown basis is not an observed zero. |
102
+ | `Cache %` | Estimated context-reuse share: `Cached Tokens / (Fresh Tokens + Cached Tokens)`, rounded to the nearest whole percent. This is a **heuristic estimate** of how much context the step reused — it is not observed provider cache behavior (task 0913, R3). An `~unknown` row carries `—`, never `0%` — unknown basis is not an observed zero. |
103
103
  | `Basis` | Observable basis for the estimate: command output, prior file read reused, generated text, etc. |
104
104
  | `Wall-clock` | Elapsed time for the step. |
105
105
 
106
- ## Token + cache estimation heuristic
106
+ ## Token + context-reuse estimation heuristic
107
107
 
108
108
  A skill **cannot read its own exact token meter** — derive an estimate and label every number
109
- `~estimate`. The accepted methodology is deterministic from the ledger rows:
109
+ `~estimate`. These chars/4 figures estimate **context volume and reuse within the driver's own
110
+ session**; they do not measure provider cache hits, cache cost, or realized savings (task 0913,
111
+ R3). The accepted methodology is deterministic from the ledger rows:
110
112
 
111
113
  1. Estimate **Fresh Tokens** from new material consumed or produced by the step:
112
114
  - text read from files or command output: `ceil(characters / 4)`, rounded to the nearest 100;
@@ -122,8 +124,9 @@ A skill **cannot read its own exact token meter** — derive an estimate and lab
122
124
  both sums (or surfaced as a separate unknown bucket), never folded in as `Cached ~0`:
123
125
  `aggregate cache% = round(sum(Cached Tokens) / sum(Fresh Tokens + Cached Tokens) * 100)`.
124
126
 
125
- The **trend across runs** is the signal, not the absolute value: rising cache% = the testee is
126
- reusing context efficiently; falling cache% = context bloat creeping in.
127
+ The **trend across runs** is the signal, not the absolute value: a falling estimated reuse share
128
+ suggests context bloat creeping in; a rising one, efficient reference reuse. Trend claims from
129
+ estimates stay labeled as such — they never become causal waste findings on their own.
127
130
 
128
131
  > Never print a precise token number you cannot substantiate. The numbers exist to show a *trend*,
129
132
  > not to bill anyone.
@@ -169,27 +172,58 @@ missing, mark the row `~unknown`, exclude it from the aggregate (or surface it a
169
172
  bucket), and explain the missing basis — never fold it in as `Cached Tokens = ~0`: unknown cache use
170
173
  is not an observed zero-percent hit rate, and a low-cache diagnosis needs observed data.
171
174
 
172
- ## Cache-health finding rule
173
-
174
- Cache% is the operational signal for testee-tuning:
175
-
176
- - Any **individual step with cache% < 40%** → it is re-reading files or re-sending prompt context
177
- unnecessarily. Emit a **P3** finding naming that step, **even if the step succeeded**.
178
- - A run with **aggregate cache% < 50%** → the testee is a tuning candidate regardless of the
179
- PASS/PARTIAL/FAIL verdict. Emit a **P3** finding: "Low cache hit rate — candidate for
180
- context-window or prompt trimming."
181
-
182
- These feed the report's §6 Findings (see [report-template.md](report-template.md)).
175
+ ## Context-reuse observation rule (task 0913, R3 — replaces the fixed-threshold cache-health rule)
176
+
177
+ The estimated reuse share is a **heuristic** (chars/4 over the driver's own session); it cannot
178
+ prove provider cache behavior, unnecessary re-reading, or realized savings. Threshold crossings
179
+ (40%, 50%, or any other fixed line) therefore **never auto-generate a causal waste finding**.
180
+ Instead:
181
+
182
+ - **Estimated-only evidence** (chars/4 ledger, no external meter): a low estimated reuse share may
183
+ be reported only as a **labeled hypothesis** finding carrying the `[unverifiable]` tag, naming
184
+ the confirmation it needs (per-step provider telemetry, a real meter, or a measured reread
185
+ trace). Example shape: `P3 — estimated reuse share 36%, below the run's own trend — hypothesis:
186
+ repeated task-list refetching; confirm with per-step metering` `[unverifiable]`.
187
+ - **Measured evidence** (ccusage session/day delta, or agent usage fields in tool results): report
188
+ the measured figures with their scope label and confidence, and keep them out of per-step ledger
189
+ cells. A waste finding may then name the measured evidence — still without claiming causality
190
+ the meter does not show.
191
+ - **Unavailable evidence**: print `Meter: n/a` and totals `n/a` rather than a fabricated share;
192
+ route the missing-measurement gap to its owner as an improvement proposal.
193
+
194
+ Unobservable chained-step cost keeps its **mandatory** P3 finding (`P3 — chained-step cost not
195
+ observable`, task 0278 R3) — that rule reports missing evidence, not causal waste, and stays.
196
+
197
+ These findings feed the report's §6 Findings (see [report-template.md](report-template.md)).
198
+
199
+ ## Pipeline-run provenance observations (task 0913, R5)
200
+
201
+ When the testee drives a task pipeline, the driver may adopt the supported findings from the
202
+ 0912 workflow baseline (`docs/reports/i31/0912-workflow-baseline.md`) as bounded observations —
203
+ with artifact anchors and owner handoffs, never as invented performance conclusions:
204
+
205
+ - **Run-row closure / structured emission (supported, pilot-selected):** if the observed pipeline
206
+ run leaves a non-terminal run row or emits no `action_runs` rows, record the observation with
207
+ the run id as anchor and hand off to the owners named in the baseline (driver adoption: D62;
208
+ row-closure defect: P). These are the baseline's pilot-eligible findings (F1/F2).
209
+ - **Performance conclusions (INSUFFICIENT_EVIDENCE in the baseline):** no token/USD, percentage,
210
+ speedup, or fleet claim may be derived from a dogfood run. Record the limitation explicitly
211
+ (what was unmeasured and why) and route the missing evidence to the baseline's named gaps
212
+ (session/cost joins E6; scoped-gate experiment F3/F4; fleet serve session S5).
213
+ - Historical comparison, recurrence, and cohort trends stay with history-anatomy and existing
214
+ doctor tooling — the dogfood report owns only what its own ledger observed.
183
215
 
184
216
  ## Cache-conservation discipline (how to keep cache% high)
185
217
 
186
- The cache-health rule above *detects* waste; this section is the mitigation. The dogfooding driver
218
+ The context-reuse observation rule above interprets the signal; this section is the mitigation. The
219
+ dogfooding driver
187
220
  (the agent running Phase 2/3) controls most of the cache% it later reports — low cache% is usually
188
221
  the driver re-fetching data it already holds. Apply these while monitoring each step:
189
222
 
190
223
  ### Driver cache checklist (task 0278 R7)
191
224
 
192
- When aggregate cache% risks falling under 50%, apply this checklist **before** re-reading:
225
+ When the estimated reuse share looks low (or trends down across runs), apply this checklist
226
+ **before** re-reading:
193
227
 
194
228
  | # | Action | Why |
195
229
  | --- | -------- | ----- |
@@ -198,12 +232,13 @@ When aggregate cache% risks falling under 50%, apply this checklist **before** r
198
232
  | 3 | Prefer `--json` CLI over re-parsing freeform prose | Smaller, stable payloads |
199
233
  | 4 | Dual-write ledger rows without re-reading the whole report each step | Append/patch; don't full-file re-load |
200
234
  | 5 | Skip redundant `bun test` full suite between steps when a focused file suite already green | Run the broad suite once at the end |
201
- | 6 | For batch testees (`verifyall` / `runall` / `refineall`): freeze `task list --json` once at resolve | Re-listing the set per task is the #1 sub-50% cache pattern on feature dogfoods |
202
- | 7 | On re-verify of done tasks: re-read only cited `file:line` anchors, not full Solution blobs | Anchor-first re-verify keeps cache% above the 50% floor |
235
+ | 6 | For batch testees (`verifyall` / `runall` / `refineall`): freeze `task list --json` once at resolve | Re-listing the set per task is the #1 low-reuse-estimate pattern on feature dogfoods |
236
+ | 7 | On re-verify of done tasks: re-read only cited `file:line` anchors, not full Solution blobs | Anchor-first re-verify keeps the reuse estimate honest and high |
203
237
 
204
238
  1. **Reuse CLI output already in context.** If a prior step (or a prior tool call this step)
205
239
  captured `spur task show`/`check`/`list` output, do **not** re-invoke the same command for that
206
- data — reference the prior result. Re-invocation is the #1 cause of sub-40% steps. Only re-fetch
240
+ data — reference the prior result. Re-invocation is the #1 cause of low estimated-reuse steps.
241
+ Only re-fetch
207
242
  when the underlying state *changed* (e.g. you just wrote a section and need the new
208
243
  `requiredSections`).
209
244
  2. **Don't re-ground shared scaffolding per step.** Command docs, the skill preamble, and the
@@ -217,7 +252,8 @@ When aggregate cache% risks falling under 50%, apply this checklist **before** r
217
252
  useful as a trend if it reflects what actually happened.
218
253
 
219
254
  The point is not to game the number — it is to drive the testee (and your own monitoring) toward
220
- reusing context, which is the real cost saving the cache% signal stands for.
255
+ reusing context, which is the context-efficiency gain the estimated reuse share stands for. The
256
+ share itself remains a heuristic trend signal (task 0913, R3), not a measured saving.
221
257
 
222
258
  ## Worked ledger example
223
259
 
@@ -230,6 +266,7 @@ reusing context, which is the real cost saving the cache% signal stands for.
230
266
  | 4 profile | 1 | PASS | — | — | ~500 | ~350 | 41% | command output + prior profile reused | ~2s |
231
267
  ```
232
268
 
233
- Aggregate: total = `3700 + 2050 = 5750`; cached = `2050`; cache% =
234
- `round(2050 / 5750 * 100) = 36%` `[~estimate]` — below the 50% floor, so emit the P3 cache-health
235
- finding.
269
+ Aggregate: total = `3700 + 2050 = 5750`; cached = `2050`; estimated reuse share =
270
+ `round(2050 / 5750 * 100) = 36%` `[~estimate]`. A low estimated reuse share is reported only as a
271
+ labeled-hypothesis `[unverifiable]` finding (or backed by a real meter) — never as an automatic
272
+ causal waste claim (task 0913, R3).
@@ -96,8 +96,21 @@ Skeleton is written in Phase 1; filled as the run progresses; finalized in Phase
96
96
  - **Mode:** `observe-only (--max-retry 0)` | `fix (--max-retry N)`
97
97
  - **Task under test:** WBS + title (if applicable)
98
98
  - **Run id:** `<run_id>` · **Live:** `<live_path>` · **Report:** `<report_path>`
99
+ - **Source evidence:** `<commit sha>` + dirty-state (`clean` | `porcelain hash` | `unrecorded`)
100
+ - **Definition:** resolved definition digest, or `path+hash`; `unknown` when not resolvable
101
+ - **Execution mode:** `<observe-only | fix>` · driver/testee separation noted for chained legs
102
+ - **Measurement scope:** `<driver ledger (chars/4, per-step) | ccusage <day\|session> | agent usage fields>`; external meter name or `none`
103
+ - **Sample coverage:** `<k>/<n>` executed steps observed; testee workflow/run/session IDs `<ids | unknown>`
99
104
  ```
100
105
 
106
+ **Provenance and comparability (task 0913, R4 — additive).** The provenance lines above are
107
+ optional in `@1.2`: legacy reports without them remain readable and valid, but are **not
108
+ automatically eligible for comparison**. Before treating two reports as comparable, require
109
+ compatible measurement scope, execution mode, and resolved definition identity; otherwise emit
110
+ **`not comparable`** and route the missing evidence to its owner. Unsupported identifiers are
111
+ recorded as `unknown` — never invented. Driver cost and testee/chained-leg cost stay separate
112
+ rows; pilot comparisons must not mix them.
113
+
101
114
  When `status` is `aborted` or the report is partial, add under §1:
102
115
 
103
116
  ```
@@ -122,14 +135,22 @@ When `status` is `aborted` or the report is partial, add under §1:
122
135
 
123
136
  **Cost honesty rules:**
124
137
 
125
- - Always include ledger-derived `~estimate` total / cached / cache% with a **Method** line and
138
+ - Separate the three evidence classes in every Cost block (task 0913, R3): **measured** usage
139
+ (real meter, scope-labeled), **heuristic estimates** (chars/4 ledger, labeled `~estimate`,
140
+ confidence `LOW`), and **unavailable** values rendered `n/a` — never a fabricated number or a
141
+ share computed over unknown rows.
142
+ - Always include ledger-derived `~estimate` total / cached / reuse share with a **Method** line and
126
143
  **confidence** (`LOW` when estimate-only; `MEDIUM` when a real meter is also present).
127
144
  - Optional meters when available (never invent):
128
145
  - `ccusage` session/daily delta — label scope (`day` / `session`), **not** per-step
129
146
  - agent usage fields if present in tool results
130
- - If no meter: print `Meter: n/a` explicitly.
147
+ - If no meter: print `Meter: n/a` explicitly; if no ledger row is observable, print totals and
148
+ share as `n/a` rather than folding unknowns into zero.
131
149
  - Never present an unsubstantiated precise integer as billed/metered cost.
132
- - Aggregate cache% MUST equal the ledger formula (see §3); otherwise the report is invalid.
150
+ - Observable totals MUST satisfy `total = fresh + cached` and the share MUST equal the ledger
151
+ formula over observable rows within display rounding (±1 point); the validator (task 0913)
152
+ rejects violations. The chars/4 figures estimate the driver session's context reuse — they are
153
+ **not** provider cache measurements and establish no realized savings.
133
154
  - **Chained-step segmentation (@1.2):** when a derived step is implement-heavy (the step runs a
134
155
  pipeline leg, writes code, or mutates more than its own arguments), its cost MUST be a separate
135
156
  ledger row tagged `chained:<step>` and kept out of the driver's row. If the chained leg ran in a
@@ -251,8 +272,8 @@ from the trailing feasibility tag:
251
272
  ```
252
273
 
253
274
  Omitting the class preserves the current line shape; untagged findings remain valid and the
254
- protocol stays `sp:dogfood-testing@1.2` (the validator gains no required field; the cache-health
255
- P3 above needs no class).
275
+ protocol stays `sp:dogfood-testing@1.2` (the validator gains no required field; an
276
+ evidence-based reuse observation needs no class).
256
277
 
257
278
  - `testee` — a defect in the testee's contract (the protocol the run grades). Bounded fix-mode
258
279
  may repair it, unchanged.
@@ -278,13 +299,15 @@ Severity scale:
278
299
  finding:** when a drift row (`drift:external`) is present in the ledger, a P2 finding naming the
279
300
  drifted paths is mandatory in the report (not optional). The finding states the run's evidence is
280
301
  degraded, not voided. See [SKILL.md §Workspace-drift guard](../SKILL.md#workspace-drift-guard-r2--task-0296).
281
- - **P3** — efficiency / DX / observation (includes the cache-health rule below).
302
+ - **P3** — efficiency / DX / observation (includes evidence-based context-reuse observations below).
282
303
  - **P4** — nice-to-have, cosmetic, or speculative.
283
304
 
284
- **Cache-health rule** (from [monitor-ledger.md](monitor-ledger.md)): if aggregate cache% < 50% or any
285
- step < 40%, emit a **P3** — "Low cache hit rate — candidate for context-window or prompt trimming"
286
- with the offending step(s). Absolute token totals from the heuristic are trend-only (`[unverifiable]`
287
- as billable cost proof is expected).
305
+ **Context-reuse observation rule** (task 0913 — replaces the fixed-threshold cache-health rule;
306
+ see [monitor-ledger.md](monitor-ledger.md)): estimated reuse shares (chars/4) never auto-generate
307
+ causal waste findings. Report a low estimated share only as a labeled-hypothesis finding carrying
308
+ `[unverifiable]` and naming the confirmation needed, or cite measured meter evidence with its
309
+ scope. Absolute token totals from the heuristic are trend-only (`[unverifiable]` as billable cost
310
+ proof is expected).
288
311
 
289
312
  **Migration grep rule.** When dogfooding migrations or retired surfaces, distinguish intentional
290
313
  legacy-term mentions in guidance from live routed surfaces. Pair any broad grep for old skill or
@@ -1,7 +1,9 @@
1
1
  # Mode contract — `sp:history-anatomy` (HA-S1, 0658)
2
2
 
3
3
  The skill resolves exactly two modes. Everything else fails loud. This matrix is the enforcement
4
- surface the workflow (0660) and the skill share; keep the vocabulary frozen.
4
+ surface the workflow (0660) and the skill share; keep the vocabulary frozen. Execution note
5
+ (0920): argument validation runs deterministically in the `history-anatomy-cache` helper `paths`
6
+ command before the workflow starts; the model hop no longer performs it.
5
7
 
6
8
  ## Mode vocabulary (frozen)
7
9
 
@@ -26,7 +28,7 @@ Rejected arguments (each fails loud, naming the offending argument):
26
28
  | --- | --- |
27
29
  | focus text (positional) | Daily mode has no focus string. |
28
30
  | `--since` / `--until` | Daily always uses the calendar-day window. |
29
- | `--output` | Daily always writes to the run directory (see 0660). |
31
+ | `--output` | Daily always writes to `docs/report/<date>-history-anatomy.md` (the executing helper's default; 0660). |
30
32
 
31
33
  A daily invocation must print the normalized **inclusive ISO bounds** and the timezone used, so the
32
34
  wall-clock window is auditable.
@@ -40,7 +42,7 @@ Requires a **non-empty focus** and **two ordered inclusive bounds**.
40
42
  | focus (positional) | Required; a missing or empty focus fails loud. |
41
43
  | `--since <iso>` | Required; the inclusive lower bound. |
42
44
  | `--until <iso>` | Required; must be present with `--since`; must not be earlier than `--since`. |
43
- | `--output <path>` | Optional; when present, writes to that explicit path. When absent, writes to the run directory. |
45
+ | `--output <path>` | Optional; when present, writes to that explicit path. When absent, writes to `docs/report/<date>-history-anatomy.md` (the executing helper's default). |
44
46
 
45
47
  Rejected arguments (each fails loud, naming the offending argument):
46
48
 
@@ -40,7 +40,7 @@ and no command-derived number or `file:line` citation is a defect in the report,
40
40
  ## The gated list
41
41
 
42
42
  Gated features are listed **separately, never ranked**, each with its gate reason from the
43
- actionability pass (`blocked: 0142 — external trigger` / `no open tasks` / `all tasks terminal`).
43
+ actionability pass (`blocked: <wbs> — external trigger` / `no open tasks` / `all tasks terminal`).
44
44
  Features whose gate reason is "all tasks terminal" are T4 candidates — say so once, in the sync-first
45
45
  block, rather than repeating per row.
46
46
 
@@ -113,6 +113,13 @@ token, deterministic).
113
113
  | C3 | A5/A6 when Testing empty/N/A **and** verify would fail for missing tests — only if prior implement claims code exists | Coverage/test signal: `bun test` fail attributed to task paths OR explicit "insufficient tests" in prior verify verdict artifact `.spur/run/<wbs>-verdict.json` | test fail / coverage gap | `/sp:dev-unit <wbs> --auto` | continue |
114
114
  | C4 | A3/A5/A6 when operator or task tags mention rules, OR `spur rule run` last report dirty in `.spur/` if present | `spur rule run` (default project preset) non-zero with findings | rule findings | **HITL STOP** — print rule summary; suggest `/sp:rule-scan` or `rule-add`/`rule-refine` (do not auto-author rules) | continue |
115
115
  | C5 | A6 only | Existing `.spur/run/<wbs>-verdict.json` with FAIL and findings pointing at coverage | verdict artifact | `/sp:dev-unit <wbs>` then re-verify on next invocation (`--once` friendly) | `/sp:dev-verify …` |
116
+ | C6 | A4/A5 | `.spur/run/<wbs>-verdict.json` has a failing (PARTIAL/FAIL) `residual-sweep` check — the bounded remediation loop already ran and failed (F96) | folded verdict artifact | **HITL STOP** — print `.spur/run/<wbs>-residual-report.md` and the recovery command `/sp:dev-run <wbs>`; never auto-dispatch a fix (repeating the loop unattended burns quota without new information) | continue |
117
+
118
+ **C-row precedence note (F96):** C6 outranks C2/C3/C5 for the same task — a residual-sweep failure
119
+ means the workspace holds unfinished task residue, so fixall/unit reruns would either clean it
120
+ unintentionally or re-certify a folded verdict. C6 fires before any fix/unit dispatch; recovery is
121
+ the operator re-running the pipeline (`/sp:dev-run <wbs>`), which re-enters the normal
122
+ verify → test-fix remediation budget.
116
123
 
117
124
  **Explicit non-probes in v1:** no freeform chat history; no always-on full `bun run test` for every
118
125
  call; no git dirtiness as a route (optional advisory print only).
@@ -37,7 +37,7 @@ native subagent.
37
37
  | 1 | **Different model or coding agent required** | The step needs a model or a coding agent the host session cannot provide (`--model`, `--agent`). | "verify on o3" where the host is Claude Code; "run this through omp" from a non-omp host. |
38
38
  | 2 | **Headless or unattended step** | The step must run without a live session - scheduled, detached, or driven by a non-interactive caller. | A batch launched by `spur workflow run --async` with no operator attached. |
39
39
  | 3 | **Durable auditable run record required** | The dispatch must produce a persisted run record (cost ledger, trace, exit code) for after-the-fact audit. | `spur agent run` writes `.spur/run/` artifacts; a native subagent does not. |
40
- | 4 | **Workspace or credential isolation required** | The step must run in a separate workspace, worktree, or credential scope from the orchestrating session. | A destructive step isolated to a throwaway worktree; a step that must not inherit the session's `cwd` secrets. |
40
+ | 4 | **Workspace or credential isolation required** | The step must run in a separate workspace, worktree, or credential scope from the orchestrating session. | A destructive step isolated to a throwaway worktree; a step that must not inherit the session's `cwd` secrets; `--mode parallel` batch fan-out, one worktree per task ([execution-batch.md § Parallel isolation](../../spur-dev/references/execution-batch.md#parallel-isolation---mode-parallel)). |
41
41
 
42
42
  ## The naming requirement
43
43
 
@@ -84,7 +84,14 @@ three buckets — never skip triage and start fixing from the raw findings list.
84
84
  Apply the placement rule in
85
85
  [the environment-improvement mapping](../../references/environment-lens.md): automate with a
86
86
  check when possible, place coding standards on the review path, and keep always-loaded steering
87
- as navigation pointers.
87
+ as navigation pointers. When the session drove a task pipeline, at most three bounded
88
+ diagnostic questions may be drawn from the supported observations (F1/F2) and measured
89
+ overhead candidates (F3/F4) in the 0912 workflow baseline
90
+ (`docs/reports/i31/0912-workflow-baseline.md`) — e.g. repeated gate runs (F4), full-loss
91
+ test-fix timeouts (F3), or unemitted/non-terminal run rows (F1/F2) — each citing its artifact
92
+ anchor and owner handoff (driver adoption D62, row-closure defect P). Areas the baseline marks
93
+ INSUFFICIENT_EVIDENCE (token/USD, percentages, fleet) stay excluded; record the limitation
94
+ instead of a performance conclusion.
88
95
  5. **Render the report.** Use the exact compact output contract below. Omit empty table rows, not
89
96
  headings; write `None observed` when a section has no supported entry.
90
97
 
@@ -123,7 +130,10 @@ session. Do not list ordinary implementation steps as issues.
123
130
  ### Process and environment improvements
124
131
 
125
132
  For each supported proposal, name its owner surface, expected impact, verification method, and
126
- reversibility. Proposals remain report-only: apply no change and create no task.
133
+ reversibility. Proposals remain report-only: apply no change and create no task. A pipeline
134
+ observation adopted from the 0912 workflow baseline cites its anchor
135
+ (`docs/reports/i31/0912-workflow-baseline.md`) and owner handoff, and carries no unsupported
136
+ performance claim.
127
137
 
128
138
  ### Triage (only when `--triage` was passed)
129
139