devflow-kit 3.1.0 → 3.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (138) hide show
  1. package/CHANGELOG.md +52 -0
  2. package/README.md +2 -2
  3. package/dist/cli/agents-view/render.js +69 -15
  4. package/dist/cli/agents-view/state.js +40 -14
  5. package/dist/cli/commands/agents.js +135 -45
  6. package/dist/cli/commands/init.js +128 -53
  7. package/dist/cli/commands/learning.js +61 -13
  8. package/dist/cli/commands/memory.js +35 -14
  9. package/dist/cli/commands/uninstall.js +163 -39
  10. package/dist/commands/code-review.md +1 -3
  11. package/dist/commands/debug.md +15 -12
  12. package/dist/commands/dynamic-build.md +172 -135
  13. package/dist/commands/dynamic-plan.md +9 -3
  14. package/dist/commands/explore.md +10 -4
  15. package/dist/commands/implement.md +149 -145
  16. package/dist/commands/plan.md +13 -9
  17. package/dist/commands/release.md +8 -2
  18. package/dist/commands/research.md +8 -2
  19. package/dist/commands/resolve.md +28 -19
  20. package/dist/commands/self-review.md +16 -13
  21. package/dist/core/agent-frontmatter.js +25 -0
  22. package/dist/core/agent-models.js +201 -36
  23. package/dist/core/agent-state.js +27 -5
  24. package/dist/core/assets.js +1 -1
  25. package/dist/core/feature-config.js +68 -10
  26. package/dist/core/flags.js +24 -0
  27. package/dist/core/learning-queue-cleanup.js +10 -11
  28. package/dist/core/learning-tuning-config.js +8 -0
  29. package/dist/core/linked-path.js +46 -0
  30. package/dist/core/plugins.js +16 -5
  31. package/dist/core/queue-drain.js +31 -0
  32. package/dist/hud/components/learning-counts.js +54 -8
  33. package/dist/skills/git/references/tracker/github/create-release.md +2 -2
  34. package/dist/skills/git/references/tracker/github/gather-release-evidence.md +1 -1
  35. package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
  36. package/dist/skills/git/references/tracker/jira/gather-release-evidence.md +1 -1
  37. package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
  38. package/dist/skills/git/references/tracker/linear/gather-release-evidence.md +1 -1
  39. package/dist/targets/claude-code/installer.js +36 -9
  40. package/dist/targets/claude-code/post-install.js +128 -38
  41. package/package.json +1 -1
  42. package/src/assets/agents/code.md +85 -35
  43. package/src/assets/agents/design.md +12 -0
  44. package/src/assets/agents/diagnose.md +18 -11
  45. package/src/assets/agents/evaluate.md +17 -24
  46. package/src/assets/agents/knowledge.md +7 -3
  47. package/src/assets/agents/learning.md +4 -6
  48. package/src/assets/agents/research.md +21 -0
  49. package/src/assets/agents/review.md +12 -0
  50. package/src/assets/agents/scrutinize.md +37 -9
  51. package/src/assets/agents/simplify.md +24 -0
  52. package/src/assets/agents/skim.md +6 -2
  53. package/src/assets/agents/synthesize.md +18 -0
  54. package/src/assets/agents/test.md +19 -11
  55. package/src/assets/agents/triage.md +8 -0
  56. package/src/assets/agents/validate.md +20 -11
  57. package/src/assets/commands/_partials/_engine.mds +36 -55
  58. package/src/assets/commands/_partials/_knowledge.mds +1 -3
  59. package/src/assets/commands/_partials/_plan_contract.mds +1 -1
  60. package/src/assets/commands/_partials/_tracker.mds +1 -1
  61. package/src/assets/commands/_partials/_wave.mds +8 -6
  62. package/src/assets/commands/code-review.mds +1 -3
  63. package/src/assets/commands/debug.mds +13 -8
  64. package/src/assets/commands/dynamic-build.mds +126 -72
  65. package/src/assets/commands/dynamic-plan.mds +7 -1
  66. package/src/assets/commands/explore.mds +9 -1
  67. package/src/assets/commands/implement.mds +147 -141
  68. package/src/assets/commands/plan.mds +12 -8
  69. package/src/assets/commands/release.md +8 -2
  70. package/src/assets/commands/research.mds +8 -2
  71. package/src/assets/commands/resolve.mds +27 -16
  72. package/src/assets/commands/self-review.mds +15 -10
  73. package/src/assets/mds/tracker/_common.mds +1 -1
  74. package/src/assets/mds/tracker/_github.mds +2 -2
  75. package/src/assets/mds/tracker/_jira.mds +2 -2
  76. package/src/assets/mds/tracker/_linear.mds +2 -2
  77. package/src/assets/scripts/ci-wait.cjs +636 -0
  78. package/src/assets/scripts/hooks/assets/orchestrator-charter.md +4 -3
  79. package/src/assets/scripts/hooks/background-memory-update +356 -17
  80. package/src/assets/scripts/hooks/capture-prompt +4 -3
  81. package/src/assets/scripts/hooks/capture-question +4 -3
  82. package/src/assets/scripts/hooks/capture-turn +4 -3
  83. package/src/assets/scripts/hooks/ensure-devflow-init +13 -1
  84. package/src/assets/scripts/hooks/ensure-root-gitignore +122 -10
  85. package/src/assets/scripts/hooks/git-marker +71 -0
  86. package/src/assets/scripts/hooks/json-helper.cjs +12 -145
  87. package/src/assets/scripts/hooks/json-parse +24 -129
  88. package/src/assets/scripts/hooks/lib/learning-store.cjs +169 -64
  89. package/src/assets/scripts/hooks/lib/render-decisions.cjs +1 -1
  90. package/src/assets/scripts/hooks/memory-worker +10 -0
  91. package/src/assets/scripts/hooks/pre-compact-memory +66 -14
  92. package/src/assets/scripts/hooks/preamble +9 -1
  93. package/src/assets/scripts/hooks/queue-append +53 -21
  94. package/src/assets/scripts/hooks/session-start-context +108 -29
  95. package/src/assets/scripts/hooks/session-start-memory +33 -11
  96. package/src/assets/scripts/release-trace.cjs +27 -10
  97. package/src/assets/skills/accessibility/SKILL.md +1 -1
  98. package/src/assets/skills/apply-decisions/SKILL.md +12 -82
  99. package/src/assets/skills/apply-feature-knowledge/SKILL.md +8 -42
  100. package/src/assets/skills/architecture/SKILL.md +1 -1
  101. package/src/assets/skills/boundary-validation/SKILL.md +1 -1
  102. package/src/assets/skills/complexity/SKILL.md +1 -1
  103. package/src/assets/skills/compliance/SKILL.md +1 -1
  104. package/src/assets/skills/consistency/SKILL.md +1 -1
  105. package/src/assets/skills/database/SKILL.md +1 -1
  106. package/src/assets/skills/dependencies/SKILL.md +1 -1
  107. package/src/assets/skills/dependency-research/SKILL.md +3 -6
  108. package/src/assets/skills/design-review/SKILL.md +1 -1
  109. package/src/assets/skills/docs-framework/SKILL.md +1 -1
  110. package/src/assets/skills/documentation/SKILL.md +1 -1
  111. package/src/assets/skills/gap-analysis/SKILL.md +1 -1
  112. package/src/assets/skills/git/SKILL.md +1 -1
  113. package/src/assets/skills/go/SKILL.md +1 -1
  114. package/src/assets/skills/java/SKILL.md +1 -1
  115. package/src/assets/skills/patterns/SKILL.md +1 -1
  116. package/src/assets/skills/performance/SKILL.md +1 -1
  117. package/src/assets/skills/python/SKILL.md +1 -1
  118. package/src/assets/skills/qa/SKILL.md +1 -3
  119. package/src/assets/skills/quality-gates/SKILL.md +9 -12
  120. package/src/assets/skills/quality-gates/references/report-template.md +20 -20
  121. package/src/assets/skills/react/SKILL.md +1 -1
  122. package/src/assets/skills/regression/SKILL.md +1 -1
  123. package/src/assets/skills/reliability/SKILL.md +1 -1
  124. package/src/assets/skills/research-codebase/SKILL.md +1 -1
  125. package/src/assets/skills/research-competitor/SKILL.md +1 -1
  126. package/src/assets/skills/research-external/SKILL.md +1 -1
  127. package/src/assets/skills/research-technology/SKILL.md +1 -1
  128. package/src/assets/skills/review-methodology/SKILL.md +1 -1
  129. package/src/assets/skills/rust/SKILL.md +1 -1
  130. package/src/assets/skills/security/SKILL.md +1 -1
  131. package/src/assets/skills/software-design/SKILL.md +1 -1
  132. package/src/assets/skills/test-driven-development/SKILL.md +15 -33
  133. package/src/assets/skills/testing/SKILL.md +1 -1
  134. package/src/assets/skills/typescript/SKILL.md +1 -1
  135. package/src/assets/skills/ui-design/SKILL.md +1 -1
  136. package/src/assets/skills/worktree-support/SKILL.md +3 -55
  137. package/src/assets/skills/worktree-support/references/discovery.md +48 -0
  138. package/src/assets/skills/worktree-support/references/roots.md +2 -2
@@ -19,7 +19,13 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
19
19
 
20
20
  ## Input
21
21
 
22
- `$ARGUMENTS` contains whatever follows `/debug`:
22
+ What follows `/debug` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
23
+
24
+ <command-input>
25
+ $ARGUMENTS
26
+ </command-input>
27
+
28
+ `COMMAND_INPUT` is one of:
23
29
  - Bug description: "login fails after session timeout"
24
30
  - Issue reference: "#42"
25
31
  - Empty: use conversation context
@@ -30,8 +36,6 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
30
36
 
31
37
  **Produces:** DECISIONS_CONTEXT
32
38
 
33
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:software-design`, `devflow:testing`. If a skill fails to load, continue without it.
34
-
35
39
  {{decisions_load()}}
36
40
 
37
41
  The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Phase 2) — prior pitfalls and decisions can suggest specific root causes to investigate. Follow `devflow:apply-decisions` to Read full entry bodies on demand. **Do NOT pass `DECISIONS_CONTEXT` to Explore investigators** — decisions context stays in the orchestrator; investigators examine code directly.
@@ -43,7 +47,7 @@ The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Ph
43
47
  **Produces:** HYPOTHESES, BUG_CONTEXT
44
48
  **Requires:** DECISIONS_CONTEXT
45
49
 
46
- If `$ARGUMENTS` opens with a candidate issue reference, fetch the issue:
50
+ If `COMMAND_INPUT` opens with a candidate issue reference, fetch the issue:
47
51
 
48
52
  {{issue_ref_grammar()}}
49
53
 
@@ -88,7 +92,8 @@ Return a structured report:
88
92
  - Status: CONFIRMED / DISPROVED / PARTIAL
89
93
  - Evidence FOR: [list with file:line refs]
90
94
  - Evidence AGAINST: [list with file:line refs]
91
- - Key finding: {one-sentence summary}"
95
+ - Key finding: {one-sentence summary}
96
+ Keep the whole report to at most about 1,500 tokens."
92
97
 
93
98
  Agent(subagent_type="Explore"):
94
99
  "Investigate this bug: {bug_description}
@@ -96,7 +101,7 @@ Agent(subagent_type="Explore"):
96
101
  Hypothesis: {hypothesis B description}
97
102
  Focus area: {specific code area, mechanism, or condition}
98
103
 
99
- [same steps and return format]"
104
+ [same steps and return format; the whole report at most about 1,500 tokens]"
100
105
 
101
106
  Agent(subagent_type="Explore"):
102
107
  "Investigate this bug: {bug_description}
@@ -104,7 +109,7 @@ Agent(subagent_type="Explore"):
104
109
  Hypothesis: {hypothesis C description}
105
110
  Focus area: {specific code area, mechanism, or condition}
106
111
 
107
- [same steps and return format]"
112
+ [same steps and return format; the whole report at most about 1,500 tokens]"
108
113
 
109
114
  (Add more investigators if bug complexity warrants 4-5 hypotheses)
110
115
  ```
@@ -116,7 +121,7 @@ Focus area: {specific code area, mechanism, or condition}
116
121
 
117
122
  Evaluate investigation verdicts before synthesis:
118
123
 
119
- - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias)
124
+ - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias), each reporting at most about 1,500 tokens
120
125
  - **Multiple PARTIAL**: Look for a unifying root cause that explains all partial evidence
121
126
  - **All DISPROVED**: Report honestly — "No root cause identified from initial hypotheses." Generate 2-3 second-round hypotheses if conversation context suggests avenues not yet explored. Loop back to Phase 3.
122
127
 
@@ -79,7 +79,13 @@ Note the `budget` value from the Workflow tool context (or default to "medium" i
79
79
 
80
80
  **3. Detect mode: SINGLE or WAVE**
81
81
 
82
- - **SINGLE mode:** input is one ticket, one issue, one task description, or one plan document
82
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
83
+
84
+ <command-input>
85
+ $ARGUMENTS
86
+ </command-input>
87
+
88
+ - **SINGLE mode:** `COMMAND_INPUT` is one ticket, one issue, one task description, or one plan document
83
89
  - **WAVE mode:** input is a set of tracker issues (wave labels, milestone, issue list), or the user says "wave" / "all tickets in wave N"
84
90
  - **A `/devflow:dynamic-tickets` ticket directory** (`{worktree}/.devflow/docs/tickets/{slug}/{ts}/`) is WAVE input: in each ticket file (every `.md` there but `tracking-issue.md`), the `**Issue:**` line directly after `**Depends on:**` is one raw `ISSUE_REFS` token, forwarded to the wave's pre-fetch verbatim — never rendered, normalised or re-derived. A ticket file with no `**Issue:**` line, or more than one, contributes no token; name it in the run summary as `not filed`.
85
91
 
@@ -113,7 +119,7 @@ If none found: build proceeds Gate-1-only (Gate 2 skipped with a note). Never re
113
119
  **5. Resolve tracking-issue number (optional)**
114
120
 
115
121
  Check, in priority order:
116
- - An explicit candidate issue reference or issue URL in the user's input (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
122
+ - An explicit candidate issue reference or issue URL in `COMMAND_INPUT` (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
117
123
  - The `**Issue:**` line directly after the H1 of the ticket set's `tracking-issue.md` (written by `/devflow:dynamic-tickets`' filing step; the file is at `{worktree}/.devflow/docs/tickets/{slug}/{ts}/tracking-issue.md`), as a raw token — only when the file holds exactly one `**Issue:**` line
118
124
  - Otherwise: none
119
125
 
@@ -181,7 +187,8 @@ if (BRANCH === "(none)") {
181
187
 
182
188
  // Phase 2: Implement
183
189
  await phase("implement", () =>
184
- agent(`Implement the following ticket on branch ${BRANCH}:
190
+ agent(`OPERATION: implement
191
+ Implement the following ticket on branch ${BRANCH}:
185
192
 
186
193
  ${TICKET}
187
194
 
@@ -194,7 +201,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
194
201
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
195
202
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
196
203
 
197
- When you build or run tests to verify your work, use your "Long-running commands" discipline (background-Bash + Monitor poll) for anything that may run silent >120s, and prefer package-scoped commands.
204
+ When you build or run tests to verify your work, follow your Running commands block.
198
205
 
199
206
  After implementing, commit your changes with a conventional-commit message and report:
200
207
  - Files changed
@@ -202,39 +209,51 @@ After implementing, commit your changes with a conventional-commit message and r
202
209
  - Any open questions or blockers`, { agentType: "Code" })
203
210
  );
204
211
 
205
- // Phase 3: Gate 1 — post-code pipeline
212
+ // Phase 3: Gate 1 #1 — post-code pipeline: Simplify agent → Scrutinize agent → Validate agent.
213
+ // Validate runs last and unconditionally, so its one full run covers every commit the pass made (see gate1_postcode()).
206
214
  const gate1 = await phase("gate1", async () => {
207
- // Gate 1 #1 — runs once, immediately after the initial implementation.
215
+ await agent(`Simplify and reduce complexity of recent changes on branch ${BRANCH}. Commit any improvements.`, { agentType: "Simplify" });
216
+
217
+ const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Fix P0/P1 issues and commit the fixes.
218
+ Return: {"status": "PASS" | "FIXED" | "BLOCKED"}`, { agentType: "Scrutinize" });
219
+ // Anything but an explicit PASS or FIXED is a stop: a Scrutinize that returned no status, or one this skeleton does not know, has not accepted the code.
220
+ if (!["PASS", "FIXED"].includes(scrutiny?.status)) {
221
+ return { verdict: "ESCALATED", type: "scrutiny-blocked", reason: scrutiny?.status === "BLOCKED" ? "Scrutinize agent returned BLOCKED: a P0 issue cannot be fixed in scope" : "Scrutinize agent returned no recognised status: counted as BLOCKED" };
222
+ }
223
+
208
224
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH}.
209
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
210
- Report: PASS or FAIL with details.`, { agentType: "Validate" });
225
+ Follow your Running commands block for every build and test command.
226
+ Return: {"verdict": "PASS" | "FAIL", "details": "..."}`, { agentType: "Validate" });
211
227
 
212
- if (validation.verdict === "FAIL") {
213
- // Up to 2 fix attempts
228
+ // Anything but an explicit PASS is a failure: a Validate that returned no verdict has not validated the branch.
229
+ if (validation?.verdict !== "PASS") {
230
+ let failureDetails = validation?.details || "Validate agent returned no PASS verdict";
231
+ // Up to 2 fix attempts, each followed by a Validate re-run
214
232
  for (let attempt = 1; attempt <= 2; attempt++) {
215
- await agent(`Fix the validation failures on branch ${BRANCH}:
216
- ${validation.details}
233
+ await agent(`OPERATION: validation-fix
234
+ Fix the validation failures on branch ${BRANCH}:
235
+ ${failureDetails}
217
236
  ISSUE_NUMBER: ${ISSUE_NUMBER}
218
237
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
219
238
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
220
239
  Commit fixes with conventional-commit message.`, { agentType: "Code" });
221
- const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH}. Report: PASS or FAIL.`, { agentType: "Validate" });
222
- if (recheck.verdict === "PASS") break;
223
- if (attempt === 2) return { verdict: "ESCALATED", reason: "Validation exhausted after 2 Code agent fix attempts" };
240
+ const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} after fix attempt ${attempt} (follow your Running commands block).
241
+ Return: {"verdict": "PASS" | "FAIL", "details": "..."}`, { agentType: "Validate" });
242
+ if (recheck?.verdict === "PASS") break;
243
+ failureDetails = recheck?.details || failureDetails;
244
+ if (attempt === 2) return { verdict: "ESCALATED", type: "validation-exhausted", reason: "Validation exhausted after 2 Code agent fix attempts" };
224
245
  }
225
246
  }
226
247
 
227
- await agent(`Simplify and reduce complexity of recent changes on branch ${BRANCH}. Commit any improvements.`, { agentType: "Simplify" });
228
-
229
- const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
230
-
231
- if (scrutiny.codeChanged) {
232
- await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes). Report: PASS or FAIL.`, { agentType: "Validate" });
233
- }
234
-
235
248
  return { verdict: "PASS" };
236
249
  });
237
250
 
251
+ // Gate 1 #1 stop, before Gate 2: an exhausted Validate or a BLOCKED Scrutinize leaves a broken or unfinished build,
252
+ // so Gate 2 and the review pass never run on it — a correctness stop, like the ticket-link and branch stops.
253
+ if (gate1.verdict === "ESCALATED") {
254
+ return { ticket: TICKET, branch: BRANCH, verdict: "ESCALATED", issueId: ISSUE_NUMBER, issuePrLink: ISSUE_PR_LINK, escalations: [{ type: gate1.type, description: gate1.reason }] };
255
+ }
256
+
238
257
  // Phase 4: Gate 2 — acceptance gate (once, before review pass)
239
258
  const gate2 = await phase("gate2", async () => {
240
259
  if (!PLAN && !CRITERIA && !TEST_PLAN) {
@@ -243,25 +262,23 @@ const gate2 = await phase("gate2", async () => {
243
262
 
244
263
  let evalVerdict = "SKIPPED";
245
264
  if (PLAN) {
246
- const panel = await parallel([
247
- () => agent(`Acceptance-criteria evaluator: does the implementation on branch ${BRANCH} satisfy each numbered criterion, INCLUDING negative criteria?
265
+ const evaluation = await agent(`Acceptance and scope check on branch ${BRANCH}. Answer both questions:
266
+ 1. Does the implementation satisfy each numbered criterion, INCLUDING negative criteria (what it must NOT do)?
267
+ 2. Did the Code agent introduce any unplanned changes, smuggled anti-features, or drift from the plan's intent?
248
268
  Plan: ${PLAN}
249
269
  Criteria: ${CRITERIA || "see plan"}
250
- Report: PASS or FAIL per criterion with rationale.`, { agentType: "Evaluate" }),
251
- () => agent(`Scope/intent-drift evaluator: did the Code agent on branch ${BRANCH} introduce any unplanned changes, smuggled anti-features, or drift from the plan's intent?
252
- Plan: ${PLAN}
253
- Report: PASS or FAIL with rationale.`, { agentType: "Evaluate" }),
254
- ]);
255
- evalVerdict = panel.every(p => p.verdict === "PASS") ? "PASS" : "FAIL";
270
+ Return: {"verdict": "PASS" | "FAIL", "rationale": "..."} — FAIL if either answer is no, naming the failing criterion or the drift.`, { agentType: "Evaluate" });
271
+ evalVerdict = evaluation?.verdict === "PASS" ? "PASS" : "FAIL";
256
272
  if (evalVerdict === "FAIL") {
257
273
  // fix-and-continue — no re-evaluate, no inline Gate 1. The final Gate 1 (#2)
258
274
  // after the review pass is the build gate before the branch is handed back.
259
- await agent(`Fix the alignment issues identified by the Evaluate agent panel on branch ${BRANCH}:
260
- ${panel.filter(p => p.verdict === "FAIL").map(p => p.rationale).join("\n")}
275
+ await agent(`OPERATION: alignment-fix
276
+ Fix the alignment issues identified by the Evaluate agent on branch ${BRANCH}:
277
+ ${evaluation?.rationale || "see the Evaluate agent's report"}
261
278
  ISSUE_NUMBER: ${ISSUE_NUMBER}
262
279
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
263
280
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
264
- Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit fixes.`, { agentType: "Code" });
281
+ Self-verify your fix compiles, following your Running commands block. Commit fixes.`, { agentType: "Code" });
265
282
  evalVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
266
283
  }
267
284
  }
@@ -271,17 +288,18 @@ Self-verify your fix compiles (background-Bash + Monitor for any build >120s —
271
288
  const testResult = await agent(`Run scenario-based acceptance tests on branch ${BRANCH} against these criteria:
272
289
  ${CRITERIA || "(none)"}
273
290
  TEST_PLAN: ${TEST_PLAN || "(none)"}
274
- For any test/build command that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog.
291
+ Follow your Running commands block for every scenario command.
275
292
  Cover: functionality, API contracts, performance, and cover every TEST_PLAN scenario. Report: PASS or FAIL per scenario.`, { agentType: "Test" });
276
293
  testVerdict = testResult.verdict;
277
294
  if (testVerdict === "FAIL") {
278
295
  // fix-and-continue — no re-test, no inline Gate 1.
279
- await agent(`Fix the failing acceptance test scenarios on branch ${BRANCH}:
296
+ await agent(`OPERATION: qa-fix
297
+ Fix the failing acceptance test scenarios on branch ${BRANCH}:
280
298
  ${testResult.failures}
281
299
  ISSUE_NUMBER: ${ISSUE_NUMBER}
282
300
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
283
301
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
284
- Self-verify your fix compiles and the scenarios pass (background-Bash + Monitor for any build/test >120s). Commit fixes.`, { agentType: "Code" });
302
+ Self-verify your fix compiles and the scenarios pass, following your Running commands block. Commit fixes.`, { agentType: "Code" });
285
303
  testVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
286
304
  }
287
305
  }
@@ -303,7 +321,7 @@ const reviewResult = await phase("review", async () => {
303
321
  const chunkSize = 5;
304
322
  const reviewThunks = reviewFocuses.map(focus =>
305
323
  () => agent(`${focus.charAt(0).toUpperCase() + focus.slice(1)} review. ${reviewScope}
306
- Focus discipline: ${focus}. Return: {"focus": "${focus}", "reviewed": true, "filesExamined": ["<path>"], "findings": [{"file": "<path-or-null>", "description": "<one-sentence issue>", "severity": "critical|high|medium|low"}]}`, { agentType: "Review" })
324
+ Focus discipline: ${focus}. Return: {"focus": "${focus}", "reviewed": true, "filesExamined": ["<path>"], "findings": [{"file": "<path-or-null>", "line": <number-or-null>, "description": "<one-sentence issue>", "severity": "critical|high|medium|low"}]}`, { agentType: "Review" })
307
325
  );
308
326
 
309
327
  let rawReviews = [];
@@ -320,7 +338,7 @@ Focus discipline: ${focus}. Return: {"focus": "${focus}", "reviewed": true, "fil
320
338
  const focus = reviewFocuses[i];
321
339
  if (!r || r.reviewed !== true) {
322
340
  // Dead Review agent: retry once sequentially
323
- const retry = await agent(`Retry ${focus} review. ${reviewScope} Focus: ${focus} discipline. Return: {"focus": "${focus}", "reviewed": true, "filesExamined": ["<path>"], "findings": [{"file": "<path-or-null>", "description": "<one-sentence issue>", "severity": "critical|high|medium|low"}]}`, { agentType: "Review" });
341
+ const retry = await agent(`Retry ${focus} review. ${reviewScope} Focus: ${focus} discipline. Return: {"focus": "${focus}", "reviewed": true, "filesExamined": ["<path>"], "findings": [{"file": "<path-or-null>", "line": <number-or-null>, "description": "<one-sentence issue>", "severity": "critical|high|medium|low"}]}`, { agentType: "Review" });
324
342
  if (!retry || retry.reviewed !== true) {
325
343
  coverageGaps.push(focus);
326
344
  log(`review coverage incomplete: ${focus}`);
@@ -385,17 +403,22 @@ ${JSON.stringify(allFindings.map((f, i) => ({ index: i, description: f.descripti
385
403
  // Distinct-file groups are paced in chunks of FIX_CHUNK — same pacing bar as the Review spawn
386
404
  // path — to avoid launching all Code agents at once on a large branch (reliability: explicit bound).
387
405
  const FIX_CHUNK = 5;
388
- const fixThunks = Object.entries(fileGroups).map(([fileKey, findings]) => async () => {
406
+ // issue-fix SCOPE per finding: CRITICAL or HIGH takes the Careful protocol (failing regression test first); anything else is Standard.
407
+ const CAREFUL_SEVERITIES = new Set(["critical", "high"]);
408
+ const fixThunks = Object.values(fileGroups).map(findings => async () => {
389
409
  const chunkResults = [];
390
410
  for (let i = 0; i < findings.length; i += MAX_BATCH) {
391
411
  const chunk = findings.slice(i, i + MAX_BATCH);
392
- const r = await agent(`Fix the following confirmed review findings on branch ${BRANCH}${fileKey.startsWith("__nofile") ? "" : ` in ${fileKey}`}:
393
- ${chunk.map(f => `- ${f.description} (${f.severity})`).join("\n")}
394
-
412
+ const r = await agent(`OPERATION: issue-fix
413
+ Fix the following confirmed review findings on branch ${BRANCH}:
414
+ ISSUES:
415
+ ${chunk.map((f, n) => `${n + 1}. ${f.file ? `${f.file}${f.line ? `:${f.line}` : ""} — ` : ""}${f.description} (${f.severity})`).join("\n")}
416
+ SCOPE: ${chunk.map((f, n) => `${n + 1}=${CAREFUL_SEVERITIES.has(String(f.severity).toLowerCase()) ? "Careful" : "Standard"}`).join(", ")}
417
+ PUSH: false
395
418
  ISSUE_NUMBER: ${ISSUE_NUMBER}
396
419
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
397
420
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
398
- Fix all findings in this batch. Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit with conventional-commit message.
421
+ Fix all findings in this batch. Self-verify your fix compiles, following your Running commands block. Commit with conventional-commit message.
399
422
  Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<description of any finding that could not be fixed>"]}`, { agentType: "Code" });
400
423
  chunkResults.push({ chunk, result: r });
401
424
  }
@@ -430,38 +453,40 @@ Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<
430
453
  return { survivingFindings: notAddressed, fixedFindings, coverageGaps };
431
454
  });
432
455
 
433
- // Phase 5.5: Gate 1 #2 — FINAL post-fix gate (full Validate agent → Simplify agent → Scrutinize agent).
456
+ // Phase 5.5: Gate 1 #2 — FINAL post-fix gate (Simplify agent → Scrutinize agent → full Validate agent, Validate last and unconditional).
434
457
  // Runs ONCE, after the review pass. Inside the pass the Code agent self-verified its own fixes;
435
458
  // this is the build gate before the branch is handed back. See gate1_postcode() cadence.
436
459
  const gate1Final = await phase("gate1-final", async () => {
460
+ await agent(`Simplify and reduce complexity of recent changes on branch ${BRANCH}. Commit any improvements.`, { agentType: "Simplify" });
461
+
462
+ const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Fix P0/P1 issues and commit the fixes.
463
+ Return: {"status": "PASS" | "FIXED" | "BLOCKED"}`, { agentType: "Scrutinize" });
464
+ if (!["PASS", "FIXED"].includes(scrutiny?.status)) {
465
+ return { verdict: "ESCALATED", type: "scrutiny-blocked", reason: scrutiny?.status === "BLOCKED" ? "Final Gate 1 Scrutinize agent returned BLOCKED: a P0 issue cannot be fixed in scope" : "Final Gate 1 Scrutinize agent returned no recognised status: counted as BLOCKED" };
466
+ }
467
+
437
468
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH} (final gate after all fixing).
438
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
439
- Report: PASS or FAIL with details.`, { agentType: "Validate" });
469
+ Follow your Running commands block for every build and test command.
470
+ Return: {"verdict": "PASS" | "FAIL", "details": "..."}`, { agentType: "Validate" });
440
471
 
441
- if (validation.verdict === "FAIL") {
442
- let failureDetails = validation.details;
472
+ if (validation?.verdict !== "PASS") {
473
+ let failureDetails = validation?.details || "Validate agent returned no PASS verdict";
443
474
  for (let attempt = 1; attempt <= 2; attempt++) {
444
- await agent(`Fix the final validation failures on branch ${BRANCH}:
475
+ await agent(`OPERATION: validation-fix
476
+ Fix the final validation failures on branch ${BRANCH}:
445
477
  ${failureDetails}
446
478
  ISSUE_NUMBER: ${ISSUE_NUMBER}
447
479
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
448
480
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
449
481
  Self-verify your fix compiles. Commit fixes with conventional-commit message.`, { agentType: "Code" });
450
- const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
451
- if (recheck.verdict === "PASS") break;
452
- failureDetails = recheck.details || failureDetails;
453
- if (attempt === 2) return { verdict: "ESCALATED", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
482
+ const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (final gate, after fix attempt ${attempt}; follow your Running commands block).
483
+ Return: {"verdict": "PASS" | "FAIL", "details": "..."}`, { agentType: "Validate" });
484
+ if (recheck?.verdict === "PASS") break;
485
+ failureDetails = recheck?.details || failureDetails;
486
+ if (attempt === 2) return { verdict: "ESCALATED", type: "validation-exhausted", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
454
487
  }
455
488
  }
456
489
 
457
- await agent(`Simplify and reduce complexity of recent changes on branch ${BRANCH}. Commit any improvements.`, { agentType: "Simplify" });
458
-
459
- const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
460
-
461
- if (scrutiny.codeChanged) {
462
- await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
463
- }
464
-
465
490
  return { verdict: "PASS" };
466
491
  });
467
492
 
@@ -490,7 +515,7 @@ Write a concise report. Present only surviving findings as outstanding — never
490
515
  survivingFindings: reviewResult.survivingFindings || [], fixedFindings: reviewResult.fixedFindings || [],
491
516
  reviewCoverage: { failedFocuses: coverageGaps, complete: coverageGaps.length === 0 },
492
517
  escalations: [
493
- ...(gate1Final.verdict === "ESCALATED" ? [{ type: "validation-exhausted", description: gate1Final.reason || "final Gate 1 escalated" }] : []),
518
+ ...(gate1Final.verdict === "ESCALATED" ? [{ type: gate1Final.type || "validation-exhausted", description: gate1Final.reason || "final Gate 1 escalated" }] : []),
494
519
  ...coverageGaps.map(focus => ({ type: "review-coverage-incomplete", description: `review coverage incomplete: ${focus}` })),
495
520
  ],
496
521
  gate2, report,
@@ -534,7 +559,7 @@ When WAVE mode is detected, author a workflow that wraps the single-ticket engin
534
559
 
535
560
  **Wave workflow structure (author after the SINGLE engine blocks above):**
536
561
 
537
- The wave workflow uses the same phases as SINGLE but wraps them in a wave loop. The integration branch is `wave/<initiative>` — the initiative slug, referenced below as `{slug}` — (or the user's current branch). Each ticket works on the branch setup-task created. The Git agent manages worktrees for parallel-eligible tickets. After every merge: Validate agent (build + test). Escalations accumulate in a list; the final report lists all of them.
562
+ The wave workflow uses the same phases as SINGLE but wraps them in a wave loop. The integration branch is `wave/<initiative>` — the initiative slug, referenced below as `{slug}` — (or the user's current branch). Each ticket works on the branch setup-task created. The Git agent manages worktrees for parallel-eligible tickets. After every merge the workflow spawns the Validate agent (build + test) and undoes a red merge. Escalations accumulate in a list; the final report lists all of them.
538
563
 
539
564
  Wave skeleton — compact reference (see `wave_loop()` doctrine for full semantics):
540
565
 
@@ -545,11 +570,12 @@ Wave skeleton — compact reference (see `wave_loop()` doctrine for full semanti
545
570
  const WAVE_TICKETS = [...remainingTickets]; // the pre-fetch's refs, in input order — the wave block's row order
546
571
  const results = {}; // ticketId → its row of the workflow's return; a ticket with no entry never ran
547
572
  // A row from an engine result: the fields step 3 after the workflow renders
548
- const rowOf = (ticketId, r, merged) => ({ ticket: ticketId, ran: true, verdict: r?.verdict || r?.overallVerdict || null, merged, issuePrLink: r?.issuePrLink || "(none)", evaluateVerdict: r?.gate2?.evaluateVerdict, testVerdict: r?.gate2?.testVerdict, surviving: r?.survivingFindings?.length, coverageComplete: r?.reviewCoverage?.complete });
573
+ const rowOf = (ticketId, r, merged, treeEqual = null) => ({ ticket: ticketId, ran: true, verdict: r?.verdict || r?.overallVerdict || null, merged, treeEqual, issuePrLink: r?.issuePrLink || "(none)", evaluateVerdict: r?.gate2?.evaluateVerdict, testVerdict: r?.gate2?.testVerdict, surviving: r?.survivingFindings?.length, coverageComplete: r?.reviewCoverage?.complete });
549
574
  const MAX_ROUNDS = Math.max(10, remainingTickets.length * 2 + 5); // heuristic; always finite
550
- const waveState = { quarantined: [], round: 0 };
575
+ const MERGE_SHA = /^[0-9a-f]{40}$/; // a merge commit as git prints it: the only Git-reported value that reaches a later prompt
576
+ const waveState = { quarantined: [], round: 0, halted: null };
551
577
  let reAskedThisDeadlock = false; // re-ask guard: re-ask once on empty ready-set, then escalate
552
- while (remainingTickets.length > 0 && waveState.round < MAX_ROUNDS) {
578
+ while (remainingTickets.length > 0 && waveState.round < MAX_ROUNDS && waveState.halted === null) {
553
579
  waveState.round++;
554
580
 
555
581
  // Design agent reader (opus) applies vacuous-truth rule: a ticket with no unmet dependencies is ALWAYS ready.
@@ -582,10 +608,36 @@ Return: {"ready": [...ticket-ids], "blocked": [{"ticket": "id", "namedBlocker":
582
608
  // Check both verdict (engine_output_schema) and overallVerdict (SINGLE skeleton alias).
583
609
  // PASS and UNVERIFIED merge; PARTIAL, FAIL, ESCALATED (the ticket-link and branch stops included) or no verdict quarantine.
584
610
  if (["PASS", "UNVERIFIED"].includes(engineResult.verdict || engineResult.overallVerdict)) {
585
- const merge = await agent(`Merge ${engineResult.branch} to ${INTEGRATION_BRANCH}. Include ticket ID ${ticketId} in the merge commit message. Run Validate agent (build + test) after merge.
586
- Return: {"merged": true} — or {"merged": false, "reason": "<why>"} when the merge or the post-merge build failed and the merge was not kept.`, { agentType: "Git" });
587
- results[ticketId] = rowOf(ticketId, engineResult, merge?.merged === true);
588
- if (merge?.merged !== true) waveState.quarantined.push({ ticket: ticketId, reason: merge?.reason || "merge or post-merge build failed" });
611
+ // Git merges locally and validates nothing: the workflow owns the post-merge Validate (D-WAVE-MERGE-VALIDATE).
612
+ const merge = await agent(`Merge ${engineResult.branch} to ${INTEGRATION_BRANCH} locally. Include ticket ID ${ticketId} in the merge commit message. Do not push, and run no build or test.
613
+ Return: {"merged": true, "mergeSha": "<40-hex merge commit>", "treeEqual": <true when mergeSha^{tree} equals the tree of ${engineResult.branch}'s head>} — or {"merged": false, "reason": "<why>"} when no merge is kept.`, { agentType: "Git" });
614
+ let kept = false;
615
+ let reason = merge?.reason || "merge failed";
616
+ if (merge?.merged === true && !MERGE_SHA.test(String(merge.mergeSha))) {
617
+ // A merge may sit on the integration branch with no SHA to validate or undo it by: stop taking merges.
618
+ waveState.halted = { ticket: ticketId, mergeSha: null, reason: "merge reported without a 40-hex mergeSha; the integration branch may hold an unvalidated merge" };
619
+ reason = waveState.halted.reason;
620
+ } else if (merge?.merged === true) {
621
+ // A merge counts as kept only when this Validate returns PASS. It always runs: treeEqual is recorded, never a reason to skip.
622
+ const post = await agent(`Run build, typecheck, lint, and tests on ${INTEGRATION_BRANCH} after merging ${engineResult.branch} (merge commit ${merge.mergeSha}).
623
+ Follow your Running commands block for every build and test command.
624
+ Return: {"verdict": "PASS" | "FAIL", "details": "..."}`, { agentType: "Validate" });
625
+ if (post?.verdict === "PASS") {
626
+ kept = true;
627
+ } else {
628
+ const undo = await agent(`Undo the merge ${merge.mergeSha} on ${INTEGRATION_BRANCH}, only when no remote branch contains ${merge.mergeSha} and the HEAD of ${INTEGRATION_BRANCH} still equals ${merge.mergeSha}. On the integration branch run exactly this one command and no other that changes a branch or a remote: git reset --keep ${merge.mergeSha}^1
629
+ Return: {"undone": true} — or {"undone": false, "reason": "<why>"} when you did not undo it.`, { agentType: "Git" });
630
+ if (undo?.undone === true) {
631
+ reason = "post-merge build red; merge undone";
632
+ } else {
633
+ // A merge that stays red would poison every later ticket: stop taking merges and rounds, and name the red HEAD.
634
+ reason = `post-merge build red; merge not undone (${undo?.reason || "no reason given"})`;
635
+ waveState.halted = { ticket: ticketId, mergeSha: merge.mergeSha, reason };
636
+ }
637
+ }
638
+ }
639
+ results[ticketId] = rowOf(ticketId, engineResult, kept, merge?.merged === true && typeof merge.treeEqual === "boolean" ? merge.treeEqual : null);
640
+ if (!kept) waveState.quarantined.push({ ticket: ticketId, reason });
589
641
  } else {
590
642
  results[ticketId] = rowOf(ticketId, engineResult, false);
591
643
  waveState.quarantined.push({ ticket: ticketId, reason: engineResult.escalations?.[0]?.description || "engine fail/escalated" });
@@ -596,11 +648,12 @@ Return: {"merged": true} — or {"merged": false, "reason": "<why>"} when the me
596
648
  waveState.quarantined.push({ ticket: ticketId, reason: `engine crash: ${String(err)}` });
597
649
  }
598
650
  remainingTickets = remainingTickets.filter(t => t !== ticketId);
651
+ if (waveState.halted !== null) break; // a red integration HEAD takes no further merge in this round
599
652
  }
600
653
  }
601
654
 
602
655
  // After the wave report is written: one row per wave ticket, in input order. No entry ⇒ never ran (cascade, deadlock, MAX_ROUNDS).
603
- return { tickets: WAVE_TICKETS.map(t => results[t] || { ticket: t, ran: false, verdict: null, merged: false, issuePrLink: "(none)" }), quarantined: waveState.quarantined };
656
+ return { tickets: WAVE_TICKETS.map(t => results[t] || { ticket: t, ran: false, verdict: null, merged: false, treeEqual: null, issuePrLink: "(none)" }), quarantined: waveState.quarantined, halted: waveState.halted };
604
657
  ```
605
658
 
606
659
  ---
@@ -624,6 +677,7 @@ A workflow cannot pause mid-run. After the build/wave workflow returns, you (the
624
677
  In WAVE mode, if no tracking-issue number was resolved in Pre-authoring step 5: state `TRACEABILITY: DEGRADED (no tracking issue for this run)` in the run summary and skip — never skip silently.
625
678
  3. **Compose the wave PR inputs** (WAVE mode only — skip this step entirely in SINGLE mode). The workflow returned `tickets`: one entry per wave ticket, in its input order, each `{ticket, ran, verdict, merged, issuePrLink, evaluateVerdict, testVerdict, surviving, coverageComplete}`. Every value in it is agent-reported and treated as untrusted: nothing below is repaired, and no reference is ever composed from a number.
626
679
  - **Branch check.** Run `git -C "{integration worktree root}" branch --show-current`. Only a name matching `^wave/[a-z0-9][a-z0-9-]{0,59}$` opens a wave PR, and its part after `wave/` is the `{slug}` below. Any other name: record `TRACEABILITY: DEGRADED (not a wave branch)` and go to step 4 with no wave PR.
680
+ - **Halted** (the workflow returned a non-null `halted`: a post-merge build it could not undo, or a merge reported with no SHA): the integration HEAD may be red. Record `Wave PR: skipped (integration HEAD red, wave halted at <halted.ticket>)`, name `halted.mergeSha` (or `none reported`) and `halted.reason` as the red integration HEAD in the run summary, and go to step 4 with no wave PR.
627
681
  - **Nothing merged** (no entry has `merged: true`): record `Wave PR: skipped (nothing merged)` and go to step 4.
628
682
  - Otherwise compose (a) and then (b).
629
683
 
@@ -59,7 +59,13 @@ If present, note its contents as `PREFERENCE_PROFILE`. This will be used to auto
59
59
 
60
60
  **3. Resolve ticket input**
61
61
 
62
- Determine the ticket source (in priority order):
62
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
63
+
64
+ <command-input>
65
+ $ARGUMENTS
66
+ </command-input>
67
+
68
+ Determine the ticket source from `COMMAND_INPUT` (in priority order):
63
69
  - A directory of ticket `.md` files (from `/devflow:dynamic-tickets` output)
64
70
  - A list of candidate issue references or issue URLs
65
71
  - Inline ticket descriptions passed as args
@@ -18,7 +18,13 @@ Explore a codebase area by spawning parallel agents for flow tracing, dependency
18
18
 
19
19
  ## Input
20
20
 
21
- `$ARGUMENTS` contains whatever follows `/explore`:
21
+ What follows `/explore` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
22
+
23
+ <command-input>
24
+ $ARGUMENTS
25
+ </command-input>
26
+
27
+ `COMMAND_INPUT` is one of:
22
28
  - Area description: "how does the auth system work"
23
29
  - Flow question: "trace the request lifecycle"
24
30
  - Empty: use conversation context
@@ -58,6 +64,8 @@ Based on Skim agent findings, spawn 2-3 `Agent(subagent_type="Explore")` agents
58
64
 
59
65
  Adjust explorer focus based on the specific exploration question.
60
66
 
67
+ Ask each explorer for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
68
+
61
69
  ### Phase 4: Synthesize
62
70
 
63
71
  **Produces:** MERGED_FINDINGS