devflow-kit 3.1.0 → 3.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/CHANGELOG.md +36 -0
  2. package/README.md +1 -1
  3. package/dist/cli/agents-view/render.js +69 -15
  4. package/dist/cli/agents-view/state.js +40 -14
  5. package/dist/cli/commands/agents.js +135 -45
  6. package/dist/cli/commands/init.js +128 -53
  7. package/dist/cli/commands/learning.js +36 -8
  8. package/dist/cli/commands/memory.js +35 -14
  9. package/dist/cli/commands/uninstall.js +163 -39
  10. package/dist/commands/code-review.md +0 -2
  11. package/dist/commands/debug.md +14 -11
  12. package/dist/commands/dynamic-build.md +33 -43
  13. package/dist/commands/dynamic-plan.md +8 -2
  14. package/dist/commands/explore.md +9 -3
  15. package/dist/commands/implement.md +20 -16
  16. package/dist/commands/plan.md +13 -9
  17. package/dist/commands/release.md +8 -2
  18. package/dist/commands/research.md +8 -2
  19. package/dist/commands/resolve.md +1 -3
  20. package/dist/commands/self-review.md +0 -2
  21. package/dist/core/agent-frontmatter.js +25 -0
  22. package/dist/core/agent-models.js +198 -36
  23. package/dist/core/agent-state.js +27 -5
  24. package/dist/core/assets.js +1 -1
  25. package/dist/core/feature-config.js +68 -10
  26. package/dist/core/flags.js +24 -0
  27. package/dist/core/learning-queue-cleanup.js +10 -11
  28. package/dist/core/linked-path.js +46 -0
  29. package/dist/core/plugins.js +9 -3
  30. package/dist/core/queue-drain.js +31 -0
  31. package/dist/hud/components/learning-counts.js +54 -8
  32. package/dist/skills/git/references/tracker/github/create-release.md +2 -2
  33. package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
  34. package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
  35. package/dist/targets/claude-code/installer.js +36 -9
  36. package/dist/targets/claude-code/post-install.js +128 -38
  37. package/package.json +1 -1
  38. package/src/assets/agents/code.md +14 -17
  39. package/src/assets/agents/design.md +2 -0
  40. package/src/assets/agents/diagnose.md +2 -0
  41. package/src/assets/agents/evaluate.md +4 -0
  42. package/src/assets/agents/knowledge.md +2 -0
  43. package/src/assets/agents/research.md +2 -0
  44. package/src/assets/agents/review.md +2 -0
  45. package/src/assets/agents/scrutinize.md +4 -0
  46. package/src/assets/agents/simplify.md +4 -0
  47. package/src/assets/agents/skim.md +3 -1
  48. package/src/assets/agents/synthesize.md +6 -0
  49. package/src/assets/agents/test.md +18 -10
  50. package/src/assets/agents/triage.md +2 -0
  51. package/src/assets/agents/validate.md +14 -10
  52. package/src/assets/commands/_partials/_engine.mds +15 -31
  53. package/src/assets/commands/_partials/_knowledge.mds +0 -2
  54. package/src/assets/commands/_partials/_tracker.mds +1 -1
  55. package/src/assets/commands/code-review.mds +0 -2
  56. package/src/assets/commands/debug.mds +13 -8
  57. package/src/assets/commands/dynamic-build.mds +17 -11
  58. package/src/assets/commands/dynamic-plan.mds +7 -1
  59. package/src/assets/commands/explore.mds +9 -1
  60. package/src/assets/commands/implement.mds +19 -13
  61. package/src/assets/commands/plan.mds +12 -8
  62. package/src/assets/commands/release.md +8 -2
  63. package/src/assets/commands/research.mds +8 -2
  64. package/src/assets/commands/resolve.mds +1 -1
  65. package/src/assets/mds/tracker/_github.mds +2 -2
  66. package/src/assets/mds/tracker/_jira.mds +2 -2
  67. package/src/assets/mds/tracker/_linear.mds +2 -2
  68. package/src/assets/scripts/hooks/assets/orchestrator-charter.md +3 -2
  69. package/src/assets/scripts/hooks/background-memory-update +69 -11
  70. package/src/assets/scripts/hooks/capture-prompt +4 -3
  71. package/src/assets/scripts/hooks/capture-question +4 -3
  72. package/src/assets/scripts/hooks/capture-turn +4 -3
  73. package/src/assets/scripts/hooks/ensure-devflow-init +13 -1
  74. package/src/assets/scripts/hooks/ensure-root-gitignore +122 -10
  75. package/src/assets/scripts/hooks/git-marker +71 -0
  76. package/src/assets/scripts/hooks/json-helper.cjs +12 -145
  77. package/src/assets/scripts/hooks/json-parse +24 -129
  78. package/src/assets/scripts/hooks/lib/learning-store.cjs +169 -64
  79. package/src/assets/scripts/hooks/lib/render-decisions.cjs +1 -1
  80. package/src/assets/scripts/hooks/memory-worker +10 -0
  81. package/src/assets/scripts/hooks/pre-compact-memory +66 -14
  82. package/src/assets/scripts/hooks/preamble +9 -1
  83. package/src/assets/scripts/hooks/queue-append +53 -21
  84. package/src/assets/scripts/hooks/session-start-context +108 -29
  85. package/src/assets/scripts/hooks/session-start-memory +33 -11
  86. package/src/assets/skills/test-driven-development/SKILL.md +6 -4
@@ -15,7 +15,13 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
15
15
 
16
16
  ## Input
17
17
 
18
- `$ARGUMENTS` contains whatever follows `/debug`:
18
+ What follows `/debug` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
19
+
20
+ <command-input>
21
+ $ARGUMENTS
22
+ </command-input>
23
+
24
+ `COMMAND_INPUT` is one of:
19
25
  - Bug description: "login fails after session timeout"
20
26
  - Issue reference: "#42"
21
27
  - Empty: use conversation context
@@ -26,8 +32,6 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
26
32
 
27
33
  **Produces:** DECISIONS_CONTEXT
28
34
 
29
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:software-design`, `devflow:testing`. If a skill fails to load, continue without it.
30
-
31
35
  ### Load DECISIONS_CONTEXT
32
36
 
33
37
  The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
@@ -66,9 +70,9 @@ The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Ph
66
70
  **Produces:** HYPOTHESES, BUG_CONTEXT
67
71
  **Requires:** DECISIONS_CONTEXT
68
72
 
69
- If `$ARGUMENTS` opens with a candidate issue reference, fetch the issue:
73
+ If `COMMAND_INPUT` opens with a candidate issue reference, fetch the issue:
70
74
 
71
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
75
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
72
76
 
73
77
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
74
78
 
@@ -119,7 +123,8 @@ Return a structured report:
119
123
  - Status: CONFIRMED / DISPROVED / PARTIAL
120
124
  - Evidence FOR: [list with file:line refs]
121
125
  - Evidence AGAINST: [list with file:line refs]
122
- - Key finding: {one-sentence summary}"
126
+ - Key finding: {one-sentence summary}
127
+ Keep the whole report to at most about 1,500 tokens."
123
128
 
124
129
  Agent(subagent_type="Explore"):
125
130
  "Investigate this bug: {bug_description}
@@ -127,7 +132,7 @@ Agent(subagent_type="Explore"):
127
132
  Hypothesis: {hypothesis B description}
128
133
  Focus area: {specific code area, mechanism, or condition}
129
134
 
130
- [same steps and return format]"
135
+ [same steps and return format; the whole report at most about 1,500 tokens]"
131
136
 
132
137
  Agent(subagent_type="Explore"):
133
138
  "Investigate this bug: {bug_description}
@@ -135,7 +140,7 @@ Agent(subagent_type="Explore"):
135
140
  Hypothesis: {hypothesis C description}
136
141
  Focus area: {specific code area, mechanism, or condition}
137
142
 
138
- [same steps and return format]"
143
+ [same steps and return format; the whole report at most about 1,500 tokens]"
139
144
 
140
145
  (Add more investigators if bug complexity warrants 4-5 hypotheses)
141
146
  ```
@@ -147,7 +152,7 @@ Focus area: {specific code area, mechanism, or condition}
147
152
 
148
153
  Evaluate investigation verdicts before synthesis:
149
154
 
150
- - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias)
155
+ - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias), each reporting at most about 1,500 tokens
151
156
  - **Multiple PARTIAL**: Look for a unifying root cause that explains all partial evidence
152
157
  - **All DISPROVED**: Report honestly — "No root cause identified from initial hypotheses." Generate 2-3 second-round hypotheses if conversation context suggests avenues not yet explored. Loop back to Phase 3.
153
158
 
@@ -256,8 +261,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
256
261
  FILES_CHANGED: {list of files changed}
257
262
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
258
263
 
259
- Load the devflow:feature-knowledge skill and follow its authoring process.
260
-
261
264
  Write the knowledge base to:
262
265
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
263
266
 
@@ -206,7 +206,13 @@ Note the `budget` value from the Workflow tool context (or default to "medium" i
206
206
 
207
207
  **3. Detect mode: SINGLE or WAVE**
208
208
 
209
- - **SINGLE mode:** input is one ticket, one issue, one task description, or one plan document
209
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
210
+
211
+ <command-input>
212
+ $ARGUMENTS
213
+ </command-input>
214
+
215
+ - **SINGLE mode:** `COMMAND_INPUT` is one ticket, one issue, one task description, or one plan document
210
216
  - **WAVE mode:** input is a set of tracker issues (wave labels, milestone, issue list), or the user says "wave" / "all tickets in wave N"
211
217
  - **A `/devflow:dynamic-tickets` ticket directory** (`{worktree}/.devflow/docs/tickets/{slug}/{ts}/`) is WAVE input: in each ticket file (every `.md` there but `tracking-issue.md`), the `**Issue:**` line directly after `**Depends on:**` is one raw `ISSUE_REFS` token, forwarded to the wave's pre-fetch verbatim — never rendered, normalised or re-derived. A ticket file with no `**Issue:**` line, or more than one, contributes no token; name it in the run summary as `not filed`.
212
218
 
@@ -244,11 +250,11 @@ If none found: build proceeds Gate-1-only (Gate 2 skipped with a note). Never re
244
250
  **5. Resolve tracking-issue number (optional)**
245
251
 
246
252
  Check, in priority order:
247
- - An explicit candidate issue reference or issue URL in the user's input (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
253
+ - An explicit candidate issue reference or issue URL in `COMMAND_INPUT` (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
248
254
  - The `**Issue:**` line directly after the H1 of the ticket set's `tracking-issue.md` (written by `/devflow:dynamic-tickets`' filing step; the file is at `{worktree}/.devflow/docs/tickets/{slug}/{ts}/tracking-issue.md`), as a raw token — only when the file holds exactly one `**Issue:**` line
249
255
  - Otherwise: none
250
256
 
251
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
257
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
252
258
 
253
259
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
254
260
 
@@ -329,7 +335,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
329
335
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
330
336
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
331
337
 
332
- When you build or run tests to verify your work, use your "Long-running commands" discipline (background-Bash + Monitor poll) for anything that may run silent >120s, and prefer package-scoped commands.
338
+ When you build or run tests to verify your work, follow your Running commands block.
333
339
 
334
340
  After implementing, commit your changes with a conventional-commit message and report:
335
341
  - Files changed
@@ -341,7 +347,7 @@ After implementing, commit your changes with a conventional-commit message and r
341
347
  const gate1 = await phase("gate1", async () => {
342
348
  // Gate 1 #1 — runs once, immediately after the initial implementation.
343
349
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH}.
344
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
350
+ Follow your Running commands block for every build and test command.
345
351
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
346
352
 
347
353
  if (validation.verdict === "FAIL") {
@@ -396,7 +402,7 @@ ${panel.filter(p => p.verdict === "FAIL").map(p => p.rationale).join("\n")}
396
402
  ISSUE_NUMBER: ${ISSUE_NUMBER}
397
403
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
398
404
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
399
- Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit fixes.`, { agentType: "Code" });
405
+ Self-verify your fix compiles, following your Running commands block. Commit fixes.`, { agentType: "Code" });
400
406
  evalVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
401
407
  }
402
408
  }
@@ -406,7 +412,7 @@ Self-verify your fix compiles (background-Bash + Monitor for any build >120s —
406
412
  const testResult = await agent(`Run scenario-based acceptance tests on branch ${BRANCH} against these criteria:
407
413
  ${CRITERIA || "(none)"}
408
414
  TEST_PLAN: ${TEST_PLAN || "(none)"}
409
- For any test/build command that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog.
415
+ Follow your Running commands block for every scenario command.
410
416
  Cover: functionality, API contracts, performance, and cover every TEST_PLAN scenario. Report: PASS or FAIL per scenario.`, { agentType: "Test" });
411
417
  testVerdict = testResult.verdict;
412
418
  if (testVerdict === "FAIL") {
@@ -416,7 +422,7 @@ ${testResult.failures}
416
422
  ISSUE_NUMBER: ${ISSUE_NUMBER}
417
423
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
418
424
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
419
- Self-verify your fix compiles and the scenarios pass (background-Bash + Monitor for any build/test >120s). Commit fixes.`, { agentType: "Code" });
425
+ Self-verify your fix compiles and the scenarios pass, following your Running commands block. Commit fixes.`, { agentType: "Code" });
420
426
  testVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
421
427
  }
422
428
  }
@@ -530,7 +536,7 @@ ${chunk.map(f => `- ${f.description} (${f.severity})`).join("\n")}
530
536
  ISSUE_NUMBER: ${ISSUE_NUMBER}
531
537
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
532
538
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
533
- Fix all findings in this batch. Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit with conventional-commit message.
539
+ Fix all findings in this batch. Self-verify your fix compiles, following your Running commands block. Commit with conventional-commit message.
534
540
  Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<description of any finding that could not be fixed>"]}`, { agentType: "Code" });
535
541
  chunkResults.push({ chunk, result: r });
536
542
  }
@@ -570,7 +576,7 @@ Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<
570
576
  // this is the build gate before the branch is handed back. See gate1_postcode() cadence.
571
577
  const gate1Final = await phase("gate1-final", async () => {
572
578
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH} (final gate after all fixing).
573
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
579
+ Follow your Running commands block for every build and test command.
574
580
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
575
581
 
576
582
  if (validation.verdict === "FAIL") {
@@ -582,7 +588,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
582
588
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
583
589
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
584
590
  Self-verify your fix compiles. Commit fixes with conventional-commit message.`, { agentType: "Code" });
585
- const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
591
+ const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
586
592
  if (recheck.verdict === "PASS") break;
587
593
  failureDetails = recheck.details || failureDetails;
588
594
  if (attempt === 2) return { verdict: "ESCALATED", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
@@ -594,7 +600,7 @@ Self-verify your fix compiles. Commit fixes with conventional-commit message.`,
594
600
  const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
595
601
 
596
602
  if (scrutiny.codeChanged) {
597
- await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
603
+ await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
598
604
  }
599
605
 
600
606
  return { verdict: "PASS" };
@@ -647,42 +653,26 @@ Write a concise report. Present only surviving findings as outstanding — never
647
653
 
648
654
  This applies to both: multiple Code agents working on a single ticket AND multi-ticket scheduling in a wave.
649
655
 
650
- ### Build execution doctrine — long-running commands (LOAD-BEARING)
651
-
652
- The Workflow runtime KILLS any sub-agent that emits no output for 180 seconds. A cold `cargo build`, `cargo test`, a large `tsc`, `gradle build`, `go build ./...`, etc. routinely runs silent far longer and trips this watchdog (the failure reads `agent stalled on all N attempts`). Plain foreground `Bash` also defaults to a 120s timeout.
653
-
654
- **RULE: any agent (Validate agent, Code agent, Test agent) running a build / test / compile / install that may run silent for more than ~120s MUST run it in the BACKGROUND and POLL — never as a single silent foreground command.**
655
-
656
- Build commands are NEVER wrapped in `sh -c`, `bash -c`, or inline interpreters (`python3 -c`, `node -e`). Invoke commands directly with the step-1 redirect form — permission systems deny wrapper-invoked commands that would be allowed directly.
656
+ ### Build execution doctrine — the Running commands block (LOAD-BEARING)
657
657
 
658
- Mechanical procedure (spike-verified — a workflow sub-agent survived a 253s job this way):
658
+ The Validate, Code and Test agents each carry this `## Running commands` block in their bodies, so each prompt in this engine names the block ("your Running commands block") instead of copying it:
659
659
 
660
- 0. **Pre-load Monitor:** before launching any background task, load the `Monitor` tool via ToolSearch (`select:Monitor`).
661
- 1. Choose ONE unique base path for this run and reuse it verbatim in steps 1–3, e.g. `BASE=/tmp/df-build-<ticket-slug>`. Launch the command with the Bash tool using `run_in_background: true`:
662
- ```
663
- <build/test command> > <BASE>.log 2>&1; echo "EXIT=$?" > <BASE>.done
664
- ```
665
- This returns immediately with a background task id — do NOT block on it.
666
- 2. Arm ONE Monitor that emits a heartbeat well under 180s AND exits when the job finishes:
667
- - description: short, e.g. `await <build cmd>`
668
- - persistent: false
669
- - timeout_ms: comfortably ABOVE the expected job time (e.g. 600000)
670
- - command: `until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
671
-
672
- The `building` heartbeat every 25s (≪ 180s) is delivered as a notification that re-invokes you, so the watchdog never sees a >180s gap. `BUILD_DONE` + the `EXIT=` line signal completion.
660
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
673
661
 
674
- **Exit-code honesty:** the background task's own exit status is meaningless (the trailing `echo` always exits 0). ALWAYS read the `EXIT=` value written inside `<BASE>.done` — that is the authoritative result.
675
-
676
- **Bounded polling:** arm ONE Monitor then stop acting. On Monitor timeout, re-arm at most 2× (never more than 3 total Monitor calls per build). If the build has not finished after 3 Monitor calls: record the state and escalate — never babysit. A Code agent burned 241k tokens polling one build.
677
- 3. When the monitor reports `BUILD_DONE`: the job PASSES iff `<BASE>.done` contains `EXIT=0`. Read `<BASE>.log` for output/failure detail.
662
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
663
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
664
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
665
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
666
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
667
+ - Never re-run a command when nothing it reads has changed.
668
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
669
+ - The same rules hold inside a dynamic Workflow sub-agent.
678
670
 
679
671
  **Cheapest-sufficient validation:** iterate with the fastest check that proves the change — `tsc --noEmit`, `cargo check`, `go vet`, a single test file — the ecosystem's cheapest sufficient signal. Reserve expensive full/optimized builds and whole-suite runs for the final gates only.
680
672
 
681
673
  **One build gate per phase:** batch related fixes, validate once. A fix-pass Code agent runs ONE light check over its whole batch — never several invocations per small fix. Do NOT validate after every individual mutation; validate once after the batch is complete.
682
674
 
683
- **Scope commands to stay short.** During the engine, PREFER crate/package-scoped builds and tests — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — over the whole workspace. The full-workspace regression is the human's job after the wave (the wave already hands the integrated branch back to the user). Scoping keeps most commands under the watchdog window and under budget.
684
-
685
- **Invariants:** heartbeat interval MUST stay well under 180s (25–30s is the tested value); Monitor `timeout_ms` MUST exceed the expected job duration (a too-short timeout kills the poll, not the build). Never substitute a single silent long command for this procedure.
675
+ **Scope commands to stay short.** During the engine, prefer the scoped command for the change over the whole workspace — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — and, when the whole set is needed, one workspace-level command over a per-package loop. After the wave, the full-workspace regression is the human's job (the wave already hands the integrated branch back to the user).
686
676
 
687
677
  ### Implement bundle
688
678
 
@@ -703,7 +693,7 @@ Gate 2 runs at implementation acceptance — this matches devflow's deliberate p
703
693
  ORDER IS LOAD-BEARING. Run exactly in this sequence:
704
694
 
705
695
  1. **Validate agent** — build / typecheck / lint / test
706
- - Build/test commands that may run silent for >~120s MUST follow `build_execution_doctrine()` (background Bash + Monitor poll), or they trip the 180s workflow watchdog.
696
+ - Build/test commands follow `build_execution_doctrine()`: the Validate agent's Running commands block, foreground under an explicit Bash timeout.
707
697
  - FAIL → Code agent fix (max 2 retries) → re-run Validate agent
708
698
  - If still FAIL after 2 retries → escalate (do not loop endlessly)
709
699
  2. **Simplify agent** — reduce complexity, remove duplication
@@ -718,7 +708,7 @@ Depth scales to change size + budget: a trivial one-line fix warrants a lighter
718
708
  1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`).
719
709
  2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed.
720
710
 
721
- It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's "Long-running commands" discipline). The final Gate 1 #2 is the invariant that all written code passes before merge.
711
+ It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's Running commands block). The final Gate 1 #2 is the invariant that all written code passes before merge.
722
712
 
723
713
  ### GATE 2 — Acceptance gate (per ticket, plan-scoped — fires ONCE at implementation acceptance)
724
714
 
@@ -850,7 +840,7 @@ Majority-survives: a finding needs >50% of verification lenses to confirm it. St
850
840
 
851
841
  If no surviving findings: return early (no fixes needed). Any coverageGaps are carried in the return — they block a PASS verdict downstream, not the early exit.
852
842
 
853
- If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's "Long-running commands" discipline). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
843
+ If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's Running commands block, once over its whole batch). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
854
844
 
855
845
  ### Engine invariants (non-negotiable)
856
846
 
@@ -171,12 +171,18 @@ If present, note its contents as `PREFERENCE_PROFILE`. This will be used to auto
171
171
 
172
172
  **3. Resolve ticket input**
173
173
 
174
- Determine the ticket source (in priority order):
174
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
175
+
176
+ <command-input>
177
+ $ARGUMENTS
178
+ </command-input>
179
+
180
+ Determine the ticket source from `COMMAND_INPUT` (in priority order):
175
181
  - A directory of ticket `.md` files (from `/devflow:dynamic-tickets` output)
176
182
  - A list of candidate issue references or issue URLs
177
183
  - Inline ticket descriptions passed as args
178
184
 
179
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
185
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
180
186
 
181
187
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
182
188
 
@@ -15,7 +15,13 @@ Explore a codebase area by spawning parallel agents for flow tracing, dependency
15
15
 
16
16
  ## Input
17
17
 
18
- `$ARGUMENTS` contains whatever follows `/explore`:
18
+ What follows `/explore` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
19
+
20
+ <command-input>
21
+ $ARGUMENTS
22
+ </command-input>
23
+
24
+ `COMMAND_INPUT` is one of:
19
25
  - Area description: "how does the auth system work"
20
26
  - Flow question: "trace the request lifecycle"
21
27
  - Empty: use conversation context
@@ -82,6 +88,8 @@ Based on Skim agent findings, spawn 2-3 `Agent(subagent_type="Explore")` agents
82
88
 
83
89
  Adjust explorer focus based on the specific exploration question.
84
90
 
91
+ Ask each explorer for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
92
+
85
93
  ### Phase 4: Synthesize
86
94
 
87
95
  **Produces:** MERGED_FINDINGS
@@ -159,8 +167,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
159
167
  FILES_CHANGED: {list of files changed}
160
168
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
161
169
 
162
- Load the devflow:feature-knowledge skill and follow its authoring process.
163
-
164
170
  Write the knowledge base to:
165
171
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
166
172
 
@@ -16,13 +16,19 @@ Orchestrate a single task through implementation by spawning specialized agents.
16
16
 
17
17
  ## Input
18
18
 
19
- `$ARGUMENTS` contains whatever follows `/implement`:
19
+ What follows `/implement` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
20
+
21
+ <command-input>
22
+ $ARGUMENTS
23
+ </command-input>
24
+
25
+ `COMMAND_INPUT` is one of:
20
26
  - Plan document path: `.devflow/docs/design/42-jwt-auth.2026-04-07_1430.md` (path to an existing `.md` file)
21
27
  - Issue reference: `#42`
22
28
  - Task description: "implement JWT auth"
23
29
  - Empty: use conversation context
24
30
 
25
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
31
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
26
32
 
27
33
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
28
34
 
@@ -48,8 +54,6 @@ If the user prompt does NOT match re-validation, proceed with the full pipeline
48
54
 
49
55
  **Produces:** TASK_ID, BASE_BRANCH, EXECUTION_PLAN, DECISIONS_CONTEXT, FEATURE_KNOWLEDGE, PR_DESCRIPTION_GUIDANCE, ISSUE_NUMBER, EVIDENCE_POLICY, ISSUE_REQUIRED, APPLY_CONVENTIONS, REQUIRE_NON_AUTHOR_APPROVAL, PR_EXCEPTIONS, TEST_PLAN, EVIDENCE_FILE, PR_TEST_PLAN_BLOCK, REVIEW_PUBLICATION
50
56
 
51
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:patterns`, `devflow:dependency-research`. If a skill fails to load, continue without it.
52
-
53
57
  Record the current branch name as `BASE_BRANCH` - this will be the PR target.
54
58
 
55
59
  **Docs root (D-DOCS-ROOT).** Every `.devflow/docs/` path this command reads or writes lives at the checkout's toplevel, never under the directory the session started in. Resolve `{worktree}` from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`) — by running
@@ -70,7 +74,7 @@ Accept the output only when it is exactly two lines: `exit=0` last and, before i
70
74
 
71
75
  Set `EVIDENCE_POLICY`, `ISSUE_REQUIRED`, `APPLY_CONVENTIONS` and `REQUIRE_NON_AUTHOR_APPROVAL` from the accepted line. Pass agents only the three mechanism inputs, never `EVIDENCE_POLICY`. Report `Evidence policy: {EVIDENCE_POLICY} (source: {SOURCE})`, plus any `WARN` tokens as advisory, once in the final report.
72
76
 
73
- **Plan Document Handling** (when $ARGUMENTS is a path ending in `.md`):
77
+ **Plan Document Handling** (when `COMMAND_INPUT` is a path ending in `.md`):
74
78
  1. Read the plan document from the path provided
75
79
  2. Extract from YAML frontmatter: `execution-strategy`, `context-risk`, `issue` number
76
80
  3. Extract from body: Subtask Breakdown, Implementation Plan, Patterns to Follow, Acceptance Criteria
@@ -81,23 +85,25 @@ Set `EVIDENCE_POLICY`, `ISSUE_REQUIRED`, `APPLY_CONVENTIONS` and `REQUIRE_NON_AU
81
85
 
82
86
  If `PR_DESCRIPTION_GUIDANCE` was not set above (non-plan paths: issue input or task description), set it to `(none)`.
83
87
 
88
+ **Empty input.** When `COMMAND_INPUT` is empty — a plan handoff arrives this way, with the plan already in the conversation — there is no argument to name the branch from, and the Git agent sees none of this conversation. Write the description yourself: one line, taken from the plan's title or, with no plan, from the conversation. Send it as the setup-task `TASK_DESCRIPTION` below.
89
+
84
90
  Spawn Git agent to set up task environment. The Git agent derives the branch name automatically from the issue or task description:
85
91
 
86
92
  ```
87
93
  Agent(subagent_type="Git"):
88
94
  "OPERATION: setup-task
89
95
  BASE_BRANCH: {current branch name}
90
- ISSUE_INPUT: {$ARGUMENTS verbatim, when it is a single whitespace-delimited token that does not end in .md; when it ends in .md, the plan frontmatter's issue value verbatim unless absent or pending — otherwise omit}
91
- TASK_DESCRIPTION: {$ARGUMENTS verbatim, when it is two or more whitespace-delimited tokens — otherwise omit}
96
+ ISSUE_INPUT: {COMMAND_INPUT verbatim, when it is a single whitespace-delimited token that does not end in .md; when it ends in .md, the plan frontmatter's issue value verbatim unless absent or pending — otherwise omit}
97
+ TASK_DESCRIPTION: {COMMAND_INPUT verbatim, when it is two or more whitespace-delimited tokens; when COMMAND_INPUT is empty, the one-line description you wrote above — otherwise omit}
92
98
  ISSUE_REQUIRED: {ISSUE_REQUIRED}
93
99
  APPLY_CONVENTIONS: {APPLY_CONVENTIONS}
94
- PLAN_ARTIFACT_PATH: {path to plan document if $ARGUMENTS ends in .md, otherwise (none)}
100
+ PLAN_ARTIFACT_PATH: {path to plan document if COMMAND_INPUT ends in .md, otherwise (none)}
95
101
  Derive branch name from issue or description, create feature branch, and fetch issue if specified.
96
102
  Return the branch setup summary."
97
103
  ```
98
104
 
99
105
  The issue token is forwarded **unclassified**, and the routing is decided by
100
- SHAPE alone — how many tokens `$ARGUMENTS` has, and whether it ends in `.md`.
106
+ SHAPE alone — how many tokens `COMMAND_INPUT` has, and whether it ends in `.md`.
101
107
 
102
108
  `setup-task` is the one step that has resolved a provider, and therefore the only
103
109
  one that knows what an issue reference looks like on this machine: `#123`,
@@ -158,7 +164,7 @@ Note: the section reaches a public PR body. The Code agent re-checks every line
158
164
 
159
165
  **Test plan.** Before any Code spawn, give the task a test plan in the evidence file `{worktree}/.devflow/docs/evidence-{branch_slug}.md` (`EVIDENCE_FILE`) — unlike the handoff file, it stays after the PR exists. It holds up to three sections, in this order and nothing else: `## Test Plan`; `## Evidence Exceptions`, a byte copy of `PR_EXCEPTIONS` present only while that is not `(none)`; and `## Claims`, always last, so every claim is appended at the end of the file. Create the file if absent; if it exists, replace its `## Test Plan` and `## Evidence Exceptions` sections and keep `## Claims` byte-identical.
160
166
 
161
- Write the `## Test Plan` section: when `$ARGUMENTS` is a plan document with a `## Test Plan` section, copy that section's lines verbatim; otherwise write one TP line per acceptance criterion the plan, the issue or the task text states, numbered from `TP-1`. Never invent a criterion, and word every scenario yourself in plain words: the lines reach the PR body, so a scenario holds no `#`, `@` or `/` — no issue reference, mention, closing keyword target or URL — and the files a TP covers go in its `files:` field. Each line follows the TP-line contract:
167
+ Write the `## Test Plan` section: when `COMMAND_INPUT` is a plan document with a `## Test Plan` section, copy that section's lines verbatim; otherwise write one TP line per acceptance criterion the plan, the issue or the task text states, numbered from `TP-1`. Never invent a criterion, and word every scenario yourself in plain words: the lines reach the PR body, so a scenario holds no `#`, `@` or `/` — no issue reference, mention, closing keyword target or URL — and the files a TP covers go in its `files:` field. Each line follows the TP-line contract:
162
168
 
163
169
  **Test-plan line (TP).** Write every test-plan entry as one line in exactly this shape. `TP_LINE_RE` in `pr-evidence.cjs` parses it and refuses any other line.
164
170
 
@@ -428,7 +434,7 @@ ISSUE_PR_LINK: {ISSUE_PR_LINK captured in Phase 1, or (none)}"
428
434
  After Code agent completes, spawn Validate agent to verify correctness:
429
435
 
430
436
  ```
431
- Agent(subagent_type="Validate", model="haiku"):
437
+ Agent(subagent_type="Validate"):
432
438
  "FILES_CHANGED: {list of files from Code agent output}
433
439
  VALIDATION_SCOPE: full
434
440
  Run build, typecheck, lint, test. Report pass/fail with failure details."
@@ -508,7 +514,7 @@ If Scrutinize agent returns BLOCKED, report to user and halt.
508
514
  If Scrutinize agent made code changes (status: FIXED), spawn Validate agent to verify:
509
515
 
510
516
  ```
511
- Agent(subagent_type="Validate", model="haiku"):
517
+ Agent(subagent_type="Validate"):
512
518
  "FILES_CHANGED: {files modified by Scrutinize agent}
513
519
  VALIDATION_SCOPE: changed-only
514
520
  Verify Scrutinize agent's fixes didn't break anything."
@@ -556,7 +562,7 @@ Validate alignment with request and plan. Report ALIGNED or MISALIGNED with deta
556
562
  ```
557
563
  - Spawn Validate agent to verify fix didn't break tests:
558
564
  ```
559
- Agent(subagent_type="Validate", model="haiku"):
565
+ Agent(subagent_type="Validate"):
560
566
  "FILES_CHANGED: {files modified by fix Code agent}
561
567
  VALIDATION_SCOPE: changed-only"
562
568
  ```
@@ -604,7 +610,7 @@ After every Test agent run — PASS or FAIL, first run or retry — append its T
604
610
  ```
605
611
  - Spawn Validate agent to verify fix didn't break tests:
606
612
  ```
607
- Agent(subagent_type="Validate", model="haiku"):
613
+ Agent(subagent_type="Validate"):
608
614
  "FILES_CHANGED: {files modified by fix Code agent}
609
615
  VALIDATION_SCOPE: changed-only"
610
616
  ```
@@ -740,8 +746,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
740
746
  FILES_CHANGED: {list of files changed}
741
747
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
742
748
 
743
- Load the devflow:feature-knowledge skill and follow its authoring process.
744
-
745
749
  Write the knowledge base to:
746
750
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
747
751
 
@@ -18,13 +18,19 @@ The orchestrator only spawns agents and gates — all analytical work is done by
18
18
 
19
19
  ## Input
20
20
 
21
- `$ARGUMENTS` contains whatever follows `/plan`:
21
+ What follows `/plan` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
22
+
23
+ <command-input>
24
+ $ARGUMENTS
25
+ </command-input>
26
+
27
+ `COMMAND_INPUT` is one of:
22
28
  - Opens with a candidate issue reference → issue mode (one candidate = single-ref, more than one = multi-issue)
23
29
  - Path to existing `.md` file → **error**: "Use /implement with plan documents"
24
30
  - Other text → feature description
25
31
  - Empty → use conversation context
26
32
 
27
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
33
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
28
34
 
29
35
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
30
36
 
@@ -62,7 +68,7 @@ Explore the user's intent through focused Socratic questioning before spawning a
62
68
 
63
69
  **Step 0 — Fetch issue(s)** (issue mode only; skip for feature-description and empty modes):
64
70
 
65
- - **Single-ref** (one candidate ref in `$ARGUMENTS`):
71
+ - **Single-ref** (one candidate ref in `COMMAND_INPUT`):
66
72
 
67
73
  ```
68
74
  Agent(subagent_type="Git"):
@@ -106,8 +112,6 @@ If the user says "skip" or "just proceed" — skip remaining questions, present
106
112
  **Produces:** SKIM_CONTEXT, DECISIONS_CONTEXT, FEATURE_KNOWLEDGE
107
113
  **Requires:** CONFIRMED_SCOPE
108
114
 
109
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:patterns`, `devflow:software-design`, `devflow:security`, `devflow:design-review`. If a skill fails to load, continue without it.
110
-
111
115
  Spawn Skim agent for codebase context:
112
116
 
113
117
  ```
@@ -212,7 +216,7 @@ Pass `FEATURE_KNOWLEDGE` alongside `DECISIONS_CONTEXT` to Explore and Design age
212
216
  **Produces:** EXPLORE_OUTPUTS
213
217
  **Requires:** SKIM_CONTEXT, DECISIONS_CONTEXT
214
218
 
215
- Spawn 4 Explore agents **in a single message**, each with Skim agent context, `DECISIONS_CONTEXT` (from Phase 2), and `FEATURE_KNOWLEDGE` (from Phase 2). Include instructions: "follow `devflow:apply-decisions` for DECISIONS_CONTEXT" and "The FEATURE_KNOWLEDGE is a baseline — VALIDATE, EXTEND, and CORRECT it, don't repeat it. Focus on areas the feature knowledge doesn't cover and changes since it was last updated."
219
+ Spawn 4 Explore agents **in a single message**, each with Skim agent context, `DECISIONS_CONTEXT` (from Phase 2), and `FEATURE_KNOWLEDGE` (from Phase 2). Include instructions: "follow `devflow:apply-decisions` for DECISIONS_CONTEXT" and "The FEATURE_KNOWLEDGE is a baseline — VALIDATE, EXTEND, and CORRECT it, don't repeat it. Focus on areas the feature knowledge doesn't cover and changes since it was last updated." Ask each agent for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
216
220
 
217
221
  | Focus | Thoroughness | Find |
218
222
  |-------|-------------|------|
@@ -355,7 +359,7 @@ User can:
355
359
  **Produces:** IMPL_EXPLORE_OUTPUTS
356
360
  **Requires:** SKIM_CONTEXT, ACCEPTED_SCOPE
357
361
 
358
- Spawn 4 Explore agents **in a single message**, each with Skim agent context + accepted scope:
362
+ Spawn 4 Explore agents **in a single message**, each with Skim agent context + accepted scope. Ask each agent for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
359
363
 
360
364
  | Focus | Thoroughness | Find |
361
365
  |-------|-------------|------|
@@ -384,7 +388,7 @@ Combine into: patterns to follow, integration points, reusable code, edge cases"
384
388
  **Produces:** PLAN_OUTPUTS
385
389
  **Requires:** IMPL_EXPLORATION_SYNTHESIS, GAP_SYNTHESIS, DECISIONS_CONTEXT
386
390
 
387
- Spawn 3 Plan agents **in a single message**, each with implementation exploration synthesis:
391
+ Spawn 3 Plan agents **in a single message**, each with implementation exploration synthesis. Ask each agent for a final report of at most about 1,500 tokens: the plan itself, not a restatement of the exploration.
388
392
 
389
393
  | Focus | Output |
390
394
  |-------|--------|
@@ -573,7 +577,7 @@ Spawn a Git agent with `OPERATION: ensure-traceable-issue`:
573
577
  ```
574
578
  Agent(subagent_type="Git"):
575
579
  "OPERATION: ensure-traceable-issue
576
- ISSUE_INPUT: {the raw candidate token from $ARGUMENTS if /plan was invoked with an issue reference, else omit}
580
+ ISSUE_INPUT: {the raw candidate token from COMMAND_INPUT if /plan was invoked with an issue reference, else omit}
577
581
  TASK_DESCRIPTION: {Gate 0 confirmed scope — one-line title}
578
582
  INITIAL_REQUEST: {the Gate 0 confirmed scope statement}
579
583
  REQUIREMENTS: {discovered requirements summary from Phase 6 gap synthesis}
@@ -17,13 +17,19 @@ Release the project using adaptive learned configuration. On first run, scans th
17
17
 
18
18
  ## Input
19
19
 
20
- `$ARGUMENTS` contains whatever follows `/release`:
20
+ What follows `/release` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
21
+
22
+ <command-input>
23
+ $ARGUMENTS
24
+ </command-input>
25
+
26
+ `COMMAND_INPUT` is one of:
21
27
  - Explicit version: `v1.2.3` or `1.2.3`
22
28
  - Bump type: `patch`, `minor`, `major`
23
29
  - Flag: `--dry-run`
24
30
  - Empty: interactive mode (will ask for version)
25
31
 
26
- Parse from $ARGUMENTS:
32
+ Parse from `COMMAND_INPUT`:
27
33
  - `VERSION`: explicit version string if present (strip leading `v`)
28
34
  - `BUMP_TYPE`: `patch | minor | major` if bump type provided
29
35
  - `DRY_RUN`: true if `--dry-run` present, false otherwise
@@ -16,7 +16,13 @@ Research a topic by spawning parallel Research agents across multiple research t
16
16
 
17
17
  ## Input
18
18
 
19
- `$ARGUMENTS` contains whatever follows `/research`:
19
+ What follows `/research` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
20
+
21
+ <command-input>
22
+ $ARGUMENTS
23
+ </command-input>
24
+
25
+ `COMMAND_INPUT` is one of:
20
26
  - Research question: "best caching strategies"
21
27
  - Comparison question: "compare React vs Svelte for our use case"
22
28
  - Empty: use conversation context
@@ -183,7 +189,7 @@ If external research was skipped due to tool unavailability: inform user.
183
189
  1. If `codebase` type was not in RESEARCH_PLAN → skip
184
190
  2. Check if matching feature knowledge already exists by reading `{worktree}/.devflow/features/index.md` (or globbing frontmatter if absent). If covered → skip
185
191
  3. Use AskUserQuestion: "No feature knowledge exists for {researched area}. Create one?"
186
- 4. If user accepts: spawn `Agent(subagent_type="Knowledge")` with researched area context + worktree root, instructing it to load `devflow:feature-knowledge`, write `KNOWLEDGE.md`, and update `index.md` directly
192
+ 4. If user accepts: spawn `Agent(subagent_type="Knowledge")` with researched area context + worktree root, instructing it to write `KNOWLEDGE.md` and update `index.md` directly
187
193
  5. Set FEATURE_KNOWLEDGE_STATUS = created or skipped
188
194
 
189
195
  **Failure handling**: Non-blocking. If Knowledge agent fails, log and continue.
@@ -361,7 +361,7 @@ If no fixes were made (CODE_AGENT_RESULTS contains 0 commits) → set VERIFICATI
361
361
  Otherwise, spawn Validate agent:
362
362
 
363
363
  ```
364
- Agent(subagent_type="Validate", model="haiku"):
364
+ Agent(subagent_type="Validate"):
365
365
  "FILES_CHANGED: {list of files from Code agent output}
366
366
  VALIDATION_SCOPE: full
367
367
  Run build, typecheck, lint, test. Report pass/fail with failure details."
@@ -665,8 +665,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
665
665
  FILES_CHANGED: {list of files changed}
666
666
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
667
667
 
668
- Load the devflow:feature-knowledge skill and follow its authoring process.
669
-
670
668
  Write the knowledge base to:
671
669
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
672
670
 
@@ -225,8 +225,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
225
225
  FILES_CHANGED: {list of files changed}
226
226
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
227
227
 
228
- Load the devflow:feature-knowledge skill and follow its authoring process.
229
-
230
228
  Write the knowledge base to:
231
229
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
232
230