devflow-kit 3.0.1 → 3.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (166) hide show
  1. package/CHANGELOG.md +49 -0
  2. package/README.md +1 -1
  3. package/dist/agents/git.md +2 -2
  4. package/dist/cli/agents-view/index.js +1 -1
  5. package/dist/cli/agents-view/render.js +71 -17
  6. package/dist/cli/agents-view/state.js +42 -16
  7. package/dist/cli/agents-view/terminal.js +5 -5
  8. package/dist/cli/commands/agents.js +142 -51
  9. package/dist/cli/commands/ambient.js +1 -1
  10. package/dist/cli/commands/attribution-prompts.js +8 -8
  11. package/dist/cli/commands/capture.js +1 -1
  12. package/dist/cli/commands/compliance-prompts.js +8 -8
  13. package/dist/cli/commands/compliance.js +8 -7
  14. package/dist/cli/commands/flags.js +33 -31
  15. package/dist/cli/commands/hud.js +1 -1
  16. package/dist/cli/commands/init-seed.js +9 -9
  17. package/dist/cli/commands/init.js +162 -85
  18. package/dist/cli/commands/install-report.js +10 -10
  19. package/dist/cli/commands/learning.js +302 -136
  20. package/dist/cli/commands/memory.js +36 -15
  21. package/dist/cli/commands/proxy.js +23 -23
  22. package/dist/cli/commands/rules.js +6 -5
  23. package/dist/cli/commands/tracker-prompts.js +6 -6
  24. package/dist/cli/commands/tracker.js +9 -9
  25. package/dist/cli/commands/uninstall.js +183 -59
  26. package/dist/cli/flags-view/render.js +5 -5
  27. package/dist/cli/flags-view/state.js +9 -9
  28. package/dist/cli/flags-view/terminal.js +4 -4
  29. package/dist/cli/tui/cells.js +1 -1
  30. package/dist/cli/tui/terminal.js +6 -6
  31. package/dist/commands/code-review.md +0 -2
  32. package/dist/commands/debug.md +14 -11
  33. package/dist/commands/dynamic-build.md +51 -47
  34. package/dist/commands/dynamic-plan.md +27 -7
  35. package/dist/commands/dynamic-profile.md +17 -3
  36. package/dist/commands/dynamic-tickets.md +18 -4
  37. package/dist/commands/explore.md +9 -3
  38. package/dist/commands/implement.md +20 -16
  39. package/dist/commands/plan.md +13 -9
  40. package/dist/commands/release.md +23 -3
  41. package/dist/commands/research.md +9 -3
  42. package/dist/commands/resolve.md +9 -12
  43. package/dist/commands/self-review.md +0 -2
  44. package/dist/core/agent-frontmatter.js +28 -3
  45. package/dist/core/agent-models.js +204 -42
  46. package/dist/core/agent-state.js +28 -6
  47. package/dist/core/ansi.js +2 -2
  48. package/dist/core/assets.js +1 -1
  49. package/dist/core/cache.js +7 -8
  50. package/dist/core/codex-auth-inspect.js +4 -4
  51. package/dist/core/compliance-compose.js +3 -3
  52. package/dist/core/compliance.js +3 -4
  53. package/dist/core/evidence-policy.js +14 -13
  54. package/dist/core/external-models.js +1 -1
  55. package/dist/core/feature-config.js +71 -13
  56. package/dist/core/feature-switch.js +3 -3
  57. package/dist/core/flags.js +49 -25
  58. package/dist/core/fs-atomic.js +6 -7
  59. package/dist/core/learning-queue-cleanup.js +16 -81
  60. package/dist/core/learning-store.js +61 -0
  61. package/dist/core/linked-path.js +46 -0
  62. package/dist/core/manifest.js +5 -5
  63. package/dist/core/mds-variants.js +13 -13
  64. package/dist/core/model-discovery.js +8 -8
  65. package/dist/core/observations.js +17 -101
  66. package/dist/core/orphan-sweep.js +4 -4
  67. package/dist/core/plugins.js +13 -8
  68. package/dist/core/project-paths.js +9 -13
  69. package/dist/core/proxy-log.js +8 -8
  70. package/dist/core/proxy-state.js +3 -3
  71. package/dist/core/queue-drain.js +31 -0
  72. package/dist/core/reference-sweep.js +6 -6
  73. package/dist/core/teammate-mode-cleanup.js +1 -1
  74. package/dist/core/tracker.js +14 -14
  75. package/dist/hud/colors.js +2 -2
  76. package/dist/hud/components/learning-counts.js +54 -22
  77. package/dist/hud/components/version-badge.js +1 -1
  78. package/dist/skills/git/references/pr/resolve-review-threads.md +2 -2
  79. package/dist/skills/git/references/tracker/github/create-release.md +2 -2
  80. package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
  81. package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
  82. package/dist/targets/claude-code/compliance-install.js +17 -15
  83. package/dist/targets/claude-code/hooks.js +2 -2
  84. package/dist/targets/claude-code/installer.js +59 -32
  85. package/dist/targets/claude-code/legacy.js +1 -1
  86. package/dist/targets/claude-code/post-install.js +135 -45
  87. package/dist/targets/claude-code/tracker-install.js +2 -2
  88. package/package.json +1 -1
  89. package/src/assets/agents/code.md +15 -21
  90. package/src/assets/agents/design.md +4 -2
  91. package/src/assets/agents/diagnose.md +3 -1
  92. package/src/assets/agents/evaluate.md +4 -0
  93. package/src/assets/agents/git.mds +2 -2
  94. package/src/assets/agents/knowledge.md +5 -3
  95. package/src/assets/agents/learning.md +281 -196
  96. package/src/assets/agents/research.md +3 -1
  97. package/src/assets/agents/review.md +5 -3
  98. package/src/assets/agents/scrutinize.md +5 -1
  99. package/src/assets/agents/simplify.md +4 -0
  100. package/src/assets/agents/skim.md +4 -2
  101. package/src/assets/agents/synthesize.md +6 -0
  102. package/src/assets/agents/test.md +18 -10
  103. package/src/assets/agents/triage.md +11 -9
  104. package/src/assets/agents/validate.md +14 -10
  105. package/src/assets/commands/_partials/_decisions.mds +8 -3
  106. package/src/assets/commands/_partials/_docs_root.mds +3 -3
  107. package/src/assets/commands/_partials/_engine.mds +16 -32
  108. package/src/assets/commands/_partials/_knowledge.mds +0 -2
  109. package/src/assets/commands/_partials/_preamble.mds +6 -2
  110. package/src/assets/commands/_partials/_settings.mds +2 -2
  111. package/src/assets/commands/_partials/_tracker.mds +1 -1
  112. package/src/assets/commands/code-review.mds +0 -2
  113. package/src/assets/commands/debug.mds +13 -8
  114. package/src/assets/commands/dynamic-build.mds +18 -12
  115. package/src/assets/commands/dynamic-plan.mds +10 -4
  116. package/src/assets/commands/dynamic-profile.mds +1 -1
  117. package/src/assets/commands/dynamic-tickets.mds +2 -2
  118. package/src/assets/commands/explore.mds +9 -1
  119. package/src/assets/commands/implement.mds +19 -13
  120. package/src/assets/commands/plan.mds +12 -8
  121. package/src/assets/commands/release.md +23 -3
  122. package/src/assets/commands/research.mds +9 -3
  123. package/src/assets/commands/resolve.mds +9 -10
  124. package/src/assets/mds/git/_pr.mds +3 -3
  125. package/src/assets/mds/tracker/_common.mds +1 -1
  126. package/src/assets/mds/tracker/_github.mds +3 -3
  127. package/src/assets/mds/tracker/_jira.mds +3 -3
  128. package/src/assets/mds/tracker/_linear.mds +3 -3
  129. package/src/assets/mds/tracker/_mcp.mds +6 -5
  130. package/src/assets/scripts/hooks/assets/orchestrator-charter.md +4 -2
  131. package/src/assets/scripts/hooks/background-memory-update +97 -33
  132. package/src/assets/scripts/hooks/capture-prompt +4 -3
  133. package/src/assets/scripts/hooks/capture-question +4 -3
  134. package/src/assets/scripts/hooks/capture-turn +5 -20
  135. package/src/assets/scripts/hooks/ensure-devflow-init +14 -2
  136. package/src/assets/scripts/hooks/ensure-proxy +5 -6
  137. package/src/assets/scripts/hooks/ensure-root-gitignore +123 -11
  138. package/src/assets/scripts/hooks/git-marker +71 -0
  139. package/src/assets/scripts/hooks/is-hex-sha +1 -1
  140. package/src/assets/scripts/hooks/json-helper.cjs +345 -944
  141. package/src/assets/scripts/hooks/json-parse +25 -129
  142. package/src/assets/scripts/hooks/lib/decisions-format.cjs +205 -156
  143. package/src/assets/scripts/hooks/lib/learning-store.cjs +3207 -0
  144. package/src/assets/scripts/hooks/lib/mkdir-lock.cjs +7 -5
  145. package/src/assets/scripts/hooks/lib/project-paths.cjs +13 -19
  146. package/src/assets/scripts/hooks/lib/render-decisions.cjs +253 -226
  147. package/src/assets/scripts/hooks/memory-worker +10 -0
  148. package/src/assets/scripts/hooks/pre-compact-memory +66 -14
  149. package/src/assets/scripts/hooks/preamble +9 -1
  150. package/src/assets/scripts/hooks/queue-append +55 -23
  151. package/src/assets/scripts/hooks/resolve-project-root +3 -4
  152. package/src/assets/scripts/hooks/session-start-context +146 -45
  153. package/src/assets/scripts/hooks/session-start-memory +33 -11
  154. package/src/assets/scripts/lib/project-config.cjs +2 -2
  155. package/src/assets/scripts/pr-evidence.cjs +3 -3
  156. package/src/assets/scripts/redact-secrets.cjs +20 -20
  157. package/src/assets/scripts/release-trace.cjs +1 -1
  158. package/src/assets/scripts/resolve-evidence-policy.cjs +3 -3
  159. package/src/assets/scripts/resolve-settings.cjs +3 -3
  160. package/src/assets/scripts/verify-evidence.cjs +2 -2
  161. package/src/assets/skills/apply-decisions/SKILL.md +37 -17
  162. package/src/assets/skills/docs-framework/SKILL.md +2 -2
  163. package/src/assets/skills/feature-knowledge/SKILL.md +6 -5
  164. package/src/assets/skills/test-driven-development/SKILL.md +6 -4
  165. package/dist/core/observation-io.js +0 -50
  166. package/src/assets/scripts/hooks/decisions-usage-scan.cjs +0 -131
@@ -1,10 +1,10 @@
1
1
  /**
2
2
  * Generic TUI shell — shared by agents-view and flags-view.
3
3
  *
4
- * applies ADR-013: impure I/O shell in CLI layer; pure logic lives in state + render.
5
- * avoids PF-014: cleanup wired via Promise resolve — never process.exit() inside
6
- * a finally-guarded scope.
7
- * avoids PF-017: one generic shell, thin adapters per TUI — not copy-adapted per consumer.
4
+ * Impure I/O shell in CLI layer; pure logic lives in state + render.
5
+ * Cleanup wired via Promise resolve — never process.exit() inside
6
+ * a finally-guarded scope (it would skip the finally).
7
+ * One generic shell, thin adapters per TUI — not copy-adapted per consumer.
8
8
  *
9
9
  * Bounded: MAX_KEYPRESSES = 50_000 hard limit (reliability rule — every loop bounded).
10
10
  *
@@ -268,8 +268,8 @@ export async function runTui(spec) {
268
268
  * A throw inside an EventEmitter listener does NOT reject the enclosing
269
269
  * promise — it escapes as an uncaughtException and kills the process with
270
270
  * cleanup() never having run, leaving raw mode and alt-screen set. Routing
271
- * every handler failure through here keeps the PF-014 invariant (cleanup
272
- * always runs) while still surfacing the error rather than swallowing it.
271
+ * every handler failure through here keeps the invariant that cleanup
272
+ * always runs, while still surfacing the error rather than swallowing it.
273
273
  */
274
274
  function fail(err) {
275
275
  cleanup();
@@ -228,8 +228,6 @@ Only `exit=0` means the skill is installed. On any other result, do NOT add that
228
228
 
229
229
  **Produces:** DECISIONS_CONTEXT, FEATURE_KNOWLEDGE, COMPLIANCE_FRAMEWORKS (carried from Step 0b)
230
230
 
231
- **Load Companion Skills** — Load via Skill tool: `devflow:quality-gates`, `devflow:software-design`. If a skill fails to load, continue without it.
232
-
233
231
  ### Load DECISIONS_CONTEXT
234
232
 
235
233
  The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
@@ -15,7 +15,13 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
15
15
 
16
16
  ## Input
17
17
 
18
- `$ARGUMENTS` contains whatever follows `/debug`:
18
+ What follows `/debug` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
19
+
20
+ <command-input>
21
+ $ARGUMENTS
22
+ </command-input>
23
+
24
+ `COMMAND_INPUT` is one of:
19
25
  - Bug description: "login fails after session timeout"
20
26
  - Issue reference: "#42"
21
27
  - Empty: use conversation context
@@ -26,8 +32,6 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
26
32
 
27
33
  **Produces:** DECISIONS_CONTEXT
28
34
 
29
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:software-design`, `devflow:testing`. If a skill fails to load, continue without it.
30
-
31
35
  ### Load DECISIONS_CONTEXT
32
36
 
33
37
  The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
@@ -66,9 +70,9 @@ The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Ph
66
70
  **Produces:** HYPOTHESES, BUG_CONTEXT
67
71
  **Requires:** DECISIONS_CONTEXT
68
72
 
69
- If `$ARGUMENTS` opens with a candidate issue reference, fetch the issue:
73
+ If `COMMAND_INPUT` opens with a candidate issue reference, fetch the issue:
70
74
 
71
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
75
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
72
76
 
73
77
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
74
78
 
@@ -119,7 +123,8 @@ Return a structured report:
119
123
  - Status: CONFIRMED / DISPROVED / PARTIAL
120
124
  - Evidence FOR: [list with file:line refs]
121
125
  - Evidence AGAINST: [list with file:line refs]
122
- - Key finding: {one-sentence summary}"
126
+ - Key finding: {one-sentence summary}
127
+ Keep the whole report to at most about 1,500 tokens."
123
128
 
124
129
  Agent(subagent_type="Explore"):
125
130
  "Investigate this bug: {bug_description}
@@ -127,7 +132,7 @@ Agent(subagent_type="Explore"):
127
132
  Hypothesis: {hypothesis B description}
128
133
  Focus area: {specific code area, mechanism, or condition}
129
134
 
130
- [same steps and return format]"
135
+ [same steps and return format; the whole report at most about 1,500 tokens]"
131
136
 
132
137
  Agent(subagent_type="Explore"):
133
138
  "Investigate this bug: {bug_description}
@@ -135,7 +140,7 @@ Agent(subagent_type="Explore"):
135
140
  Hypothesis: {hypothesis C description}
136
141
  Focus area: {specific code area, mechanism, or condition}
137
142
 
138
- [same steps and return format]"
143
+ [same steps and return format; the whole report at most about 1,500 tokens]"
139
144
 
140
145
  (Add more investigators if bug complexity warrants 4-5 hypotheses)
141
146
  ```
@@ -147,7 +152,7 @@ Focus area: {specific code area, mechanism, or condition}
147
152
 
148
153
  Evaluate investigation verdicts before synthesis:
149
154
 
150
- - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias)
155
+ - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias), each reporting at most about 1,500 tokens
151
156
  - **Multiple PARTIAL**: Look for a unifying root cause that explains all partial evidence
152
157
  - **All DISPROVED**: Report honestly — "No root cause identified from initial hypotheses." Generate 2-3 second-round hypotheses if conversation context suggests avenues not yet explored. Loop back to Phase 3.
153
158
 
@@ -256,8 +261,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
256
261
  FILES_CHANGED: {list of files changed}
257
262
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
258
263
 
259
- Load the devflow:feature-knowledge skill and follow its authoring process.
260
-
261
264
  Write the knowledge base to:
262
265
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
263
266
 
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
65
65
 
66
66
  ### DECISIONS_CONTEXT — obtain BEFORE authoring
67
67
 
68
- Before you author the workflow script, read `.devflow/learning/index.md` for the current worktree. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
68
+ The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
69
+
70
+ ```bash
71
+ git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
72
+ ```
73
+
74
+ Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
75
+
76
+ 1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
77
+ 2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
78
+ 3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
79
+
80
+ This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
81
+
82
+ Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
69
83
 
70
84
  The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
71
85
 
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
73
87
 
74
88
  When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
75
89
 
76
- ### IRON RULE (ADR-008: LLM-vs-plumbing)
90
+ ### IRON RULE (LLM-vs-plumbing)
77
91
 
78
92
  **Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
79
93
 
@@ -192,7 +206,13 @@ Note the `budget` value from the Workflow tool context (or default to "medium" i
192
206
 
193
207
  **3. Detect mode: SINGLE or WAVE**
194
208
 
195
- - **SINGLE mode:** input is one ticket, one issue, one task description, or one plan document
209
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
210
+
211
+ <command-input>
212
+ $ARGUMENTS
213
+ </command-input>
214
+
215
+ - **SINGLE mode:** `COMMAND_INPUT` is one ticket, one issue, one task description, or one plan document
196
216
  - **WAVE mode:** input is a set of tracker issues (wave labels, milestone, issue list), or the user says "wave" / "all tickets in wave N"
197
217
  - **A `/devflow:dynamic-tickets` ticket directory** (`{worktree}/.devflow/docs/tickets/{slug}/{ts}/`) is WAVE input: in each ticket file (every `.md` there but `tracking-issue.md`), the `**Issue:**` line directly after `**Depends on:**` is one raw `ISSUE_REFS` token, forwarded to the wave's pre-fetch verbatim — never rendered, normalised or re-derived. A ticket file with no `**Issue:**` line, or more than one, contributes no token; name it in the run summary as `not filed`.
198
218
 
@@ -230,11 +250,11 @@ If none found: build proceeds Gate-1-only (Gate 2 skipped with a note). Never re
230
250
  **5. Resolve tracking-issue number (optional)**
231
251
 
232
252
  Check, in priority order:
233
- - An explicit candidate issue reference or issue URL in the user's input (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
253
+ - An explicit candidate issue reference or issue URL in `COMMAND_INPUT` (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
234
254
  - The `**Issue:**` line directly after the H1 of the ticket set's `tracking-issue.md` (written by `/devflow:dynamic-tickets`' filing step; the file is at `{worktree}/.devflow/docs/tickets/{slug}/{ts}/tracking-issue.md`), as a raw token — only when the file holds exactly one `**Issue:**` line
235
255
  - Otherwise: none
236
256
 
237
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
257
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
238
258
 
239
259
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
240
260
 
@@ -315,7 +335,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
315
335
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
316
336
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
317
337
 
318
- When you build or run tests to verify your work, use your "Long-running commands" discipline (background-Bash + Monitor poll) for anything that may run silent >120s, and prefer package-scoped commands.
338
+ When you build or run tests to verify your work, follow your Running commands block.
319
339
 
320
340
  After implementing, commit your changes with a conventional-commit message and report:
321
341
  - Files changed
@@ -327,7 +347,7 @@ After implementing, commit your changes with a conventional-commit message and r
327
347
  const gate1 = await phase("gate1", async () => {
328
348
  // Gate 1 #1 — runs once, immediately after the initial implementation.
329
349
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH}.
330
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
350
+ Follow your Running commands block for every build and test command.
331
351
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
332
352
 
333
353
  if (validation.verdict === "FAIL") {
@@ -382,7 +402,7 @@ ${panel.filter(p => p.verdict === "FAIL").map(p => p.rationale).join("\n")}
382
402
  ISSUE_NUMBER: ${ISSUE_NUMBER}
383
403
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
384
404
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
385
- Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit fixes.`, { agentType: "Code" });
405
+ Self-verify your fix compiles, following your Running commands block. Commit fixes.`, { agentType: "Code" });
386
406
  evalVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
387
407
  }
388
408
  }
@@ -392,7 +412,7 @@ Self-verify your fix compiles (background-Bash + Monitor for any build >120s —
392
412
  const testResult = await agent(`Run scenario-based acceptance tests on branch ${BRANCH} against these criteria:
393
413
  ${CRITERIA || "(none)"}
394
414
  TEST_PLAN: ${TEST_PLAN || "(none)"}
395
- For any test/build command that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog.
415
+ Follow your Running commands block for every scenario command.
396
416
  Cover: functionality, API contracts, performance, and cover every TEST_PLAN scenario. Report: PASS or FAIL per scenario.`, { agentType: "Test" });
397
417
  testVerdict = testResult.verdict;
398
418
  if (testVerdict === "FAIL") {
@@ -402,7 +422,7 @@ ${testResult.failures}
402
422
  ISSUE_NUMBER: ${ISSUE_NUMBER}
403
423
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
404
424
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
405
- Self-verify your fix compiles and the scenarios pass (background-Bash + Monitor for any build/test >120s). Commit fixes.`, { agentType: "Code" });
425
+ Self-verify your fix compiles and the scenarios pass, following your Running commands block. Commit fixes.`, { agentType: "Code" });
406
426
  testVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
407
427
  }
408
428
  }
@@ -516,7 +536,7 @@ ${chunk.map(f => `- ${f.description} (${f.severity})`).join("\n")}
516
536
  ISSUE_NUMBER: ${ISSUE_NUMBER}
517
537
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
518
538
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
519
- Fix all findings in this batch. Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit with conventional-commit message.
539
+ Fix all findings in this batch. Self-verify your fix compiles, following your Running commands block. Commit with conventional-commit message.
520
540
  Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<description of any finding that could not be fixed>"]}`, { agentType: "Code" });
521
541
  chunkResults.push({ chunk, result: r });
522
542
  }
@@ -556,7 +576,7 @@ Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<
556
576
  // this is the build gate before the branch is handed back. See gate1_postcode() cadence.
557
577
  const gate1Final = await phase("gate1-final", async () => {
558
578
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH} (final gate after all fixing).
559
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
579
+ Follow your Running commands block for every build and test command.
560
580
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
561
581
 
562
582
  if (validation.verdict === "FAIL") {
@@ -568,7 +588,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
568
588
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
569
589
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
570
590
  Self-verify your fix compiles. Commit fixes with conventional-commit message.`, { agentType: "Code" });
571
- const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
591
+ const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
572
592
  if (recheck.verdict === "PASS") break;
573
593
  failureDetails = recheck.details || failureDetails;
574
594
  if (attempt === 2) return { verdict: "ESCALATED", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
@@ -580,7 +600,7 @@ Self-verify your fix compiles. Commit fixes with conventional-commit message.`,
580
600
  const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
581
601
 
582
602
  if (scrutiny.codeChanged) {
583
- await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
603
+ await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
584
604
  }
585
605
 
586
606
  return { verdict: "PASS" };
@@ -633,42 +653,26 @@ Write a concise report. Present only surviving findings as outstanding — never
633
653
 
634
654
  This applies to both: multiple Code agents working on a single ticket AND multi-ticket scheduling in a wave.
635
655
 
636
- ### Build execution doctrine — long-running commands (LOAD-BEARING)
656
+ ### Build execution doctrine — the Running commands block (LOAD-BEARING)
637
657
 
638
- The Workflow runtime KILLS any sub-agent that emits no output for 180 seconds. A cold `cargo build`, `cargo test`, a large `tsc`, `gradle build`, `go build ./...`, etc. routinely runs silent far longer and trips this watchdog (the failure reads `agent stalled on all N attempts`). Plain foreground `Bash` also defaults to a 120s timeout.
658
+ The Validate, Code and Test agents each carry this `## Running commands` block in their bodies, so each prompt in this engine names the block ("your Running commands block") instead of copying it:
639
659
 
640
- **RULE: any agent (Validate agent, Code agent, Test agent) running a build / test / compile / install that may run silent for more than ~120s MUST run it in the BACKGROUND and POLL — never as a single silent foreground command.**
660
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
641
661
 
642
- Build commands are NEVER wrapped in `sh -c`, `bash -c`, or inline interpreters (`python3 -c`, `node -e`). Invoke commands directly with the step-1 redirect form — permission systems deny wrapper-invoked commands that would be allowed directly.
643
-
644
- Mechanical procedure (spike-verified — a workflow sub-agent survived a 253s job this way):
645
-
646
- 0. **Pre-load Monitor:** before launching any background task, load the `Monitor` tool via ToolSearch (`select:Monitor`).
647
- 1. Choose ONE unique base path for this run and reuse it verbatim in steps 1–3, e.g. `BASE=/tmp/df-build-<ticket-slug>`. Launch the command with the Bash tool using `run_in_background: true`:
648
- ```
649
- <build/test command> > <BASE>.log 2>&1; echo "EXIT=$?" > <BASE>.done
650
- ```
651
- This returns immediately with a background task id — do NOT block on it.
652
- 2. Arm ONE Monitor that emits a heartbeat well under 180s AND exits when the job finishes:
653
- - description: short, e.g. `await <build cmd>`
654
- - persistent: false
655
- - timeout_ms: comfortably ABOVE the expected job time (e.g. 600000)
656
- - command: `until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
657
-
658
- The `building` heartbeat every 25s (≪ 180s) is delivered as a notification that re-invokes you, so the watchdog never sees a >180s gap. `BUILD_DONE` + the `EXIT=` line signal completion.
659
-
660
- **Exit-code honesty:** the background task's own exit status is meaningless (the trailing `echo` always exits 0). ALWAYS read the `EXIT=` value written inside `<BASE>.done` — that is the authoritative result.
661
-
662
- **Bounded polling:** arm ONE Monitor then stop acting. On Monitor timeout, re-arm at most 2× (never more than 3 total Monitor calls per build). If the build has not finished after 3 Monitor calls: record the state and escalate — never babysit. A Code agent burned 241k tokens polling one build.
663
- 3. When the monitor reports `BUILD_DONE`: the job PASSES iff `<BASE>.done` contains `EXIT=0`. Read `<BASE>.log` for output/failure detail.
662
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
663
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
664
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
665
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
666
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
667
+ - Never re-run a command when nothing it reads has changed.
668
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
669
+ - The same rules hold inside a dynamic Workflow sub-agent.
664
670
 
665
671
  **Cheapest-sufficient validation:** iterate with the fastest check that proves the change — `tsc --noEmit`, `cargo check`, `go vet`, a single test file — the ecosystem's cheapest sufficient signal. Reserve expensive full/optimized builds and whole-suite runs for the final gates only.
666
672
 
667
673
  **One build gate per phase:** batch related fixes, validate once. A fix-pass Code agent runs ONE light check over its whole batch — never several invocations per small fix. Do NOT validate after every individual mutation; validate once after the batch is complete.
668
674
 
669
- **Scope commands to stay short.** During the engine, PREFER crate/package-scoped builds and tests — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — over the whole workspace. The full-workspace regression is the human's job after the wave (the wave already hands the integrated branch back to the user). Scoping keeps most commands under the watchdog window and under budget.
670
-
671
- **Invariants:** heartbeat interval MUST stay well under 180s (25–30s is the tested value); Monitor `timeout_ms` MUST exceed the expected job duration (a too-short timeout kills the poll, not the build). Never substitute a single silent long command for this procedure.
675
+ **Scope commands to stay short.** During the engine, prefer the scoped command for the change over the whole workspace — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — and, when the whole set is needed, one workspace-level command over a per-package loop. After the wave, the full-workspace regression is the human's job (the wave already hands the integrated branch back to the user).
672
676
 
673
677
  ### Implement bundle
674
678
 
@@ -680,7 +684,7 @@ Code(agentType:"Code", prompt: full task + plan + DECISIONS_CONTEXT + handoff if
680
684
  → gate2_acceptance() ← Gate 2 runs HERE — before the review pass, not after
681
685
  ```
682
686
 
683
- The Code agent prompt must include: task description, implementation plan (if one exists), relevant DECISIONS_CONTEXT (from `.devflow/learning/index.md`), the compliance lens (`COMPLIANCE_FRAMEWORKS` — every Code prompt carries it, fix prompts included), and any PRIOR_PHASE_SUMMARY / HANDOFF_FILE for sequential multi-phase tickets.
687
+ The Code agent prompt must include: task description, implementation plan (if one exists), relevant DECISIONS_CONTEXT (the index you loaded before authoring), the compliance lens (`COMPLIANCE_FRAMEWORKS` — every Code prompt carries it, fix prompts included), and any PRIOR_PHASE_SUMMARY / HANDOFF_FILE for sequential multi-phase tickets.
684
688
 
685
689
  Gate 2 runs at implementation acceptance — this matches devflow's deliberate placement: "evaluation is part of implementation acceptance, not post-review" (§6.1).
686
690
 
@@ -689,7 +693,7 @@ Gate 2 runs at implementation acceptance — this matches devflow's deliberate p
689
693
  ORDER IS LOAD-BEARING. Run exactly in this sequence:
690
694
 
691
695
  1. **Validate agent** — build / typecheck / lint / test
692
- - Build/test commands that may run silent for >~120s MUST follow `build_execution_doctrine()` (background Bash + Monitor poll), or they trip the 180s workflow watchdog.
696
+ - Build/test commands follow `build_execution_doctrine()`: the Validate agent's Running commands block, foreground under an explicit Bash timeout.
693
697
  - FAIL → Code agent fix (max 2 retries) → re-run Validate agent
694
698
  - If still FAIL after 2 retries → escalate (do not loop endlessly)
695
699
  2. **Simplify agent** — reduce complexity, remove duplication
@@ -704,7 +708,7 @@ Depth scales to change size + budget: a trivial one-line fix warrants a lighter
704
708
  1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`).
705
709
  2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed.
706
710
 
707
- It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's "Long-running commands" discipline). The final Gate 1 #2 is the invariant that all written code passes before merge.
711
+ It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's Running commands block). The final Gate 1 #2 is the invariant that all written code passes before merge.
708
712
 
709
713
  ### GATE 2 — Acceptance gate (per ticket, plan-scoped — fires ONCE at implementation acceptance)
710
714
 
@@ -836,7 +840,7 @@ Majority-survives: a finding needs >50% of verification lenses to confirm it. St
836
840
 
837
841
  If no surviving findings: return early (no fixes needed). Any coverageGaps are carried in the return — they block a PASS verdict downstream, not the early exit.
838
842
 
839
- If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's "Long-running commands" discipline). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
843
+ If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's Running commands block, once over its whole batch). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
840
844
 
841
845
  ### Engine invariants (non-negotiable)
842
846
 
@@ -1240,4 +1244,4 @@ Update the wave PR's test-plan block and post its evidence comment."
1240
1244
 
1241
1245
  ### Maintenance note
1242
1246
 
1243
- This recipe encodes the current `/implement` + `/code-review` + `/resolve` orchestration shape as of the authoring date (2026-06-12). When those base commands change their orchestration, update this recipe to match. No tooling detects drift — by design (ADR-008 Iron Rule). The reminder lives in the design doc §16.
1247
+ This recipe encodes the current `/implement` + `/code-review` + `/resolve` orchestration shape as of the authoring date (2026-06-12). When those base commands change their orchestration, update this recipe to match. No tooling detects drift — by design (the LLM-vs-plumbing Iron Rule). The reminder lives in the design doc §16.
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
65
65
 
66
66
  ### DECISIONS_CONTEXT — obtain BEFORE authoring
67
67
 
68
- Before you author the workflow script, read `.devflow/learning/index.md` for the current worktree. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
68
+ The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
69
+
70
+ ```bash
71
+ git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
72
+ ```
73
+
74
+ Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
75
+
76
+ 1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
77
+ 2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
78
+ 3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
79
+
80
+ This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
81
+
82
+ Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
69
83
 
70
84
  The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
71
85
 
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
73
87
 
74
88
  When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
75
89
 
76
- ### IRON RULE (ADR-008: LLM-vs-plumbing)
90
+ ### IRON RULE (LLM-vs-plumbing)
77
91
 
78
92
  **Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
79
93
 
@@ -157,12 +171,18 @@ If present, note its contents as `PREFERENCE_PROFILE`. This will be used to auto
157
171
 
158
172
  **3. Resolve ticket input**
159
173
 
160
- Determine the ticket source (in priority order):
174
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
175
+
176
+ <command-input>
177
+ $ARGUMENTS
178
+ </command-input>
179
+
180
+ Determine the ticket source from `COMMAND_INPUT` (in priority order):
161
181
  - A directory of ticket `.md` files (from `/devflow:dynamic-tickets` output)
162
182
  - A list of candidate issue references or issue URLs
163
183
  - Inline ticket descriptions passed as args
164
184
 
165
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
185
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
166
186
 
167
187
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
168
188
 
@@ -233,7 +253,7 @@ const plans = await phase("plan-parallel", () =>
233
253
  parallel((tickets || []).map(ticket => () =>
234
254
  agent(`Write an implementation plan for this ticket.
235
255
  Ticket: ${JSON.stringify(ticket)}
236
- Decisions context (apply devflow:apply-decisions; cite ADR/PF IDs): ${DECISIONS_CONTEXT}
256
+ Decisions context (apply devflow:apply-decisions; a plan file can be posted to the tracker, so state each decision in words, never by ID): ${DECISIONS_CONTEXT}
237
257
  The plan must cover: approach overview, affected files and modules, key design decisions, implementation sequence (what to build first), risks and mitigations, and any open questions you cannot resolve from the ticket alone.
238
258
  Write a thorough but tight plan — every section must earn its place for a Code agent who has no other context.
239
259
  Return: { ticketTitle, planMarkdown, openDecisions (array of genuine unknowns requiring user input) }.`, { agentType: "Design" })
@@ -324,7 +344,7 @@ For each ticket, write ${OUTDIR}/{ticket-slug}-plan.md containing:
324
344
  - ## Acceptance Criteria (numbered, positive + negative)
325
345
  - ## Test Plan — the challenger's testPlan lines, verbatim, one per line, and nothing else: no prose, no blank line between them, no setup or outcome
326
346
  - ## Test Scenarios — one line per TP, in TP order: TP-n: its setup, then its expected outcome, from testScenarios
327
- - ## Auto-Resolved Decisions (if any — list each as: decision → resolution → source)
347
+ - ## Auto-Resolved Decisions (if any — list each as: decision → resolution → source, naming a recorded decision in words, never by ID)
328
348
 
329
349
  Then write ${OUTDIR}/DECISIONS-NEEDED.md:
330
350
  - ## Auto-Resolved Decisions — list each silently-resolved decision as: decision → resolution → source (preference profile / ADR-NNN), so auto-resolution is auditable and reversible. If none, write "None."
@@ -430,4 +450,4 @@ After the workflow returns: check each plan's test plan (step 1 of the F4 list a
430
450
 
431
451
  ### Maintenance note
432
452
 
433
- This recipe encodes the planning pipeline as of the authoring date (2026-06-12). The plan-challenge verbatim intent (§5.1) is load-bearing — do not paraphrase it when authoring the challenger agent prompt. The "Acceptance criteria + test plan contract" section above is the shared shape with `/devflow:dynamic-build` Gate 2; any change must be kept in sync. No tooling detects drift — by design (ADR-008 Iron Rule).
453
+ This recipe encodes the planning pipeline as of the authoring date (2026-06-12). The plan-challenge verbatim intent (§5.1) is load-bearing — do not paraphrase it when authoring the challenger agent prompt. The "Acceptance criteria + test plan contract" section above is the shared shape with `/devflow:dynamic-build` Gate 2; any change must be kept in sync. No tooling detects drift — by design (the LLM-vs-plumbing Iron Rule).
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
65
65
 
66
66
  ### DECISIONS_CONTEXT — obtain BEFORE authoring
67
67
 
68
- Before you author the workflow script, read `.devflow/learning/index.md` for the current worktree. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
68
+ The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
69
+
70
+ ```bash
71
+ git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
72
+ ```
73
+
74
+ Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
75
+
76
+ 1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
77
+ 2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
78
+ 3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
79
+
80
+ This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
81
+
82
+ Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
69
83
 
70
84
  The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
71
85
 
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
73
87
 
74
88
  When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
75
89
 
76
- ### IRON RULE (ADR-008: LLM-vs-plumbing)
90
+ ### IRON RULE (LLM-vs-plumbing)
77
91
 
78
92
  **Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
79
93
 
@@ -217,4 +231,4 @@ Next steps:
217
231
 
218
232
  ### Maintenance note
219
233
 
220
- This command mines ALL projects' history on this machine. The bounded-reading discipline (grep/rg + sample — never full-read) is mandatory and must be preserved in every revision. The agent writes prose; no extraction or clustering algorithm is authored here. Per ADR-008 Iron Rule: the agent does the reading and summarizing — not a script we maintain.
234
+ This command mines ALL projects' history on this machine. The bounded-reading discipline (grep/rg + sample — never full-read) is mandatory and must be preserved in every revision. The agent writes prose; no extraction or clustering algorithm is authored here. Per the LLM-vs-plumbing Iron Rule: the agent does the reading and summarizing — not a script we maintain.
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
65
65
 
66
66
  ### DECISIONS_CONTEXT — obtain BEFORE authoring
67
67
 
68
- Before you author the workflow script, read `.devflow/learning/index.md` for the current worktree. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
68
+ The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
69
+
70
+ ```bash
71
+ git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
72
+ ```
73
+
74
+ Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
75
+
76
+ 1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
77
+ 2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
78
+ 3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
79
+
80
+ This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
81
+
82
+ Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
69
83
 
70
84
  The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
71
85
 
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
73
87
 
74
88
  When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
75
89
 
76
- ### IRON RULE (ADR-008: LLM-vs-plumbing)
90
+ ### IRON RULE (LLM-vs-plumbing)
77
91
 
78
92
  **Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
79
93
 
@@ -210,7 +224,7 @@ const drafts = await phase("draft", () =>
210
224
  agent(`Draft ticket for initiative: "${initiative}"
211
225
  Ticket: ${JSON.stringify(c)}
212
226
  Constraints: ${constraints}
213
- Decisions context (apply devflow:apply-decisions; cite ADR/PF IDs): ${DECISIONS_CONTEXT}
227
+ Decisions context (apply devflow:apply-decisions; the ticket is filed to the tracker, so state each decision in words, never by ID): ${DECISIONS_CONTEXT}
214
228
  Write the ticket body following the ticket_body_template structure (Wave/Depends-on header, Summary, Scope with In/Out + anti-features, Invariants, numbered Acceptance Criteria with at least one negative criterion, Open Questions).
215
229
  Return a JSON object with: title (string), summary (string), wave (number), dependsOn (array), bodyMarkdown (string), openQuestions (array).`, { agentType: "Design" })
216
230
  ))
@@ -641,4 +655,4 @@ The tracking-issue path and any open questions are the primary handoff to `/devf
641
655
 
642
656
  ### Maintenance note
643
657
 
644
- This recipe encodes the ticket-factory shape as of the authoring date (2026-06-12). The pipeline structure (`draft → [2-lens review] → revise → whole-set critic → amend → tracking-issue`) is the load-bearing invariant. Per ADR-008, no deterministic ticket-parsing logic is added — ticket slates are proposed by the model and confirmed by the user. When the devflow agent roster changes, update the `agentType` values above. No tooling detects drift — by design (ADR-008 Iron Rule).
658
+ This recipe encodes the ticket-factory shape as of the authoring date (2026-06-12). The pipeline structure (`draft → [2-lens review] → revise → whole-set critic → amend → tracking-issue`) is the load-bearing invariant. Per the LLM-vs-plumbing Iron Rule, no deterministic ticket-parsing logic is added — ticket slates are proposed by the model and confirmed by the user. When the devflow agent roster changes, update the `agentType` values above. No tooling detects drift — by design, under the same rule.
@@ -15,7 +15,13 @@ Explore a codebase area by spawning parallel agents for flow tracing, dependency
15
15
 
16
16
  ## Input
17
17
 
18
- `$ARGUMENTS` contains whatever follows `/explore`:
18
+ What follows `/explore` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
19
+
20
+ <command-input>
21
+ $ARGUMENTS
22
+ </command-input>
23
+
24
+ `COMMAND_INPUT` is one of:
19
25
  - Area description: "how does the auth system work"
20
26
  - Flow question: "trace the request lifecycle"
21
27
  - Empty: use conversation context
@@ -82,6 +88,8 @@ Based on Skim agent findings, spawn 2-3 `Agent(subagent_type="Explore")` agents
82
88
 
83
89
  Adjust explorer focus based on the specific exploration question.
84
90
 
91
+ Ask each explorer for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
92
+
85
93
  ### Phase 4: Synthesize
86
94
 
87
95
  **Produces:** MERGED_FINDINGS
@@ -159,8 +167,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
159
167
  FILES_CHANGED: {list of files changed}
160
168
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
161
169
 
162
- Load the devflow:feature-knowledge skill and follow its authoring process.
163
-
164
170
  Write the knowledge base to:
165
171
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
166
172