devflow-kit 3.0.1 → 3.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +49 -0
- package/README.md +1 -1
- package/dist/agents/git.md +2 -2
- package/dist/cli/agents-view/index.js +1 -1
- package/dist/cli/agents-view/render.js +71 -17
- package/dist/cli/agents-view/state.js +42 -16
- package/dist/cli/agents-view/terminal.js +5 -5
- package/dist/cli/commands/agents.js +142 -51
- package/dist/cli/commands/ambient.js +1 -1
- package/dist/cli/commands/attribution-prompts.js +8 -8
- package/dist/cli/commands/capture.js +1 -1
- package/dist/cli/commands/compliance-prompts.js +8 -8
- package/dist/cli/commands/compliance.js +8 -7
- package/dist/cli/commands/flags.js +33 -31
- package/dist/cli/commands/hud.js +1 -1
- package/dist/cli/commands/init-seed.js +9 -9
- package/dist/cli/commands/init.js +162 -85
- package/dist/cli/commands/install-report.js +10 -10
- package/dist/cli/commands/learning.js +302 -136
- package/dist/cli/commands/memory.js +36 -15
- package/dist/cli/commands/proxy.js +23 -23
- package/dist/cli/commands/rules.js +6 -5
- package/dist/cli/commands/tracker-prompts.js +6 -6
- package/dist/cli/commands/tracker.js +9 -9
- package/dist/cli/commands/uninstall.js +183 -59
- package/dist/cli/flags-view/render.js +5 -5
- package/dist/cli/flags-view/state.js +9 -9
- package/dist/cli/flags-view/terminal.js +4 -4
- package/dist/cli/tui/cells.js +1 -1
- package/dist/cli/tui/terminal.js +6 -6
- package/dist/commands/code-review.md +0 -2
- package/dist/commands/debug.md +14 -11
- package/dist/commands/dynamic-build.md +51 -47
- package/dist/commands/dynamic-plan.md +27 -7
- package/dist/commands/dynamic-profile.md +17 -3
- package/dist/commands/dynamic-tickets.md +18 -4
- package/dist/commands/explore.md +9 -3
- package/dist/commands/implement.md +20 -16
- package/dist/commands/plan.md +13 -9
- package/dist/commands/release.md +23 -3
- package/dist/commands/research.md +9 -3
- package/dist/commands/resolve.md +9 -12
- package/dist/commands/self-review.md +0 -2
- package/dist/core/agent-frontmatter.js +28 -3
- package/dist/core/agent-models.js +204 -42
- package/dist/core/agent-state.js +28 -6
- package/dist/core/ansi.js +2 -2
- package/dist/core/assets.js +1 -1
- package/dist/core/cache.js +7 -8
- package/dist/core/codex-auth-inspect.js +4 -4
- package/dist/core/compliance-compose.js +3 -3
- package/dist/core/compliance.js +3 -4
- package/dist/core/evidence-policy.js +14 -13
- package/dist/core/external-models.js +1 -1
- package/dist/core/feature-config.js +71 -13
- package/dist/core/feature-switch.js +3 -3
- package/dist/core/flags.js +49 -25
- package/dist/core/fs-atomic.js +6 -7
- package/dist/core/learning-queue-cleanup.js +16 -81
- package/dist/core/learning-store.js +61 -0
- package/dist/core/linked-path.js +46 -0
- package/dist/core/manifest.js +5 -5
- package/dist/core/mds-variants.js +13 -13
- package/dist/core/model-discovery.js +8 -8
- package/dist/core/observations.js +17 -101
- package/dist/core/orphan-sweep.js +4 -4
- package/dist/core/plugins.js +13 -8
- package/dist/core/project-paths.js +9 -13
- package/dist/core/proxy-log.js +8 -8
- package/dist/core/proxy-state.js +3 -3
- package/dist/core/queue-drain.js +31 -0
- package/dist/core/reference-sweep.js +6 -6
- package/dist/core/teammate-mode-cleanup.js +1 -1
- package/dist/core/tracker.js +14 -14
- package/dist/hud/colors.js +2 -2
- package/dist/hud/components/learning-counts.js +54 -22
- package/dist/hud/components/version-badge.js +1 -1
- package/dist/skills/git/references/pr/resolve-review-threads.md +2 -2
- package/dist/skills/git/references/tracker/github/create-release.md +2 -2
- package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
- package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
- package/dist/targets/claude-code/compliance-install.js +17 -15
- package/dist/targets/claude-code/hooks.js +2 -2
- package/dist/targets/claude-code/installer.js +59 -32
- package/dist/targets/claude-code/legacy.js +1 -1
- package/dist/targets/claude-code/post-install.js +135 -45
- package/dist/targets/claude-code/tracker-install.js +2 -2
- package/package.json +1 -1
- package/src/assets/agents/code.md +15 -21
- package/src/assets/agents/design.md +4 -2
- package/src/assets/agents/diagnose.md +3 -1
- package/src/assets/agents/evaluate.md +4 -0
- package/src/assets/agents/git.mds +2 -2
- package/src/assets/agents/knowledge.md +5 -3
- package/src/assets/agents/learning.md +281 -196
- package/src/assets/agents/research.md +3 -1
- package/src/assets/agents/review.md +5 -3
- package/src/assets/agents/scrutinize.md +5 -1
- package/src/assets/agents/simplify.md +4 -0
- package/src/assets/agents/skim.md +4 -2
- package/src/assets/agents/synthesize.md +6 -0
- package/src/assets/agents/test.md +18 -10
- package/src/assets/agents/triage.md +11 -9
- package/src/assets/agents/validate.md +14 -10
- package/src/assets/commands/_partials/_decisions.mds +8 -3
- package/src/assets/commands/_partials/_docs_root.mds +3 -3
- package/src/assets/commands/_partials/_engine.mds +16 -32
- package/src/assets/commands/_partials/_knowledge.mds +0 -2
- package/src/assets/commands/_partials/_preamble.mds +6 -2
- package/src/assets/commands/_partials/_settings.mds +2 -2
- package/src/assets/commands/_partials/_tracker.mds +1 -1
- package/src/assets/commands/code-review.mds +0 -2
- package/src/assets/commands/debug.mds +13 -8
- package/src/assets/commands/dynamic-build.mds +18 -12
- package/src/assets/commands/dynamic-plan.mds +10 -4
- package/src/assets/commands/dynamic-profile.mds +1 -1
- package/src/assets/commands/dynamic-tickets.mds +2 -2
- package/src/assets/commands/explore.mds +9 -1
- package/src/assets/commands/implement.mds +19 -13
- package/src/assets/commands/plan.mds +12 -8
- package/src/assets/commands/release.md +23 -3
- package/src/assets/commands/research.mds +9 -3
- package/src/assets/commands/resolve.mds +9 -10
- package/src/assets/mds/git/_pr.mds +3 -3
- package/src/assets/mds/tracker/_common.mds +1 -1
- package/src/assets/mds/tracker/_github.mds +3 -3
- package/src/assets/mds/tracker/_jira.mds +3 -3
- package/src/assets/mds/tracker/_linear.mds +3 -3
- package/src/assets/mds/tracker/_mcp.mds +6 -5
- package/src/assets/scripts/hooks/assets/orchestrator-charter.md +4 -2
- package/src/assets/scripts/hooks/background-memory-update +97 -33
- package/src/assets/scripts/hooks/capture-prompt +4 -3
- package/src/assets/scripts/hooks/capture-question +4 -3
- package/src/assets/scripts/hooks/capture-turn +5 -20
- package/src/assets/scripts/hooks/ensure-devflow-init +14 -2
- package/src/assets/scripts/hooks/ensure-proxy +5 -6
- package/src/assets/scripts/hooks/ensure-root-gitignore +123 -11
- package/src/assets/scripts/hooks/git-marker +71 -0
- package/src/assets/scripts/hooks/is-hex-sha +1 -1
- package/src/assets/scripts/hooks/json-helper.cjs +345 -944
- package/src/assets/scripts/hooks/json-parse +25 -129
- package/src/assets/scripts/hooks/lib/decisions-format.cjs +205 -156
- package/src/assets/scripts/hooks/lib/learning-store.cjs +3207 -0
- package/src/assets/scripts/hooks/lib/mkdir-lock.cjs +7 -5
- package/src/assets/scripts/hooks/lib/project-paths.cjs +13 -19
- package/src/assets/scripts/hooks/lib/render-decisions.cjs +253 -226
- package/src/assets/scripts/hooks/memory-worker +10 -0
- package/src/assets/scripts/hooks/pre-compact-memory +66 -14
- package/src/assets/scripts/hooks/preamble +9 -1
- package/src/assets/scripts/hooks/queue-append +55 -23
- package/src/assets/scripts/hooks/resolve-project-root +3 -4
- package/src/assets/scripts/hooks/session-start-context +146 -45
- package/src/assets/scripts/hooks/session-start-memory +33 -11
- package/src/assets/scripts/lib/project-config.cjs +2 -2
- package/src/assets/scripts/pr-evidence.cjs +3 -3
- package/src/assets/scripts/redact-secrets.cjs +20 -20
- package/src/assets/scripts/release-trace.cjs +1 -1
- package/src/assets/scripts/resolve-evidence-policy.cjs +3 -3
- package/src/assets/scripts/resolve-settings.cjs +3 -3
- package/src/assets/scripts/verify-evidence.cjs +2 -2
- package/src/assets/skills/apply-decisions/SKILL.md +37 -17
- package/src/assets/skills/docs-framework/SKILL.md +2 -2
- package/src/assets/skills/feature-knowledge/SKILL.md +6 -5
- package/src/assets/skills/test-driven-development/SKILL.md +6 -4
- package/dist/core/observation-io.js +0 -50
- package/src/assets/scripts/hooks/decisions-usage-scan.cjs +0 -131
package/dist/cli/tui/terminal.js
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* Generic TUI shell — shared by agents-view and flags-view.
|
|
3
3
|
*
|
|
4
|
-
*
|
|
5
|
-
*
|
|
6
|
-
* a finally-guarded scope.
|
|
7
|
-
*
|
|
4
|
+
* Impure I/O shell in CLI layer; pure logic lives in state + render.
|
|
5
|
+
* Cleanup wired via Promise resolve — never process.exit() inside
|
|
6
|
+
* a finally-guarded scope (it would skip the finally).
|
|
7
|
+
* One generic shell, thin adapters per TUI — not copy-adapted per consumer.
|
|
8
8
|
*
|
|
9
9
|
* Bounded: MAX_KEYPRESSES = 50_000 hard limit (reliability rule — every loop bounded).
|
|
10
10
|
*
|
|
@@ -268,8 +268,8 @@ export async function runTui(spec) {
|
|
|
268
268
|
* A throw inside an EventEmitter listener does NOT reject the enclosing
|
|
269
269
|
* promise — it escapes as an uncaughtException and kills the process with
|
|
270
270
|
* cleanup() never having run, leaving raw mode and alt-screen set. Routing
|
|
271
|
-
* every handler failure through here keeps the
|
|
272
|
-
* always runs
|
|
271
|
+
* every handler failure through here keeps the invariant that cleanup
|
|
272
|
+
* always runs, while still surfacing the error rather than swallowing it.
|
|
273
273
|
*/
|
|
274
274
|
function fail(err) {
|
|
275
275
|
cleanup();
|
|
@@ -228,8 +228,6 @@ Only `exit=0` means the skill is installed. On any other result, do NOT add that
|
|
|
228
228
|
|
|
229
229
|
**Produces:** DECISIONS_CONTEXT, FEATURE_KNOWLEDGE, COMPLIANCE_FRAMEWORKS (carried from Step 0b)
|
|
230
230
|
|
|
231
|
-
**Load Companion Skills** — Load via Skill tool: `devflow:quality-gates`, `devflow:software-design`. If a skill fails to load, continue without it.
|
|
232
|
-
|
|
233
231
|
### Load DECISIONS_CONTEXT
|
|
234
232
|
|
|
235
233
|
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
package/dist/commands/debug.md
CHANGED
|
@@ -15,7 +15,13 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
|
|
|
15
15
|
|
|
16
16
|
## Input
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
What follows `/debug` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
|
|
19
|
+
|
|
20
|
+
<command-input>
|
|
21
|
+
$ARGUMENTS
|
|
22
|
+
</command-input>
|
|
23
|
+
|
|
24
|
+
`COMMAND_INPUT` is one of:
|
|
19
25
|
- Bug description: "login fails after session timeout"
|
|
20
26
|
- Issue reference: "#42"
|
|
21
27
|
- Empty: use conversation context
|
|
@@ -26,8 +32,6 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
|
|
|
26
32
|
|
|
27
33
|
**Produces:** DECISIONS_CONTEXT
|
|
28
34
|
|
|
29
|
-
**Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:software-design`, `devflow:testing`. If a skill fails to load, continue without it.
|
|
30
|
-
|
|
31
35
|
### Load DECISIONS_CONTEXT
|
|
32
36
|
|
|
33
37
|
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
|
@@ -66,9 +70,9 @@ The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Ph
|
|
|
66
70
|
**Produces:** HYPOTHESES, BUG_CONTEXT
|
|
67
71
|
**Requires:** DECISIONS_CONTEXT
|
|
68
72
|
|
|
69
|
-
If
|
|
73
|
+
If `COMMAND_INPUT` opens with a candidate issue reference, fetch the issue:
|
|
70
74
|
|
|
71
|
-
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan
|
|
75
|
+
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
|
|
72
76
|
|
|
73
77
|
**A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
|
|
74
78
|
|
|
@@ -119,7 +123,8 @@ Return a structured report:
|
|
|
119
123
|
- Status: CONFIRMED / DISPROVED / PARTIAL
|
|
120
124
|
- Evidence FOR: [list with file:line refs]
|
|
121
125
|
- Evidence AGAINST: [list with file:line refs]
|
|
122
|
-
- Key finding: {one-sentence summary}
|
|
126
|
+
- Key finding: {one-sentence summary}
|
|
127
|
+
Keep the whole report to at most about 1,500 tokens."
|
|
123
128
|
|
|
124
129
|
Agent(subagent_type="Explore"):
|
|
125
130
|
"Investigate this bug: {bug_description}
|
|
@@ -127,7 +132,7 @@ Agent(subagent_type="Explore"):
|
|
|
127
132
|
Hypothesis: {hypothesis B description}
|
|
128
133
|
Focus area: {specific code area, mechanism, or condition}
|
|
129
134
|
|
|
130
|
-
[same steps and return format]"
|
|
135
|
+
[same steps and return format; the whole report at most about 1,500 tokens]"
|
|
131
136
|
|
|
132
137
|
Agent(subagent_type="Explore"):
|
|
133
138
|
"Investigate this bug: {bug_description}
|
|
@@ -135,7 +140,7 @@ Agent(subagent_type="Explore"):
|
|
|
135
140
|
Hypothesis: {hypothesis C description}
|
|
136
141
|
Focus area: {specific code area, mechanism, or condition}
|
|
137
142
|
|
|
138
|
-
[same steps and return format]"
|
|
143
|
+
[same steps and return format; the whole report at most about 1,500 tokens]"
|
|
139
144
|
|
|
140
145
|
(Add more investigators if bug complexity warrants 4-5 hypotheses)
|
|
141
146
|
```
|
|
@@ -147,7 +152,7 @@ Focus area: {specific code area, mechanism, or condition}
|
|
|
147
152
|
|
|
148
153
|
Evaluate investigation verdicts before synthesis:
|
|
149
154
|
|
|
150
|
-
- **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias)
|
|
155
|
+
- **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias), each reporting at most about 1,500 tokens
|
|
151
156
|
- **Multiple PARTIAL**: Look for a unifying root cause that explains all partial evidence
|
|
152
157
|
- **All DISPROVED**: Report honestly — "No root cause identified from initial hypotheses." Generate 2-3 second-round hypotheses if conversation context suggests avenues not yet explored. Loop back to Phase 3.
|
|
153
158
|
|
|
@@ -256,8 +261,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
|
|
|
256
261
|
FILES_CHANGED: {list of files changed}
|
|
257
262
|
DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
|
|
258
263
|
|
|
259
|
-
Load the devflow:feature-knowledge skill and follow its authoring process.
|
|
260
|
-
|
|
261
264
|
Write the knowledge base to:
|
|
262
265
|
{worktree}/.devflow/features/{slug}/KNOWLEDGE.md
|
|
263
266
|
|
|
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
|
|
|
65
65
|
|
|
66
66
|
### DECISIONS_CONTEXT — obtain BEFORE authoring
|
|
67
67
|
|
|
68
|
-
|
|
68
|
+
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
|
|
75
|
+
|
|
76
|
+
1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
|
|
77
|
+
2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
|
|
78
|
+
3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
|
|
79
|
+
|
|
80
|
+
This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
|
|
81
|
+
|
|
82
|
+
Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
|
|
69
83
|
|
|
70
84
|
The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
|
|
71
85
|
|
|
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
|
|
|
73
87
|
|
|
74
88
|
When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
|
|
75
89
|
|
|
76
|
-
### IRON RULE (
|
|
90
|
+
### IRON RULE (LLM-vs-plumbing)
|
|
77
91
|
|
|
78
92
|
**Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
|
|
79
93
|
|
|
@@ -192,7 +206,13 @@ Note the `budget` value from the Workflow tool context (or default to "medium" i
|
|
|
192
206
|
|
|
193
207
|
**3. Detect mode: SINGLE or WAVE**
|
|
194
208
|
|
|
195
|
-
|
|
209
|
+
What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
|
|
210
|
+
|
|
211
|
+
<command-input>
|
|
212
|
+
$ARGUMENTS
|
|
213
|
+
</command-input>
|
|
214
|
+
|
|
215
|
+
- **SINGLE mode:** `COMMAND_INPUT` is one ticket, one issue, one task description, or one plan document
|
|
196
216
|
- **WAVE mode:** input is a set of tracker issues (wave labels, milestone, issue list), or the user says "wave" / "all tickets in wave N"
|
|
197
217
|
- **A `/devflow:dynamic-tickets` ticket directory** (`{worktree}/.devflow/docs/tickets/{slug}/{ts}/`) is WAVE input: in each ticket file (every `.md` there but `tracking-issue.md`), the `**Issue:**` line directly after `**Depends on:**` is one raw `ISSUE_REFS` token, forwarded to the wave's pre-fetch verbatim — never rendered, normalised or re-derived. A ticket file with no `**Issue:**` line, or more than one, contributes no token; name it in the run summary as `not filed`.
|
|
198
218
|
|
|
@@ -230,11 +250,11 @@ If none found: build proceeds Gate-1-only (Gate 2 skipped with a note). Never re
|
|
|
230
250
|
**5. Resolve tracking-issue number (optional)**
|
|
231
251
|
|
|
232
252
|
Check, in priority order:
|
|
233
|
-
- An explicit candidate issue reference or issue URL in
|
|
253
|
+
- An explicit candidate issue reference or issue URL in `COMMAND_INPUT` (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
|
|
234
254
|
- The `**Issue:**` line directly after the H1 of the ticket set's `tracking-issue.md` (written by `/devflow:dynamic-tickets`' filing step; the file is at `{worktree}/.devflow/docs/tickets/{slug}/{ts}/tracking-issue.md`), as a raw token — only when the file holds exactly one `**Issue:**` line
|
|
235
255
|
- Otherwise: none
|
|
236
256
|
|
|
237
|
-
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan
|
|
257
|
+
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
|
|
238
258
|
|
|
239
259
|
**A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
|
|
240
260
|
|
|
@@ -315,7 +335,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
|
|
|
315
335
|
ISSUE_PR_LINK: ${ISSUE_PR_LINK}
|
|
316
336
|
COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
|
|
317
337
|
|
|
318
|
-
When you build or run tests to verify your work,
|
|
338
|
+
When you build or run tests to verify your work, follow your Running commands block.
|
|
319
339
|
|
|
320
340
|
After implementing, commit your changes with a conventional-commit message and report:
|
|
321
341
|
- Files changed
|
|
@@ -327,7 +347,7 @@ After implementing, commit your changes with a conventional-commit message and r
|
|
|
327
347
|
const gate1 = await phase("gate1", async () => {
|
|
328
348
|
// Gate 1 #1 — runs once, immediately after the initial implementation.
|
|
329
349
|
const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH}.
|
|
330
|
-
|
|
350
|
+
Follow your Running commands block for every build and test command.
|
|
331
351
|
Report: PASS or FAIL with details.`, { agentType: "Validate" });
|
|
332
352
|
|
|
333
353
|
if (validation.verdict === "FAIL") {
|
|
@@ -382,7 +402,7 @@ ${panel.filter(p => p.verdict === "FAIL").map(p => p.rationale).join("\n")}
|
|
|
382
402
|
ISSUE_NUMBER: ${ISSUE_NUMBER}
|
|
383
403
|
ISSUE_PR_LINK: ${ISSUE_PR_LINK}
|
|
384
404
|
COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
|
|
385
|
-
Self-verify your fix compiles
|
|
405
|
+
Self-verify your fix compiles, following your Running commands block. Commit fixes.`, { agentType: "Code" });
|
|
386
406
|
evalVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
|
|
387
407
|
}
|
|
388
408
|
}
|
|
@@ -392,7 +412,7 @@ Self-verify your fix compiles (background-Bash + Monitor for any build >120s —
|
|
|
392
412
|
const testResult = await agent(`Run scenario-based acceptance tests on branch ${BRANCH} against these criteria:
|
|
393
413
|
${CRITERIA || "(none)"}
|
|
394
414
|
TEST_PLAN: ${TEST_PLAN || "(none)"}
|
|
395
|
-
|
|
415
|
+
Follow your Running commands block for every scenario command.
|
|
396
416
|
Cover: functionality, API contracts, performance, and cover every TEST_PLAN scenario. Report: PASS or FAIL per scenario.`, { agentType: "Test" });
|
|
397
417
|
testVerdict = testResult.verdict;
|
|
398
418
|
if (testVerdict === "FAIL") {
|
|
@@ -402,7 +422,7 @@ ${testResult.failures}
|
|
|
402
422
|
ISSUE_NUMBER: ${ISSUE_NUMBER}
|
|
403
423
|
ISSUE_PR_LINK: ${ISSUE_PR_LINK}
|
|
404
424
|
COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
|
|
405
|
-
Self-verify your fix compiles and the scenarios pass
|
|
425
|
+
Self-verify your fix compiles and the scenarios pass, following your Running commands block. Commit fixes.`, { agentType: "Code" });
|
|
406
426
|
testVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
|
|
407
427
|
}
|
|
408
428
|
}
|
|
@@ -516,7 +536,7 @@ ${chunk.map(f => `- ${f.description} (${f.severity})`).join("\n")}
|
|
|
516
536
|
ISSUE_NUMBER: ${ISSUE_NUMBER}
|
|
517
537
|
ISSUE_PR_LINK: ${ISSUE_PR_LINK}
|
|
518
538
|
COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
|
|
519
|
-
Fix all findings in this batch. Self-verify your fix compiles
|
|
539
|
+
Fix all findings in this batch. Self-verify your fix compiles, following your Running commands block. Commit with conventional-commit message.
|
|
520
540
|
Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<description of any finding that could not be fixed>"]}`, { agentType: "Code" });
|
|
521
541
|
chunkResults.push({ chunk, result: r });
|
|
522
542
|
}
|
|
@@ -556,7 +576,7 @@ Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<
|
|
|
556
576
|
// this is the build gate before the branch is handed back. See gate1_postcode() cadence.
|
|
557
577
|
const gate1Final = await phase("gate1-final", async () => {
|
|
558
578
|
const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH} (final gate after all fixing).
|
|
559
|
-
|
|
579
|
+
Follow your Running commands block for every build and test command.
|
|
560
580
|
Report: PASS or FAIL with details.`, { agentType: "Validate" });
|
|
561
581
|
|
|
562
582
|
if (validation.verdict === "FAIL") {
|
|
@@ -568,7 +588,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
|
|
|
568
588
|
ISSUE_PR_LINK: ${ISSUE_PR_LINK}
|
|
569
589
|
COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
|
|
570
590
|
Self-verify your fix compiles. Commit fixes with conventional-commit message.`, { agentType: "Code" });
|
|
571
|
-
const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (
|
|
591
|
+
const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
|
|
572
592
|
if (recheck.verdict === "PASS") break;
|
|
573
593
|
failureDetails = recheck.details || failureDetails;
|
|
574
594
|
if (attempt === 2) return { verdict: "ESCALATED", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
|
|
@@ -580,7 +600,7 @@ Self-verify your fix compiles. Commit fixes with conventional-commit message.`,
|
|
|
580
600
|
const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
|
|
581
601
|
|
|
582
602
|
if (scrutiny.codeChanged) {
|
|
583
|
-
await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes;
|
|
603
|
+
await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
|
|
584
604
|
}
|
|
585
605
|
|
|
586
606
|
return { verdict: "PASS" };
|
|
@@ -633,42 +653,26 @@ Write a concise report. Present only surviving findings as outstanding — never
|
|
|
633
653
|
|
|
634
654
|
This applies to both: multiple Code agents working on a single ticket AND multi-ticket scheduling in a wave.
|
|
635
655
|
|
|
636
|
-
### Build execution doctrine —
|
|
656
|
+
### Build execution doctrine — the Running commands block (LOAD-BEARING)
|
|
637
657
|
|
|
638
|
-
The
|
|
658
|
+
The Validate, Code and Test agents each carry this `## Running commands` block in their bodies, so each prompt in this engine names the block ("your Running commands block") instead of copying it:
|
|
639
659
|
|
|
640
|
-
|
|
660
|
+
Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
|
|
641
661
|
|
|
642
|
-
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
|
|
647
|
-
|
|
648
|
-
|
|
649
|
-
|
|
650
|
-
```
|
|
651
|
-
This returns immediately with a background task id — do NOT block on it.
|
|
652
|
-
2. Arm ONE Monitor that emits a heartbeat well under 180s AND exits when the job finishes:
|
|
653
|
-
- description: short, e.g. `await <build cmd>`
|
|
654
|
-
- persistent: false
|
|
655
|
-
- timeout_ms: comfortably ABOVE the expected job time (e.g. 600000)
|
|
656
|
-
- command: `until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
|
|
657
|
-
|
|
658
|
-
The `building` heartbeat every 25s (≪ 180s) is delivered as a notification that re-invokes you, so the watchdog never sees a >180s gap. `BUILD_DONE` + the `EXIT=` line signal completion.
|
|
659
|
-
|
|
660
|
-
**Exit-code honesty:** the background task's own exit status is meaningless (the trailing `echo` always exits 0). ALWAYS read the `EXIT=` value written inside `<BASE>.done` — that is the authoritative result.
|
|
661
|
-
|
|
662
|
-
**Bounded polling:** arm ONE Monitor then stop acting. On Monitor timeout, re-arm at most 2× (never more than 3 total Monitor calls per build). If the build has not finished after 3 Monitor calls: record the state and escalate — never babysit. A Code agent burned 241k tokens polling one build.
|
|
663
|
-
3. When the monitor reports `BUILD_DONE`: the job PASSES iff `<BASE>.done` contains `EXIT=0`. Read `<BASE>.log` for output/failure detail.
|
|
662
|
+
- Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
|
|
663
|
+
- Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
|
|
664
|
+
- Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
|
|
665
|
+
- A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
|
|
666
|
+
- A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
|
|
667
|
+
- Never re-run a command when nothing it reads has changed.
|
|
668
|
+
- Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
|
|
669
|
+
- The same rules hold inside a dynamic Workflow sub-agent.
|
|
664
670
|
|
|
665
671
|
**Cheapest-sufficient validation:** iterate with the fastest check that proves the change — `tsc --noEmit`, `cargo check`, `go vet`, a single test file — the ecosystem's cheapest sufficient signal. Reserve expensive full/optimized builds and whole-suite runs for the final gates only.
|
|
666
672
|
|
|
667
673
|
**One build gate per phase:** batch related fixes, validate once. A fix-pass Code agent runs ONE light check over its whole batch — never several invocations per small fix. Do NOT validate after every individual mutation; validate once after the batch is complete.
|
|
668
674
|
|
|
669
|
-
**Scope commands to stay short.** During the engine,
|
|
670
|
-
|
|
671
|
-
**Invariants:** heartbeat interval MUST stay well under 180s (25–30s is the tested value); Monitor `timeout_ms` MUST exceed the expected job duration (a too-short timeout kills the poll, not the build). Never substitute a single silent long command for this procedure.
|
|
675
|
+
**Scope commands to stay short.** During the engine, prefer the scoped command for the change over the whole workspace — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — and, when the whole set is needed, one workspace-level command over a per-package loop. After the wave, the full-workspace regression is the human's job (the wave already hands the integrated branch back to the user).
|
|
672
676
|
|
|
673
677
|
### Implement bundle
|
|
674
678
|
|
|
@@ -680,7 +684,7 @@ Code(agentType:"Code", prompt: full task + plan + DECISIONS_CONTEXT + handoff if
|
|
|
680
684
|
→ gate2_acceptance() ← Gate 2 runs HERE — before the review pass, not after
|
|
681
685
|
```
|
|
682
686
|
|
|
683
|
-
The Code agent prompt must include: task description, implementation plan (if one exists), relevant DECISIONS_CONTEXT (
|
|
687
|
+
The Code agent prompt must include: task description, implementation plan (if one exists), relevant DECISIONS_CONTEXT (the index you loaded before authoring), the compliance lens (`COMPLIANCE_FRAMEWORKS` — every Code prompt carries it, fix prompts included), and any PRIOR_PHASE_SUMMARY / HANDOFF_FILE for sequential multi-phase tickets.
|
|
684
688
|
|
|
685
689
|
Gate 2 runs at implementation acceptance — this matches devflow's deliberate placement: "evaluation is part of implementation acceptance, not post-review" (§6.1).
|
|
686
690
|
|
|
@@ -689,7 +693,7 @@ Gate 2 runs at implementation acceptance — this matches devflow's deliberate p
|
|
|
689
693
|
ORDER IS LOAD-BEARING. Run exactly in this sequence:
|
|
690
694
|
|
|
691
695
|
1. **Validate agent** — build / typecheck / lint / test
|
|
692
|
-
- Build/test commands
|
|
696
|
+
- Build/test commands follow `build_execution_doctrine()`: the Validate agent's Running commands block, foreground under an explicit Bash timeout.
|
|
693
697
|
- FAIL → Code agent fix (max 2 retries) → re-run Validate agent
|
|
694
698
|
- If still FAIL after 2 retries → escalate (do not loop endlessly)
|
|
695
699
|
2. **Simplify agent** — reduce complexity, remove duplication
|
|
@@ -704,7 +708,7 @@ Depth scales to change size + budget: a trivial one-line fix warrants a lighter
|
|
|
704
708
|
1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`).
|
|
705
709
|
2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed.
|
|
706
710
|
|
|
707
|
-
It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's
|
|
711
|
+
It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's Running commands block). The final Gate 1 #2 is the invariant that all written code passes before merge.
|
|
708
712
|
|
|
709
713
|
### GATE 2 — Acceptance gate (per ticket, plan-scoped — fires ONCE at implementation acceptance)
|
|
710
714
|
|
|
@@ -836,7 +840,7 @@ Majority-survives: a finding needs >50% of verification lenses to confirm it. St
|
|
|
836
840
|
|
|
837
841
|
If no surviving findings: return early (no fixes needed). Any coverageGaps are carried in the return — they block a PASS verdict downstream, not the early exit.
|
|
838
842
|
|
|
839
|
-
If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's
|
|
843
|
+
If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's Running commands block, once over its whole batch). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
|
|
840
844
|
|
|
841
845
|
### Engine invariants (non-negotiable)
|
|
842
846
|
|
|
@@ -1240,4 +1244,4 @@ Update the wave PR's test-plan block and post its evidence comment."
|
|
|
1240
1244
|
|
|
1241
1245
|
### Maintenance note
|
|
1242
1246
|
|
|
1243
|
-
This recipe encodes the current `/implement` + `/code-review` + `/resolve` orchestration shape as of the authoring date (2026-06-12). When those base commands change their orchestration, update this recipe to match. No tooling detects drift — by design (
|
|
1247
|
+
This recipe encodes the current `/implement` + `/code-review` + `/resolve` orchestration shape as of the authoring date (2026-06-12). When those base commands change their orchestration, update this recipe to match. No tooling detects drift — by design (the LLM-vs-plumbing Iron Rule). The reminder lives in the design doc §16.
|
|
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
|
|
|
65
65
|
|
|
66
66
|
### DECISIONS_CONTEXT — obtain BEFORE authoring
|
|
67
67
|
|
|
68
|
-
|
|
68
|
+
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
|
|
75
|
+
|
|
76
|
+
1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
|
|
77
|
+
2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
|
|
78
|
+
3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
|
|
79
|
+
|
|
80
|
+
This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
|
|
81
|
+
|
|
82
|
+
Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
|
|
69
83
|
|
|
70
84
|
The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
|
|
71
85
|
|
|
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
|
|
|
73
87
|
|
|
74
88
|
When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
|
|
75
89
|
|
|
76
|
-
### IRON RULE (
|
|
90
|
+
### IRON RULE (LLM-vs-plumbing)
|
|
77
91
|
|
|
78
92
|
**Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
|
|
79
93
|
|
|
@@ -157,12 +171,18 @@ If present, note its contents as `PREFERENCE_PROFILE`. This will be used to auto
|
|
|
157
171
|
|
|
158
172
|
**3. Resolve ticket input**
|
|
159
173
|
|
|
160
|
-
|
|
174
|
+
What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
|
|
175
|
+
|
|
176
|
+
<command-input>
|
|
177
|
+
$ARGUMENTS
|
|
178
|
+
</command-input>
|
|
179
|
+
|
|
180
|
+
Determine the ticket source from `COMMAND_INPUT` (in priority order):
|
|
161
181
|
- A directory of ticket `.md` files (from `/devflow:dynamic-tickets` output)
|
|
162
182
|
- A list of candidate issue references or issue URLs
|
|
163
183
|
- Inline ticket descriptions passed as args
|
|
164
184
|
|
|
165
|
-
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan
|
|
185
|
+
**Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
|
|
166
186
|
|
|
167
187
|
**A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
|
|
168
188
|
|
|
@@ -233,7 +253,7 @@ const plans = await phase("plan-parallel", () =>
|
|
|
233
253
|
parallel((tickets || []).map(ticket => () =>
|
|
234
254
|
agent(`Write an implementation plan for this ticket.
|
|
235
255
|
Ticket: ${JSON.stringify(ticket)}
|
|
236
|
-
Decisions context (apply devflow:apply-decisions;
|
|
256
|
+
Decisions context (apply devflow:apply-decisions; a plan file can be posted to the tracker, so state each decision in words, never by ID): ${DECISIONS_CONTEXT}
|
|
237
257
|
The plan must cover: approach overview, affected files and modules, key design decisions, implementation sequence (what to build first), risks and mitigations, and any open questions you cannot resolve from the ticket alone.
|
|
238
258
|
Write a thorough but tight plan — every section must earn its place for a Code agent who has no other context.
|
|
239
259
|
Return: { ticketTitle, planMarkdown, openDecisions (array of genuine unknowns requiring user input) }.`, { agentType: "Design" })
|
|
@@ -324,7 +344,7 @@ For each ticket, write ${OUTDIR}/{ticket-slug}-plan.md containing:
|
|
|
324
344
|
- ## Acceptance Criteria (numbered, positive + negative)
|
|
325
345
|
- ## Test Plan — the challenger's testPlan lines, verbatim, one per line, and nothing else: no prose, no blank line between them, no setup or outcome
|
|
326
346
|
- ## Test Scenarios — one line per TP, in TP order: TP-n: its setup, then its expected outcome, from testScenarios
|
|
327
|
-
- ## Auto-Resolved Decisions (if any — list each as: decision → resolution → source)
|
|
347
|
+
- ## Auto-Resolved Decisions (if any — list each as: decision → resolution → source, naming a recorded decision in words, never by ID)
|
|
328
348
|
|
|
329
349
|
Then write ${OUTDIR}/DECISIONS-NEEDED.md:
|
|
330
350
|
- ## Auto-Resolved Decisions — list each silently-resolved decision as: decision → resolution → source (preference profile / ADR-NNN), so auto-resolution is auditable and reversible. If none, write "None."
|
|
@@ -430,4 +450,4 @@ After the workflow returns: check each plan's test plan (step 1 of the F4 list a
|
|
|
430
450
|
|
|
431
451
|
### Maintenance note
|
|
432
452
|
|
|
433
|
-
This recipe encodes the planning pipeline as of the authoring date (2026-06-12). The plan-challenge verbatim intent (§5.1) is load-bearing — do not paraphrase it when authoring the challenger agent prompt. The "Acceptance criteria + test plan contract" section above is the shared shape with `/devflow:dynamic-build` Gate 2; any change must be kept in sync. No tooling detects drift — by design (
|
|
453
|
+
This recipe encodes the planning pipeline as of the authoring date (2026-06-12). The plan-challenge verbatim intent (§5.1) is load-bearing — do not paraphrase it when authoring the challenger agent prompt. The "Acceptance criteria + test plan contract" section above is the shared shape with `/devflow:dynamic-build` Gate 2; any change must be kept in sync. No tooling detects drift — by design (the LLM-vs-plumbing Iron Rule).
|
|
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
|
|
|
65
65
|
|
|
66
66
|
### DECISIONS_CONTEXT — obtain BEFORE authoring
|
|
67
67
|
|
|
68
|
-
|
|
68
|
+
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
|
|
75
|
+
|
|
76
|
+
1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
|
|
77
|
+
2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
|
|
78
|
+
3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
|
|
79
|
+
|
|
80
|
+
This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
|
|
81
|
+
|
|
82
|
+
Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
|
|
69
83
|
|
|
70
84
|
The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
|
|
71
85
|
|
|
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
|
|
|
73
87
|
|
|
74
88
|
When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
|
|
75
89
|
|
|
76
|
-
### IRON RULE (
|
|
90
|
+
### IRON RULE (LLM-vs-plumbing)
|
|
77
91
|
|
|
78
92
|
**Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
|
|
79
93
|
|
|
@@ -217,4 +231,4 @@ Next steps:
|
|
|
217
231
|
|
|
218
232
|
### Maintenance note
|
|
219
233
|
|
|
220
|
-
This command mines ALL projects' history on this machine. The bounded-reading discipline (grep/rg + sample — never full-read) is mandatory and must be preserved in every revision. The agent writes prose; no extraction or clustering algorithm is authored here. Per
|
|
234
|
+
This command mines ALL projects' history on this machine. The bounded-reading discipline (grep/rg + sample — never full-read) is mandatory and must be preserved in every revision. The agent writes prose; no extraction or clustering algorithm is authored here. Per the LLM-vs-plumbing Iron Rule: the agent does the reading and summarizing — not a script we maintain.
|
|
@@ -65,7 +65,21 @@ The `budget` global governs depth. Scale Review agent roster and verification vo
|
|
|
65
65
|
|
|
66
66
|
### DECISIONS_CONTEXT — obtain BEFORE authoring
|
|
67
67
|
|
|
68
|
-
|
|
68
|
+
The decisions ledger belongs to the repository, not to one checkout: in a linked worktree it lives in the main worktree, and a session started in a subdirectory reads the copy at the repository root. Locate it with ONE git call, run from the start directory — `WORKTREE_PATH` if provided, otherwise cwd (`devflow:worktree-support`):
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
Line 1 is the checkout's toplevel, line 2 the repository's common git directory. A git older than 2.31 echoes `--path-format=absolute` back as a line of its own first. `{ledger}` is the first of these that applies:
|
|
75
|
+
|
|
76
|
+
1. **The main worktree** — line 2 without its trailing `/.git`, when the output is exactly two lines each beginning with `/`, line 2 ends in `/.git`, and the directory left once it is removed is not your home directory and contains a `.devflow/` directory.
|
|
77
|
+
2. **The toplevel** — line 1, or on an older git the line after the echoed flag.
|
|
78
|
+
3. **The start directory itself** — when the command failed or printed no absolute toplevel (outside a git repository).
|
|
79
|
+
|
|
80
|
+
This is the rule the learning hooks apply (D-LEDGER-MAIN-WORKTREE, D-PROMPT-ROOT), so you read the index the Learning agent writes.
|
|
81
|
+
|
|
82
|
+
Before you author the workflow script, read `{ledger}/.devflow/learning/index.md`. If the file is absent or empty, set `DECISIONS_CONTEXT` to `(none)`; otherwise use the file content as `DECISIONS_CONTEXT`.
|
|
69
83
|
|
|
70
84
|
The script body cannot perform this read — you (the main model) do it before authoring. Then inject the relevant DECISIONS_CONTEXT into agent prompts using the `devflow:apply-decisions` consumption algorithm (scan index → Read relevant entries → cite verbatim IDs in agent prompts). Only agents that need architectural context (Code agent, Evaluate agent, Review agent, Scrutinize agent) need DECISIONS_CONTEXT injected; lightweight agents (Validate agent, Simplify agent) do not.
|
|
71
85
|
|
|
@@ -73,7 +87,7 @@ The script body cannot perform this read — you (the main model) do it before a
|
|
|
73
87
|
|
|
74
88
|
When a ticket requires multiple sequential Code agent phases, each Code agent writes `{toplevel}/.devflow/docs/handoff-{branch_slug}.md` (branch-scoped to prevent concurrent session clobber), `{toplevel}` being `git rev-parse --show-toplevel` in the checkout the ticket's branch is in — never a subdirectory. The next Code agent reads it via HANDOFF_FILE input. PRIOR_PHASE_SUMMARY is the compact in-context form; the handoff file is the durable form that survives context compaction. Always read the handoff file directly — code is authoritative, summaries are supplementary.
|
|
75
89
|
|
|
76
|
-
### IRON RULE (
|
|
90
|
+
### IRON RULE (LLM-vs-plumbing)
|
|
77
91
|
|
|
78
92
|
**Author ZERO deterministic feature code.** No parsers, no schedulers, no topological-sort, no dependency-graph helpers, no confidence formulas. ALL issue reading, dependency reasoning, and scheduling decisions are LLM judgment at runtime, performed by the workflow's agents. The recipe is instructions. The workflow script Claude authors IS the runtime logic — keep it free of hand-coded feature algorithms.
|
|
79
93
|
|
|
@@ -210,7 +224,7 @@ const drafts = await phase("draft", () =>
|
|
|
210
224
|
agent(`Draft ticket for initiative: "${initiative}"
|
|
211
225
|
Ticket: ${JSON.stringify(c)}
|
|
212
226
|
Constraints: ${constraints}
|
|
213
|
-
Decisions context (apply devflow:apply-decisions;
|
|
227
|
+
Decisions context (apply devflow:apply-decisions; the ticket is filed to the tracker, so state each decision in words, never by ID): ${DECISIONS_CONTEXT}
|
|
214
228
|
Write the ticket body following the ticket_body_template structure (Wave/Depends-on header, Summary, Scope with In/Out + anti-features, Invariants, numbered Acceptance Criteria with at least one negative criterion, Open Questions).
|
|
215
229
|
Return a JSON object with: title (string), summary (string), wave (number), dependsOn (array), bodyMarkdown (string), openQuestions (array).`, { agentType: "Design" })
|
|
216
230
|
))
|
|
@@ -641,4 +655,4 @@ The tracking-issue path and any open questions are the primary handoff to `/devf
|
|
|
641
655
|
|
|
642
656
|
### Maintenance note
|
|
643
657
|
|
|
644
|
-
This recipe encodes the ticket-factory shape as of the authoring date (2026-06-12). The pipeline structure (`draft → [2-lens review] → revise → whole-set critic → amend → tracking-issue`) is the load-bearing invariant. Per
|
|
658
|
+
This recipe encodes the ticket-factory shape as of the authoring date (2026-06-12). The pipeline structure (`draft → [2-lens review] → revise → whole-set critic → amend → tracking-issue`) is the load-bearing invariant. Per the LLM-vs-plumbing Iron Rule, no deterministic ticket-parsing logic is added — ticket slates are proposed by the model and confirmed by the user. When the devflow agent roster changes, update the `agentType` values above. No tooling detects drift — by design, under the same rule.
|
package/dist/commands/explore.md
CHANGED
|
@@ -15,7 +15,13 @@ Explore a codebase area by spawning parallel agents for flow tracing, dependency
|
|
|
15
15
|
|
|
16
16
|
## Input
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
What follows `/explore` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
|
|
19
|
+
|
|
20
|
+
<command-input>
|
|
21
|
+
$ARGUMENTS
|
|
22
|
+
</command-input>
|
|
23
|
+
|
|
24
|
+
`COMMAND_INPUT` is one of:
|
|
19
25
|
- Area description: "how does the auth system work"
|
|
20
26
|
- Flow question: "trace the request lifecycle"
|
|
21
27
|
- Empty: use conversation context
|
|
@@ -82,6 +88,8 @@ Based on Skim agent findings, spawn 2-3 `Agent(subagent_type="Explore")` agents
|
|
|
82
88
|
|
|
83
89
|
Adjust explorer focus based on the specific exploration question.
|
|
84
90
|
|
|
91
|
+
Ask each explorer for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
|
|
92
|
+
|
|
85
93
|
### Phase 4: Synthesize
|
|
86
94
|
|
|
87
95
|
**Produces:** MERGED_FINDINGS
|
|
@@ -159,8 +167,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
|
|
|
159
167
|
FILES_CHANGED: {list of files changed}
|
|
160
168
|
DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
|
|
161
169
|
|
|
162
|
-
Load the devflow:feature-knowledge skill and follow its authoring process.
|
|
163
|
-
|
|
164
170
|
Write the knowledge base to:
|
|
165
171
|
{worktree}/.devflow/features/{slug}/KNOWLEDGE.md
|
|
166
172
|
|