devflow-kit 3.1.0 → 3.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/CHANGELOG.md +36 -0
  2. package/README.md +1 -1
  3. package/dist/cli/agents-view/render.js +69 -15
  4. package/dist/cli/agents-view/state.js +40 -14
  5. package/dist/cli/commands/agents.js +135 -45
  6. package/dist/cli/commands/init.js +128 -53
  7. package/dist/cli/commands/learning.js +36 -8
  8. package/dist/cli/commands/memory.js +35 -14
  9. package/dist/cli/commands/uninstall.js +163 -39
  10. package/dist/commands/code-review.md +0 -2
  11. package/dist/commands/debug.md +14 -11
  12. package/dist/commands/dynamic-build.md +33 -43
  13. package/dist/commands/dynamic-plan.md +8 -2
  14. package/dist/commands/explore.md +9 -3
  15. package/dist/commands/implement.md +20 -16
  16. package/dist/commands/plan.md +13 -9
  17. package/dist/commands/release.md +8 -2
  18. package/dist/commands/research.md +8 -2
  19. package/dist/commands/resolve.md +1 -3
  20. package/dist/commands/self-review.md +0 -2
  21. package/dist/core/agent-frontmatter.js +25 -0
  22. package/dist/core/agent-models.js +198 -36
  23. package/dist/core/agent-state.js +27 -5
  24. package/dist/core/assets.js +1 -1
  25. package/dist/core/feature-config.js +68 -10
  26. package/dist/core/flags.js +24 -0
  27. package/dist/core/learning-queue-cleanup.js +10 -11
  28. package/dist/core/linked-path.js +46 -0
  29. package/dist/core/plugins.js +9 -3
  30. package/dist/core/queue-drain.js +31 -0
  31. package/dist/hud/components/learning-counts.js +54 -8
  32. package/dist/skills/git/references/tracker/github/create-release.md +2 -2
  33. package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
  34. package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
  35. package/dist/targets/claude-code/installer.js +36 -9
  36. package/dist/targets/claude-code/post-install.js +128 -38
  37. package/package.json +1 -1
  38. package/src/assets/agents/code.md +14 -17
  39. package/src/assets/agents/design.md +2 -0
  40. package/src/assets/agents/diagnose.md +2 -0
  41. package/src/assets/agents/evaluate.md +4 -0
  42. package/src/assets/agents/knowledge.md +2 -0
  43. package/src/assets/agents/research.md +2 -0
  44. package/src/assets/agents/review.md +2 -0
  45. package/src/assets/agents/scrutinize.md +4 -0
  46. package/src/assets/agents/simplify.md +4 -0
  47. package/src/assets/agents/skim.md +3 -1
  48. package/src/assets/agents/synthesize.md +6 -0
  49. package/src/assets/agents/test.md +18 -10
  50. package/src/assets/agents/triage.md +2 -0
  51. package/src/assets/agents/validate.md +14 -10
  52. package/src/assets/commands/_partials/_engine.mds +15 -31
  53. package/src/assets/commands/_partials/_knowledge.mds +0 -2
  54. package/src/assets/commands/_partials/_tracker.mds +1 -1
  55. package/src/assets/commands/code-review.mds +0 -2
  56. package/src/assets/commands/debug.mds +13 -8
  57. package/src/assets/commands/dynamic-build.mds +17 -11
  58. package/src/assets/commands/dynamic-plan.mds +7 -1
  59. package/src/assets/commands/explore.mds +9 -1
  60. package/src/assets/commands/implement.mds +19 -13
  61. package/src/assets/commands/plan.mds +12 -8
  62. package/src/assets/commands/release.md +8 -2
  63. package/src/assets/commands/research.mds +8 -2
  64. package/src/assets/commands/resolve.mds +1 -1
  65. package/src/assets/mds/tracker/_github.mds +2 -2
  66. package/src/assets/mds/tracker/_jira.mds +2 -2
  67. package/src/assets/mds/tracker/_linear.mds +2 -2
  68. package/src/assets/scripts/hooks/assets/orchestrator-charter.md +3 -2
  69. package/src/assets/scripts/hooks/background-memory-update +69 -11
  70. package/src/assets/scripts/hooks/capture-prompt +4 -3
  71. package/src/assets/scripts/hooks/capture-question +4 -3
  72. package/src/assets/scripts/hooks/capture-turn +4 -3
  73. package/src/assets/scripts/hooks/ensure-devflow-init +13 -1
  74. package/src/assets/scripts/hooks/ensure-root-gitignore +122 -10
  75. package/src/assets/scripts/hooks/git-marker +71 -0
  76. package/src/assets/scripts/hooks/json-helper.cjs +12 -145
  77. package/src/assets/scripts/hooks/json-parse +24 -129
  78. package/src/assets/scripts/hooks/lib/learning-store.cjs +169 -64
  79. package/src/assets/scripts/hooks/lib/render-decisions.cjs +1 -1
  80. package/src/assets/scripts/hooks/memory-worker +10 -0
  81. package/src/assets/scripts/hooks/pre-compact-memory +66 -14
  82. package/src/assets/scripts/hooks/preamble +9 -1
  83. package/src/assets/scripts/hooks/queue-append +53 -21
  84. package/src/assets/scripts/hooks/session-start-context +108 -29
  85. package/src/assets/scripts/hooks/session-start-memory +33 -11
  86. package/src/assets/skills/test-driven-development/SKILL.md +6 -4
@@ -110,6 +110,8 @@ Write findings to OUTPUT_PATH using the Write tool:
110
110
  {What was not investigated, scope boundaries, data freshness concerns}
111
111
  ```
112
112
 
113
+ Report cap: final message at most about 1,500 tokens; the findings document is the file at the output path, other longer material goes to a `mktemp` file (via Bash or Write), and the message gives its path. Exempt: none.
114
+
113
115
  ## Token Budget
114
116
 
115
117
  Target output: ~4K–8K tokens. Prioritize structured tables and key findings over exhaustive lists.
@@ -178,6 +178,8 @@ Report format for `{output_path}`:
178
178
  **Recommendation**: {BLOCK | CHANGES_REQUESTED | APPROVED_WITH_CONDITIONS | APPROVED}
179
179
  ```
180
180
 
181
+ Report cap: final message at most about 1,500 tokens; the report is the file at `{output_path}`, other longer material goes to a `mktemp` file (via Bash or Write), and the message gives its path. Exempt, inline in full: in a `/code-review` spawn, the report path, counts and recommendation; in a Workflow spawn, the structured result (`focus`, `reviewed`, `filesExamined`, `findings`).
182
+
181
183
  ## Secret Handling in Findings
182
184
 
183
185
  When a finding involves a secret or credential value, cite `file:line` and the secret TYPE
@@ -44,6 +44,8 @@ Follow the `devflow:apply-decisions` skill to scan the index, Read full bodies o
44
44
 
45
45
  7. **Report status**: Return structured report with pillar evaluations and changes made.
46
46
 
47
+ **Gate ownership:** Run only a test file you added or changed, once. Only Validate runs the full suite.
48
+
47
49
  ## Principles
48
50
 
49
51
  1. **Fix, don't report** - Self-review means fixing issues, not generating reports
@@ -81,6 +83,8 @@ Return structured completion status:
81
83
  - {sha} fix: address self-review issues
82
84
  ```
83
85
 
86
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `### Status` line and `### Files Modified` (`/self-review` reads `changes_made` from them).
87
+
84
88
  ## Boundaries
85
89
 
86
90
  **Escalate to orchestrator (BLOCKED):**
@@ -76,6 +76,8 @@ Your refinement process:
76
76
 
77
77
  You operate autonomously and proactively, refining code immediately after it's written or modified without requiring explicit requests. Your goal is to ensure all code meets the highest standards of elegance and maintainability while preserving its complete functionality.
78
78
 
79
+ **Gate ownership:** Run no build, test or lint command. Reading files and git diffs is not verification. Only Validate runs the full suite.
80
+
79
81
  ## Output
80
82
 
81
83
  Return structured completion status:
@@ -93,6 +95,8 @@ Return structured completion status:
93
95
  - {file} ({change description})
94
96
  ```
95
97
 
98
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the completion status, `### Files Modified` and your commits.
99
+
96
100
  ## Boundaries
97
101
 
98
102
  **Escalate to orchestrator:**
@@ -70,7 +70,7 @@ If you already know a file needs content, go straight to Read — don't skim it
70
70
 
71
71
  ### Step 6: Project Knowledge
72
72
 
73
- If `.devflow/learning/decisions.md` exists, Read its `<!-- TL;DR: ... -->` first-line comment and include active decision count under "### Active Decisions". Only the TL;DR — intentional for token efficiency.
73
+ Read the decisions TL;DR at the repository's main worktree, where the ledger lives. Run `git -C "{start}" rev-parse --path-format=absolute --show-toplevel --git-common-dir`, `{start}` being `WORKTREE_PATH` if provided, otherwise cwd. `{ledger}` is the first that applies: the main worktree — when the output is two absolute lines and line 2 ends in `/.git`, its parent, provided that directory contains `.devflow/` and is not your home directory; else the toplevel, line 1 (on a git older than 2.31, the line after the echoed flag); else `{start}`, when the command failed. If `{ledger}/.devflow/learning/decisions.md` exists, Read its first line, `<!-- TL;DR: N decisions -->`, and report N under "### Active Decisions". Only the TL;DR — intentional for token efficiency.
74
74
 
75
75
  ### Step 7: Generate Summary
76
76
 
@@ -127,6 +127,8 @@ skim also handles prose/config files (`.md`, `.json`, `.yaml`, `.toml`) — the
127
127
  {Brief recommendation based on codebase structure}
128
128
  ```
129
129
 
130
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `### Relevant Files for Task` table.
131
+
130
132
  ## Principles
131
133
 
132
134
  1. **Speed and focus** — Get oriented quickly on what's relevant; task-focused exploration only
@@ -378,6 +378,12 @@ If CYCLE_NUMBER >= 3 and prior FP ratio > 70%: append "Note: High false-positive
378
378
 
379
379
  ---
380
380
 
381
+ ## Output
382
+
383
+ Each mode's template above is its output.
384
+
385
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the synthesis in exploration, planning and design modes. Review, bug-analysis and research modes write the summary to disk and return its path.
386
+
381
387
  ## Principles
382
388
 
383
389
  1. **No new research** - Only synthesize what agents found
@@ -72,18 +72,24 @@ For each scenario:
72
72
 
73
73
  If a previous run failed (PREVIOUS_FAILURES provided), prioritize re-testing those scenarios first.
74
74
 
75
- ## Long-running commands (test/build commands that may run >120s)
75
+ **Gate ownership:** Run the scenario commands. Never the full suite. Only Validate runs the full suite.
76
76
 
77
- A plain `Bash` call defaults to a 120s timeout, and inside a dynamic Workflow a sub-agent that emits no output for 180s is KILLED ("agent stalled"). For any scenario whose command may run silent longer than ~120s (a full `cargo test` / `go test ./...`, a build step, a slow integration suite), do NOT run it as one silent foreground command. Instead:
77
+ ## Running commands
78
78
 
79
- 1. Run it in the BACKGROUND with the Bash tool (`run_in_background: true`), capturing output + exit code under a unique `<slug>` reused in step 2:
80
- `<command> > /tmp/df-test-<slug>.log 2>&1; echo "EXIT=$?" > /tmp/df-test-<slug>.done`
81
- 2. Poll with the `Monitor` tool (load it via ToolSearch `select:Monitor` if it is not available): set `persistent: false`, `timeout_ms` above the expected run time (e.g. 600000), and
82
- `command: until [ -f /tmp/df-test-<slug>.done ]; do echo running; sleep 25; done; echo DONE; cat /tmp/df-test-<slug>.done`
83
- The 25s heartbeat (≪ 180s) is delivered as a notification that keeps you alive past the watchdog.
84
- 3. When the monitor reports `DONE`: the scenario's command PASSED iff the `.done` file contains `EXIT=0`. Read the `.log` for evidence.
79
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
85
80
 
86
- For a foreground command that exceeds the 120s default but stays under 180s, pass an explicit higher `timeout` to the Bash tool (up to 600000ms). Prefer scoping to the changed package/path where possible.
81
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
82
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
83
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
84
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
85
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
86
+ - Never re-run a command when nothing it reads has changed.
87
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
88
+ - The same rules hold inside a dynamic Workflow sub-agent.
89
+
90
+ ## Dev server (test.md only)
91
+
92
+ When a scenario needs the dev server, follow the lifecycle in `devflow:qa/references/browser-testing.md`: start the server in the background, check readiness with one bounded loop inside a single Bash call, and kill the server before you finish. Never end the turn while it runs.
87
93
 
88
94
  ## Output
89
95
 
@@ -139,9 +145,11 @@ One row per TP line, in TP order. PASS only when every scenario covering the TP
139
145
  - **Remediation**: {what Code agent should fix}
140
146
 
141
147
  ### Evidence Log
142
- {Raw command outputs for traceability}
148
+ {Each command run's `LOG=` path, for traceability}
143
149
  ```
144
150
 
151
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: `### Test Plan Evidence` (HEAD line and TP table), the Scenario Results table (`| ID | TP | Type |`) and `### Failed Scenarios`. `### Evidence Log` lists each run's `LOG=` path, never raw output.
152
+
145
153
  ## Principles
146
154
 
147
155
  1. **User perspective** - Test what the user asked for, not implementation internals
@@ -145,6 +145,8 @@ Return the verdict ledger grouped by disposition:
145
145
  - DUPLICATE: {n}
146
146
  ```
147
147
 
148
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the whole ledger, since `/resolve` checks that every issue id appears in it.
149
+
148
150
  ## Boundaries
149
151
 
150
152
  **You are TRIAGE ONLY — read and judge, never write:**
@@ -38,18 +38,20 @@ Execute in this order, stopping on first failure:
38
38
  | 3 | Lint | `npm run lint`, `cargo clippy`, `make lint` |
39
39
  | 4 | Test | `npm test`, `cargo test`, `make test` |
40
40
 
41
- ## Long-running commands (builds/tests that may run >120s)
41
+ **Gate ownership:** Run the full suite once per HEAD. You are the only agent that does.
42
42
 
43
- A plain `Bash` call defaults to a 120s timeout, and inside a dynamic Workflow a sub-agent that emits no output for 180s is KILLED ("agent stalled"). For any build/test that may run silent longer than ~120s (cold `cargo build`/`cargo test`, large `tsc`, `gradle`, `go build ./...`), do NOT run it as one silent foreground command. Instead:
43
+ ## Running commands
44
44
 
45
- 1. Run it in the BACKGROUND, capturing output + exit code. With the Bash tool set `run_in_background: true` and pick a unique `<slug>` (reuse the same paths in step 2):
46
- `<command> > /tmp/df-val-<slug>.log 2>&1; echo "EXIT=$?" > /tmp/df-val-<slug>.done`
47
- 2. Poll with the `Monitor` tool (load it via ToolSearch `select:Monitor` if it is not available): set `persistent: false`, `timeout_ms` above the expected run time (e.g. 600000), and
48
- `command: until [ -f /tmp/df-val-<slug>.done ]; do echo running; sleep 25; done; echo DONE; cat /tmp/df-val-<slug>.done`
49
- The 25s heartbeat (≪ 180s) is delivered as a notification that keeps you alive past the watchdog.
50
- 3. When the monitor reports `DONE`: the command PASSED iff the `.done` file contains `EXIT=0`. Read the `.log` for failure details to parse.
45
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
51
46
 
52
- For a foreground command that merely exceeds the 120s default but stays well under 180s, simply pass an explicit higher `timeout` to the Bash tool (up to 600000ms). Prefer package-scoped commands (`cargo build -p <crate>`, `cargo test -p <crate>`) when the project supports them.
47
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
48
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
49
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
50
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
51
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
52
+ - Never re-run a command when nothing it reads has changed.
53
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
54
+ - The same rules hold inside a dynamic Workflow sub-agent.
53
55
 
54
56
  ## Principles
55
57
 
@@ -78,7 +80,7 @@ HEAD: {the 40-hex `git rev-parse HEAD`, read before the first command}
78
80
 
79
81
  ### Failures (if FAIL)
80
82
 
81
- #### typecheck
83
+ #### typecheck (first 30 lines; full output at {log path})
82
84
  ```
83
85
  src/auth/login.ts:42:15 - error TS2339: Property 'email' does not exist on type 'User'.
84
86
  src/auth/login.ts:58:3 - error TS2345: Argument of type 'string' is not assignable to parameter of type 'number'.
@@ -92,6 +94,8 @@ src/auth/login.ts:58:3 - error TS2345: Argument of type 'string' is not assignab
92
94
  {Description of why validation couldn't run - e.g., missing dependencies, broken config}
93
95
  ```
94
96
 
97
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `HEAD:` line, the `| Command | Status | Exit | Duration |` table and Parsed References. Failure output: at most 30 lines per failing command, then its log path.
98
+
95
99
  ## Boundaries
96
100
 
97
101
  **Escalate to orchestrator (BLOCKED):**
@@ -4,7 +4,7 @@
4
4
  ORDER IS LOAD-BEARING. Run exactly in this sequence:
5
5
 
6
6
  1. **Validate agent** — build / typecheck / lint / test
7
- - Build/test commands that may run silent for >~120s MUST follow `build_execution_doctrine()` (background Bash + Monitor poll), or they trip the 180s workflow watchdog.
7
+ - Build/test commands follow `build_execution_doctrine()`: the Validate agent's Running commands block, foreground under an explicit Bash timeout.
8
8
  - FAIL → Code agent fix (max 2 retries) → re-run Validate agent
9
9
  - If still FAIL after 2 retries → escalate (do not loop endlessly)
10
10
  2. **Simplify agent** — reduce complexity, remove duplication
@@ -19,7 +19,7 @@ Depth scales to change size + budget: a trivial one-line fix warrants a lighter
19
19
  1. **Gate 1 #1** — immediately after the initial Code agent implementation (inside `implement_bundle()`).
20
20
  2. **Gate 1 #2** — the FINAL gate, after ALL Gate-2 fixes AND the entire review pass have completed.
21
21
 
22
- It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's "Long-running commands" discipline). The final Gate 1 #2 is the invariant that all written code passes before merge.
22
+ It does NOT run inside the review pass, nor after each individual Gate-2 / review / QA fix. At those points the fixing Code agent self-verifies its OWN build compiles (see `review_pass()` and the Code agent's Running commands block). The final Gate 1 #2 is the invariant that all written code passes before merge.
23
23
  @end
24
24
 
25
25
  @define gate2_acceptance():
@@ -115,7 +115,7 @@ Majority-survives: a finding needs >50% of verification lenses to confirm it. St
115
115
 
116
116
  If no surviving findings: return early (no fixes needed). Any coverageGaps are carried in the return — they block a PASS verdict downstream, not the early exit.
117
117
 
118
- If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's "Long-running commands" discipline). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
118
+ If survivors remain: batch the confirmed findings for fixing: group findings by file — one file per set of sub-batches, chunked at max 5 findings per sub-batch; never mix two files in one batch. A finding with no `file` field is its own singleton batch. Sub-batches for the SAME file run sequentially (never two Code agents editing the same file concurrently — same-file edits in `parallel()` cause index contention and lost fixes); sub-batches for DISTINCT files run via `parallel()` in staggered chunks of ~5, same pacing bar as the Review spawn path (different code areas — safe per concurrency doctrine). Each Code agent's prompt pins a return contract: `{"status": "fixed"|"blocked", "commitShas": [...], "unresolved": [...]}` — a chunk is FIXED only when `result.status === "fixed"` AND `commitShas` is non-empty AND `result.unresolved` is empty; never decide disposition from status alone. A non-empty `unresolved` list means the agent named work it could not complete — carry the whole chunk into `survivingFindings` rather than guessing which findings the strings map to. `survivingFindings` = findings NOT addressed: fix Code agent dead/failed/blocked/deferred OR committed but left work named in `unresolved`. The fixing Code agent **self-verifies its own fix builds** (build/typecheck per the Code agent's Running commands block, once over its whole batch). Do **NOT** run Gate 1 or Gate 2 inside the pass (no Validate agent, no Simplify agent, no Scrutinize agent, no Evaluate agent, no Test agent). The engine runs ONE final Gate 1 after the pass exits — see the `gate1_postcode()` cadence (Gate 1 #2).
119
119
  @end
120
120
 
121
121
  @define concurrency_doctrine():
@@ -135,42 +135,26 @@ This applies to both: multiple Code agents working on a single ticket AND multi-
135
135
  @end
136
136
 
137
137
  @define build_execution_doctrine():
138
- ### Build execution doctrine — long-running commands (LOAD-BEARING)
138
+ ### Build execution doctrine — the Running commands block (LOAD-BEARING)
139
139
 
140
- The Workflow runtime KILLS any sub-agent that emits no output for 180 seconds. A cold `cargo build`, `cargo test`, a large `tsc`, `gradle build`, `go build ./...`, etc. routinely runs silent far longer and trips this watchdog (the failure reads `agent stalled on all N attempts`). Plain foreground `Bash` also defaults to a 120s timeout.
140
+ The Validate, Code and Test agents each carry this `## Running commands` block in their bodies, so each prompt in this engine names the block ("your Running commands block") instead of copying it:
141
141
 
142
- **RULE: any agent (Validate agent, Code agent, Test agent) running a build / test / compile / install that may run silent for more than ~120s MUST run it in the BACKGROUND and POLL — never as a single silent foreground command.**
142
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
143
143
 
144
- Build commands are NEVER wrapped in `sh -c`, `bash -c`, or inline interpreters (`python3 -c`, `node -e`). Invoke commands directly with the step-1 redirect form — permission systems deny wrapper-invoked commands that would be allowed directly.
145
-
146
- Mechanical procedure (spike-verified — a workflow sub-agent survived a 253s job this way):
147
-
148
- 0. **Pre-load Monitor:** before launching any background task, load the `Monitor` tool via ToolSearch (`select:Monitor`).
149
- 1. Choose ONE unique base path for this run and reuse it verbatim in steps 1–3, e.g. `BASE=/tmp/df-build-<ticket-slug>`. Launch the command with the Bash tool using `run_in_background: true`:
150
- ```
151
- <build/test command> > <BASE>.log 2>&1; echo "EXIT=$?" > <BASE>.done
152
- ```
153
- This returns immediately with a background task id — do NOT block on it.
154
- 2. Arm ONE Monitor that emits a heartbeat well under 180s AND exits when the job finishes:
155
- - description: short, e.g. `await <build cmd>`
156
- - persistent: false
157
- - timeout_ms: comfortably ABOVE the expected job time (e.g. 600000)
158
- - command: `until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
159
-
160
- The `building` heartbeat every 25s (≪ 180s) is delivered as a notification that re-invokes you, so the watchdog never sees a >180s gap. `BUILD_DONE` + the `EXIT=` line signal completion.
161
-
162
- **Exit-code honesty:** the background task's own exit status is meaningless (the trailing `echo` always exits 0). ALWAYS read the `EXIT=` value written inside `<BASE>.done` — that is the authoritative result.
163
-
164
- **Bounded polling:** arm ONE Monitor then stop acting. On Monitor timeout, re-arm at most 2× (never more than 3 total Monitor calls per build). If the build has not finished after 3 Monitor calls: record the state and escalate — never babysit. A Code agent burned 241k tokens polling one build.
165
- 3. When the monitor reports `BUILD_DONE`: the job PASSES iff `<BASE>.done` contains `EXIT=0`. Read `<BASE>.log` for output/failure detail.
144
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
145
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
146
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
147
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
148
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
149
+ - Never re-run a command when nothing it reads has changed.
150
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
151
+ - The same rules hold inside a dynamic Workflow sub-agent.
166
152
 
167
153
  **Cheapest-sufficient validation:** iterate with the fastest check that proves the change — `tsc --noEmit`, `cargo check`, `go vet`, a single test file — the ecosystem's cheapest sufficient signal. Reserve expensive full/optimized builds and whole-suite runs for the final gates only.
168
154
 
169
155
  **One build gate per phase:** batch related fixes, validate once. A fix-pass Code agent runs ONE light check over its whole batch — never several invocations per small fix. Do NOT validate after every individual mutation; validate once after the batch is complete.
170
156
 
171
- **Scope commands to stay short.** During the engine, PREFER crate/package-scoped builds and tests — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — over the whole workspace. The full-workspace regression is the human's job after the wave (the wave already hands the integrated branch back to the user). Scoping keeps most commands under the watchdog window and under budget.
172
-
173
- **Invariants:** heartbeat interval MUST stay well under 180s (25–30s is the tested value); Monitor `timeout_ms` MUST exceed the expected job duration (a too-short timeout kills the poll, not the build). Never substitute a single silent long command for this procedure.
157
+ **Scope commands to stay short.** During the engine, prefer the scoped command for the change over the whole workspace — `cargo build -p <crate>`, `cargo test -p <crate>`, `npm test -- <path>`, `go test ./pkg/...` — and, when the whole set is needed, one workspace-level command over a per-package loop. After the wave, the full-workspace regression is the human's job (the wave already hands the integrated branch back to the user).
174
158
  @end
175
159
 
176
160
  @define engine_output_schema():
@@ -86,8 +86,6 @@ DIRECTORIES: {list of primary directories touched by this workflow}
86
86
  FILES_CHANGED: {list of files changed}
87
87
  DECISIONS_CONTEXT: {DECISIONS_CONTEXT if available, else (none)}
88
88
 
89
- Load the devflow:feature-knowledge skill and follow its authoring process.
90
-
91
89
  Write the knowledge base to:
92
90
  {worktree}/.devflow/features/{slug}/KNOWLEDGE.md
93
91
 
@@ -1,5 +1,5 @@
1
1
  @define issue_ref_grammar():
2
- **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `$ARGUMENTS` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
2
+ **Issue-reference grammar (L1 — command layer, permissive and provider-blind):** scan `COMMAND_INPUT` for candidate issue references — a `#`-prefixed token and a bare digit run are both candidates — and collect them in source order as the raw token list `ISSUE_REFS`. Forward that list to the Git agent **verbatim**: the command never renders, normalises, pads, strips or coerces a token, and never rules a candidate out. Under `github` a token matching `^#?[1-9][0-9]{0,8}$` **is** a reference and the Git agent renders it as `#{n}`.
3
3
 
4
4
  **A token of any other shape is neither coerced nor dropped silently — and no producer-side grammar check rejects it before the fetch.** Adjudication belongs to the operation that runs, and each one answers in its own Output block: `fetch-issue` strips a leading `#` and takes the text branch, so a non-numeric token is used as a **search term** and the operation returns the first open match or nothing; `fetch-issues-batch` resolves each token to an issue number, drops the ones it cannot resolve, and names them in `NOT_FOUND ({refs})` beside the issues it did fetch. Read the outcome from the operation that ran — a token's shape is a verdict nowhere, and there is nothing upstream holding it back.
5
5
 
@@ -198,8 +198,6 @@ Only `exit=0` means the skill is installed. On any other result, do NOT add that
198
198
 
199
199
  **Produces:** DECISIONS_CONTEXT, FEATURE_KNOWLEDGE, COMPLIANCE_FRAMEWORKS (carried from Step 0b)
200
200
 
201
- **Load Companion Skills** — Load via Skill tool: `devflow:quality-gates`, `devflow:software-design`. If a skill fails to load, continue without it.
202
-
203
201
  {{decisions_load()}}
204
202
 
205
203
  This produces a compact index of active ADR/PF entries. Pass `DECISIONS_CONTEXT` to all Review agents. Review agents use `devflow:apply-decisions` to Read full entry bodies on demand.
@@ -19,7 +19,13 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
19
19
 
20
20
  ## Input
21
21
 
22
- `$ARGUMENTS` contains whatever follows `/debug`:
22
+ What follows `/debug` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
23
+
24
+ <command-input>
25
+ $ARGUMENTS
26
+ </command-input>
27
+
28
+ `COMMAND_INPUT` is one of:
23
29
  - Bug description: "login fails after session timeout"
24
30
  - Issue reference: "#42"
25
31
  - Empty: use conversation context
@@ -30,8 +36,6 @@ Investigate bugs by spawning parallel agents, each pursuing a different hypothes
30
36
 
31
37
  **Produces:** DECISIONS_CONTEXT
32
38
 
33
- **Load Companion Skills** — Load via Skill tool: `devflow:test-driven-development`, `devflow:software-design`, `devflow:testing`. If a skill fails to load, continue without it.
34
-
35
39
  {{decisions_load()}}
36
40
 
37
41
  The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Phase 2) — prior pitfalls and decisions can suggest specific root causes to investigate. Follow `devflow:apply-decisions` to Read full entry bodies on demand. **Do NOT pass `DECISIONS_CONTEXT` to Explore investigators** — decisions context stays in the orchestrator; investigators examine code directly.
@@ -43,7 +47,7 @@ The orchestrator uses `DECISIONS_CONTEXT` locally when generating hypotheses (Ph
43
47
  **Produces:** HYPOTHESES, BUG_CONTEXT
44
48
  **Requires:** DECISIONS_CONTEXT
45
49
 
46
- If `$ARGUMENTS` opens with a candidate issue reference, fetch the issue:
50
+ If `COMMAND_INPUT` opens with a candidate issue reference, fetch the issue:
47
51
 
48
52
  {{issue_ref_grammar()}}
49
53
 
@@ -88,7 +92,8 @@ Return a structured report:
88
92
  - Status: CONFIRMED / DISPROVED / PARTIAL
89
93
  - Evidence FOR: [list with file:line refs]
90
94
  - Evidence AGAINST: [list with file:line refs]
91
- - Key finding: {one-sentence summary}"
95
+ - Key finding: {one-sentence summary}
96
+ Keep the whole report to at most about 1,500 tokens."
92
97
 
93
98
  Agent(subagent_type="Explore"):
94
99
  "Investigate this bug: {bug_description}
@@ -96,7 +101,7 @@ Agent(subagent_type="Explore"):
96
101
  Hypothesis: {hypothesis B description}
97
102
  Focus area: {specific code area, mechanism, or condition}
98
103
 
99
- [same steps and return format]"
104
+ [same steps and return format; the whole report at most about 1,500 tokens]"
100
105
 
101
106
  Agent(subagent_type="Explore"):
102
107
  "Investigate this bug: {bug_description}
@@ -104,7 +109,7 @@ Agent(subagent_type="Explore"):
104
109
  Hypothesis: {hypothesis C description}
105
110
  Focus area: {specific code area, mechanism, or condition}
106
111
 
107
- [same steps and return format]"
112
+ [same steps and return format; the whole report at most about 1,500 tokens]"
108
113
 
109
114
  (Add more investigators if bug complexity warrants 4-5 hypotheses)
110
115
  ```
@@ -116,7 +121,7 @@ Focus area: {specific code area, mechanism, or condition}
116
121
 
117
122
  Evaluate investigation verdicts before synthesis:
118
123
 
119
- - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias)
124
+ - **One CONFIRMED**: Spawn 1-2 additional `Agent(subagent_type="Explore")` agents to validate from different angles (prevent confirmation bias), each reporting at most about 1,500 tokens
120
125
  - **Multiple PARTIAL**: Look for a unifying root cause that explains all partial evidence
121
126
  - **All DISPROVED**: Report honestly — "No root cause identified from initial hypotheses." Generate 2-3 second-round hypotheses if conversation context suggests avenues not yet explored. Loop back to Phase 3.
122
127
 
@@ -79,7 +79,13 @@ Note the `budget` value from the Workflow tool context (or default to "medium" i
79
79
 
80
80
  **3. Detect mode: SINGLE or WAVE**
81
81
 
82
- - **SINGLE mode:** input is one ticket, one issue, one task description, or one plan document
82
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
83
+
84
+ <command-input>
85
+ $ARGUMENTS
86
+ </command-input>
87
+
88
+ - **SINGLE mode:** `COMMAND_INPUT` is one ticket, one issue, one task description, or one plan document
83
89
  - **WAVE mode:** input is a set of tracker issues (wave labels, milestone, issue list), or the user says "wave" / "all tickets in wave N"
84
90
  - **A `/devflow:dynamic-tickets` ticket directory** (`{worktree}/.devflow/docs/tickets/{slug}/{ts}/`) is WAVE input: in each ticket file (every `.md` there but `tracking-issue.md`), the `**Issue:**` line directly after `**Depends on:**` is one raw `ISSUE_REFS` token, forwarded to the wave's pre-fetch verbatim — never rendered, normalised or re-derived. A ticket file with no `**Issue:**` line, or more than one, contributes no token; name it in the run summary as `not filed`.
85
91
 
@@ -113,7 +119,7 @@ If none found: build proceeds Gate-1-only (Gate 2 skipped with a note). Never re
113
119
  **5. Resolve tracking-issue number (optional)**
114
120
 
115
121
  Check, in priority order:
116
- - An explicit candidate issue reference or issue URL in the user's input (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
122
+ - An explicit candidate issue reference or issue URL in `COMMAND_INPUT` (e.g. `#42`, `42`, or `https://github.com/…/issues/42`)
117
123
  - The `**Issue:**` line directly after the H1 of the ticket set's `tracking-issue.md` (written by `/devflow:dynamic-tickets`' filing step; the file is at `{worktree}/.devflow/docs/tickets/{slug}/{ts}/tracking-issue.md`), as a raw token — only when the file holds exactly one `**Issue:**` line
118
124
  - Otherwise: none
119
125
 
@@ -194,7 +200,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
194
200
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
195
201
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
196
202
 
197
- When you build or run tests to verify your work, use your "Long-running commands" discipline (background-Bash + Monitor poll) for anything that may run silent >120s, and prefer package-scoped commands.
203
+ When you build or run tests to verify your work, follow your Running commands block.
198
204
 
199
205
  After implementing, commit your changes with a conventional-commit message and report:
200
206
  - Files changed
@@ -206,7 +212,7 @@ After implementing, commit your changes with a conventional-commit message and r
206
212
  const gate1 = await phase("gate1", async () => {
207
213
  // Gate 1 #1 — runs once, immediately after the initial implementation.
208
214
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH}.
209
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
215
+ Follow your Running commands block for every build and test command.
210
216
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
211
217
 
212
218
  if (validation.verdict === "FAIL") {
@@ -261,7 +267,7 @@ ${panel.filter(p => p.verdict === "FAIL").map(p => p.rationale).join("\n")}
261
267
  ISSUE_NUMBER: ${ISSUE_NUMBER}
262
268
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
263
269
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
264
- Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit fixes.`, { agentType: "Code" });
270
+ Self-verify your fix compiles, following your Running commands block. Commit fixes.`, { agentType: "Code" });
265
271
  evalVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
266
272
  }
267
273
  }
@@ -271,7 +277,7 @@ Self-verify your fix compiles (background-Bash + Monitor for any build >120s —
271
277
  const testResult = await agent(`Run scenario-based acceptance tests on branch ${BRANCH} against these criteria:
272
278
  ${CRITERIA || "(none)"}
273
279
  TEST_PLAN: ${TEST_PLAN || "(none)"}
274
- For any test/build command that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog.
280
+ Follow your Running commands block for every scenario command.
275
281
  Cover: functionality, API contracts, performance, and cover every TEST_PLAN scenario. Report: PASS or FAIL per scenario.`, { agentType: "Test" });
276
282
  testVerdict = testResult.verdict;
277
283
  if (testVerdict === "FAIL") {
@@ -281,7 +287,7 @@ ${testResult.failures}
281
287
  ISSUE_NUMBER: ${ISSUE_NUMBER}
282
288
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
283
289
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
284
- Self-verify your fix compiles and the scenarios pass (background-Bash + Monitor for any build/test >120s). Commit fixes.`, { agentType: "Code" });
290
+ Self-verify your fix compiles and the scenarios pass, following your Running commands block. Commit fixes.`, { agentType: "Code" });
285
291
  testVerdict = "FAIL-FIXED"; // issues found, fixes applied, not re-evaluated by design
286
292
  }
287
293
  }
@@ -395,7 +401,7 @@ ${chunk.map(f => `- ${f.description} (${f.severity})`).join("\n")}
395
401
  ISSUE_NUMBER: ${ISSUE_NUMBER}
396
402
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
397
403
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
398
- Fix all findings in this batch. Self-verify your fix compiles (background-Bash + Monitor for any build >120s — see your "Long-running commands" discipline). Commit with conventional-commit message.
404
+ Fix all findings in this batch. Self-verify your fix compiles, following your Running commands block. Commit with conventional-commit message.
399
405
  Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<description of any finding that could not be fixed>"]}`, { agentType: "Code" });
400
406
  chunkResults.push({ chunk, result: r });
401
407
  }
@@ -435,7 +441,7 @@ Return: {"status": "fixed"|"blocked", "commitShas": ["<sha>"], "unresolved": ["<
435
441
  // this is the build gate before the branch is handed back. See gate1_postcode() cadence.
436
442
  const gate1Final = await phase("gate1-final", async () => {
437
443
  const validation = await agent(`Run build, typecheck, lint, and tests on branch ${BRANCH} (final gate after all fixing).
438
- For any build/test that may run silent >120s, use the background-Bash + Monitor poll procedure (your "Long-running commands" discipline) so you never trip the 180s watchdog; prefer package-scoped commands.
444
+ Follow your Running commands block for every build and test command.
439
445
  Report: PASS or FAIL with details.`, { agentType: "Validate" });
440
446
 
441
447
  if (validation.verdict === "FAIL") {
@@ -447,7 +453,7 @@ ISSUE_NUMBER: ${ISSUE_NUMBER}
447
453
  ISSUE_PR_LINK: ${ISSUE_PR_LINK}
448
454
  COMPLIANCE_FRAMEWORKS: ${COMPLIANCE_FRAMEWORKS}
449
455
  Self-verify your fix compiles. Commit fixes with conventional-commit message.`, { agentType: "Code" });
450
- const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
456
+ const recheck = await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
451
457
  if (recheck.verdict === "PASS") break;
452
458
  failureDetails = recheck.details || failureDetails;
453
459
  if (attempt === 2) return { verdict: "ESCALATED", reason: "Final Gate 1 validation exhausted after 2 Code agent fix attempts" };
@@ -459,7 +465,7 @@ Self-verify your fix compiles. Commit fixes with conventional-commit message.`,
459
465
  const scrutiny = await agent(`9-pillar self-review of recent changes on branch ${BRANCH}. Report any code you changed and your findings.`, { agentType: "Scrutinize" });
460
466
 
461
467
  if (scrutiny.codeChanged) {
462
- await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; background+Monitor for long commands). Report: PASS or FAIL.`, { agentType: "Validate" });
468
+ await agent(`Re-run build, typecheck, lint, tests on branch ${BRANCH} (Scrutinize agent made changes; follow your Running commands block). Report: PASS or FAIL.`, { agentType: "Validate" });
463
469
  }
464
470
 
465
471
  return { verdict: "PASS" };
@@ -59,7 +59,13 @@ If present, note its contents as `PREFERENCE_PROFILE`. This will be used to auto
59
59
 
60
60
  **3. Resolve ticket input**
61
61
 
62
- Determine the ticket source (in priority order):
62
+ What follows the command is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
63
+
64
+ <command-input>
65
+ $ARGUMENTS
66
+ </command-input>
67
+
68
+ Determine the ticket source from `COMMAND_INPUT` (in priority order):
63
69
  - A directory of ticket `.md` files (from `/devflow:dynamic-tickets` output)
64
70
  - A list of candidate issue references or issue URLs
65
71
  - Inline ticket descriptions passed as args
@@ -18,7 +18,13 @@ Explore a codebase area by spawning parallel agents for flow tracing, dependency
18
18
 
19
19
  ## Input
20
20
 
21
- `$ARGUMENTS` contains whatever follows `/explore`:
21
+ What follows `/explore` is bound once, here. Every later step names it `COMMAND_INPUT` and never restates it:
22
+
23
+ <command-input>
24
+ $ARGUMENTS
25
+ </command-input>
26
+
27
+ `COMMAND_INPUT` is one of:
22
28
  - Area description: "how does the auth system work"
23
29
  - Flow question: "trace the request lifecycle"
24
30
  - Empty: use conversation context
@@ -58,6 +64,8 @@ Based on Skim agent findings, spawn 2-3 `Agent(subagent_type="Explore")` agents
58
64
 
59
65
  Adjust explorer focus based on the specific exploration question.
60
66
 
67
+ Ask each explorer for a final report of at most about 1,500 tokens: findings with file:line references, not file dumps.
68
+
61
69
  ### Phase 4: Synthesize
62
70
 
63
71
  **Produces:** MERGED_FINDINGS