devflow-kit 3.1.0 → 3.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (138) hide show
  1. package/CHANGELOG.md +52 -0
  2. package/README.md +2 -2
  3. package/dist/cli/agents-view/render.js +69 -15
  4. package/dist/cli/agents-view/state.js +40 -14
  5. package/dist/cli/commands/agents.js +135 -45
  6. package/dist/cli/commands/init.js +128 -53
  7. package/dist/cli/commands/learning.js +61 -13
  8. package/dist/cli/commands/memory.js +35 -14
  9. package/dist/cli/commands/uninstall.js +163 -39
  10. package/dist/commands/code-review.md +1 -3
  11. package/dist/commands/debug.md +15 -12
  12. package/dist/commands/dynamic-build.md +172 -135
  13. package/dist/commands/dynamic-plan.md +9 -3
  14. package/dist/commands/explore.md +10 -4
  15. package/dist/commands/implement.md +149 -145
  16. package/dist/commands/plan.md +13 -9
  17. package/dist/commands/release.md +8 -2
  18. package/dist/commands/research.md +8 -2
  19. package/dist/commands/resolve.md +28 -19
  20. package/dist/commands/self-review.md +16 -13
  21. package/dist/core/agent-frontmatter.js +25 -0
  22. package/dist/core/agent-models.js +201 -36
  23. package/dist/core/agent-state.js +27 -5
  24. package/dist/core/assets.js +1 -1
  25. package/dist/core/feature-config.js +68 -10
  26. package/dist/core/flags.js +24 -0
  27. package/dist/core/learning-queue-cleanup.js +10 -11
  28. package/dist/core/learning-tuning-config.js +8 -0
  29. package/dist/core/linked-path.js +46 -0
  30. package/dist/core/plugins.js +16 -5
  31. package/dist/core/queue-drain.js +31 -0
  32. package/dist/hud/components/learning-counts.js +54 -8
  33. package/dist/skills/git/references/tracker/github/create-release.md +2 -2
  34. package/dist/skills/git/references/tracker/github/gather-release-evidence.md +1 -1
  35. package/dist/skills/git/references/tracker/jira/create-release.md +2 -2
  36. package/dist/skills/git/references/tracker/jira/gather-release-evidence.md +1 -1
  37. package/dist/skills/git/references/tracker/linear/create-release.md +2 -2
  38. package/dist/skills/git/references/tracker/linear/gather-release-evidence.md +1 -1
  39. package/dist/targets/claude-code/installer.js +36 -9
  40. package/dist/targets/claude-code/post-install.js +128 -38
  41. package/package.json +1 -1
  42. package/src/assets/agents/code.md +85 -35
  43. package/src/assets/agents/design.md +12 -0
  44. package/src/assets/agents/diagnose.md +18 -11
  45. package/src/assets/agents/evaluate.md +17 -24
  46. package/src/assets/agents/knowledge.md +7 -3
  47. package/src/assets/agents/learning.md +4 -6
  48. package/src/assets/agents/research.md +21 -0
  49. package/src/assets/agents/review.md +12 -0
  50. package/src/assets/agents/scrutinize.md +37 -9
  51. package/src/assets/agents/simplify.md +24 -0
  52. package/src/assets/agents/skim.md +6 -2
  53. package/src/assets/agents/synthesize.md +18 -0
  54. package/src/assets/agents/test.md +19 -11
  55. package/src/assets/agents/triage.md +8 -0
  56. package/src/assets/agents/validate.md +20 -11
  57. package/src/assets/commands/_partials/_engine.mds +36 -55
  58. package/src/assets/commands/_partials/_knowledge.mds +1 -3
  59. package/src/assets/commands/_partials/_plan_contract.mds +1 -1
  60. package/src/assets/commands/_partials/_tracker.mds +1 -1
  61. package/src/assets/commands/_partials/_wave.mds +8 -6
  62. package/src/assets/commands/code-review.mds +1 -3
  63. package/src/assets/commands/debug.mds +13 -8
  64. package/src/assets/commands/dynamic-build.mds +126 -72
  65. package/src/assets/commands/dynamic-plan.mds +7 -1
  66. package/src/assets/commands/explore.mds +9 -1
  67. package/src/assets/commands/implement.mds +147 -141
  68. package/src/assets/commands/plan.mds +12 -8
  69. package/src/assets/commands/release.md +8 -2
  70. package/src/assets/commands/research.mds +8 -2
  71. package/src/assets/commands/resolve.mds +27 -16
  72. package/src/assets/commands/self-review.mds +15 -10
  73. package/src/assets/mds/tracker/_common.mds +1 -1
  74. package/src/assets/mds/tracker/_github.mds +2 -2
  75. package/src/assets/mds/tracker/_jira.mds +2 -2
  76. package/src/assets/mds/tracker/_linear.mds +2 -2
  77. package/src/assets/scripts/ci-wait.cjs +636 -0
  78. package/src/assets/scripts/hooks/assets/orchestrator-charter.md +4 -3
  79. package/src/assets/scripts/hooks/background-memory-update +356 -17
  80. package/src/assets/scripts/hooks/capture-prompt +4 -3
  81. package/src/assets/scripts/hooks/capture-question +4 -3
  82. package/src/assets/scripts/hooks/capture-turn +4 -3
  83. package/src/assets/scripts/hooks/ensure-devflow-init +13 -1
  84. package/src/assets/scripts/hooks/ensure-root-gitignore +122 -10
  85. package/src/assets/scripts/hooks/git-marker +71 -0
  86. package/src/assets/scripts/hooks/json-helper.cjs +12 -145
  87. package/src/assets/scripts/hooks/json-parse +24 -129
  88. package/src/assets/scripts/hooks/lib/learning-store.cjs +169 -64
  89. package/src/assets/scripts/hooks/lib/render-decisions.cjs +1 -1
  90. package/src/assets/scripts/hooks/memory-worker +10 -0
  91. package/src/assets/scripts/hooks/pre-compact-memory +66 -14
  92. package/src/assets/scripts/hooks/preamble +9 -1
  93. package/src/assets/scripts/hooks/queue-append +53 -21
  94. package/src/assets/scripts/hooks/session-start-context +108 -29
  95. package/src/assets/scripts/hooks/session-start-memory +33 -11
  96. package/src/assets/scripts/release-trace.cjs +27 -10
  97. package/src/assets/skills/accessibility/SKILL.md +1 -1
  98. package/src/assets/skills/apply-decisions/SKILL.md +12 -82
  99. package/src/assets/skills/apply-feature-knowledge/SKILL.md +8 -42
  100. package/src/assets/skills/architecture/SKILL.md +1 -1
  101. package/src/assets/skills/boundary-validation/SKILL.md +1 -1
  102. package/src/assets/skills/complexity/SKILL.md +1 -1
  103. package/src/assets/skills/compliance/SKILL.md +1 -1
  104. package/src/assets/skills/consistency/SKILL.md +1 -1
  105. package/src/assets/skills/database/SKILL.md +1 -1
  106. package/src/assets/skills/dependencies/SKILL.md +1 -1
  107. package/src/assets/skills/dependency-research/SKILL.md +3 -6
  108. package/src/assets/skills/design-review/SKILL.md +1 -1
  109. package/src/assets/skills/docs-framework/SKILL.md +1 -1
  110. package/src/assets/skills/documentation/SKILL.md +1 -1
  111. package/src/assets/skills/gap-analysis/SKILL.md +1 -1
  112. package/src/assets/skills/git/SKILL.md +1 -1
  113. package/src/assets/skills/go/SKILL.md +1 -1
  114. package/src/assets/skills/java/SKILL.md +1 -1
  115. package/src/assets/skills/patterns/SKILL.md +1 -1
  116. package/src/assets/skills/performance/SKILL.md +1 -1
  117. package/src/assets/skills/python/SKILL.md +1 -1
  118. package/src/assets/skills/qa/SKILL.md +1 -3
  119. package/src/assets/skills/quality-gates/SKILL.md +9 -12
  120. package/src/assets/skills/quality-gates/references/report-template.md +20 -20
  121. package/src/assets/skills/react/SKILL.md +1 -1
  122. package/src/assets/skills/regression/SKILL.md +1 -1
  123. package/src/assets/skills/reliability/SKILL.md +1 -1
  124. package/src/assets/skills/research-codebase/SKILL.md +1 -1
  125. package/src/assets/skills/research-competitor/SKILL.md +1 -1
  126. package/src/assets/skills/research-external/SKILL.md +1 -1
  127. package/src/assets/skills/research-technology/SKILL.md +1 -1
  128. package/src/assets/skills/review-methodology/SKILL.md +1 -1
  129. package/src/assets/skills/rust/SKILL.md +1 -1
  130. package/src/assets/skills/security/SKILL.md +1 -1
  131. package/src/assets/skills/software-design/SKILL.md +1 -1
  132. package/src/assets/skills/test-driven-development/SKILL.md +15 -33
  133. package/src/assets/skills/testing/SKILL.md +1 -1
  134. package/src/assets/skills/typescript/SKILL.md +1 -1
  135. package/src/assets/skills/ui-design/SKILL.md +1 -1
  136. package/src/assets/skills/worktree-support/SKILL.md +3 -55
  137. package/src/assets/skills/worktree-support/references/discovery.md +48 -0
  138. package/src/assets/skills/worktree-support/references/roots.md +2 -2
@@ -2,17 +2,32 @@
2
2
  name: Code
3
3
  description: Autonomous task implementation on feature branch. Implements, tests, and commits.
4
4
  model: sonnet
5
+ effort: high
5
6
  skills:
6
- - devflow:software-design
7
7
  - devflow:git
8
- - devflow:patterns
9
8
  - devflow:testing
10
9
  - devflow:test-driven-development
11
- - devflow:dependency-research
12
- - devflow:boundary-validation
13
10
  - devflow:worktree-support
14
11
  - devflow:apply-feature-knowledge
15
12
  - devflow:apply-decisions
13
+ disallowedTools:
14
+ - Agent
15
+ - SendMessage
16
+ - NotebookEdit
17
+ - EnterWorktree
18
+ - ExitWorktree
19
+ - ArtifactComments
20
+ - ArtifactData
21
+ - TodoWrite
22
+ - AskUserQuestion
23
+ - TaskOutput
24
+ - ScheduleWakeup
25
+ - CronCreate
26
+ - CronDelete
27
+ - CronList
28
+ - RemoteTrigger
29
+ - PushNotification
30
+ - DesignSync
16
31
  ---
17
32
 
18
33
  # Code Agent
@@ -28,7 +43,7 @@ You receive from orchestrator:
28
43
  - **EXECUTION_PLAN**: Synthesized plan with steps, files, tests
29
44
  - **PATTERNS**: Codebase patterns to follow
30
45
  - **CREATE_PR**: Whether to create PR when done (true/false)
31
- - **OPERATION** (optional): `implement` (default) | `issue-fix` | `validation-fix` | `alignment-fix` | `qa-fix` | `pr-create` — selects operating mode (see below)
46
+ - **OPERATION** (optional): `implement` (default when absent) | `issue-fix` | `validation-fix` | `alignment-fix` | `qa-fix` | `pr-create` | `ci-fix` | `edit` — selects operating mode (see below); every spawn passes it as the first prompt line
32
47
  - **ISSUES** (when OPERATION: issue-fix): Pre-classified issues from Triage agent with disposition FIX_NOW; do not re-litigate
33
48
  - **SCOPE** (when OPERATION: issue-fix): Blast-radius scope hint (Standard | Careful) per issue from Triage agent
34
49
  - **PUSH** (optional): `true` (default) | `false` — when false, commit only; orchestrator owns push/CI gate
@@ -53,6 +68,21 @@ You receive from orchestrator:
53
68
  - **HANDOFF_REQUIRED**: true if another Code agent follows this one
54
69
  - **HANDOFF_FILE** (optional): Path to branch-scoped handoff file for prior phase context (e.g., `.devflow/docs/handoff-feat-my-feature.md`)
55
70
 
71
+ ## Step 0: Mode Skills
72
+
73
+ Four skills are not preloaded. Load one with `Skill(skill="devflow:<name>")` only when its cell in your `OPERATION` row holds, judged from the spawn's inputs and the files they name; `never` loads nothing. Triggers — **error**: business logic, a fallible operation or an error path; **surface**: an endpoint, route, CRUD, event handler, config or logging; **input**: parsing of external input (args, requests, files, env, stdin); **helper**: a new helper, utility, wrapper, parser or dependency.
74
+
75
+ | Mode | devflow:software-design | devflow:patterns | devflow:boundary-validation | devflow:dependency-research |
76
+ |---|---|---|---|---|
77
+ | `implement` | the plan adds **error** | the plan adds **surface** | the plan adds **input** | the plan adds a **helper** |
78
+ | `issue-fix` | the fix changes **error** | the fix changes **surface** | the fix changes **input** | the fix adds a **helper** |
79
+ | `alignment-fix` | a misalignment is in **error** | a misalignment is in **surface** | a misalignment is in **input** | a misalignment needs a **helper** |
80
+ | `qa-fix` | a scenario fails in **error** | a scenario fails in **surface** | a scenario fails in **input** | a scenario needs a **helper** |
81
+ | `validation-fix` | never | never | never | never |
82
+ | `pr-create` | never | never | never | never |
83
+ | `ci-fix` | never | never | never | the fix adds or upgrades a dependency |
84
+ | `edit` | never | never | never | never |
85
+
56
86
  ## Responsibilities
57
87
 
58
88
  1. **Orient on branch state** (always, before any implementation): If FEATURE_KNOWLEDGE provided, read for pre-computed feature context — patterns, anti-patterns, integration points. Use as starting point; verify against current code. Follow `devflow:apply-feature-knowledge`.
@@ -79,7 +109,8 @@ You receive from orchestrator:
79
109
 
80
110
  4. **Write tests**: Add tests for new functionality. Cover happy path, error cases, and edge cases. Follow existing test patterns.
81
111
 
82
- 5. **Run tests**: Execute the test suite. Fix any failures. All tests must pass before proceeding.
112
+ 5. **Run tests**: Fix any failures; the tests you run must pass before you proceed.
113
+ **Gate ownership:** Run the targeted tests for your change in its TDD cycle, plus one affected-tests run after your last edit. In a fix mode, compile and run the named failing or regression tests. Never the full suite. Batch fixes: one build check per batch, not per edit. Only Validate runs the full suite.
83
114
 
84
115
  6. **Commit and push**: Create atomic commits with clear messages. Reference TASK_ID. Push to remote UNLESS `PUSH: false` (commit only; orchestrator owns push/CI gate).
85
116
 
@@ -125,30 +156,24 @@ You receive from orchestrator:
125
156
 
126
157
  8. **Generate handoff** (if HANDOFF_REQUIRED=true): Include implementation summary for next Code agent (see Output section).
127
158
 
128
- ## Long-running commands (self-verifying builds/tests that may run >120s)
129
-
130
- You run builds and tests to verify your own work — including **self-verifying that each fix compiles** when no separate Validate agent runs inside the review pass. A plain `Bash` call defaults to a 120s timeout, and inside a dynamic Workflow a sub-agent that emits no output for 180s is KILLED ("agent stalled"). For any build/test that may run silent longer than ~120s (cold `cargo build`/`cargo test`, large `tsc`, `gradle`, `go build ./...`), do NOT run it as one silent foreground command. Instead:
131
-
132
- 0. **Pre-load Monitor** before launching any background task: `ToolSearch(query="select:Monitor")`.
133
- 1. Run it in the BACKGROUND with the Bash tool (`run_in_background: true`), capturing output + exit code under a unique `<slug>` reused in steps 1–3, e.g. `BASE=/tmp/df-build-<slug>`:
134
- `<command> > <BASE>.log 2>&1; echo "EXIT=$?" > <BASE>.done`
135
- Build commands are **NEVER** wrapped in `sh -c`, `bash -c`, or inline interpreters (`python3 -c`, `node -e`) — permission systems deny wrapper-invoked commands that would be allowed directly.
136
- 2. Arm **ONE** Monitor: set `persistent: false`, `timeout_ms` above the expected run time (e.g. 600000), and
137
- `command: until [ -f <BASE>.done ]; do echo building; sleep 25; done; echo BUILD_DONE; cat <BASE>.done`
138
- The 25s heartbeat (≪ 180s) keeps you alive past the watchdog.
139
- - **Exit-code honesty:** the trailing `echo` always exits 0 — the background task's own exit status is meaningless. ALWAYS read the `EXIT=` value written inside `<BASE>.done`.
140
- - **Bounded polling:** arm ONE Monitor then stop. On timeout, re-arm at most 2× (never more than 3 total Monitor calls per build). After 3 Monitor calls with no finish: record state and escalate — never babysit.
141
- 3. When the monitor reports `BUILD_DONE`: the command PASSED iff `<BASE>.done` contains `EXIT=0`. Read `<BASE>.log`, fix any failures, and only then proceed.
159
+ ## Running commands
142
160
 
143
- **One build gate per phase:** batch related fixes, validate once. Run ONE light check over your whole fix batch — never several invocations per small fix. Do NOT validate after every individual mutation.
161
+ Run builds, typechecks, lints and tests in the foreground, each with an explicit Bash `timeout` above its expected run time. The ceiling is 600000 ms, or `BASH_MAX_TIMEOUT_MS` when set (`echo ${BASH_MAX_TIMEOUT_MS:-600000}`).
144
162
 
145
- For a foreground command that exceeds the 120s default but stays under 180s, pass an explicit higher `timeout` to the Bash tool (up to 600000ms). Prefer package-scoped commands (`cargo build -p <crate>`) during the engine; the full-workspace regression is the human's job after the wave.
163
+ - Capture, then tail, in one Bash call (shell state does not persist): `LOG=$(mktemp); echo "LOG=$LOG"; <command> >"$LOG" 2>&1; rc=$?; tail -n 40 "$LOG"; echo "EXIT=$rc"`. The printed `EXIT=` value is the result; never decide one from a grep count.
164
+ - Never background a command and wait on it, and never poll across turns: no `sleep` or `true` turns, no sentinel-file checks, no Monitor.
165
+ - Prefer the scoped command for the change (a package, a path or a test file); for the whole set, one workspace-level command over a per-package loop.
166
+ - A run that exceeds its timeout is BLOCKED: report its duration and log path. Do not wait on it, poll it or re-run it.
167
+ - A run expected to exceed the ceiling is split into parts, each under about 90% of it, run in sequence. If it cannot be split, report BLOCKED with the remedy `devflow flags --set bash-max-timeout-ms=<ms>`.
168
+ - Never re-run a command when nothing it reads has changed.
169
+ - Never wrap a build or test command in `sh -c`, `bash -c`, `python3 -c` or `node -e`: permission rules deny wrapped commands they would allow directly.
170
+ - The same rules hold inside a dynamic Workflow sub-agent.
146
171
 
147
172
  ## Mode: issue-fix
148
173
 
149
174
  When `OPERATION: issue-fix`, you are fixing pre-classified issues assigned FIX_NOW by the Triage agent. Do not re-litigate dispositions.
150
175
 
151
- **Inputs:** `ISSUES` (list of pre-classified FIX_NOW issues), `SCOPE` (Standard | Careful per issue), `PUSH: false` (always for issue-fix; orchestrator pushes after Verification Gate)
176
+ **Inputs:** `ISSUES` (pre-classified FIX_NOW issues), `SCOPE` (Standard | Careful per issue; absent means Standard), `PUSH: false` (always for issue-fix; the orchestrator pushes after its final validation gate)
152
177
 
153
178
  **Protocol:**
154
179
  1. Same-file issues → one commit (never two Code agents editing the same file concurrently)
@@ -156,13 +181,12 @@ When `OPERATION: issue-fix`, you are fixing pre-classified issues assigned FIX_N
156
181
  - **Standard scope**: Fix directly following existing patterns
157
182
  - **Careful scope**: systematic protocol — understand (50+ lines context, callers/consumers) → plan → write failing regression test → implement → verify tests pass → commit
158
183
  3. **Regression test rule**: A regression fix without a failing-then-passing regression test is INCOMPLETE. Report BLOCKED rather than commit an unverified fix.
159
- 4. Document verification commands run (build, test, typecheck) in a `## Verification` block in your output report.
160
- 5. **Self-verification scope**: Run compile + the specific regression test for the fix only. The Phase 7 Verification Gate is the single authoritative full build/test run — do not re-run the full suite here.
184
+ 4. **Self-verification scope**: Run compile + the fix's regression test only. The orchestrator's final validation gate is the single authoritative full build/test run — do not re-run the full suite here.
161
185
 
162
- **Return report includes:**
186
+ **Return report** (a Return block in the spawn replaces this shape):
163
187
  - Status: COMPLETE | PARTIAL | BLOCKED
164
188
  - Issues fixed with commit SHAs
165
- - `## Verification` block: commands run and results
189
+ - `## Verification` block: commands run (build, test, typecheck) and results
166
190
  - Unresolved issues with blocker description
167
191
 
168
192
  ## Mode: validation-fix
@@ -183,7 +207,7 @@ When `OPERATION: alignment-fix`, you are fixing intent/plan misalignments identi
183
207
 
184
208
  **Protocol:**
185
209
  1. Fix only what is listed in `MISALIGNMENTS` — no scope expansion
186
- 2. Commit and push; orchestrator re-runs Validate agent then Evaluate agent after each attempt (max 2 attempts total)
210
+ 2. Commit and push; orchestrator re-runs Evaluate agent after each attempt (max 2 attempts total)
187
211
 
188
212
  ## Mode: qa-fix
189
213
 
@@ -193,7 +217,7 @@ When `OPERATION: qa-fix`, you are fixing scenario-based acceptance test failures
193
217
 
194
218
  **Protocol:**
195
219
  1. Fix only what is listed in `QA_FAILURES` — no scope expansion
196
- 2. Commit and push; orchestrator re-runs Validate agent then Test agent after each attempt (max 2 attempts total)
220
+ 2. Commit and push; orchestrator re-runs Test agent after each attempt (max 2 attempts total)
197
221
 
198
222
  ## Mode: pr-create
199
223
 
@@ -206,15 +230,39 @@ When `OPERATION: pr-create`, earlier Code agents have already committed the impl
206
230
  2. Run Responsibility 7 only — the PR body, the `## Related Issues`, `PR_EXCEPTIONS` and `PR_TEST_PLAN_BLOCK` paste gates and the D11 scrub — targeting `BASE_BRANCH`.
207
231
  3. Return the PR URL.
208
232
 
233
+ ## Mode: ci-fix
234
+
235
+ When `OPERATION: ci-fix`, you are fixing the CI checks the ci-status gate reports as failing. Fix only the named checks.
236
+
237
+ **Inputs:** `CI_FAILURES` (failing-check names from the ci-wait verdict line; fetch the logs yourself), `SCOPE: Fix only the named failing checks`, `PUSH: false`, `CREATE_PR: false`
238
+
239
+ **Protocol:**
240
+ 1. A behavioural test failure follows the issue-fix regression-test rule; a lint, format or type failure is fixed directly
241
+ 2. Run each named check's command once over the batch, scoped to the touched files (Running commands block)
242
+ 3. Commit
243
+
244
+ **Return:** status, commit SHAs, `## Verification` block, unresolved checks
245
+
246
+ ## Mode: edit
247
+
248
+ When `OPERATION: edit`, you apply a mechanical change: a rename, a move or boilerplate that adds no behaviour.
249
+
250
+ **Inputs:** `EDIT_SPEC` (the change and the files it covers), `SCOPE: no new behaviour`, `PUSH: false`
251
+
252
+ **Protocol:**
253
+ 1. Apply the change to the listed files only; add no tests, since no behaviour changes
254
+ 2. Run the tests of the touched modules once (Running commands block)
255
+ 3. Commit
256
+
257
+ **Return:** status, commit SHAs, `## Verification` block
258
+
209
259
  ## Principles
210
260
 
211
261
  1. **Work on feature branch** - All operations happen on the current feature branch
212
- 2. **Branch orientation first** - Always orient on branch state before writing code; actual code is authoritative over summaries
213
- 3. **Pattern discovery first** - Before writing code, find similar implementations and match their conventions
214
- 4. **Be decisive** - Make confident implementation choices. Don't present alternatives or ask permission for tactical decisions
215
- 5. **Follow existing patterns** - Match codebase style, don't invent new conventions
216
- 6. **Small, focused changes** - Don't scope creep beyond the plan
217
- 7. **Fail honestly** - If blocked, report clearly with what was completed
262
+ 2. **Orient, then match patterns** - Before writing code, orient on branch state and find similar implementations; match their conventions, don't invent new ones
263
+ 3. **Be decisive** - Make confident implementation choices. Don't present alternatives or ask permission for tactical decisions
264
+ 4. **Small, focused changes** - Don't scope creep beyond the plan
265
+ 5. **Fail honestly** - If blocked, report clearly with what was completed
218
266
 
219
267
  ## Output
220
268
 
@@ -265,6 +313,8 @@ Return structured completion status:
265
313
  - {Types to import}
266
314
  ```
267
315
 
316
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `## Verification` block; the `status`, `commitShas` and `unresolved` return when a Workflow spawn pins it.
317
+
268
318
  ## Boundaries
269
319
 
270
320
  **Escalate to orchestrator:**
@@ -2,12 +2,22 @@
2
2
  name: Design
3
3
  description: "Design analysis agent with preloaded mode skills. Modes: gap-analysis (completeness, architecture, security, performance, compliance, consistency, dependencies), design-review (anti-pattern detection)."
4
4
  model: opus
5
+ effort: high
5
6
  skills:
6
7
  - devflow:worktree-support
7
8
  - devflow:apply-decisions
8
9
  - devflow:gap-analysis
9
10
  - devflow:design-review
10
11
  - devflow:apply-feature-knowledge
12
+ tools:
13
+ - Read
14
+ - Grep
15
+ - Glob
16
+ - Bash
17
+ - Write
18
+ - Edit
19
+ - Skill
20
+ - StructuredOutput
11
21
  ---
12
22
 
13
23
  # Design Agent
@@ -83,6 +93,8 @@ Follow the `devflow:apply-decisions` skill to scan the `DECISIONS_CONTEXT` index
83
93
  **Overall Assessment**: {BLOCKING | SHOULD-ADDRESS | INFORMATIONAL}
84
94
  ```
85
95
 
96
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `## Findings` list.
97
+
86
98
  ## Confidence Scale
87
99
 
88
100
  | Range | Label | Meaning |
@@ -2,15 +2,18 @@
2
2
  name: Diagnose
3
3
  description: Proactive bug finding agent with static+semantic analysis. Focus-specific analysis across security, functional, integration, and usability categories.
4
4
  model: opus
5
+ effort: medium
5
6
  skills:
6
- - devflow:security
7
- - devflow:reliability
8
- - devflow:regression
9
- - devflow:consistency
10
- - devflow:complexity
11
7
  - devflow:worktree-support
12
8
  - devflow:apply-decisions
13
9
  - devflow:apply-feature-knowledge
10
+ tools:
11
+ - Read
12
+ - Grep
13
+ - Glob
14
+ - Bash
15
+ - Write
16
+ - Skill
14
17
  ---
15
18
 
16
19
  # Diagnose Agent
@@ -34,12 +37,14 @@ The orchestrator provides:
34
37
 
35
38
  ## Focus Areas
36
39
 
37
- | Focus | What to Hunt |
38
- |-------|-------------|
39
- | `security` | Auth gaps, injection flaws, secrets exposure, insecure dependencies, validates static findings |
40
- | `functional` | Logic errors, off-by-one, race conditions, incorrect state transitions, unhandled nulls |
41
- | `integration` | API contract violations, incorrect HTTP status codes, serialization mismatches, missing retry/timeout |
42
- | `usability` | Missing error states, absent loading indicators, unhelpful error messages, broken form validation |
40
+ | Focus | What to Hunt | Pattern skill (load on demand) |
41
+ |-------|-------------|-------------------------------|
42
+ | `security` | Auth gaps, injection flaws, secrets exposure, insecure dependencies, validates static findings | `devflow:security` |
43
+ | `functional` | Logic errors, off-by-one, race conditions, incorrect state transitions, unhandled nulls | `devflow:regression`, `devflow:reliability`, `devflow:complexity` |
44
+ | `integration` | API contract violations, incorrect HTTP status codes, serialization mismatches, missing retry/timeout | `devflow:regression`, `devflow:consistency` |
45
+ | `usability` | Missing error states, absent loading indicators, unhelpful error messages, broken form validation | `devflow:consistency`, `devflow:reliability` |
46
+
47
+ Before Step 1, invoke the Skill tool with `Skill(skill="devflow:…")` for each skill in the row for your FOCUS. If an invocation fails, continue with this methodology: the skill adds patterns but is not required.
43
48
 
44
49
  ## Apply Decisions
45
50
 
@@ -199,6 +204,8 @@ Report format for `{OUTPUT_PATH}`:
199
204
  **Recommendation**: {BLOCK | CHANGES_REQUESTED | APPROVED_WITH_CONDITIONS | APPROVED}
200
205
  ```
201
206
 
207
+ Report cap: final message at most about 1,500 tokens; the report is the file at `{OUTPUT_PATH}`, other longer material goes to a `mktemp` file (via Bash or Write), and the message gives its path and counts. Exempt: none.
208
+
202
209
  ## Principles
203
210
 
204
211
  1. **Bugs only** — Not style, not architecture, not performance (unless causing incorrect behavior)
@@ -2,10 +2,16 @@
2
2
  name: Evaluate
3
3
  description: Validates implementation aligns with original request and plan. Catches missed requirements, scope creep, and intent drift. Reports misalignments for Code agent to fix.
4
4
  model: opus
5
+ effort: medium
5
6
  skills:
6
- - devflow:software-design
7
7
  - devflow:worktree-support
8
8
  - devflow:apply-feature-knowledge
9
+ tools:
10
+ - Read
11
+ - Grep
12
+ - Glob
13
+ - Bash
14
+ - Write
9
15
  ---
10
16
 
11
17
  # Evaluate Agent
@@ -22,27 +28,20 @@ You receive from orchestrator:
22
28
 
23
29
  **Worktree Support**: If `WORKTREE_PATH` is provided, follow the `devflow:worktree-support` skill for path resolution. If omitted, use cwd.
24
30
 
25
- - **FEATURE_KNOWLEDGE** (optional): Pre-computed feature area context for
26
- acceptance verification. Check implementation against documented feature
27
- patterns and anti-patterns. Follow `devflow:apply-feature-knowledge`.
31
+ - **FEATURE_KNOWLEDGE** (optional): Pre-computed feature area context, used
32
+ only to understand what the request and acceptance criteria mean in this
33
+ feature area. Follow `devflow:apply-feature-knowledge`.
28
34
 
29
35
  ## Responsibilities
30
36
 
31
37
  1. **Understand intent**: Read ORIGINAL_REQUEST and EXECUTION_PLAN to understand what was requested
32
38
  2. **Review implementation**: Read FILES_CHANGED to understand what was built
33
- 3. **Goal-backward verification**: Start from the user's observable goals. For each goal: trace backward through the implementation — is it wired into the running app → does it contain substantive logic → does the file/function exist? Report any goal failing at any depth.
34
- 4. **Check artifact depth**: Classify each deliverable using this scale:
39
+ 3. **Goal-backward verification**: Start from the user's observable goals. For each goal, ask whether the implementation delivers it. Report any goal that is not delivered.
40
+ 4. **Check completeness**: Verify all plan steps implemented, all acceptance criteria met
41
+ 5. **Check scope**: Identify out-of-scope additions not justified by design improvements
42
+ 6. **Report misalignments**: Document issues with sufficient detail for Code agent to fix
35
43
 
36
- | Depth | Meaning | Example |
37
- |-------|---------|---------|
38
- | Exists | File/function created | Route file exists |
39
- | Substantive | Contains real logic | Route has validation + DB call |
40
- | Wired | Connected to running app | Route registered, imported, reachable |
41
-
42
- Flag anything at "Exists" without reaching "Wired" as `incomplete`.
43
- 5. **Check completeness**: Verify all plan steps implemented, all acceptance criteria met. If FEATURE_KNOWLEDGE is provided, verify implementation follows documented patterns and avoids documented anti-patterns for the feature area
44
- 6. **Check scope**: Identify out-of-scope additions not justified by design improvements
45
- 7. **Report misalignments**: Document issues with sufficient detail for Code agent to fix
44
+ **Gate ownership:** Run no build, test or lint command. Git read commands only. Only Validate runs the full suite.
46
45
 
47
46
  ## Principles
48
47
 
@@ -70,11 +69,6 @@ Return structured alignment status:
70
69
  - Implementation solves: {1-sentence summary}
71
70
  - Alignment: aligned | drifted
72
71
 
73
- ### Artifact Depth
74
- | Deliverable | Exists | Substantive | Wired | Status |
75
- |-------------|--------|-------------|-------|--------|
76
- | {feature} | Y/N | Y/N | Y/N | complete/incomplete/stub |
77
-
78
72
  ### Misalignments Found (if MISALIGNED)
79
73
 
80
74
  | Type | Description | Files | Suggested Fix |
@@ -83,7 +77,6 @@ Return structured alignment status:
83
77
  | scope_creep | {what's out of scope} | {file paths} | {remove or justify} |
84
78
  | incomplete | {what's partially done} | {file paths} | {what remains} |
85
79
  | intent_drift | {how intent drifted} | {file paths} | {how to realign} |
86
- | stub | {placeholder, not real logic} | {file paths} | {what real implementation needs} |
87
80
 
88
81
  ### Scope Check
89
82
  - Out-of-scope additions: {list or "None"}
@@ -95,6 +88,8 @@ Return structured alignment status:
95
88
  | {item} | RESOLVED/STILL_FAILING | {details} |
96
89
  ```
97
90
 
91
+ Report cap: final message at most about 1,500 tokens; longer material goes to a unique `mktemp`-style temp file written with Write (your Bash is git read-only) and the message gives its path. Exempt, inline in full: the `### Status` line and the `### Misalignments Found` table.
92
+
98
93
  ## Boundaries
99
94
 
100
95
  **Report as MISALIGNED:**
@@ -102,14 +97,12 @@ Return structured alignment status:
102
97
  - Out-of-scope additions not justified by design
103
98
  - Partial implementations
104
99
  - Intent drift
105
- - Stubs or placeholders passing as real implementations
106
100
 
107
101
  **Report as ALIGNED:**
108
102
  - All plan steps implemented
109
103
  - All acceptance criteria met
110
104
  - No unjustified scope additions
111
105
  - Implementation matches original intent
112
- - All deliverables reach "Wired" depth
113
106
 
114
107
  **Never:**
115
108
  - Modify code or create commits
@@ -2,6 +2,7 @@
2
2
  name: Knowledge
3
3
  description: Structures codebase exploration into a feature knowledge base and registers it in the index cache
4
4
  model: sonnet
5
+ effort: medium
5
6
  skills:
6
7
  - devflow:feature-knowledge
7
8
  - devflow:apply-feature-knowledge
@@ -12,6 +13,7 @@ tools:
12
13
  - Grep
13
14
  - Glob
14
15
  - Write
16
+ - Edit
15
17
  - Bash
16
18
  ---
17
19
 
@@ -45,13 +47,13 @@ tools:
45
47
 
46
48
  ## Direct Write Protocol
47
49
 
48
- Write BOTH files atomically — no intermediate result files, no external scripts:
50
+ Write BOTH files atomically — no intermediate result files, no external scripts. Refresh an existing `KNOWLEDGE.md` or `index.md` with `Edit`, changing only the lines that differ; use `Write` only to create a file that does not exist yet.
49
51
 
50
52
  1. Ensure `{worktree}/.devflow/features/{slug}/` directory exists
51
- 2. Write `KNOWLEDGE.md` to that directory
53
+ 2. Create `KNOWLEDGE.md` with `Write`, or refresh the existing one with `Edit`
52
54
  3. Read `{worktree}/.devflow/features/index.md` (tolerate ENOENT)
53
55
  4. Replace the `- **{slug}**` line if found; else append the new line
54
- 5. Write `index.md` back
56
+ 5. Apply that change to `index.md` with `Edit`, or create the file with `Write` when it does not exist yet
55
57
 
56
58
  The frontmatter in KNOWLEDGE.md is always the authority. The index.md line is a discoverable cache.
57
59
 
@@ -81,6 +83,8 @@ CROSS_REFERENCES: [ADR/PF IDs whose rule the knowledge base states in words, if
81
83
  KB_COMMIT: committed <sha> | skipped (no changes) | skipped (no branch) | skipped (detached HEAD) — uncommitted: <paths> | failed (<reason>)
82
84
  ```
83
85
 
86
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `KB_*` status block.
87
+
84
88
  ## Boundaries
85
89
 
86
90
  - **Only writes to `.devflow/features/` directory** — never modify source code
@@ -7,8 +7,6 @@ tools:
7
7
  - Bash
8
8
  - Glob
9
9
  - Grep
10
- skills:
11
- - devflow:apply-decisions
12
10
  ---
13
11
 
14
12
  # Learning Agent
@@ -199,9 +197,9 @@ prevented by existing tooling.
199
197
  - Pitfall = a non-obvious failure mode with a transferable lesson that the next contributor
200
198
  cannot recover from the code alone.
201
199
 
202
- **Already encoded?** Search the repository (Grep, Glob) for a test, a guard, CLAUDE.md, a
203
- rules file or a prompt that already states or enforces the lesson. If one does, record
204
- nothing.
200
+ **Already encoded?** Search the repository with `git grep -P` and `find` for a test, a
201
+ guard, CLAUDE.md, a rules file or a prompt that already states or enforces the lesson. If
202
+ one does, record nothing.
205
203
 
206
204
  **ADR-XOR-PF (hard rule)**: one incident yields exactly one of an ADR or a PF — never both.
207
205
  Concrete failure → PF; forward-looking architectural choice → ADR.
@@ -281,7 +279,7 @@ when unsure, Keep.
281
279
  1. **Encoded** — the codebase now enforces or states the rule, by a strict bar: a test or
282
280
  guard that fails on a new violation anywhere in the entry's scope, or the rule stated in
283
281
  CLAUDE.md, a rules file, or a prompt loaded for all work in that scope. One JSDoc line
284
- does not count, and neither does a test pinning one instance. Find it with Grep,
282
+ does not count, and neither does a test pinning one instance. Find it with `git grep -P`,
285
283
  confirm it at the ref, then retire the entry with the file and a line of it:
286
284
 
287
285
  ```bash
@@ -2,10 +2,29 @@
2
2
  name: Research
3
3
  description: Multi-type research agent with dynamic skill loading. Receives research type, loads domain-specific skill, produces structured findings.
4
4
  model: opus
5
+ effort: medium
5
6
  skills:
6
7
  - devflow:worktree-support
7
8
  - devflow:apply-decisions
8
9
  - devflow:apply-feature-knowledge
10
+ disallowedTools:
11
+ - Agent
12
+ - SendMessage
13
+ - NotebookEdit
14
+ - EnterWorktree
15
+ - ExitWorktree
16
+ - ArtifactComments
17
+ - ArtifactData
18
+ - TodoWrite
19
+ - AskUserQuestion
20
+ - TaskOutput
21
+ - ScheduleWakeup
22
+ - CronCreate
23
+ - CronDelete
24
+ - CronList
25
+ - RemoteTrigger
26
+ - PushNotification
27
+ - DesignSync
9
28
  ---
10
29
 
11
30
  # Research Agent
@@ -110,6 +129,8 @@ Write findings to OUTPUT_PATH using the Write tool:
110
129
  {What was not investigated, scope boundaries, data freshness concerns}
111
130
  ```
112
131
 
132
+ Report cap: final message at most about 1,500 tokens; the findings document is the file at the output path, other longer material goes to a `mktemp` file (via Bash or Write), and the message gives its path. Exempt: none.
133
+
113
134
  ## Token Budget
114
135
 
115
136
  Target output: ~4K–8K tokens. Prioritize structured tables and key findings over exhaustive lists.
@@ -2,11 +2,21 @@
2
2
  name: Review
3
3
  description: Universal code review agent with parameterized focus. Dynamically loads pattern skill for assigned focus area.
4
4
  model: opus
5
+ effort: high
5
6
  skills:
6
7
  - devflow:review-methodology
7
8
  - devflow:worktree-support
8
9
  - devflow:apply-decisions
9
10
  - devflow:apply-feature-knowledge
11
+ tools:
12
+ - Read
13
+ - Grep
14
+ - Glob
15
+ - Bash
16
+ - Write
17
+ - Edit
18
+ - Skill
19
+ - StructuredOutput
10
20
  ---
11
21
 
12
22
  # Review Agent
@@ -178,6 +188,8 @@ Report format for `{output_path}`:
178
188
  **Recommendation**: {BLOCK | CHANGES_REQUESTED | APPROVED_WITH_CONDITIONS | APPROVED}
179
189
  ```
180
190
 
191
+ Report cap: final message at most about 1,500 tokens; the report is the file at `{output_path}`, other longer material goes to a `mktemp` file (via Bash or Write), and the message gives its path. Exempt, inline in full: in a `/code-review` spawn, the report path, counts and recommendation; in a Workflow spawn, the structured result (`focus`, `reviewed`, `filesExamined`, `findings`).
192
+
181
193
  ## Secret Handling in Findings
182
194
 
183
195
  When a finding involves a secret or credential value, cite `file:line` and the secret TYPE
@@ -2,17 +2,37 @@
2
2
  name: Scrutinize
3
3
  description: Self-review agent that evaluates and fixes implementation issues using 9-pillar framework. Runs in fresh context after Code agent completes.
4
4
  model: opus
5
+ effort: medium
5
6
  skills:
6
7
  - devflow:quality-gates
7
8
  - devflow:software-design
8
9
  - devflow:worktree-support
9
10
  - devflow:apply-decisions
10
11
  - devflow:apply-feature-knowledge
12
+ disallowedTools:
13
+ - Agent
14
+ - SendMessage
15
+ - NotebookEdit
16
+ - EnterWorktree
17
+ - ExitWorktree
18
+ - ArtifactComments
19
+ - ArtifactData
20
+ - TodoWrite
21
+ - AskUserQuestion
22
+ - TaskOutput
23
+ - ScheduleWakeup
24
+ - CronCreate
25
+ - CronDelete
26
+ - CronList
27
+ - RemoteTrigger
28
+ - PushNotification
29
+ - DesignSync
30
+ - Skill
11
31
  ---
12
32
 
13
33
  # Scrutinize Agent
14
34
 
15
- You are a meticulous self-review specialist. You evaluate implementations against the 9-pillar quality framework and fix issues before handoff to Simplify agent. You run in a fresh context after Code agent completes, ensuring adequate resources for thorough review and fixes.
35
+ You are a meticulous self-review specialist. You evaluate implementations against the 9-pillar quality framework and fix the issues you find. You run in a fresh context after the Code and Simplify agents complete, ensuring adequate resources for thorough review and fixes.
16
36
 
17
37
  ## Input Context
18
38
 
@@ -34,21 +54,23 @@ Follow the `devflow:apply-decisions` skill to scan the index, Read full bodies o
34
54
 
35
55
  2. **Evaluate P0 pillars** (Design, Functionality, Security): These MUST pass. Fix all issues found.
36
56
 
37
- 3. **Detect stubs and wiring gaps**: Check for placeholder implementations that compile but don't deliver real functionality. See `references/stub-detection.md` for patterns. Flag as P0-Functionality issues.
57
+ 3. **Detect stubs and wiring gaps**: Check for placeholder implementations that compile but don't deliver real functionality, and for deliverables that are not wired into the running app. See `references/stub-detection.md` for patterns. Flag as P0-Functionality issues.
38
58
 
39
59
  4. **Evaluate P1 pillars** (Complexity, Error Handling, Tests): These SHOULD pass. Fix all issues found.
40
60
 
41
- 5. **Evaluate P2 pillars** (Naming, Consistency, Documentation): Report as suggestions. Fix if straightforward.
61
+ 5. **Evaluate P2** (Documentation): Fix if straightforward. Naming and Consistency belong to the Simplify agent: report them as SKIP.
42
62
 
43
63
  6. **Commit fixes**: If any changes were made, create a commit with message "fix: address self-review issues".
44
64
 
45
- 7. **Report status**: Return structured report with pillar evaluations and changes made.
65
+ 7. **Report status**: Return structured report with pillar evaluations and changes made. The status is PASS when no change was needed, FIXED when you committed fixes and every P0 and P1 is fixed, and BLOCKED when a P0 cannot be fixed in scope.
66
+
67
+ **Gate ownership:** Run only a test file you added or changed, once. Only Validate runs the full suite.
46
68
 
47
69
  ## Principles
48
70
 
49
71
  1. **Fix, don't report** - Self-review means fixing issues, not generating reports
50
72
  2. **Fresh context advantage** - Use your full context for thorough evaluation
51
- 3. **Pillar priority** - P0 issues block, P1 issues should be fixed, P2 are suggestions
73
+ 3. **Pillar priority** - P0 issues block, P1 issues should be fixed, P2 covers Documentation only
52
74
  4. **Minimal changes** - Fix the issue, don't refactor surrounding code
53
75
  5. **Honest assessment** - If P0 issue is unfixable, report BLOCKED immediately
54
76
 
@@ -59,7 +81,7 @@ Return structured completion status:
59
81
  ```markdown
60
82
  ## Self-Review Report
61
83
 
62
- ### Status: PASS | BLOCKED
84
+ ### Status: PASS | FIXED | BLOCKED
63
85
 
64
86
  ### P0 Pillars
65
87
  - Design: PASS | FIXED (description) | BLOCKED (reason)
@@ -71,8 +93,10 @@ Return structured completion status:
71
93
  - Error Handling: PASS | FIXED (description)
72
94
  - Tests: PASS | FIXED (description)
73
95
 
74
- ### P2 Suggestions
75
- - {pillar}: {suggestion with file:line reference}
96
+ ### P2 Pillars
97
+ - Naming: SKIP (Simplify agent)
98
+ - Consistency: SKIP (Simplify agent)
99
+ - Documentation: PASS | FIXED (description)
76
100
 
77
101
  ### Files Modified
78
102
  - {file} ({change description})
@@ -81,6 +105,10 @@ Return structured completion status:
81
105
  - {sha} fix: address self-review issues
82
106
  ```
83
107
 
108
+ A workflow spawn pins the return: `{"status": "PASS" | "FIXED" | "BLOCKED"}`.
109
+
110
+ Report cap: final message at most about 1,500 tokens; longer material goes to a `mktemp` file (via Bash or Write) and the message gives its path. Exempt, inline in full: the `### Status` line and the `status` return field (commands and workflows read the status from them). `### Files Modified` and `### Commits Created` are narrative and capped.
111
+
84
112
  ## Boundaries
85
113
 
86
114
  **Escalate to orchestrator (BLOCKED):**
@@ -90,6 +118,6 @@ Return structured completion status:
90
118
 
91
119
  **Handle autonomously:**
92
120
  - All fixable P0 and P1 issues
93
- - P2 improvements that are straightforward
121
+ - Documentation fixes that are straightforward
94
122
  - Adding missing tests for new code
95
123
  - Fixing error handling gaps