aiblueprint-cli 1.4.98 → 1.4.100

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (98) hide show
  1. package/README.md +17 -1
  2. package/agents-config/skills/agents-manager/SKILL.md +2 -2
  3. package/agents-config/skills/agents-manager/agents/openai.yaml +7 -0
  4. package/agents-config/skills/agents-manager/assets/codex-icon.svg +20 -0
  5. package/agents-config/skills/apex/SKILL.md +120 -118
  6. package/agents-config/skills/apex/agents/openai.yaml +10 -0
  7. package/agents-config/skills/apex/assets/codex-icon.svg +15 -0
  8. package/agents-config/skills/apex/scripts/apex-state.py +740 -0
  9. package/agents-config/skills/apex/scripts/setup-templates.sh +27 -145
  10. package/agents-config/skills/apex/scripts/test_apex_state.py +413 -0
  11. package/agents-config/skills/apex/scripts/update-progress.sh +17 -73
  12. package/agents-config/skills/apex/steps/step-00-init.md +85 -231
  13. package/agents-config/skills/apex/steps/step-00b-branch.md +10 -118
  14. package/agents-config/skills/apex/steps/step-00b-economy.md +12 -239
  15. package/agents-config/skills/apex/steps/step-00b-interactive.md +13 -162
  16. package/agents-config/skills/apex/steps/step-00b-save.md +13 -114
  17. package/agents-config/skills/apex/steps/step-01-analyze.md +40 -361
  18. package/agents-config/skills/apex/steps/step-02-plan.md +55 -562
  19. package/agents-config/skills/apex/steps/step-02b-tasks.md +15 -291
  20. package/agents-config/skills/apex/steps/step-03-execute-teams.md +47 -267
  21. package/agents-config/skills/apex/steps/step-03-execute.md +32 -212
  22. package/agents-config/skills/apex/steps/step-04-validate.md +42 -246
  23. package/agents-config/skills/apex/steps/step-05-examine.md +47 -371
  24. package/agents-config/skills/apex/steps/step-06-resolve.md +19 -221
  25. package/agents-config/skills/apex/steps/step-07-tests.md +19 -234
  26. package/agents-config/skills/apex/steps/step-08-run-tests.md +13 -300
  27. package/agents-config/skills/apex/steps/step-09-finish.md +36 -200
  28. package/agents-config/skills/apex/steps/step-10-verify.md +46 -264
  29. package/agents-config/skills/appstore-connect/agents/openai.yaml +7 -0
  30. package/agents-config/skills/appstore-connect/assets/codex-icon.svg +17 -0
  31. package/agents-config/skills/commit/agents/openai.yaml +10 -0
  32. package/agents-config/skills/commit/assets/codex-icon.svg +17 -0
  33. package/agents-config/skills/create-pr/agents/openai.yaml +10 -0
  34. package/agents-config/skills/create-pr/assets/codex-icon.svg +17 -0
  35. package/agents-config/skills/environments-manager/SKILL.md +1 -1
  36. package/agents-config/skills/environments-manager/agents/openai.yaml +7 -0
  37. package/agents-config/skills/environments-manager/assets/codex-icon.svg +16 -0
  38. package/agents-config/skills/environments-manager/examples/scripts/claude-worktree-remove.sh +19 -3
  39. package/agents-config/skills/environments-manager/examples/scripts/worktree-up.sh +1 -1
  40. package/agents-config/skills/environments-manager/references/claude.md +1 -1
  41. package/agents-config/skills/fix-pr-comments/agents/openai.yaml +10 -0
  42. package/agents-config/skills/fix-pr-comments/assets/codex-icon.svg +17 -0
  43. package/agents-config/skills/grill-me/SKILL.md +25 -4
  44. package/agents-config/skills/grill-me/agents/openai.yaml +8 -0
  45. package/agents-config/skills/grill-me/assets/codex-icon.svg +16 -0
  46. package/agents-config/skills/hooks-manager/SKILL.md +19 -9
  47. package/agents-config/skills/hooks-manager/assets/codex-icon.svg +15 -4
  48. package/agents-config/skills/hooks-manager/references/claude-code.md +32 -0
  49. package/agents-config/skills/hooks-manager/references/codex.md +23 -0
  50. package/agents-config/skills/hooks-manager/references/cursor.md +18 -0
  51. package/agents-config/skills/hooks-manager/references/hook-types.md +5 -3
  52. package/agents-config/skills/hooks-manager/references/input-output-schemas.md +2 -2
  53. package/agents-config/skills/hooks-manager/references/research-sources.md +25 -0
  54. package/agents-config/skills/hooks-manager/references/router.md +32 -0
  55. package/agents-config/skills/hooks-manager/references/troubleshooting.md +3 -3
  56. package/agents-config/skills/merge/agents/openai.yaml +10 -0
  57. package/agents-config/skills/merge/assets/codex-icon.svg +17 -0
  58. package/agents-config/skills/oneshot/SKILL.md +4 -0
  59. package/agents-config/skills/oneshot/agents/openai.yaml +10 -0
  60. package/agents-config/skills/oneshot/assets/codex-icon.svg +18 -0
  61. package/agents-config/skills/prompt-creator/agents/openai.yaml +7 -0
  62. package/agents-config/skills/prompt-creator/assets/codex-icon.svg +16 -0
  63. package/agents-config/skills/rules-manager/agents/openai.yaml +7 -0
  64. package/agents-config/skills/rules-manager/assets/codex-icon.svg +23 -0
  65. package/agents-config/skills/skill-manager/SKILL.md +45 -3
  66. package/agents-config/skills/skill-manager/agents/openai.yaml +7 -0
  67. package/agents-config/skills/skill-manager/assets/codex-icon.svg +23 -0
  68. package/agents-config/skills/skill-manager/references/skill-writing-glossary.md +201 -0
  69. package/agents-config/skills/skill-manager/scripts/setup-codex-icons.ts +143 -0
  70. package/agents-config/skills/ultrathink/agents/openai.yaml +10 -0
  71. package/agents-config/skills/ultrathink/assets/codex-icon.svg +20 -0
  72. package/agents-config/skills/use-artifacts/SKILL.md +102 -51
  73. package/agents-config/skills/use-artifacts/assets/local-runtime.js +299 -0
  74. package/agents-config/skills/use-artifacts/scripts/create_artifact.py +1 -1
  75. package/agents-config/skills/use-delegate/SKILL.md +4 -0
  76. package/agents-config/skills/use-delegate/agents/openai.yaml +10 -0
  77. package/agents-config/skills/use-delegate/assets/codex-icon.svg +20 -0
  78. package/agents-config/skills/use-goal/SKILL.md +70 -9
  79. package/agents-config/skills/use-goal/agents/openai.yaml +1 -1
  80. package/agents-config/skills/use-goal/assets/codex-icon.svg +17 -3
  81. package/agents-config/skills/use-goal/references/claude-code-goal.md +54 -6
  82. package/agents-config/skills/use-goal/references/codex-goal.md +59 -4
  83. package/agents-config/skills/use-goal/references/verification-harnesses.md +104 -3
  84. package/dist/cli.js +264 -12
  85. package/package.json +1 -1
  86. package/agents-config/skills/apex/templates/00-context.md +0 -55
  87. package/agents-config/skills/apex/templates/01-analyze.md +0 -10
  88. package/agents-config/skills/apex/templates/02-plan.md +0 -10
  89. package/agents-config/skills/apex/templates/03-execute.md +0 -10
  90. package/agents-config/skills/apex/templates/04-validate.md +0 -10
  91. package/agents-config/skills/apex/templates/05-examine.md +0 -10
  92. package/agents-config/skills/apex/templates/06-resolve.md +0 -10
  93. package/agents-config/skills/apex/templates/07-tests.md +0 -10
  94. package/agents-config/skills/apex/templates/08-run-tests.md +0 -10
  95. package/agents-config/skills/apex/templates/09-finish.md +0 -10
  96. package/agents-config/skills/apex/templates/10-verify.md +0 -9
  97. package/agents-config/skills/apex/templates/README.md +0 -195
  98. package/agents-config/skills/apex/templates/step-complete.md +0 -7
@@ -1,400 +1,76 @@
1
1
  ---
2
2
  name: step-05-examine
3
- description: Adversarial code review - security, logic, and quality analysis
4
- prev_step: steps/step-04-validate.md
5
- next_step: steps/step-06-resolve.md
3
+ description: Select independent APEX reviewers by change risk and domain, then validate and deduplicate their findings.
6
4
  ---
7
5
 
8
- # Step 5: Examine (Adversarial Review)
6
+ # Step 5: eXamine
9
7
 
10
- ## MANDATORY EXECUTION RULES (READ FIRST):
8
+ Independent review is mandatory for material, high-risk, or explicitly adversarial work. Review depth follows the diff, not a fixed agent count.
11
9
 
12
- - 🛑 NEVER skip security review
13
- - 🛑 NEVER dismiss findings without justification
14
- - 🛑 NEVER auto-approve without thorough review
15
- - ✅ ALWAYS check OWASP top 10 vulnerabilities
16
- - ✅ ALWAYS classify findings by severity and validity
17
- - ✅ ALWAYS present findings table to user
18
- - 📋 YOU ARE A SKEPTICAL REVIEWER, not a defender
19
- - 💬 FOCUS on "What could go wrong?"
20
- - 🚫 FORBIDDEN to approve without thorough analysis
10
+ ## 1. Build the review packet
21
11
 
22
- ## EXECUTION PROTOCOLS:
12
+ Capture:
23
13
 
24
- - 🎯 Launch 4+ parallel review sub-agents in ONE message (unless economy_mode)
25
- - 🛑 NEVER launch only 1 review agent - you MUST launch Security + Logic + Clean Code + Thermo-Nuclear as separate agents
26
- - 🛑 NEVER skip the Thermo-Nuclear quality audit - it is the final maintainability gate
27
- - 💾 Document all findings with severity
28
- - 📖 Create todos for each finding
29
- - 🚫 FORBIDDEN to skip security analysis
30
- - 🚫 FORBIDDEN to skip thermo-nuclear maintainability review
31
- - 🚫 FORBIDDEN to combine review categories into a single agent
14
+ - original task and acceptance criteria;
15
+ - intended paths and actual diff;
16
+ - relevant architecture and project rules;
17
+ - validation ledger and known baseline failures;
18
+ - unresolved risks, assumptions, and proof requirements.
32
19
 
33
- ## CONTEXT BOUNDARIES:
20
+ Review the actual uncommitted task diff, not automatically `HEAD~1`.
34
21
 
35
- - Implementation is complete and validated
36
- - All tests pass
37
- - Now looking for issues that tests miss
38
- - Adversarial mindset - assume bugs exist
39
- - **If `{teams_mode}` = true:** Agent team is still alive. Do NOT shutdown teammates - that happens in step-09-finish only.
22
+ ## 2. Select review lenses
40
23
 
41
- ## YOUR TASK:
24
+ Use only lenses relevant to the change:
42
25
 
43
- Conduct an adversarial code review to identify security vulnerabilities, logic flaws, and quality issues.
26
+ | Lens | Trigger examples |
27
+ |---|---|
28
+ | Correctness and edge cases | State transitions, concurrency, parsing, error handling |
29
+ | Security and authority | Auth, tenant boundaries, secrets, input boundaries, external actions |
30
+ | Data and migration safety | Schema changes, backfills, idempotency, rollback |
31
+ | Domain specialist | Payments, email, mobile, framework, provider, performance |
32
+ | Maintainability | Cross-cutting changes, new abstractions, large or structurally risky diffs |
33
+ | Evidence and acceptance | Runtime/provider/public claims or complex proof matrix |
44
34
 
45
- ---
46
-
47
- <available_state>
48
- From previous steps:
49
-
50
- | Variable | Description |
51
- |----------|-------------|
52
- | `{task_description}` | What was implemented |
53
- | `{task_id}` | Kebab-case identifier |
54
- | `{auto_mode}` | Auto-fix Real findings |
55
- | `{save_mode}` | Save outputs to files |
56
- | `{economy_mode}` | No subagents, direct review |
57
- | `{output_dir}` | Path to output (if save_mode) |
58
- | Files modified | From step-03 |
59
- </available_state>
60
-
61
- ---
62
-
63
- ## EXECUTION SEQUENCE:
64
-
65
- ### 1. Initialize Save Output (if save_mode)
35
+ Use a fresh independent context for each genuinely distinct lens. Combine closely related lenses when separation would only duplicate context. For low-risk changes, one focused independent reviewer may be enough. For high-risk changes, use multiple non-overlapping specialists.
66
36
 
67
- **If `{save_mode}` = true:**
37
+ Reviewers are read-only unless explicitly assigned a later resolution task.
68
38
 
69
- ```bash
70
- bash {skill_dir}/scripts/update-progress.sh "{task_id}" "05" "examine" "in_progress"
71
- ```
72
-
73
- Append findings to `{output_dir}/05-examine.md` as you work.
74
-
75
- ### 2. Gather Changes
76
-
77
- ```bash
78
- git diff --name-only HEAD~1
79
- git status --porcelain
80
- ```
39
+ ## 3. Require high-signal findings
81
40
 
82
- Group files: source, tests, config, other.
41
+ Every finding must contain:
83
42
 
84
- ### 3. Conduct Review
43
+ - stable ID, severity, and confidence;
44
+ - exact file and line or artifact reference;
45
+ - concrete failure scenario or violated contract;
46
+ - evidence that the issue is introduced or exposed by the intended diff;
47
+ - smallest safe remediation direction.
85
48
 
86
- **If `{economy_mode}` = true:**
87
- → Self-review with checklist:
49
+ Reject style preference, speculative breakage without a path, duplicated findings, and issues wholly outside scope.
88
50
 
89
- ```markdown
90
- ## Security Checklist
91
- - [ ] No SQL injection (parameterized queries)
92
- - [ ] No XSS (output encoding)
93
- - [ ] No secrets in code
94
- - [ ] Input validation present
95
- - [ ] Auth checks on protected routes
51
+ ## 4. Validate findings
96
52
 
97
- ## Logic Checklist
98
- - [ ] Error handling for all failure modes
99
- - [ ] Edge cases handled
100
- - [ ] Null/undefined checks
101
- - [ ] Race conditions considered
53
+ The coordinator independently inspects each reported issue and classifies it:
102
54
 
103
- ## Quality Checklist
104
- - [ ] Follows existing patterns
105
- - [ ] No code duplication
106
- - [ ] Clear naming
55
+ - `CONFIRMED`;
56
+ - `NOISE`;
57
+ - `PREEXISTING`;
58
+ - `OUT_OF_SCOPE`;
59
+ - `UNCERTAIN` with the exact missing evidence.
107
60
 
108
- ## Thermo-Nuclear Maintainability Checklist
109
- - [ ] No file pushed from under 1k lines to over 1k lines without strong justification
110
- - [ ] No new ad-hoc conditionals / spaghetti branches in unrelated flows
111
- - [ ] No thin abstractions, identity wrappers, or "magic" mechanisms added
112
- - [ ] No unnecessary casts / `any` / `unknown` / optional params muddying contracts
113
- - [ ] Feature logic stays in the canonical layer (no leaking into shared paths)
114
- - [ ] Reuses existing canonical helpers instead of bespoke near-duplicates
115
- - [ ] No "code judo" simplification was missed (could this be dramatically simpler?)
116
- - [ ] Orchestration is parallel / atomic where the cleaner structure is obvious
117
- ```
61
+ Only confirmed findings block completion automatically. High-severity uncertain findings require targeted investigation before disposition.
118
62
 
119
- **If `{economy_mode}` = false:**
120
- → Launch parallel review sub-agents.
63
+ ## 5. Record review ledger
121
64
 
122
- **🛑 CRITICAL: You MUST launch ALL 4 agents (or 5 if Next.js) in a SINGLE message using MULTIPLE sub-agent launches. DO NOT launch them one at a time. DO NOT use only 1 agent. Each agent reviews a DIFFERENT aspect. The Thermo-Nuclear agent is MANDATORY and runs alongside the others.**
65
+ Store reviewer lens, evidence, classification, disposition, and any invalidated validation or proof artifacts.
123
66
 
124
- First, gather the list of modified files:
125
67
  ```bash
126
- git diff --name-only HEAD~1
127
- ```
128
-
129
- Then, in **ONE message with 4+ parallel sub-agent launches**, launch:
130
-
131
- ---
132
-
133
- **Agent 1: Security Review** - sub-agent profile/type: `code-reviewer`
134
- ```
135
- prompt: |
136
- You are a SECURITY reviewer. Review ONLY the following files for security vulnerabilities:
137
- {list of modified files}
138
-
139
- Focus exclusively on:
140
- - OWASP Top 10: injection flaws (SQL, command, XSS)
141
- - Authentication and authorization issues
142
- - Sensitive data exposure (secrets, tokens, PII in logs)
143
- - Security misconfiguration
144
- - Insecure deserialization
145
- - Missing input validation at system boundaries
146
-
147
- For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
148
- If no security issues found, explicitly state "No security issues found."
149
- ```
150
-
151
- ---
152
-
153
- **Agent 2: Logic & Edge Cases Review** - sub-agent profile/type: `code-reviewer`
154
- ```
155
- prompt: |
156
- You are a LOGIC reviewer. Review ONLY the following files for logic correctness:
157
- {list of modified files}
158
-
159
- Focus exclusively on:
160
- - Edge cases not handled (empty arrays, null/undefined, boundary values)
161
- - Race conditions and concurrency issues
162
- - Incorrect conditional logic or off-by-one errors
163
- - Missing error handling for failure modes
164
- - State management bugs
165
- - Incorrect assumptions about data shape or types
166
-
167
- For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
168
- If no logic issues found, explicitly state "No logic issues found."
169
- ```
170
-
171
- ---
172
-
173
- **Agent 3: Clean Code & Quality Review** - sub-agent profile/type: `code-reviewer`
174
- ```
175
- prompt: |
176
- You are a CLEAN CODE reviewer. Review ONLY the following files for code quality:
177
- {list of modified files}
178
-
179
- Focus exclusively on:
180
- - SOLID principle violations
181
- - Code smells (long methods, god objects, feature envy)
182
- - Cyclomatic complexity > 10
183
- - Code duplication > 20 lines
184
- - Naming that doesn't communicate intent
185
- - Functions doing too many things
186
-
187
- For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
188
- If no quality issues found, explicitly state "No quality issues found."
189
- ```
190
-
191
- ---
192
-
193
- **Agent 4: Thermo-Nuclear Code Quality Review** (MANDATORY - launch alongside Agents 1-3)
194
-
195
- This is the **final verification gate**. After the other reviewers find issues, this agent performs an extremely strict maintainability audit using the `thermo-nuclear-code-quality-review` skill. It is **not optional** and must run on every examine pass.
196
-
197
- Sub-agent profile/type: `thermo-nuclear-code-quality-review`
198
- ```
199
- prompt: |
200
- Perform a Thermo-Nuclear Code Quality Review on the current branch's changes.
201
-
202
- Modified files to audit:
203
- {list of modified files}
204
-
205
- Load and strictly apply the rubric from the `thermo-nuclear-code-quality-review` skill.
206
-
207
- Verify and challenge:
208
- - Structural code-quality regressions and missed "code judo" opportunities
209
- - Any file pushed from under 1k lines to over 1k lines (presumptive blocker)
210
- - New ad-hoc conditionals or spaghetti branching bolted onto unrelated flows
211
- - Thin abstractions, identity wrappers, pass-through helpers, "magic" mechanisms
212
- - Unnecessary casts, `any`, `unknown`, optional params obscuring real contracts
213
- - Feature logic leaking into shared/canonical paths
214
- - Bespoke helpers duplicating existing canonical utilities
215
- - Unnecessary sequential orchestration or non-atomic update flows
216
- - Whether the implementation could be dramatically simpler / smaller / more direct
217
-
218
- For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and a concrete restructuring suggestion (prefer DELETING complexity over rearranging it).
219
-
220
- Be ambitious, direct, and demanding. Do not soften major maintainability issues.
221
- Apply the skill's Approval Bar strictly - flag presumptive blockers explicitly.
222
-
223
- If no significant maintainability issues found, explicitly state "No thermo-nuclear findings."
224
- ```
225
-
226
- ---
227
-
228
- **Agent 5: Vercel/Next.js Best Practices** (CONDITIONAL - launch alongside Agents 1-4)
229
-
230
- → **Detection:** Check if modified files match Next.js/Vercel patterns:
231
- ```
232
- - *.tsx, *.jsx files in app/, pages/, components/
233
- - next.config.* files
234
- - Server actions (use server)
235
- - API routes (app/api/*, pages/api/*)
236
- - Middleware (middleware.ts)
237
- - Server components, client components
238
- ```
239
-
240
- → **If Next.js/Vercel code detected:** Add a 5th parallel review sub-agent:
241
-
242
- Sub-agent profile/type: `code-reviewer`
243
- ```
244
- prompt: |
245
- You are a NEXT.JS / REACT PERFORMANCE reviewer. Review ONLY the following files:
246
- {list of modified files}
247
-
248
- Focus exclusively on:
249
- - Sequential awaits that should use Promise.all for parallel fetching
250
- - Barrel imports causing bundle bloat (import from index files)
251
- - Missing dynamic imports for heavy client components
252
- - Server-side caching opportunities (React cache, unstable_cache)
253
- - Unnecessary re-renders (missing memo, useMemo, useCallback)
254
- - Wrong Server vs Client component boundaries
255
- - Data fetching patterns (preloading, parallel fetching, waterfall detection)
256
-
257
- For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
258
- If no performance issues found, explicitly state "No performance issues found."
259
- ```
260
-
261
- → **If NOT Next.js/Vercel code:** Skip this agent (launch only Agents 1-4)
262
-
263
- ---
264
-
265
- **🛑 REMINDER: You MUST have 4+ sub-agent launches in a SINGLE response (Security + Logic + Clean Code + Thermo-Nuclear, plus Next.js if applicable). If you only launched 1-3 agents, you are doing it WRONG. Go back and launch all agents. The Thermo-Nuclear agent is MANDATORY - never skip it.**
266
-
267
- ### 4. Classify Findings
268
-
269
- For each finding:
270
-
271
- **Severity:**
272
- - CRITICAL: Security vulnerability, data loss risk
273
- - HIGH: Significant bug, will cause issues
274
- - MEDIUM: Should fix, not urgent
275
- - LOW: Minor improvement
276
-
277
- **Validity:**
278
- - Real: Definitely needs fixing
279
- - Noise: Not actually a problem
280
- - Uncertain: Needs discussion
281
-
282
- ### 5. Present Findings Table
283
-
284
- ```markdown
285
- ## Findings
286
-
287
- | ID | Severity | Category | Location | Issue | Validity |
288
- |----|----------|----------|----------|-------|----------|
289
- | F1 | CRITICAL | Security | auth.ts:42 | SQL injection | Real |
290
- | F2 | HIGH | Logic | handler.ts:78 | Missing null check | Real |
291
- | F3 | MEDIUM | Quality | utils.ts:15 | Complex function | Uncertain |
292
-
293
- **Summary:** {count} findings ({blocking} blocking)
294
- ```
295
-
296
- ### 6. Create Finding Todos
297
-
68
+ python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase examine --status complete --message "Independent findings validated and deduplicated"
298
69
  ```
299
- - [ ] F1 [CRITICAL] Fix SQL injection in auth.ts:42
300
- - [ ] F2 [HIGH] Add null check in handler.ts:78
301
- ```
302
-
303
- ### 7. Get User Approval (review → resolve/test)
304
-
305
- **If `{auto_mode}` = true:**
306
- → Proceed automatically based on findings
307
-
308
- **If `{auto_mode}` = false:**
309
-
310
- ```yaml
311
- questions:
312
- - header: "Review"
313
- question: "Review complete. How would you like to proceed?"
314
- options:
315
- - label: "Resolve findings (Recommended)"
316
- description: "Address the identified issues"
317
- - label: "Skip to tests"
318
- description: "Skip resolution, proceed to test creation"
319
- - label: "Skip resolution"
320
- description: "Accept findings, don't make changes"
321
- - label: "Discuss findings"
322
- description: "I want to discuss specific findings"
323
- multiSelect: false
324
- ```
325
-
326
- <critical>
327
- This is one of the THREE transition points that requires user confirmation:
328
- 1. plan → execute
329
- 2. validate → review
330
- 3. review → resolve/test (THIS ONE)
331
- </critical>
332
-
333
- ### 8. Complete Save Output (if save_mode)
334
-
335
- **If `{save_mode}` = true:**
336
-
337
- Append to `{output_dir}/05-examine.md`:
338
- ```markdown
339
- ---
340
- ## Step Complete
341
- **Status:** ✓ Complete
342
- **Findings:** {count}
343
- **Critical:** {count}
344
- **Next:** step-06-resolve.md
345
- **Timestamp:** {ISO timestamp}
346
- ```
347
-
348
- ---
349
-
350
- ## SUCCESS METRICS:
351
-
352
- ✅ All modified files reviewed
353
- ✅ Security checklist completed
354
- ✅ Findings classified by severity
355
- ✅ Validity assessed for each finding
356
- ✅ Findings table presented
357
- ✅ Todos created for tracking
358
- ✅ Thermo-Nuclear maintainability audit completed (mandatory final gate)
359
- ✅ Next.js/Vercel best practices checked (if applicable)
360
-
361
- ## FAILURE MODES:
362
-
363
- ❌ **CRITICAL**: Launching only 1 review agent instead of 4+ - each category (Security, Logic, Clean Code, Thermo-Nuclear) MUST be a separate agent
364
- ❌ **CRITICAL**: Skipping the Thermo-Nuclear maintainability review - it is the mandatory final verification gate
365
- ❌ Combining multiple review categories into a single agent prompt
366
- ❌ Skipping security review
367
- ❌ Not classifying by severity
368
- ❌ Auto-dismissing findings
369
- ❌ Launching agents sequentially instead of in parallel
370
- ❌ Using subagents when economy_mode
371
- ❌ Skipping Vercel/Next.js review when React/Next.js files are modified
372
- ❌ **CRITICAL**: Not using AskUserQuestion for review → resolve/test transition
373
-
374
- ## REVIEW PROTOCOLS:
375
-
376
- - Adversarial mindset - assume bugs exist
377
- - Check security FIRST
378
- - Every finding gets severity and validity
379
- - Don't dismiss without justification
380
- - Present clear summary
381
-
382
- ---
383
-
384
- ## NEXT STEP:
385
-
386
- After user confirms via AskUserQuestion (or auto-proceed):
387
-
388
- **If user chooses "Resolve findings":** → Load `./step-06-resolve.md`
389
-
390
- **If user chooses "Skip to tests" (and test_mode):** → Load `./step-07-tests.md`
391
70
 
392
- **If user chooses "Skip resolution":**
393
- - **If test_mode:** → Load `./step-07-tests.md`
394
- - **If pr_mode:** → Load `./step-09-finish.md` to create pull request
395
- - **Otherwise:** → Workflow complete - show summary
71
+ ## Routing
396
72
 
397
- <critical>
398
- Remember: Be SKEPTICAL - your job is to find problems, not approve code!
399
- This step MUST ask before proceeding (unless auto_mode).
400
- </critical>
73
+ - If confirmed findings exist, load `step-06-resolve.md`.
74
+ - If test coverage must change, load `step-07-tests.md`.
75
+ - If runtime proof is required, load `step-10-verify.md`.
76
+ - Otherwise load `step-09-finish.md`.