aiblueprint-cli 1.4.98 → 1.4.100
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -1
- package/agents-config/skills/agents-manager/SKILL.md +2 -2
- package/agents-config/skills/agents-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/agents-manager/assets/codex-icon.svg +20 -0
- package/agents-config/skills/apex/SKILL.md +120 -118
- package/agents-config/skills/apex/agents/openai.yaml +10 -0
- package/agents-config/skills/apex/assets/codex-icon.svg +15 -0
- package/agents-config/skills/apex/scripts/apex-state.py +740 -0
- package/agents-config/skills/apex/scripts/setup-templates.sh +27 -145
- package/agents-config/skills/apex/scripts/test_apex_state.py +413 -0
- package/agents-config/skills/apex/scripts/update-progress.sh +17 -73
- package/agents-config/skills/apex/steps/step-00-init.md +85 -231
- package/agents-config/skills/apex/steps/step-00b-branch.md +10 -118
- package/agents-config/skills/apex/steps/step-00b-economy.md +12 -239
- package/agents-config/skills/apex/steps/step-00b-interactive.md +13 -162
- package/agents-config/skills/apex/steps/step-00b-save.md +13 -114
- package/agents-config/skills/apex/steps/step-01-analyze.md +40 -361
- package/agents-config/skills/apex/steps/step-02-plan.md +55 -562
- package/agents-config/skills/apex/steps/step-02b-tasks.md +15 -291
- package/agents-config/skills/apex/steps/step-03-execute-teams.md +47 -267
- package/agents-config/skills/apex/steps/step-03-execute.md +32 -212
- package/agents-config/skills/apex/steps/step-04-validate.md +42 -246
- package/agents-config/skills/apex/steps/step-05-examine.md +47 -371
- package/agents-config/skills/apex/steps/step-06-resolve.md +19 -221
- package/agents-config/skills/apex/steps/step-07-tests.md +19 -234
- package/agents-config/skills/apex/steps/step-08-run-tests.md +13 -300
- package/agents-config/skills/apex/steps/step-09-finish.md +36 -200
- package/agents-config/skills/apex/steps/step-10-verify.md +46 -264
- package/agents-config/skills/appstore-connect/agents/openai.yaml +7 -0
- package/agents-config/skills/appstore-connect/assets/codex-icon.svg +17 -0
- package/agents-config/skills/commit/agents/openai.yaml +10 -0
- package/agents-config/skills/commit/assets/codex-icon.svg +17 -0
- package/agents-config/skills/create-pr/agents/openai.yaml +10 -0
- package/agents-config/skills/create-pr/assets/codex-icon.svg +17 -0
- package/agents-config/skills/environments-manager/SKILL.md +1 -1
- package/agents-config/skills/environments-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/environments-manager/assets/codex-icon.svg +16 -0
- package/agents-config/skills/environments-manager/examples/scripts/claude-worktree-remove.sh +19 -3
- package/agents-config/skills/environments-manager/examples/scripts/worktree-up.sh +1 -1
- package/agents-config/skills/environments-manager/references/claude.md +1 -1
- package/agents-config/skills/fix-pr-comments/agents/openai.yaml +10 -0
- package/agents-config/skills/fix-pr-comments/assets/codex-icon.svg +17 -0
- package/agents-config/skills/grill-me/SKILL.md +25 -4
- package/agents-config/skills/grill-me/agents/openai.yaml +8 -0
- package/agents-config/skills/grill-me/assets/codex-icon.svg +16 -0
- package/agents-config/skills/hooks-manager/SKILL.md +19 -9
- package/agents-config/skills/hooks-manager/assets/codex-icon.svg +15 -4
- package/agents-config/skills/hooks-manager/references/claude-code.md +32 -0
- package/agents-config/skills/hooks-manager/references/codex.md +23 -0
- package/agents-config/skills/hooks-manager/references/cursor.md +18 -0
- package/agents-config/skills/hooks-manager/references/hook-types.md +5 -3
- package/agents-config/skills/hooks-manager/references/input-output-schemas.md +2 -2
- package/agents-config/skills/hooks-manager/references/research-sources.md +25 -0
- package/agents-config/skills/hooks-manager/references/router.md +32 -0
- package/agents-config/skills/hooks-manager/references/troubleshooting.md +3 -3
- package/agents-config/skills/merge/agents/openai.yaml +10 -0
- package/agents-config/skills/merge/assets/codex-icon.svg +17 -0
- package/agents-config/skills/oneshot/SKILL.md +4 -0
- package/agents-config/skills/oneshot/agents/openai.yaml +10 -0
- package/agents-config/skills/oneshot/assets/codex-icon.svg +18 -0
- package/agents-config/skills/prompt-creator/agents/openai.yaml +7 -0
- package/agents-config/skills/prompt-creator/assets/codex-icon.svg +16 -0
- package/agents-config/skills/rules-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/rules-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/SKILL.md +45 -3
- package/agents-config/skills/skill-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/skill-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/references/skill-writing-glossary.md +201 -0
- package/agents-config/skills/skill-manager/scripts/setup-codex-icons.ts +143 -0
- package/agents-config/skills/ultrathink/agents/openai.yaml +10 -0
- package/agents-config/skills/ultrathink/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-artifacts/SKILL.md +102 -51
- package/agents-config/skills/use-artifacts/assets/local-runtime.js +299 -0
- package/agents-config/skills/use-artifacts/scripts/create_artifact.py +1 -1
- package/agents-config/skills/use-delegate/SKILL.md +4 -0
- package/agents-config/skills/use-delegate/agents/openai.yaml +10 -0
- package/agents-config/skills/use-delegate/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-goal/SKILL.md +70 -9
- package/agents-config/skills/use-goal/agents/openai.yaml +1 -1
- package/agents-config/skills/use-goal/assets/codex-icon.svg +17 -3
- package/agents-config/skills/use-goal/references/claude-code-goal.md +54 -6
- package/agents-config/skills/use-goal/references/codex-goal.md +59 -4
- package/agents-config/skills/use-goal/references/verification-harnesses.md +104 -3
- package/dist/cli.js +264 -12
- package/package.json +1 -1
- package/agents-config/skills/apex/templates/00-context.md +0 -55
- package/agents-config/skills/apex/templates/01-analyze.md +0 -10
- package/agents-config/skills/apex/templates/02-plan.md +0 -10
- package/agents-config/skills/apex/templates/03-execute.md +0 -10
- package/agents-config/skills/apex/templates/04-validate.md +0 -10
- package/agents-config/skills/apex/templates/05-examine.md +0 -10
- package/agents-config/skills/apex/templates/06-resolve.md +0 -10
- package/agents-config/skills/apex/templates/07-tests.md +0 -10
- package/agents-config/skills/apex/templates/08-run-tests.md +0 -10
- package/agents-config/skills/apex/templates/09-finish.md +0 -10
- package/agents-config/skills/apex/templates/10-verify.md +0 -9
- package/agents-config/skills/apex/templates/README.md +0 -195
- package/agents-config/skills/apex/templates/step-complete.md +0 -7
|
@@ -1,400 +1,76 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: step-05-examine
|
|
3
|
-
description:
|
|
4
|
-
prev_step: steps/step-04-validate.md
|
|
5
|
-
next_step: steps/step-06-resolve.md
|
|
3
|
+
description: Select independent APEX reviewers by change risk and domain, then validate and deduplicate their findings.
|
|
6
4
|
---
|
|
7
5
|
|
|
8
|
-
# Step 5:
|
|
6
|
+
# Step 5: eXamine
|
|
9
7
|
|
|
10
|
-
|
|
8
|
+
Independent review is mandatory for material, high-risk, or explicitly adversarial work. Review depth follows the diff, not a fixed agent count.
|
|
11
9
|
|
|
12
|
-
|
|
13
|
-
- 🛑 NEVER dismiss findings without justification
|
|
14
|
-
- 🛑 NEVER auto-approve without thorough review
|
|
15
|
-
- ✅ ALWAYS check OWASP top 10 vulnerabilities
|
|
16
|
-
- ✅ ALWAYS classify findings by severity and validity
|
|
17
|
-
- ✅ ALWAYS present findings table to user
|
|
18
|
-
- 📋 YOU ARE A SKEPTICAL REVIEWER, not a defender
|
|
19
|
-
- 💬 FOCUS on "What could go wrong?"
|
|
20
|
-
- 🚫 FORBIDDEN to approve without thorough analysis
|
|
10
|
+
## 1. Build the review packet
|
|
21
11
|
|
|
22
|
-
|
|
12
|
+
Capture:
|
|
23
13
|
|
|
24
|
-
-
|
|
25
|
-
-
|
|
26
|
-
-
|
|
27
|
-
-
|
|
28
|
-
-
|
|
29
|
-
- 🚫 FORBIDDEN to skip security analysis
|
|
30
|
-
- 🚫 FORBIDDEN to skip thermo-nuclear maintainability review
|
|
31
|
-
- 🚫 FORBIDDEN to combine review categories into a single agent
|
|
14
|
+
- original task and acceptance criteria;
|
|
15
|
+
- intended paths and actual diff;
|
|
16
|
+
- relevant architecture and project rules;
|
|
17
|
+
- validation ledger and known baseline failures;
|
|
18
|
+
- unresolved risks, assumptions, and proof requirements.
|
|
32
19
|
|
|
33
|
-
|
|
20
|
+
Review the actual uncommitted task diff, not automatically `HEAD~1`.
|
|
34
21
|
|
|
35
|
-
|
|
36
|
-
- All tests pass
|
|
37
|
-
- Now looking for issues that tests miss
|
|
38
|
-
- Adversarial mindset - assume bugs exist
|
|
39
|
-
- **If `{teams_mode}` = true:** Agent team is still alive. Do NOT shutdown teammates - that happens in step-09-finish only.
|
|
22
|
+
## 2. Select review lenses
|
|
40
23
|
|
|
41
|
-
|
|
24
|
+
Use only lenses relevant to the change:
|
|
42
25
|
|
|
43
|
-
|
|
26
|
+
| Lens | Trigger examples |
|
|
27
|
+
|---|---|
|
|
28
|
+
| Correctness and edge cases | State transitions, concurrency, parsing, error handling |
|
|
29
|
+
| Security and authority | Auth, tenant boundaries, secrets, input boundaries, external actions |
|
|
30
|
+
| Data and migration safety | Schema changes, backfills, idempotency, rollback |
|
|
31
|
+
| Domain specialist | Payments, email, mobile, framework, provider, performance |
|
|
32
|
+
| Maintainability | Cross-cutting changes, new abstractions, large or structurally risky diffs |
|
|
33
|
+
| Evidence and acceptance | Runtime/provider/public claims or complex proof matrix |
|
|
44
34
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
<available_state>
|
|
48
|
-
From previous steps:
|
|
49
|
-
|
|
50
|
-
| Variable | Description |
|
|
51
|
-
|----------|-------------|
|
|
52
|
-
| `{task_description}` | What was implemented |
|
|
53
|
-
| `{task_id}` | Kebab-case identifier |
|
|
54
|
-
| `{auto_mode}` | Auto-fix Real findings |
|
|
55
|
-
| `{save_mode}` | Save outputs to files |
|
|
56
|
-
| `{economy_mode}` | No subagents, direct review |
|
|
57
|
-
| `{output_dir}` | Path to output (if save_mode) |
|
|
58
|
-
| Files modified | From step-03 |
|
|
59
|
-
</available_state>
|
|
60
|
-
|
|
61
|
-
---
|
|
62
|
-
|
|
63
|
-
## EXECUTION SEQUENCE:
|
|
64
|
-
|
|
65
|
-
### 1. Initialize Save Output (if save_mode)
|
|
35
|
+
Use a fresh independent context for each genuinely distinct lens. Combine closely related lenses when separation would only duplicate context. For low-risk changes, one focused independent reviewer may be enough. For high-risk changes, use multiple non-overlapping specialists.
|
|
66
36
|
|
|
67
|
-
|
|
37
|
+
Reviewers are read-only unless explicitly assigned a later resolution task.
|
|
68
38
|
|
|
69
|
-
|
|
70
|
-
bash {skill_dir}/scripts/update-progress.sh "{task_id}" "05" "examine" "in_progress"
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
Append findings to `{output_dir}/05-examine.md` as you work.
|
|
74
|
-
|
|
75
|
-
### 2. Gather Changes
|
|
76
|
-
|
|
77
|
-
```bash
|
|
78
|
-
git diff --name-only HEAD~1
|
|
79
|
-
git status --porcelain
|
|
80
|
-
```
|
|
39
|
+
## 3. Require high-signal findings
|
|
81
40
|
|
|
82
|
-
|
|
41
|
+
Every finding must contain:
|
|
83
42
|
|
|
84
|
-
|
|
43
|
+
- stable ID, severity, and confidence;
|
|
44
|
+
- exact file and line or artifact reference;
|
|
45
|
+
- concrete failure scenario or violated contract;
|
|
46
|
+
- evidence that the issue is introduced or exposed by the intended diff;
|
|
47
|
+
- smallest safe remediation direction.
|
|
85
48
|
|
|
86
|
-
|
|
87
|
-
→ Self-review with checklist:
|
|
49
|
+
Reject style preference, speculative breakage without a path, duplicated findings, and issues wholly outside scope.
|
|
88
50
|
|
|
89
|
-
|
|
90
|
-
## Security Checklist
|
|
91
|
-
- [ ] No SQL injection (parameterized queries)
|
|
92
|
-
- [ ] No XSS (output encoding)
|
|
93
|
-
- [ ] No secrets in code
|
|
94
|
-
- [ ] Input validation present
|
|
95
|
-
- [ ] Auth checks on protected routes
|
|
51
|
+
## 4. Validate findings
|
|
96
52
|
|
|
97
|
-
|
|
98
|
-
- [ ] Error handling for all failure modes
|
|
99
|
-
- [ ] Edge cases handled
|
|
100
|
-
- [ ] Null/undefined checks
|
|
101
|
-
- [ ] Race conditions considered
|
|
53
|
+
The coordinator independently inspects each reported issue and classifies it:
|
|
102
54
|
|
|
103
|
-
|
|
104
|
-
-
|
|
105
|
-
-
|
|
106
|
-
-
|
|
55
|
+
- `CONFIRMED`;
|
|
56
|
+
- `NOISE`;
|
|
57
|
+
- `PREEXISTING`;
|
|
58
|
+
- `OUT_OF_SCOPE`;
|
|
59
|
+
- `UNCERTAIN` with the exact missing evidence.
|
|
107
60
|
|
|
108
|
-
|
|
109
|
-
- [ ] No file pushed from under 1k lines to over 1k lines without strong justification
|
|
110
|
-
- [ ] No new ad-hoc conditionals / spaghetti branches in unrelated flows
|
|
111
|
-
- [ ] No thin abstractions, identity wrappers, or "magic" mechanisms added
|
|
112
|
-
- [ ] No unnecessary casts / `any` / `unknown` / optional params muddying contracts
|
|
113
|
-
- [ ] Feature logic stays in the canonical layer (no leaking into shared paths)
|
|
114
|
-
- [ ] Reuses existing canonical helpers instead of bespoke near-duplicates
|
|
115
|
-
- [ ] No "code judo" simplification was missed (could this be dramatically simpler?)
|
|
116
|
-
- [ ] Orchestration is parallel / atomic where the cleaner structure is obvious
|
|
117
|
-
```
|
|
61
|
+
Only confirmed findings block completion automatically. High-severity uncertain findings require targeted investigation before disposition.
|
|
118
62
|
|
|
119
|
-
|
|
120
|
-
→ Launch parallel review sub-agents.
|
|
63
|
+
## 5. Record review ledger
|
|
121
64
|
|
|
122
|
-
|
|
65
|
+
Store reviewer lens, evidence, classification, disposition, and any invalidated validation or proof artifacts.
|
|
123
66
|
|
|
124
|
-
First, gather the list of modified files:
|
|
125
67
|
```bash
|
|
126
|
-
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
Then, in **ONE message with 4+ parallel sub-agent launches**, launch:
|
|
130
|
-
|
|
131
|
-
---
|
|
132
|
-
|
|
133
|
-
**Agent 1: Security Review** - sub-agent profile/type: `code-reviewer`
|
|
134
|
-
```
|
|
135
|
-
prompt: |
|
|
136
|
-
You are a SECURITY reviewer. Review ONLY the following files for security vulnerabilities:
|
|
137
|
-
{list of modified files}
|
|
138
|
-
|
|
139
|
-
Focus exclusively on:
|
|
140
|
-
- OWASP Top 10: injection flaws (SQL, command, XSS)
|
|
141
|
-
- Authentication and authorization issues
|
|
142
|
-
- Sensitive data exposure (secrets, tokens, PII in logs)
|
|
143
|
-
- Security misconfiguration
|
|
144
|
-
- Insecure deserialization
|
|
145
|
-
- Missing input validation at system boundaries
|
|
146
|
-
|
|
147
|
-
For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
|
|
148
|
-
If no security issues found, explicitly state "No security issues found."
|
|
149
|
-
```
|
|
150
|
-
|
|
151
|
-
---
|
|
152
|
-
|
|
153
|
-
**Agent 2: Logic & Edge Cases Review** - sub-agent profile/type: `code-reviewer`
|
|
154
|
-
```
|
|
155
|
-
prompt: |
|
|
156
|
-
You are a LOGIC reviewer. Review ONLY the following files for logic correctness:
|
|
157
|
-
{list of modified files}
|
|
158
|
-
|
|
159
|
-
Focus exclusively on:
|
|
160
|
-
- Edge cases not handled (empty arrays, null/undefined, boundary values)
|
|
161
|
-
- Race conditions and concurrency issues
|
|
162
|
-
- Incorrect conditional logic or off-by-one errors
|
|
163
|
-
- Missing error handling for failure modes
|
|
164
|
-
- State management bugs
|
|
165
|
-
- Incorrect assumptions about data shape or types
|
|
166
|
-
|
|
167
|
-
For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
|
|
168
|
-
If no logic issues found, explicitly state "No logic issues found."
|
|
169
|
-
```
|
|
170
|
-
|
|
171
|
-
---
|
|
172
|
-
|
|
173
|
-
**Agent 3: Clean Code & Quality Review** - sub-agent profile/type: `code-reviewer`
|
|
174
|
-
```
|
|
175
|
-
prompt: |
|
|
176
|
-
You are a CLEAN CODE reviewer. Review ONLY the following files for code quality:
|
|
177
|
-
{list of modified files}
|
|
178
|
-
|
|
179
|
-
Focus exclusively on:
|
|
180
|
-
- SOLID principle violations
|
|
181
|
-
- Code smells (long methods, god objects, feature envy)
|
|
182
|
-
- Cyclomatic complexity > 10
|
|
183
|
-
- Code duplication > 20 lines
|
|
184
|
-
- Naming that doesn't communicate intent
|
|
185
|
-
- Functions doing too many things
|
|
186
|
-
|
|
187
|
-
For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
|
|
188
|
-
If no quality issues found, explicitly state "No quality issues found."
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
---
|
|
192
|
-
|
|
193
|
-
**Agent 4: Thermo-Nuclear Code Quality Review** (MANDATORY - launch alongside Agents 1-3)
|
|
194
|
-
|
|
195
|
-
This is the **final verification gate**. After the other reviewers find issues, this agent performs an extremely strict maintainability audit using the `thermo-nuclear-code-quality-review` skill. It is **not optional** and must run on every examine pass.
|
|
196
|
-
|
|
197
|
-
Sub-agent profile/type: `thermo-nuclear-code-quality-review`
|
|
198
|
-
```
|
|
199
|
-
prompt: |
|
|
200
|
-
Perform a Thermo-Nuclear Code Quality Review on the current branch's changes.
|
|
201
|
-
|
|
202
|
-
Modified files to audit:
|
|
203
|
-
{list of modified files}
|
|
204
|
-
|
|
205
|
-
Load and strictly apply the rubric from the `thermo-nuclear-code-quality-review` skill.
|
|
206
|
-
|
|
207
|
-
Verify and challenge:
|
|
208
|
-
- Structural code-quality regressions and missed "code judo" opportunities
|
|
209
|
-
- Any file pushed from under 1k lines to over 1k lines (presumptive blocker)
|
|
210
|
-
- New ad-hoc conditionals or spaghetti branching bolted onto unrelated flows
|
|
211
|
-
- Thin abstractions, identity wrappers, pass-through helpers, "magic" mechanisms
|
|
212
|
-
- Unnecessary casts, `any`, `unknown`, optional params obscuring real contracts
|
|
213
|
-
- Feature logic leaking into shared/canonical paths
|
|
214
|
-
- Bespoke helpers duplicating existing canonical utilities
|
|
215
|
-
- Unnecessary sequential orchestration or non-atomic update flows
|
|
216
|
-
- Whether the implementation could be dramatically simpler / smaller / more direct
|
|
217
|
-
|
|
218
|
-
For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and a concrete restructuring suggestion (prefer DELETING complexity over rearranging it).
|
|
219
|
-
|
|
220
|
-
Be ambitious, direct, and demanding. Do not soften major maintainability issues.
|
|
221
|
-
Apply the skill's Approval Bar strictly - flag presumptive blockers explicitly.
|
|
222
|
-
|
|
223
|
-
If no significant maintainability issues found, explicitly state "No thermo-nuclear findings."
|
|
224
|
-
```
|
|
225
|
-
|
|
226
|
-
---
|
|
227
|
-
|
|
228
|
-
**Agent 5: Vercel/Next.js Best Practices** (CONDITIONAL - launch alongside Agents 1-4)
|
|
229
|
-
|
|
230
|
-
→ **Detection:** Check if modified files match Next.js/Vercel patterns:
|
|
231
|
-
```
|
|
232
|
-
- *.tsx, *.jsx files in app/, pages/, components/
|
|
233
|
-
- next.config.* files
|
|
234
|
-
- Server actions (use server)
|
|
235
|
-
- API routes (app/api/*, pages/api/*)
|
|
236
|
-
- Middleware (middleware.ts)
|
|
237
|
-
- Server components, client components
|
|
238
|
-
```
|
|
239
|
-
|
|
240
|
-
→ **If Next.js/Vercel code detected:** Add a 5th parallel review sub-agent:
|
|
241
|
-
|
|
242
|
-
Sub-agent profile/type: `code-reviewer`
|
|
243
|
-
```
|
|
244
|
-
prompt: |
|
|
245
|
-
You are a NEXT.JS / REACT PERFORMANCE reviewer. Review ONLY the following files:
|
|
246
|
-
{list of modified files}
|
|
247
|
-
|
|
248
|
-
Focus exclusively on:
|
|
249
|
-
- Sequential awaits that should use Promise.all for parallel fetching
|
|
250
|
-
- Barrel imports causing bundle bloat (import from index files)
|
|
251
|
-
- Missing dynamic imports for heavy client components
|
|
252
|
-
- Server-side caching opportunities (React cache, unstable_cache)
|
|
253
|
-
- Unnecessary re-renders (missing memo, useMemo, useCallback)
|
|
254
|
-
- Wrong Server vs Client component boundaries
|
|
255
|
-
- Data fetching patterns (preloading, parallel fetching, waterfall detection)
|
|
256
|
-
|
|
257
|
-
For each finding, provide: file:line, severity (CRITICAL/HIGH/MEDIUM/LOW), description, and suggested fix.
|
|
258
|
-
If no performance issues found, explicitly state "No performance issues found."
|
|
259
|
-
```
|
|
260
|
-
|
|
261
|
-
→ **If NOT Next.js/Vercel code:** Skip this agent (launch only Agents 1-4)
|
|
262
|
-
|
|
263
|
-
---
|
|
264
|
-
|
|
265
|
-
**🛑 REMINDER: You MUST have 4+ sub-agent launches in a SINGLE response (Security + Logic + Clean Code + Thermo-Nuclear, plus Next.js if applicable). If you only launched 1-3 agents, you are doing it WRONG. Go back and launch all agents. The Thermo-Nuclear agent is MANDATORY - never skip it.**
|
|
266
|
-
|
|
267
|
-
### 4. Classify Findings
|
|
268
|
-
|
|
269
|
-
For each finding:
|
|
270
|
-
|
|
271
|
-
**Severity:**
|
|
272
|
-
- CRITICAL: Security vulnerability, data loss risk
|
|
273
|
-
- HIGH: Significant bug, will cause issues
|
|
274
|
-
- MEDIUM: Should fix, not urgent
|
|
275
|
-
- LOW: Minor improvement
|
|
276
|
-
|
|
277
|
-
**Validity:**
|
|
278
|
-
- Real: Definitely needs fixing
|
|
279
|
-
- Noise: Not actually a problem
|
|
280
|
-
- Uncertain: Needs discussion
|
|
281
|
-
|
|
282
|
-
### 5. Present Findings Table
|
|
283
|
-
|
|
284
|
-
```markdown
|
|
285
|
-
## Findings
|
|
286
|
-
|
|
287
|
-
| ID | Severity | Category | Location | Issue | Validity |
|
|
288
|
-
|----|----------|----------|----------|-------|----------|
|
|
289
|
-
| F1 | CRITICAL | Security | auth.ts:42 | SQL injection | Real |
|
|
290
|
-
| F2 | HIGH | Logic | handler.ts:78 | Missing null check | Real |
|
|
291
|
-
| F3 | MEDIUM | Quality | utils.ts:15 | Complex function | Uncertain |
|
|
292
|
-
|
|
293
|
-
**Summary:** {count} findings ({blocking} blocking)
|
|
294
|
-
```
|
|
295
|
-
|
|
296
|
-
### 6. Create Finding Todos
|
|
297
|
-
|
|
68
|
+
python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase examine --status complete --message "Independent findings validated and deduplicated"
|
|
298
69
|
```
|
|
299
|
-
- [ ] F1 [CRITICAL] Fix SQL injection in auth.ts:42
|
|
300
|
-
- [ ] F2 [HIGH] Add null check in handler.ts:78
|
|
301
|
-
```
|
|
302
|
-
|
|
303
|
-
### 7. Get User Approval (review → resolve/test)
|
|
304
|
-
|
|
305
|
-
**If `{auto_mode}` = true:**
|
|
306
|
-
→ Proceed automatically based on findings
|
|
307
|
-
|
|
308
|
-
**If `{auto_mode}` = false:**
|
|
309
|
-
|
|
310
|
-
```yaml
|
|
311
|
-
questions:
|
|
312
|
-
- header: "Review"
|
|
313
|
-
question: "Review complete. How would you like to proceed?"
|
|
314
|
-
options:
|
|
315
|
-
- label: "Resolve findings (Recommended)"
|
|
316
|
-
description: "Address the identified issues"
|
|
317
|
-
- label: "Skip to tests"
|
|
318
|
-
description: "Skip resolution, proceed to test creation"
|
|
319
|
-
- label: "Skip resolution"
|
|
320
|
-
description: "Accept findings, don't make changes"
|
|
321
|
-
- label: "Discuss findings"
|
|
322
|
-
description: "I want to discuss specific findings"
|
|
323
|
-
multiSelect: false
|
|
324
|
-
```
|
|
325
|
-
|
|
326
|
-
<critical>
|
|
327
|
-
This is one of the THREE transition points that requires user confirmation:
|
|
328
|
-
1. plan → execute
|
|
329
|
-
2. validate → review
|
|
330
|
-
3. review → resolve/test (THIS ONE)
|
|
331
|
-
</critical>
|
|
332
|
-
|
|
333
|
-
### 8. Complete Save Output (if save_mode)
|
|
334
|
-
|
|
335
|
-
**If `{save_mode}` = true:**
|
|
336
|
-
|
|
337
|
-
Append to `{output_dir}/05-examine.md`:
|
|
338
|
-
```markdown
|
|
339
|
-
---
|
|
340
|
-
## Step Complete
|
|
341
|
-
**Status:** ✓ Complete
|
|
342
|
-
**Findings:** {count}
|
|
343
|
-
**Critical:** {count}
|
|
344
|
-
**Next:** step-06-resolve.md
|
|
345
|
-
**Timestamp:** {ISO timestamp}
|
|
346
|
-
```
|
|
347
|
-
|
|
348
|
-
---
|
|
349
|
-
|
|
350
|
-
## SUCCESS METRICS:
|
|
351
|
-
|
|
352
|
-
✅ All modified files reviewed
|
|
353
|
-
✅ Security checklist completed
|
|
354
|
-
✅ Findings classified by severity
|
|
355
|
-
✅ Validity assessed for each finding
|
|
356
|
-
✅ Findings table presented
|
|
357
|
-
✅ Todos created for tracking
|
|
358
|
-
✅ Thermo-Nuclear maintainability audit completed (mandatory final gate)
|
|
359
|
-
✅ Next.js/Vercel best practices checked (if applicable)
|
|
360
|
-
|
|
361
|
-
## FAILURE MODES:
|
|
362
|
-
|
|
363
|
-
❌ **CRITICAL**: Launching only 1 review agent instead of 4+ - each category (Security, Logic, Clean Code, Thermo-Nuclear) MUST be a separate agent
|
|
364
|
-
❌ **CRITICAL**: Skipping the Thermo-Nuclear maintainability review - it is the mandatory final verification gate
|
|
365
|
-
❌ Combining multiple review categories into a single agent prompt
|
|
366
|
-
❌ Skipping security review
|
|
367
|
-
❌ Not classifying by severity
|
|
368
|
-
❌ Auto-dismissing findings
|
|
369
|
-
❌ Launching agents sequentially instead of in parallel
|
|
370
|
-
❌ Using subagents when economy_mode
|
|
371
|
-
❌ Skipping Vercel/Next.js review when React/Next.js files are modified
|
|
372
|
-
❌ **CRITICAL**: Not using AskUserQuestion for review → resolve/test transition
|
|
373
|
-
|
|
374
|
-
## REVIEW PROTOCOLS:
|
|
375
|
-
|
|
376
|
-
- Adversarial mindset - assume bugs exist
|
|
377
|
-
- Check security FIRST
|
|
378
|
-
- Every finding gets severity and validity
|
|
379
|
-
- Don't dismiss without justification
|
|
380
|
-
- Present clear summary
|
|
381
|
-
|
|
382
|
-
---
|
|
383
|
-
|
|
384
|
-
## NEXT STEP:
|
|
385
|
-
|
|
386
|
-
After user confirms via AskUserQuestion (or auto-proceed):
|
|
387
|
-
|
|
388
|
-
**If user chooses "Resolve findings":** → Load `./step-06-resolve.md`
|
|
389
|
-
|
|
390
|
-
**If user chooses "Skip to tests" (and test_mode):** → Load `./step-07-tests.md`
|
|
391
70
|
|
|
392
|
-
|
|
393
|
-
- **If test_mode:** → Load `./step-07-tests.md`
|
|
394
|
-
- **If pr_mode:** → Load `./step-09-finish.md` to create pull request
|
|
395
|
-
- **Otherwise:** → Workflow complete - show summary
|
|
71
|
+
## Routing
|
|
396
72
|
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
|
|
73
|
+
- If confirmed findings exist, load `step-06-resolve.md`.
|
|
74
|
+
- If test coverage must change, load `step-07-tests.md`.
|
|
75
|
+
- If runtime proof is required, load `step-10-verify.md`.
|
|
76
|
+
- Otherwise load `step-09-finish.md`.
|