aiblueprint-cli 1.4.99 → 1.4.101
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +0 -1
- package/agents-config/skills/agents-manager/SKILL.md +2 -2
- package/agents-config/skills/agents-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/agents-manager/assets/codex-icon.svg +20 -0
- package/agents-config/skills/apex/SKILL.md +120 -118
- package/agents-config/skills/apex/agents/openai.yaml +10 -0
- package/agents-config/skills/apex/assets/codex-icon.svg +15 -0
- package/agents-config/skills/apex/scripts/apex-state.py +740 -0
- package/agents-config/skills/apex/scripts/setup-templates.sh +27 -145
- package/agents-config/skills/apex/scripts/test_apex_state.py +413 -0
- package/agents-config/skills/apex/scripts/update-progress.sh +17 -73
- package/agents-config/skills/apex/steps/step-00-init.md +85 -231
- package/agents-config/skills/apex/steps/step-00b-branch.md +10 -118
- package/agents-config/skills/apex/steps/step-00b-economy.md +12 -239
- package/agents-config/skills/apex/steps/step-00b-interactive.md +13 -162
- package/agents-config/skills/apex/steps/step-00b-save.md +13 -114
- package/agents-config/skills/apex/steps/step-01-analyze.md +40 -361
- package/agents-config/skills/apex/steps/step-02-plan.md +55 -562
- package/agents-config/skills/apex/steps/step-02b-tasks.md +15 -291
- package/agents-config/skills/apex/steps/step-03-execute-teams.md +47 -267
- package/agents-config/skills/apex/steps/step-03-execute.md +32 -212
- package/agents-config/skills/apex/steps/step-04-validate.md +42 -246
- package/agents-config/skills/apex/steps/step-05-examine.md +47 -371
- package/agents-config/skills/apex/steps/step-06-resolve.md +19 -221
- package/agents-config/skills/apex/steps/step-07-tests.md +19 -234
- package/agents-config/skills/apex/steps/step-08-run-tests.md +13 -300
- package/agents-config/skills/apex/steps/step-09-finish.md +36 -200
- package/agents-config/skills/apex/steps/step-10-verify.md +46 -264
- package/agents-config/skills/appstore-connect/agents/openai.yaml +7 -0
- package/agents-config/skills/appstore-connect/assets/codex-icon.svg +17 -0
- package/agents-config/skills/commit/agents/openai.yaml +10 -0
- package/agents-config/skills/commit/assets/codex-icon.svg +17 -0
- package/agents-config/skills/create-pr/agents/openai.yaml +10 -0
- package/agents-config/skills/create-pr/assets/codex-icon.svg +17 -0
- package/agents-config/skills/environments-manager/SKILL.md +1 -1
- package/agents-config/skills/environments-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/environments-manager/assets/codex-icon.svg +16 -0
- package/agents-config/skills/environments-manager/examples/scripts/claude-worktree-remove.sh +19 -3
- package/agents-config/skills/environments-manager/examples/scripts/worktree-up.sh +1 -1
- package/agents-config/skills/environments-manager/references/claude.md +1 -1
- package/agents-config/skills/fix-pr-comments/agents/openai.yaml +10 -0
- package/agents-config/skills/fix-pr-comments/assets/codex-icon.svg +17 -0
- package/agents-config/skills/grill-me/SKILL.md +25 -4
- package/agents-config/skills/grill-me/agents/openai.yaml +8 -0
- package/agents-config/skills/grill-me/assets/codex-icon.svg +16 -0
- package/agents-config/skills/hooks-manager/SKILL.md +19 -9
- package/agents-config/skills/hooks-manager/assets/codex-icon.svg +15 -4
- package/agents-config/skills/hooks-manager/references/claude-code.md +32 -0
- package/agents-config/skills/hooks-manager/references/codex.md +23 -0
- package/agents-config/skills/hooks-manager/references/cursor.md +18 -0
- package/agents-config/skills/hooks-manager/references/hook-types.md +5 -3
- package/agents-config/skills/hooks-manager/references/input-output-schemas.md +2 -2
- package/agents-config/skills/hooks-manager/references/research-sources.md +25 -0
- package/agents-config/skills/hooks-manager/references/router.md +32 -0
- package/agents-config/skills/hooks-manager/references/troubleshooting.md +3 -3
- package/agents-config/skills/merge/agents/openai.yaml +10 -0
- package/agents-config/skills/merge/assets/codex-icon.svg +17 -0
- package/agents-config/skills/oneshot/SKILL.md +4 -0
- package/agents-config/skills/oneshot/agents/openai.yaml +10 -0
- package/agents-config/skills/oneshot/assets/codex-icon.svg +18 -0
- package/agents-config/skills/rules-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/rules-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/SKILL.md +45 -3
- package/agents-config/skills/skill-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/skill-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/references/skill-writing-glossary.md +201 -0
- package/agents-config/skills/skill-manager/scripts/setup-codex-icons.ts +143 -0
- package/agents-config/skills/ultrathink/agents/openai.yaml +10 -0
- package/agents-config/skills/ultrathink/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-artifacts/SKILL.md +102 -51
- package/agents-config/skills/use-artifacts/assets/local-runtime.js +299 -0
- package/agents-config/skills/use-artifacts/scripts/create_artifact.py +1 -1
- package/agents-config/skills/use-delegate/SKILL.md +4 -0
- package/agents-config/skills/use-delegate/agents/openai.yaml +10 -0
- package/agents-config/skills/use-delegate/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-goal/SKILL.md +70 -9
- package/agents-config/skills/use-goal/agents/openai.yaml +1 -1
- package/agents-config/skills/use-goal/assets/codex-icon.svg +17 -3
- package/agents-config/skills/use-goal/references/claude-code-goal.md +54 -6
- package/agents-config/skills/use-goal/references/codex-goal.md +59 -4
- package/agents-config/skills/use-goal/references/verification-harnesses.md +104 -3
- package/dist/cli.js +365 -363
- package/package.json +1 -1
- package/agents-config/skills/apex/templates/00-context.md +0 -55
- package/agents-config/skills/apex/templates/01-analyze.md +0 -10
- package/agents-config/skills/apex/templates/02-plan.md +0 -10
- package/agents-config/skills/apex/templates/03-execute.md +0 -10
- package/agents-config/skills/apex/templates/04-validate.md +0 -10
- package/agents-config/skills/apex/templates/05-examine.md +0 -10
- package/agents-config/skills/apex/templates/06-resolve.md +0 -10
- package/agents-config/skills/apex/templates/07-tests.md +0 -10
- package/agents-config/skills/apex/templates/08-run-tests.md +0 -10
- package/agents-config/skills/apex/templates/09-finish.md +0 -10
- package/agents-config/skills/apex/templates/10-verify.md +0 -9
- package/agents-config/skills/apex/templates/README.md +0 -195
- package/agents-config/skills/apex/templates/step-complete.md +0 -7
- package/agents-config/skills/prompt-creator/SKILL.md +0 -285
- package/agents-config/skills/prompt-creator/references/anthropic-best-practices.md +0 -126
- package/agents-config/skills/prompt-creator/references/anti-patterns.md +0 -57
- package/agents-config/skills/prompt-creator/references/clarity-principles.md +0 -54
- package/agents-config/skills/prompt-creator/references/context-management.md +0 -389
- package/agents-config/skills/prompt-creator/references/few-shot-patterns.md +0 -47
- package/agents-config/skills/prompt-creator/references/openai-best-practices.md +0 -50
- package/agents-config/skills/prompt-creator/references/prompt-templates.md +0 -110
- package/agents-config/skills/prompt-creator/references/reasoning-techniques.md +0 -52
- package/agents-config/skills/prompt-creator/references/system-prompt-patterns.md +0 -48
- package/agents-config/skills/prompt-creator/references/xml-structure.md +0 -36
|
@@ -1,240 +1,60 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: step-03-execute
|
|
3
|
-
description:
|
|
4
|
-
|
|
5
|
-
next_step: steps/step-04-validate.md
|
|
3
|
+
description: Execute the next APEX task units adaptively with bounded attempts, scope checks, checkpoints, and re-planning.
|
|
4
|
+
next_step: step-04-validate.md
|
|
6
5
|
---
|
|
7
6
|
|
|
8
|
-
# Step 3: Execute
|
|
7
|
+
# Step 3: Execute
|
|
9
8
|
|
|
10
|
-
|
|
9
|
+
Implement the task graph, not a stale narrative plan.
|
|
11
10
|
|
|
12
|
-
|
|
13
|
-
- 🛑 NEVER add features not in the plan (scope creep)
|
|
14
|
-
- 🛑 NEVER modify files without reading them first
|
|
15
|
-
- ✅ ALWAYS follow the plan file-by-file
|
|
16
|
-
- ✅ ALWAYS mark todos complete immediately after each task
|
|
17
|
-
- ✅ ALWAYS read files BEFORE editing them
|
|
18
|
-
- 📋 YOU ARE AN IMPLEMENTER following a plan, not a designer
|
|
19
|
-
- 💬 FOCUS on executing the plan exactly as approved
|
|
20
|
-
- 🚫 FORBIDDEN to add "improvements" not in the plan
|
|
11
|
+
## 1. Re-read current state
|
|
21
12
|
|
|
22
|
-
|
|
13
|
+
Before each task unit:
|
|
23
14
|
|
|
24
|
-
-
|
|
25
|
-
-
|
|
26
|
-
-
|
|
27
|
-
-
|
|
15
|
+
- confirm dependencies are complete;
|
|
16
|
+
- compare repository state with the last checkpoint;
|
|
17
|
+
- inspect overlapping local changes;
|
|
18
|
+
- confirm the unit's write boundary, side effects, validation, and evidence;
|
|
19
|
+
- re-plan if an assumption or boundary is stale.
|
|
28
20
|
|
|
29
|
-
##
|
|
21
|
+
## 2. Choose local or delegated execution
|
|
30
22
|
|
|
31
|
-
|
|
32
|
-
- Files to modify are known from the plan
|
|
33
|
-
- Patterns to follow are documented from step-01
|
|
34
|
-
- Don't add features - stick to the plan
|
|
35
|
-
- **If context was cleared ("Execute and clear context"):** The plan file contains all APEX state variables in the "APEX Workflow Context" section. Read the plan file first and restore all variables before proceeding.
|
|
23
|
+
Keep the unit local when it is on the immediate critical path, tightly coupled to current context, small, or likely to need rapid iteration. Delegate when it is self-contained and a separate context materially helps.
|
|
36
24
|
|
|
37
|
-
|
|
25
|
+
Delegated packets must include the task contract, exact boundaries, relevant project rules, dependencies, expected output, validation, and stop condition. A worker may report a newly discovered need but may not silently widen scope.
|
|
38
26
|
|
|
39
|
-
|
|
27
|
+
## 3. Record the attempt
|
|
40
28
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
<available_state>
|
|
44
|
-
From previous steps:
|
|
45
|
-
|
|
46
|
-
| Variable | Description |
|
|
47
|
-
|----------|-------------|
|
|
48
|
-
| `{task_description}` | What to implement |
|
|
49
|
-
| `{task_id}` | Kebab-case identifier |
|
|
50
|
-
| `{auto_mode}` | Skip confirmations |
|
|
51
|
-
| `{save_mode}` | Save outputs to files |
|
|
52
|
-
| `{output_dir}` | Path to output (if save_mode) |
|
|
53
|
-
| Implementation plan | File-by-file changes from step-02 |
|
|
54
|
-
| Patterns | How to implement from step-01 |
|
|
55
|
-
</available_state>
|
|
56
|
-
|
|
57
|
-
---
|
|
58
|
-
|
|
59
|
-
## EXECUTION SEQUENCE:
|
|
60
|
-
|
|
61
|
-
### 1. Initialize Save Output (if save_mode)
|
|
62
|
-
|
|
63
|
-
**If `{save_mode}` = true:**
|
|
29
|
+
Every attempt has a stable task ID and incrementing attempt number. Record its starting revision, owner, intended paths, and status before mutation.
|
|
64
30
|
|
|
65
31
|
```bash
|
|
66
|
-
|
|
32
|
+
python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase execute --task-id "{unit_id}" --status in_progress --message "Attempt started"
|
|
67
33
|
```
|
|
68
34
|
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
### 2. Create Todos from Plan
|
|
35
|
+
## 4. Implement in a tight loop
|
|
72
36
|
|
|
73
|
-
|
|
37
|
+
1. Make the smallest coherent edit.
|
|
38
|
+
2. Inspect the changed diff immediately.
|
|
39
|
+
3. Run the shortest relevant feedback command.
|
|
40
|
+
4. Fix introduced failures within scope.
|
|
41
|
+
5. Repeat until the task's evidence contract is met or a re-plan trigger fires.
|
|
74
42
|
|
|
75
|
-
|
|
76
|
-
Plan entry:
|
|
77
|
-
#### `src/auth/handler.ts`
|
|
78
|
-
- Add `validateToken` function
|
|
79
|
-
- Handle error case: expired token
|
|
80
|
-
|
|
81
|
-
Becomes:
|
|
82
|
-
- [ ] src/auth/handler.ts: Add validateToken function
|
|
83
|
-
- [ ] src/auth/handler.ts: Handle expired token error
|
|
84
|
-
```
|
|
43
|
+
Do not opportunistically refactor unrelated code. Preserve user changes even when they complicate the implementation.
|
|
85
44
|
|
|
86
|
-
|
|
45
|
+
## 5. Close or re-plan
|
|
87
46
|
|
|
88
|
-
|
|
47
|
+
A task is complete only when its declared output exists, its write boundary is respected, relevant validation has a classified result, and required evidence is recorded.
|
|
89
48
|
|
|
90
|
-
|
|
49
|
+
If blocked, record the concrete condition, attempted alternatives, and exact input or authority needed. Continue with other independent unblocked units when useful.
|
|
91
50
|
|
|
92
|
-
|
|
93
|
-
- Only ONE todo in_progress at a time
|
|
94
|
-
|
|
95
|
-
**3.2 Read Before Edit**
|
|
96
|
-
```
|
|
97
|
-
ALWAYS read the file before modifying:
|
|
98
|
-
- Understand current structure
|
|
99
|
-
- Find exact insertion points
|
|
100
|
-
- Verify patterns match expectations
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
**3.3 Implement Changes**
|
|
104
|
-
```
|
|
105
|
-
Make changes specified in the plan:
|
|
106
|
-
- Follow patterns from step-01 analysis
|
|
107
|
-
- Use exact names from plan
|
|
108
|
-
- Handle error cases as specified
|
|
109
|
-
- NO comments unless truly necessary
|
|
110
|
-
```
|
|
111
|
-
|
|
112
|
-
**3.4 Mark Complete Immediately**
|
|
113
|
-
- Mark todo complete RIGHT AFTER finishing
|
|
114
|
-
- Don't batch completions
|
|
115
|
-
|
|
116
|
-
**3.5 Log Progress (if save_mode)**
|
|
117
|
-
```markdown
|
|
118
|
-
### ✓ src/auth/handler.ts
|
|
119
|
-
- Added `validateToken` function (lines 45-78)
|
|
120
|
-
- Added error handling for expired tokens
|
|
121
|
-
**Timestamp:** {ISO}
|
|
122
|
-
```
|
|
123
|
-
|
|
124
|
-
### 4. Handle Blockers
|
|
125
|
-
|
|
126
|
-
**If `{auto_mode}` = true:**
|
|
127
|
-
→ Make reasonable decision and continue
|
|
128
|
-
|
|
129
|
-
**If `{auto_mode}` = false:**
|
|
130
|
-
|
|
131
|
-
```yaml
|
|
132
|
-
questions:
|
|
133
|
-
- header: "Blocker"
|
|
134
|
-
question: "Encountered an issue. How should we proceed?"
|
|
135
|
-
options:
|
|
136
|
-
- label: "Use alternative approach (Recommended)"
|
|
137
|
-
description: "Description of alternative"
|
|
138
|
-
- label: "Skip this part"
|
|
139
|
-
description: "Continue without this change"
|
|
140
|
-
- label: "Stop for discussion"
|
|
141
|
-
description: "I want to discuss before continuing"
|
|
142
|
-
multiSelect: false
|
|
143
|
-
```
|
|
144
|
-
|
|
145
|
-
### 5. Verify Implementation
|
|
146
|
-
|
|
147
|
-
After completing all todos:
|
|
51
|
+
After completion:
|
|
148
52
|
|
|
149
53
|
```bash
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
Fix any errors immediately.
|
|
154
|
-
|
|
155
|
-
### 6. Implementation Summary
|
|
156
|
-
|
|
157
|
-
```
|
|
158
|
-
**Implementation Complete**
|
|
159
|
-
|
|
160
|
-
**Files Modified:**
|
|
161
|
-
- `src/auth/handler.ts` - Added validateToken, error handling
|
|
162
|
-
- `src/api/auth/route.ts` - Integrated token validation
|
|
163
|
-
|
|
164
|
-
**New Files:**
|
|
165
|
-
- `src/types/auth.ts` - Auth type definitions
|
|
166
|
-
|
|
167
|
-
**Todos:** {X}/{Y} complete
|
|
54
|
+
python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase execute --task-id "{unit_id}" --status complete --message "Task output and evidence recorded"
|
|
55
|
+
python3 "{skill_dir}/scripts/apex-state.py" checkpoint --root "$PWD" --run-id "{run_id}" --phase execute --message "Task checkpoint"
|
|
168
56
|
```
|
|
169
57
|
|
|
170
|
-
|
|
171
|
-
→ Proceed to validation
|
|
172
|
-
|
|
173
|
-
**If `{auto_mode}` = false:**
|
|
174
|
-
|
|
175
|
-
```yaml
|
|
176
|
-
questions:
|
|
177
|
-
- header: "Execute"
|
|
178
|
-
question: "Implementation complete. Ready to validate?"
|
|
179
|
-
options:
|
|
180
|
-
- label: "Proceed to validation (Recommended)"
|
|
181
|
-
description: "Run typecheck, lint, and tests"
|
|
182
|
-
- label: "Review changes"
|
|
183
|
-
description: "I want to review what was changed"
|
|
184
|
-
- label: "Make adjustments"
|
|
185
|
-
description: "I want to modify something"
|
|
186
|
-
multiSelect: false
|
|
187
|
-
```
|
|
188
|
-
|
|
189
|
-
### 7. Complete Save Output (if save_mode)
|
|
190
|
-
|
|
191
|
-
**If `{save_mode}` = true:**
|
|
192
|
-
|
|
193
|
-
Append to `{output_dir}/03-execute.md`:
|
|
194
|
-
```markdown
|
|
195
|
-
---
|
|
196
|
-
## Step Complete
|
|
197
|
-
**Status:** ✓ Complete
|
|
198
|
-
**Files modified:** {count}
|
|
199
|
-
**Todos completed:** {count}
|
|
200
|
-
**Next:** step-04-validate.md
|
|
201
|
-
**Timestamp:** {ISO timestamp}
|
|
202
|
-
```
|
|
203
|
-
|
|
204
|
-
---
|
|
205
|
-
|
|
206
|
-
## SUCCESS METRICS:
|
|
207
|
-
|
|
208
|
-
✅ All plan items implemented
|
|
209
|
-
✅ All todos marked complete
|
|
210
|
-
✅ No scope creep - only plan items
|
|
211
|
-
✅ Files read before modification
|
|
212
|
-
✅ Typecheck and lint pass
|
|
213
|
-
✅ Progress logged (if save_mode)
|
|
214
|
-
|
|
215
|
-
## FAILURE MODES:
|
|
216
|
-
|
|
217
|
-
❌ Adding features not in the plan
|
|
218
|
-
❌ Modifying files without reading first
|
|
219
|
-
❌ Not updating todos as you work
|
|
220
|
-
❌ Multiple todos in_progress simultaneously
|
|
221
|
-
❌ Ignoring type or lint errors
|
|
222
|
-
❌ **CRITICAL**: Not using AskUserQuestion for blockers
|
|
223
|
-
|
|
224
|
-
## EXECUTION PROTOCOLS:
|
|
225
|
-
|
|
226
|
-
- Follow the plan EXACTLY
|
|
227
|
-
- Read before write
|
|
228
|
-
- One file at a time
|
|
229
|
-
- Update todos in real-time
|
|
230
|
-
- Fix errors immediately
|
|
231
|
-
|
|
232
|
-
---
|
|
233
|
-
|
|
234
|
-
## NEXT STEP:
|
|
235
|
-
|
|
236
|
-
After implementation complete, load `./step-04-validate.md`
|
|
58
|
+
## Completion
|
|
237
59
|
|
|
238
|
-
|
|
239
|
-
Remember: Execution is about following the plan - don't redesign or add features!
|
|
240
|
-
</critical>
|
|
60
|
+
Proceed to `step-04-validate.md` when all required graph nodes are complete or explicitly blocked with no remaining meaningful in-scope work.
|
|
@@ -1,272 +1,68 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: step-04-validate
|
|
3
|
-
description:
|
|
4
|
-
prev_step: steps/step-03-execute.md
|
|
5
|
-
next_step: steps/step-05-examine.md
|
|
3
|
+
description: Integrate the APEX diff and classify relevant validation without confusing regressions, baseline noise, or unavailable checks.
|
|
6
4
|
---
|
|
7
5
|
|
|
8
|
-
# Step 4:
|
|
6
|
+
# Step 4: Integrate and validate
|
|
9
7
|
|
|
10
|
-
|
|
8
|
+
Validation is evidence collection, not a ritual command list.
|
|
11
9
|
|
|
12
|
-
|
|
13
|
-
- 🛑 NEVER skip any validation step
|
|
14
|
-
- ✅ ALWAYS run typecheck, lint, and tests
|
|
15
|
-
- ✅ ALWAYS verify each acceptance criterion
|
|
16
|
-
- ✅ ALWAYS fix failures before proceeding
|
|
17
|
-
- 📋 YOU ARE A VALIDATOR, not an implementer
|
|
18
|
-
- 💬 FOCUS on "Does it work correctly?"
|
|
19
|
-
- 🚫 FORBIDDEN to proceed with failing checks
|
|
10
|
+
## 1. Review the integrated scope
|
|
20
11
|
|
|
21
|
-
|
|
12
|
+
Inspect current Git status, staged and unstaged diffs, untracked files, generated artifacts, and the task graph. Confirm:
|
|
22
13
|
|
|
23
|
-
-
|
|
24
|
-
-
|
|
25
|
-
-
|
|
26
|
-
-
|
|
14
|
+
- every intended change maps to an acceptance criterion;
|
|
15
|
+
- no unrelated user change was absorbed or overwritten;
|
|
16
|
+
- no task exceeded its write boundary without a recorded re-plan;
|
|
17
|
+
- dependencies and generated outputs are consistent;
|
|
18
|
+
- formatting did not create unrelated churn.
|
|
27
19
|
|
|
28
|
-
##
|
|
20
|
+
## 2. Discover relevant checks
|
|
29
21
|
|
|
30
|
-
|
|
31
|
-
- Tests may or may not pass yet
|
|
32
|
-
- Type errors may exist
|
|
33
|
-
- Focus is on verification, not new implementation
|
|
34
|
-
- **If `{teams_mode}` = true:** The agent team is still alive. Do NOT shutdown or dismiss teammates. Team shutdown happens in step-09-finish only.
|
|
22
|
+
Read project instructions, package scripts, CI configuration, and nearby tests. Select checks from the changed surface and risk:
|
|
35
23
|
|
|
36
|
-
|
|
24
|
+
- syntax, formatting, lint, and types;
|
|
25
|
+
- targeted unit, integration, contract, or end-to-end tests;
|
|
26
|
+
- build, packaging, schema, migration, or generated-code validation;
|
|
27
|
+
- runtime, provider, or public-artifact checks when required.
|
|
37
28
|
|
|
38
|
-
|
|
29
|
+
Do not invent a command because another ecosystem commonly uses it.
|
|
39
30
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
<available_state>
|
|
43
|
-
From previous steps:
|
|
44
|
-
|
|
45
|
-
| Variable | Description |
|
|
46
|
-
|----------|-------------|
|
|
47
|
-
| `{task_description}` | What was implemented |
|
|
48
|
-
| `{task_id}` | Kebab-case identifier |
|
|
49
|
-
| `{acceptance_criteria}` | Success criteria |
|
|
50
|
-
| `{auto_mode}` | Skip confirmations |
|
|
51
|
-
| `{save_mode}` | Save outputs to files |
|
|
52
|
-
| `{test_mode}` | Include test steps |
|
|
53
|
-
| `{examine_mode}` | Auto-proceed to review |
|
|
54
|
-
| `{output_dir}` | Path to output (if save_mode) |
|
|
55
|
-
| Implementation | Completed in step-03 |
|
|
56
|
-
</available_state>
|
|
57
|
-
|
|
58
|
-
---
|
|
59
|
-
|
|
60
|
-
## EXECUTION SEQUENCE:
|
|
61
|
-
|
|
62
|
-
### 1. Initialize Save Output (if save_mode)
|
|
63
|
-
|
|
64
|
-
**If `{save_mode}` = true:**
|
|
65
|
-
|
|
66
|
-
```bash
|
|
67
|
-
bash {skill_dir}/scripts/update-progress.sh "{task_id}" "04" "validate" "in_progress"
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
Append results to `{output_dir}/04-validate.md` as you work.
|
|
71
|
-
|
|
72
|
-
### 2. Discover Available Commands
|
|
73
|
-
|
|
74
|
-
Check `package.json` for exact command names:
|
|
75
|
-
```bash
|
|
76
|
-
cat package.json | grep -A 20 '"scripts"'
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
Look for: `typecheck`, `lint`, `test`, `build`, `format`
|
|
80
|
-
|
|
81
|
-
### 3. Run Validation Suite
|
|
31
|
+
## 3. Establish baseline when needed
|
|
82
32
|
|
|
83
|
-
|
|
84
|
-
```bash
|
|
85
|
-
pnpm run typecheck # or npm run typecheck
|
|
86
|
-
```
|
|
87
|
-
|
|
88
|
-
**MUST PASS.** If fails:
|
|
89
|
-
1. Read error messages
|
|
90
|
-
2. Fix type issues
|
|
91
|
-
3. Re-run until passing
|
|
92
|
-
|
|
93
|
-
**3.2 Lint**
|
|
94
|
-
```bash
|
|
95
|
-
pnpm run lint
|
|
96
|
-
```
|
|
97
|
-
|
|
98
|
-
**MUST PASS.** If fails:
|
|
99
|
-
1. Try auto-fix: `pnpm run lint --fix`
|
|
100
|
-
2. Manually fix remaining
|
|
101
|
-
3. Re-run until passing
|
|
102
|
-
|
|
103
|
-
**3.3 Tests**
|
|
104
|
-
```bash
|
|
105
|
-
pnpm run test -- --filter={affected-area}
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
**MUST PASS.** If fails:
|
|
109
|
-
1. Identify failing test
|
|
110
|
-
2. Determine if code bug or test bug
|
|
111
|
-
3. Fix the root cause
|
|
112
|
-
4. Re-run until passing
|
|
113
|
-
|
|
114
|
-
**If `{save_mode}` = true:** Log each result
|
|
33
|
+
When a broad check fails and causality is unclear, compare against the pre-task revision or use targeted diagnostics that preserve user changes. Classify each result:
|
|
115
34
|
|
|
116
|
-
|
|
35
|
+
| Status | Meaning |
|
|
36
|
+
|---|---|
|
|
37
|
+
| PASS | Check ran and passed on the current intended state |
|
|
38
|
+
| FAIL_INTRODUCED | Current APEX changes caused the failure |
|
|
39
|
+
| FAIL_PREEXISTING | Failure is reproduced outside the intended change or predates it |
|
|
40
|
+
| FAIL_UNRELATED | Failure belongs to unrelated local changes or an out-of-scope area |
|
|
41
|
+
| UNAVAILABLE | Required service, dependency, credential, command, or environment is absent |
|
|
42
|
+
| NOT_RUN | Check was intentionally omitted with a concrete reason |
|
|
117
43
|
|
|
118
|
-
|
|
44
|
+
Never turn `UNAVAILABLE`, `NOT_RUN`, or an unproven baseline inference into PASS.
|
|
119
45
|
|
|
120
|
-
|
|
121
|
-
- [ ] All todos from step-03 marked complete
|
|
122
|
-
- [ ] No tasks skipped without reason
|
|
123
|
-
- [ ] Any blocked tasks have explanation
|
|
46
|
+
## 4. Resolve introduced failures
|
|
124
47
|
|
|
125
|
-
|
|
126
|
-
- [ ] All existing tests pass
|
|
127
|
-
- [ ] New tests written for new functionality
|
|
128
|
-
- [ ] No skipped tests without reason
|
|
48
|
+
Fix `FAIL_INTRODUCED` within scope and re-run every invalidated check. Do not repair pre-existing or unrelated failures unless the user expands scope.
|
|
129
49
|
|
|
130
|
-
|
|
131
|
-
- [ ] Each AC demonstrably met
|
|
132
|
-
- [ ] Can explain how implementation satisfies AC
|
|
133
|
-
- [ ] Edge cases considered
|
|
50
|
+
If a failure exposes a flawed plan or interface, record a re-plan event and return to execution.
|
|
134
51
|
|
|
135
|
-
|
|
136
|
-
- [ ] Code follows existing patterns
|
|
137
|
-
- [ ] Error handling consistent
|
|
138
|
-
- [ ] Naming conventions match
|
|
52
|
+
## 5. Record validation ledger
|
|
139
53
|
|
|
140
|
-
|
|
54
|
+
For each check, record command/tool, environment, timestamp, revision, exit status, concise result, classification, and artifact path when useful.
|
|
141
55
|
|
|
142
|
-
If format command available:
|
|
143
56
|
```bash
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
### 6. Final Verification
|
|
148
|
-
|
|
149
|
-
Re-run all checks:
|
|
150
|
-
```bash
|
|
151
|
-
pnpm run typecheck && pnpm run lint
|
|
152
|
-
```
|
|
153
|
-
|
|
154
|
-
Both MUST pass.
|
|
155
|
-
|
|
156
|
-
### 7. Present Validation Results
|
|
157
|
-
|
|
158
|
-
```
|
|
159
|
-
**Validation Complete**
|
|
160
|
-
|
|
161
|
-
**Typecheck:** ✓ Passed
|
|
162
|
-
**Lint:** ✓ Passed
|
|
163
|
-
**Tests:** ✓ {X}/{X} passing
|
|
164
|
-
**Format:** ✓ Applied
|
|
165
|
-
|
|
166
|
-
**Acceptance Criteria:**
|
|
167
|
-
- [✓] AC1: Verified by [how]
|
|
168
|
-
- [✓] AC2: Verified by [how]
|
|
169
|
-
|
|
170
|
-
**Files Modified:** {list}
|
|
171
|
-
|
|
172
|
-
**Summary:** All checks passing, ready for next step.
|
|
173
|
-
```
|
|
174
|
-
|
|
175
|
-
### 8. Determine Next Step
|
|
176
|
-
|
|
177
|
-
**Decision tree:**
|
|
178
|
-
|
|
179
|
-
```
|
|
180
|
-
IF {test_mode} = true:
|
|
181
|
-
→ Load step-07-tests.md (test analysis and creation)
|
|
182
|
-
|
|
183
|
-
ELSE IF {examine_mode} = true:
|
|
184
|
-
→ Load step-05-examine.md (adversarial review)
|
|
185
|
-
|
|
186
|
-
ELSE IF {verify_mode} = true:
|
|
187
|
-
→ Load step-10-verify.md (feature verification)
|
|
188
|
-
|
|
189
|
-
ELSE IF {auto_mode} = false:
|
|
190
|
-
→ Ask user:
|
|
191
|
-
```
|
|
192
|
-
|
|
193
|
-
```yaml
|
|
194
|
-
questions:
|
|
195
|
-
- header: "Next"
|
|
196
|
-
question: "Validation complete. What would you like to do?"
|
|
197
|
-
options:
|
|
198
|
-
- label: "Run adversarial review"
|
|
199
|
-
description: "Deep review for security, logic, and quality"
|
|
200
|
-
- label: "Verify feature"
|
|
201
|
-
description: "Launch app and test feature works"
|
|
202
|
-
- label: "Complete workflow"
|
|
203
|
-
description: "Skip review and finalize"
|
|
204
|
-
- label: "Add tests"
|
|
205
|
-
description: "Create additional tests first"
|
|
206
|
-
multiSelect: false
|
|
207
|
-
```
|
|
208
|
-
|
|
209
|
-
```
|
|
210
|
-
ELSE:
|
|
211
|
-
→ Complete workflow (show final summary)
|
|
57
|
+
python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase validate --status complete --message "Validation ledger classified"
|
|
58
|
+
python3 "{skill_dir}/scripts/apex-state.py" checkpoint --root "$PWD" --run-id "{run_id}" --phase validate --message "Integrated validation checkpoint"
|
|
212
59
|
```
|
|
213
60
|
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
**If `{save_mode}` = true:**
|
|
217
|
-
|
|
218
|
-
Append to `{output_dir}/04-validate.md`:
|
|
219
|
-
```markdown
|
|
220
|
-
---
|
|
221
|
-
## Step Complete
|
|
222
|
-
**Status:** ✓ Complete
|
|
223
|
-
**Typecheck:** ✓
|
|
224
|
-
**Lint:** ✓
|
|
225
|
-
**Tests:** ✓
|
|
226
|
-
**Next:** {next step based on flags}
|
|
227
|
-
**Timestamp:** {ISO timestamp}
|
|
228
|
-
```
|
|
229
|
-
|
|
230
|
-
---
|
|
231
|
-
|
|
232
|
-
## SUCCESS METRICS:
|
|
233
|
-
|
|
234
|
-
✅ Typecheck passes
|
|
235
|
-
✅ Lint passes
|
|
236
|
-
✅ All tests pass
|
|
237
|
-
✅ All AC verified
|
|
238
|
-
✅ Code formatted
|
|
239
|
-
✅ User informed of status
|
|
240
|
-
|
|
241
|
-
## FAILURE MODES:
|
|
242
|
-
|
|
243
|
-
❌ Claiming checks pass when they don't
|
|
244
|
-
❌ Not running all validation commands
|
|
245
|
-
❌ Skipping tests for modified code
|
|
246
|
-
❌ Missing AC verification
|
|
247
|
-
❌ Proceeding with failures
|
|
248
|
-
❌ **CRITICAL**: Not using AskUserQuestion for next step
|
|
249
|
-
|
|
250
|
-
## VALIDATION PROTOCOLS:
|
|
251
|
-
|
|
252
|
-
- Run EVERY validation command
|
|
253
|
-
- Fix failures IMMEDIATELY
|
|
254
|
-
- Don't proceed until all green
|
|
255
|
-
- Verify EACH acceptance criterion
|
|
256
|
-
- Document all results
|
|
257
|
-
|
|
258
|
-
---
|
|
259
|
-
|
|
260
|
-
## NEXT STEP:
|
|
261
|
-
|
|
262
|
-
Based on flags (check in order):
|
|
263
|
-
- **If test_mode:** Load `./step-07-tests.md`
|
|
264
|
-
- **If examine_mode OR user requests:** Load `./step-05-examine.md`
|
|
265
|
-
- **If verify_mode:** Load `./step-10-verify.md` to verify feature
|
|
266
|
-
- **If pr_mode:** Load `./step-09-finish.md` to create pull request
|
|
267
|
-
- **Otherwise:** Workflow complete - show summary
|
|
61
|
+
## Routing
|
|
268
62
|
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
If
|
|
272
|
-
|
|
63
|
+
- If `{test_authoring}=on`, load `step-07-tests.md` when new tests remain to be authored.
|
|
64
|
+
- If `{test_authoring}=off`, do not author new tests; report material coverage gaps precisely.
|
|
65
|
+
- If `{test_authoring}=risk-based`, load `step-07-tests.md` only for an evidence-backed material gap.
|
|
66
|
+
- Load `step-05-examine.md` for adversarial or risk-required review.
|
|
67
|
+
- Load `step-10-verify.md` when runtime proof is required and review requirements are already satisfied.
|
|
68
|
+
- Otherwise continue to `step-09-finish.md`.
|