aiblueprint-cli 1.4.98 → 1.4.100
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +17 -1
- package/agents-config/skills/agents-manager/SKILL.md +2 -2
- package/agents-config/skills/agents-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/agents-manager/assets/codex-icon.svg +20 -0
- package/agents-config/skills/apex/SKILL.md +120 -118
- package/agents-config/skills/apex/agents/openai.yaml +10 -0
- package/agents-config/skills/apex/assets/codex-icon.svg +15 -0
- package/agents-config/skills/apex/scripts/apex-state.py +740 -0
- package/agents-config/skills/apex/scripts/setup-templates.sh +27 -145
- package/agents-config/skills/apex/scripts/test_apex_state.py +413 -0
- package/agents-config/skills/apex/scripts/update-progress.sh +17 -73
- package/agents-config/skills/apex/steps/step-00-init.md +85 -231
- package/agents-config/skills/apex/steps/step-00b-branch.md +10 -118
- package/agents-config/skills/apex/steps/step-00b-economy.md +12 -239
- package/agents-config/skills/apex/steps/step-00b-interactive.md +13 -162
- package/agents-config/skills/apex/steps/step-00b-save.md +13 -114
- package/agents-config/skills/apex/steps/step-01-analyze.md +40 -361
- package/agents-config/skills/apex/steps/step-02-plan.md +55 -562
- package/agents-config/skills/apex/steps/step-02b-tasks.md +15 -291
- package/agents-config/skills/apex/steps/step-03-execute-teams.md +47 -267
- package/agents-config/skills/apex/steps/step-03-execute.md +32 -212
- package/agents-config/skills/apex/steps/step-04-validate.md +42 -246
- package/agents-config/skills/apex/steps/step-05-examine.md +47 -371
- package/agents-config/skills/apex/steps/step-06-resolve.md +19 -221
- package/agents-config/skills/apex/steps/step-07-tests.md +19 -234
- package/agents-config/skills/apex/steps/step-08-run-tests.md +13 -300
- package/agents-config/skills/apex/steps/step-09-finish.md +36 -200
- package/agents-config/skills/apex/steps/step-10-verify.md +46 -264
- package/agents-config/skills/appstore-connect/agents/openai.yaml +7 -0
- package/agents-config/skills/appstore-connect/assets/codex-icon.svg +17 -0
- package/agents-config/skills/commit/agents/openai.yaml +10 -0
- package/agents-config/skills/commit/assets/codex-icon.svg +17 -0
- package/agents-config/skills/create-pr/agents/openai.yaml +10 -0
- package/agents-config/skills/create-pr/assets/codex-icon.svg +17 -0
- package/agents-config/skills/environments-manager/SKILL.md +1 -1
- package/agents-config/skills/environments-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/environments-manager/assets/codex-icon.svg +16 -0
- package/agents-config/skills/environments-manager/examples/scripts/claude-worktree-remove.sh +19 -3
- package/agents-config/skills/environments-manager/examples/scripts/worktree-up.sh +1 -1
- package/agents-config/skills/environments-manager/references/claude.md +1 -1
- package/agents-config/skills/fix-pr-comments/agents/openai.yaml +10 -0
- package/agents-config/skills/fix-pr-comments/assets/codex-icon.svg +17 -0
- package/agents-config/skills/grill-me/SKILL.md +25 -4
- package/agents-config/skills/grill-me/agents/openai.yaml +8 -0
- package/agents-config/skills/grill-me/assets/codex-icon.svg +16 -0
- package/agents-config/skills/hooks-manager/SKILL.md +19 -9
- package/agents-config/skills/hooks-manager/assets/codex-icon.svg +15 -4
- package/agents-config/skills/hooks-manager/references/claude-code.md +32 -0
- package/agents-config/skills/hooks-manager/references/codex.md +23 -0
- package/agents-config/skills/hooks-manager/references/cursor.md +18 -0
- package/agents-config/skills/hooks-manager/references/hook-types.md +5 -3
- package/agents-config/skills/hooks-manager/references/input-output-schemas.md +2 -2
- package/agents-config/skills/hooks-manager/references/research-sources.md +25 -0
- package/agents-config/skills/hooks-manager/references/router.md +32 -0
- package/agents-config/skills/hooks-manager/references/troubleshooting.md +3 -3
- package/agents-config/skills/merge/agents/openai.yaml +10 -0
- package/agents-config/skills/merge/assets/codex-icon.svg +17 -0
- package/agents-config/skills/oneshot/SKILL.md +4 -0
- package/agents-config/skills/oneshot/agents/openai.yaml +10 -0
- package/agents-config/skills/oneshot/assets/codex-icon.svg +18 -0
- package/agents-config/skills/prompt-creator/agents/openai.yaml +7 -0
- package/agents-config/skills/prompt-creator/assets/codex-icon.svg +16 -0
- package/agents-config/skills/rules-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/rules-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/SKILL.md +45 -3
- package/agents-config/skills/skill-manager/agents/openai.yaml +7 -0
- package/agents-config/skills/skill-manager/assets/codex-icon.svg +23 -0
- package/agents-config/skills/skill-manager/references/skill-writing-glossary.md +201 -0
- package/agents-config/skills/skill-manager/scripts/setup-codex-icons.ts +143 -0
- package/agents-config/skills/ultrathink/agents/openai.yaml +10 -0
- package/agents-config/skills/ultrathink/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-artifacts/SKILL.md +102 -51
- package/agents-config/skills/use-artifacts/assets/local-runtime.js +299 -0
- package/agents-config/skills/use-artifacts/scripts/create_artifact.py +1 -1
- package/agents-config/skills/use-delegate/SKILL.md +4 -0
- package/agents-config/skills/use-delegate/agents/openai.yaml +10 -0
- package/agents-config/skills/use-delegate/assets/codex-icon.svg +20 -0
- package/agents-config/skills/use-goal/SKILL.md +70 -9
- package/agents-config/skills/use-goal/agents/openai.yaml +1 -1
- package/agents-config/skills/use-goal/assets/codex-icon.svg +17 -3
- package/agents-config/skills/use-goal/references/claude-code-goal.md +54 -6
- package/agents-config/skills/use-goal/references/codex-goal.md +59 -4
- package/agents-config/skills/use-goal/references/verification-harnesses.md +104 -3
- package/dist/cli.js +264 -12
- package/package.json +1 -1
- package/agents-config/skills/apex/templates/00-context.md +0 -55
- package/agents-config/skills/apex/templates/01-analyze.md +0 -10
- package/agents-config/skills/apex/templates/02-plan.md +0 -10
- package/agents-config/skills/apex/templates/03-execute.md +0 -10
- package/agents-config/skills/apex/templates/04-validate.md +0 -10
- package/agents-config/skills/apex/templates/05-examine.md +0 -10
- package/agents-config/skills/apex/templates/06-resolve.md +0 -10
- package/agents-config/skills/apex/templates/07-tests.md +0 -10
- package/agents-config/skills/apex/templates/08-run-tests.md +0 -10
- package/agents-config/skills/apex/templates/09-finish.md +0 -10
- package/agents-config/skills/apex/templates/10-verify.md +0 -9
- package/agents-config/skills/apex/templates/README.md +0 -195
- package/agents-config/skills/apex/templates/step-complete.md +0 -7
|
@@ -1,296 +1,78 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: step-10-verify
|
|
3
|
-
description:
|
|
4
|
-
|
|
5
|
-
next_step: steps/step-09-finish.md
|
|
3
|
+
description: Prove required APEX acceptance criteria through the real user, API, provider, artifact, or deployment surface with current evidence.
|
|
4
|
+
next_step: step-09-finish.md
|
|
6
5
|
---
|
|
7
6
|
|
|
8
|
-
# Step 10:
|
|
7
|
+
# Step 10: Runtime proof
|
|
9
8
|
|
|
10
|
-
|
|
9
|
+
Runtime proof is a hard gate when requested by the user, required by project rules, or selected by risk. Tests and code inspection support proof but do not replace a stronger required surface.
|
|
11
10
|
|
|
12
|
-
|
|
13
|
-
- 🛑 NEVER claim feature works without actually testing it
|
|
14
|
-
- 📐 ALWAYS read the project's recommended verification rules FIRST (`.agents/rules/`, `AGENTS.md`, `CLAUDE.md`) and follow them
|
|
15
|
-
- ✅ ALWAYS launch a verifier agent to test through real surface
|
|
16
|
-
- ✅ ALWAYS compare results against original user request
|
|
17
|
-
- 📸 ALWAYS capture 1+ screenshots proving the feature works - verification is INCOMPLETE without visual proof
|
|
18
|
-
- ✅ ALWAYS fix issues before proceeding (unless user skips)
|
|
19
|
-
- 📋 YOU ARE A VERIFIER, not an implementer
|
|
20
|
-
- 💬 FOCUS on "Does the feature actually work for a real user?"
|
|
21
|
-
- 🚫 FORBIDDEN to proceed with FAIL verdict without user confirmation
|
|
11
|
+
## 1. Set proof state
|
|
22
12
|
|
|
23
|
-
|
|
13
|
+
Set `{proof_gate}=NOT_PROVEN`. Record environment, revision, target surface, authentication state, fixture identity, and evidence directory.
|
|
24
14
|
|
|
25
|
-
|
|
26
|
-
- 💾 Log verification results (if save_mode)
|
|
27
|
-
- 📖 Compare each AC against real behavior
|
|
28
|
-
- 🚫 FORBIDDEN to fake verification
|
|
15
|
+
Read project verification rules and use the approved local server, browser, simulator, CLI, API, provider, or release workflow. Reuse healthy managed services instead of starting duplicates.
|
|
29
16
|
|
|
30
|
-
##
|
|
17
|
+
## 2. Build the proof matrix
|
|
31
18
|
|
|
32
|
-
|
|
33
|
-
- All typecheck/lint/tests pass
|
|
34
|
-
- Now testing the REAL user experience
|
|
35
|
-
- **If `{teams_mode}` = true:** Agent team is still alive. Do NOT shutdown teammates - that happens in step-09-finish only.
|
|
19
|
+
Create one row per observable contract:
|
|
36
20
|
|
|
37
|
-
|
|
21
|
+
| ID | Acceptance criterion | Starting state | Action | Expected result | Evidence layer | Artifact | Status |
|
|
22
|
+
|---|---|---|---|---|---|---|---|
|
|
38
23
|
|
|
39
|
-
|
|
24
|
+
Include the initial state, meaningful transitions, final outcome, and relevant negative, persistence, refresh, permission, or regression paths implied by the request.
|
|
40
25
|
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
<available_state>
|
|
44
|
-
From previous steps:
|
|
45
|
-
|
|
46
|
-
| Variable | Description |
|
|
47
|
-
|----------|-------------|
|
|
48
|
-
| `{task_description}` | Original user request |
|
|
49
|
-
| `{task_id}` | Kebab-case identifier |
|
|
50
|
-
| `{acceptance_criteria}` | Success criteria from analysis |
|
|
51
|
-
| `{auto_mode}` | Skip confirmations |
|
|
52
|
-
| `{save_mode}` | Save outputs to files |
|
|
53
|
-
| `{economy_mode}` | No subagents |
|
|
54
|
-
| `{output_dir}` | Path to output (if save_mode) |
|
|
55
|
-
| Files modified | From step-03 |
|
|
56
|
-
</available_state>
|
|
57
|
-
|
|
58
|
-
---
|
|
59
|
-
|
|
60
|
-
## EXECUTION SEQUENCE:
|
|
61
|
-
|
|
62
|
-
### 1. Initialize Save Output (if save_mode)
|
|
63
|
-
|
|
64
|
-
**If `{save_mode}` = true:**
|
|
65
|
-
|
|
66
|
-
```bash
|
|
67
|
-
bash {skill_dir}/scripts/update-progress.sh "{task_id}" "10" "verify" "in_progress"
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
Append results to `{output_dir}/10-verify.md` as you work.
|
|
71
|
-
|
|
72
|
-
### 2. Gather Context for Verifier
|
|
73
|
-
|
|
74
|
-
Collect all information the verifier agent needs:
|
|
75
|
-
|
|
76
|
-
```
|
|
77
|
-
1. Original task description: {task_description}
|
|
78
|
-
2. Acceptance criteria: {acceptance_criteria}
|
|
79
|
-
3. Modified files (from git):
|
|
80
|
-
git diff --name-only HEAD~1
|
|
81
|
-
git status --porcelain
|
|
82
|
-
4. Implementation summary from step-03/04
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
### 2b. Load Recommended Verification Rules
|
|
86
|
-
|
|
87
|
-
Look for project-specific verification conventions and **follow them**. They override the defaults below. Check in order:
|
|
88
|
-
|
|
89
|
-
1. `.agents/rules/` — any rule file about verification, testing, QA, or e2e (e.g. `verify.md`, `testing.md`)
|
|
90
|
-
2. `AGENTS.md` — verification / commands sections
|
|
91
|
-
3. `CLAUDE.md` — verification preferences (e.g. "Prefer /dev-browser for verification")
|
|
92
|
-
4. `README.md` + `package.json` scripts — how to launch / test
|
|
93
|
-
|
|
94
|
-
```bash
|
|
95
|
-
ls .agents/rules/ 2>/dev/null
|
|
96
|
-
cat AGENTS.md CLAUDE.md 2>/dev/null
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
Capture the recommended launch command, auth/login steps, test URLs/credentials, and any required evidence into `{verification_rules}` and pass it to the verifier. If nothing is documented, fall back to detecting the dev/start command from `package.json`.
|
|
100
|
-
|
|
101
|
-
### 3. Launch Verifier Agent
|
|
102
|
-
|
|
103
|
-
**If `{economy_mode}` = true:**
|
|
104
|
-
-> Self-verify: manually test the feature using available tools (dev-browser, curl, shell commands). FOLLOW the recommended verification rules from step 2b, capture 1+ screenshots as proof, and include the screenshots directly in the chat using Markdown image syntax with absolute local paths. Follow the verification process below directly instead of launching an agent.
|
|
105
|
-
|
|
106
|
-
**If `{economy_mode}` = false:**
|
|
107
|
-
-> Launch a verifier sub-agent:
|
|
108
|
-
|
|
109
|
-
```
|
|
110
|
-
Sub-agent:
|
|
111
|
-
profile/type: "verifier"
|
|
112
|
-
prompt: |
|
|
113
|
-
## Feature Verification
|
|
114
|
-
|
|
115
|
-
**Original Request:** {task_description}
|
|
116
|
-
|
|
117
|
-
**Acceptance Criteria:**
|
|
118
|
-
{acceptance_criteria - list each one}
|
|
119
|
-
|
|
120
|
-
**Files Modified:**
|
|
121
|
-
{list of modified files}
|
|
122
|
-
|
|
123
|
-
**Implementation Summary:**
|
|
124
|
-
{brief summary of what was done}
|
|
26
|
+
Choose evidence that matches the surface:
|
|
125
27
|
|
|
126
|
-
|
|
127
|
-
|
|
28
|
+
- visual step: current screenshot;
|
|
29
|
+
- CLI/API: raw command or response artifact;
|
|
30
|
+
- persistence: reload, relaunch, or authoritative state read-back;
|
|
31
|
+
- provider: provider/API read-back, not local configuration alone;
|
|
32
|
+
- public artifact/deployment: independently fetch the public surface or artifact;
|
|
33
|
+
- authenticated live flow: controlled real interaction and final observable result.
|
|
128
34
|
|
|
129
|
-
|
|
130
|
-
1. FOLLOW the recommended verification rules above (launch command, auth, URLs)
|
|
131
|
-
2. Launch the app (rules first, else package.json dev/start command)
|
|
132
|
-
3. Navigate to the relevant pages/endpoints
|
|
133
|
-
4. Test each acceptance criterion through real interaction
|
|
134
|
-
5. Use /dev-browser for web UI (capture screenshots), curl for APIs, shell for CLI tools
|
|
135
|
-
6. 📸 MANDATORY: Capture 1+ screenshots proving the finished feature works. For web UI use dev-browser saveScreenshot; for non-visual surfaces capture the terminal output / response as evidence
|
|
136
|
-
7. Output the screenshot(s) directly in the chat using Markdown image syntax with absolute local paths, e.g. ``. Do NOT only list paths.
|
|
137
|
-
8. Report your findings with PASS/FAIL for each AC, listing the screenshot paths below the inline images
|
|
138
|
-
9. List anything missing compared to the original request
|
|
139
|
-
```
|
|
140
|
-
|
|
141
|
-
### 4. Process Verification Results
|
|
142
|
-
|
|
143
|
-
Parse the verifier agent's report:
|
|
35
|
+
## 3. Use an independent verifier when valuable
|
|
144
36
|
|
|
145
|
-
|
|
146
|
-
-> Verification is INCOMPLETE. Re-run capture (or send the verifier back) before accepting any PASS verdict.
|
|
37
|
+
A fresh verifier context is useful for material user-facing, high-risk, or disputed flows. Give it the original request, acceptance criteria, verification rules, current revision, and proof matrix. The coordinator inspects every returned artifact before accepting it.
|
|
147
38
|
|
|
148
|
-
|
|
149
|
-
-> Display success summary and proceed
|
|
39
|
+
## 4. Exercise the real flow
|
|
150
40
|
|
|
151
|
-
|
|
152
|
-
-> Display failure details
|
|
41
|
+
For each row:
|
|
153
42
|
|
|
154
|
-
|
|
43
|
+
1. Establish the documented starting state.
|
|
44
|
+
2. Perform the action through the intended surface.
|
|
45
|
+
3. Wait for and inspect the observable result.
|
|
46
|
+
4. Check relevant errors, failed requests, crashes, logs, and persistent state.
|
|
47
|
+
5. Capture evidence immediately with ordered artifact names.
|
|
48
|
+
6. Record timestamp, environment, revision, action, observed result, and artifact path.
|
|
49
|
+
7. Mark PASS only when the expected result is directly visible in current evidence.
|
|
155
50
|
|
|
156
|
-
|
|
157
|
-
## Verification Results
|
|
51
|
+
Do not reuse evidence invalidated by a later code, configuration, environment, or data change.
|
|
158
52
|
|
|
159
|
-
|
|
53
|
+
## 5. Evaluate and continue
|
|
160
54
|
|
|
161
|
-
|
|
162
|
-
|----|--------|---------|
|
|
163
|
-
| AC1 | ✓ PASS | Working as expected |
|
|
164
|
-
| AC2 | ✗ FAIL | {what went wrong} |
|
|
55
|
+
Set `{proof_gate}=PASS` only when all required criteria and rows pass, all artifacts exist and are current, and no observed error invalidates the flow.
|
|
165
56
|
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
**Verdict:** {PASS or FAIL}
|
|
169
|
-
```
|
|
57
|
+
While the gate is not PASS:
|
|
170
58
|
|
|
171
|
-
|
|
59
|
+
- identify the exact missing proof or failing behavior;
|
|
60
|
+
- use the shortest real feedback loop to diagnose it;
|
|
61
|
+
- return to planning or execution for in-scope fixes;
|
|
62
|
+
- re-run affected validation and independent review;
|
|
63
|
+
- reset the verification state and recapture every invalidated row.
|
|
172
64
|
|
|
173
|
-
|
|
174
|
-
-> Proceed to next step
|
|
65
|
+
There is no arbitrary retry limit while attempts produce meaningful progress. If a genuine external dependency blocks progress after safe alternatives are exhausted, report `BLOCKED — NOT PROVEN` with the exact condition and required input. Never relabel it as completion.
|
|
175
66
|
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
**If `{auto_mode}` = true:**
|
|
179
|
-
-> Auto-fix: Apply fixes for each failing AC, then re-verify (max 2 retry rounds)
|
|
180
|
-
|
|
181
|
-
**If `{auto_mode}` = false:**
|
|
182
|
-
|
|
183
|
-
```yaml
|
|
184
|
-
questions:
|
|
185
|
-
- header: "Verify"
|
|
186
|
-
question: "Verification found issues. How would you like to proceed?"
|
|
187
|
-
options:
|
|
188
|
-
- label: "Fix and re-verify (Recommended)"
|
|
189
|
-
description: "Apply fixes for failing ACs and verify again"
|
|
190
|
-
- label: "Fix without re-verify"
|
|
191
|
-
description: "Apply fixes but skip second verification"
|
|
192
|
-
- label: "Skip verification"
|
|
193
|
-
description: "Accept current state, proceed anyway"
|
|
194
|
-
- label: "Discuss issues"
|
|
195
|
-
description: "Let me review the findings first"
|
|
196
|
-
multiSelect: false
|
|
197
|
-
```
|
|
67
|
+
## 6. Present evidence
|
|
198
68
|
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
```
|
|
202
|
-
max_verify_rounds = 3
|
|
203
|
-
round = 0
|
|
204
|
-
|
|
205
|
-
WHILE round < max_verify_rounds AND verdict = FAIL:
|
|
206
|
-
round += 1
|
|
207
|
-
|
|
208
|
-
1. Read each failing AC
|
|
209
|
-
2. Identify the root cause
|
|
210
|
-
3. Apply the fix
|
|
211
|
-
4. Run typecheck + lint to ensure no regressions
|
|
212
|
-
5. Re-launch verifier agent (or self-verify if economy_mode)
|
|
213
|
-
6. Update verdict
|
|
214
|
-
```
|
|
215
|
-
|
|
216
|
-
**If still failing after 3 rounds:**
|
|
217
|
-
|
|
218
|
-
```yaml
|
|
219
|
-
questions:
|
|
220
|
-
- header: "Stuck"
|
|
221
|
-
question: "Verification still failing after {round} fix rounds. How proceed?"
|
|
222
|
-
options:
|
|
223
|
-
- label: "Keep trying"
|
|
224
|
-
description: "Continue fixing"
|
|
225
|
-
- label: "Accept current state"
|
|
226
|
-
description: "Proceed with known issues"
|
|
227
|
-
- label: "I'll fix manually"
|
|
228
|
-
description: "Let me investigate"
|
|
229
|
-
multiSelect: false
|
|
230
|
-
```
|
|
69
|
+
Show the proof matrix in flow order. Render visual artifacts inline with absolute local paths and link non-visual artifacts. State the exact local/static, provider, public-artifact/deployment, and authenticated-live boundaries proven.
|
|
231
70
|
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
**If `{save_mode}` = true:**
|
|
235
|
-
|
|
236
|
-
Append to `{output_dir}/10-verify.md`:
|
|
237
|
-
```markdown
|
|
238
|
-
---
|
|
239
|
-
## Step Complete
|
|
240
|
-
**Status:** ✓ Complete
|
|
241
|
-
**Verdict:** {PASS/FAIL}
|
|
242
|
-
**Rounds:** {count}
|
|
243
|
-
**Screenshots:** {inline Markdown image(s) first, then paths - at least 1}
|
|
244
|
-
**Failing ACs:** {list or "none"}
|
|
245
|
-
**Timestamp:** {ISO timestamp}
|
|
246
|
-
```
|
|
71
|
+
When `{proof_gate}=PASS`:
|
|
247
72
|
|
|
248
73
|
```bash
|
|
249
|
-
|
|
74
|
+
python3 "{skill_dir}/scripts/apex-state.py" event --root "$PWD" --run-id "{run_id}" --phase verify --status complete --message "Runtime proof gate passed"
|
|
75
|
+
python3 "{skill_dir}/scripts/apex-state.py" checkpoint --root "$PWD" --run-id "{run_id}" --phase verify --message "Current proof artifacts recorded"
|
|
250
76
|
```
|
|
251
77
|
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
## SUCCESS METRICS:
|
|
255
|
-
|
|
256
|
-
✅ Recommended verification rules followed
|
|
257
|
-
✅ Feature tested through real user surface
|
|
258
|
-
✅ Each AC verified with actual interaction
|
|
259
|
-
✅ 1+ screenshots captured as visual proof
|
|
260
|
-
✅ Failures identified with clear descriptions
|
|
261
|
-
✅ Fixes applied and re-verified (if needed)
|
|
262
|
-
✅ User informed of verification status
|
|
263
|
-
|
|
264
|
-
## FAILURE MODES:
|
|
265
|
-
|
|
266
|
-
❌ Claiming verification passed without actually testing
|
|
267
|
-
❌ Ignoring the project's recommended verification rules
|
|
268
|
-
❌ No screenshot / visual proof of the feature working
|
|
269
|
-
❌ Not launching the app
|
|
270
|
-
❌ Testing only happy path, ignoring ACs
|
|
271
|
-
❌ Auto-proceeding with FAIL verdict (without user OK)
|
|
272
|
-
❌ Not comparing against original user request
|
|
273
|
-
❌ **CRITICAL**: Not using AskUserQuestion for failure decisions
|
|
274
|
-
|
|
275
|
-
## VERIFICATION PROTOCOLS:
|
|
276
|
-
|
|
277
|
-
- Follow the project's recommended verification rules first (`.agents/rules/`, `AGENTS.md`, `CLAUDE.md`)
|
|
278
|
-
- Test through REAL user surface (browser, CLI, API)
|
|
279
|
-
- Capture 1+ screenshots as proof - no PASS verdict without evidence
|
|
280
|
-
- Compare against ORIGINAL request, not just code
|
|
281
|
-
- Fix and re-verify, don't just fix and hope
|
|
282
|
-
- Be honest about failures
|
|
283
|
-
|
|
284
|
-
---
|
|
285
|
-
|
|
286
|
-
## NEXT STEP:
|
|
287
|
-
|
|
288
|
-
Based on flags:
|
|
289
|
-
- **If pr_mode:** Load `./step-09-finish.md` to create pull request
|
|
290
|
-
- **Otherwise:** Workflow complete - show summary
|
|
291
|
-
|
|
292
|
-
<critical>
|
|
293
|
-
Remember: The whole point of verify is to catch things that code review and tests miss.
|
|
294
|
-
Test the feature like a real user would - don't just check the code!
|
|
295
|
-
If teams_mode is active: NEVER shutdown teammates - they stay alive until step-09-finish!
|
|
296
|
-
</critical>
|
|
78
|
+
Then load `step-09-finish.md`.
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Appstore Connect"
|
|
3
|
+
short_description: "Interact with App Store Connect via the asc CLI - apps,..."
|
|
4
|
+
icon_small: "./assets/codex-icon.svg"
|
|
5
|
+
icon_large: "./assets/codex-icon.svg"
|
|
6
|
+
brand_color: "#01C1C6"
|
|
7
|
+
default_prompt: "Use $appstore-connect to help with this task."
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="appstore-connect skill icon"
|
|
3
|
+
class="lucide lucide-cloud-upload"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<path d="M12 13v8" />
|
|
15
|
+
<path d="M4 14.899A7 7 0 1 1 15.71 8h1.79a4.5 4.5 0 0 1 2.5 8.242" />
|
|
16
|
+
<path d="m8 17 4-4 4 4" />
|
|
17
|
+
</svg>
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Commit"
|
|
3
|
+
short_description: "Quick commit and push with minimal, clean messages"
|
|
4
|
+
icon_small: "./assets/codex-icon.svg"
|
|
5
|
+
icon_large: "./assets/codex-icon.svg"
|
|
6
|
+
brand_color: "#4BD0FC"
|
|
7
|
+
default_prompt: "Use $commit to help with this task."
|
|
8
|
+
|
|
9
|
+
policy:
|
|
10
|
+
allow_implicit_invocation: false
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="commit skill icon"
|
|
3
|
+
class="lucide lucide-git-commit-horizontal"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<circle cx="12" cy="12" r="3" />
|
|
15
|
+
<line x1="3" x2="9" y1="12" y2="12" />
|
|
16
|
+
<line x1="15" x2="21" y1="12" y2="12" />
|
|
17
|
+
</svg>
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Create Pr"
|
|
3
|
+
short_description: "Create and push PR with auto-generated title and description"
|
|
4
|
+
icon_small: "./assets/codex-icon.svg"
|
|
5
|
+
icon_large: "./assets/codex-icon.svg"
|
|
6
|
+
brand_color: "#951556"
|
|
7
|
+
default_prompt: "Use $create-pr to help with this task."
|
|
8
|
+
|
|
9
|
+
policy:
|
|
10
|
+
allow_implicit_invocation: false
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="create-pr skill icon"
|
|
3
|
+
class="lucide lucide-git-merge"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<circle cx="18" cy="18" r="3" />
|
|
15
|
+
<circle cx="6" cy="6" r="3" />
|
|
16
|
+
<path d="M6 21V9a9 9 0 0 0 9 9" />
|
|
17
|
+
</svg>
|
|
@@ -108,7 +108,7 @@ Base template, before backend-specific additions:
|
|
|
108
108
|
set -euo pipefail
|
|
109
109
|
|
|
110
110
|
WORKTREE_PATH="${CODEX_WORKTREE_PATH:-${CURSOR_WORKTREE_PATH:-$(pwd)}}"
|
|
111
|
-
SOURCE_PATH="${CODEX_SOURCE_TREE_PATH:-${ROOT_WORKTREE_PATH:-$HOME/Developer/
|
|
111
|
+
SOURCE_PATH="${CODEX_SOURCE_TREE_PATH:-${ROOT_WORKTREE_PATH:-$HOME/Developer/projects/<repo>}}"
|
|
112
112
|
|
|
113
113
|
cd "$WORKTREE_PATH"
|
|
114
114
|
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Environments Manager"
|
|
3
|
+
short_description: "Set up per-worktree environments for Claude Code, Cursor, or..."
|
|
4
|
+
icon_small: "./assets/codex-icon.svg"
|
|
5
|
+
icon_large: "./assets/codex-icon.svg"
|
|
6
|
+
brand_color: "#444449"
|
|
7
|
+
default_prompt: "Use $environments-manager to help with this task."
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="environments-manager skill icon"
|
|
3
|
+
class="lucide lucide-settings"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<path d="M9.671 4.136a2.34 2.34 0 0 1 4.659 0 2.34 2.34 0 0 0 3.319 1.915 2.34 2.34 0 0 1 2.33 4.033 2.34 2.34 0 0 0 0 3.831 2.34 2.34 0 0 1-2.33 4.033 2.34 2.34 0 0 0-3.319 1.915 2.34 2.34 0 0 1-4.659 0 2.34 2.34 0 0 0-3.32-1.915 2.34 2.34 0 0 1-2.33-4.033 2.34 2.34 0 0 0 0-3.831A2.34 2.34 0 0 1 6.35 6.051a2.34 2.34 0 0 0 3.319-1.915" />
|
|
15
|
+
<circle cx="12" cy="12" r="3" />
|
|
16
|
+
</svg>
|
package/agents-config/skills/environments-manager/examples/scripts/claude-worktree-remove.sh
CHANGED
|
@@ -13,7 +13,7 @@ CLEANUP_TIMEOUT_SEC="${CLAUDE_WORKTREE_CLEANUP_TIMEOUT:-300}"
|
|
|
13
13
|
|
|
14
14
|
[[ ! -d "$WORKTREE" ]] && exit 0
|
|
15
15
|
|
|
16
|
-
LOG_FILE="$
|
|
16
|
+
LOG_FILE="${TMPDIR:-/tmp}/claude-worktree-cleanup-$(basename "$WORKTREE").log"
|
|
17
17
|
: > "$LOG_FILE" 2>/dev/null || LOG_FILE=""
|
|
18
18
|
|
|
19
19
|
log_to() {
|
|
@@ -23,6 +23,18 @@ log_to() {
|
|
|
23
23
|
fi
|
|
24
24
|
}
|
|
25
25
|
|
|
26
|
+
assert_clean_worktree() {
|
|
27
|
+
local changes
|
|
28
|
+
changes=$(git -C "$WORKTREE" status --porcelain --untracked-files=all 2>/dev/null || true)
|
|
29
|
+
if [[ -n "$changes" ]]; then
|
|
30
|
+
log_to "[claude-worktree-remove] REFUSED: worktree has uncommitted changes"
|
|
31
|
+
printf '%s\n' "$changes" >&2
|
|
32
|
+
exit 1
|
|
33
|
+
fi
|
|
34
|
+
}
|
|
35
|
+
|
|
36
|
+
assert_clean_worktree
|
|
37
|
+
|
|
26
38
|
# Cleanup may call CLIs (e.g. convex env remove) that need the project's Node version.
|
|
27
39
|
if [[ -s "$HOME/.nvm/nvm.sh" ]]; then
|
|
28
40
|
# shellcheck disable=SC1090,SC1091
|
|
@@ -61,6 +73,10 @@ if [[ -x "${REPO}/scripts/worktree-down.sh" ]]; then
|
|
|
61
73
|
wait "$WATCHDOG_PID" 2>/dev/null || true
|
|
62
74
|
fi
|
|
63
75
|
|
|
76
|
+
assert_clean_worktree
|
|
64
77
|
BRANCH=$(git -C "$WORKTREE" rev-parse --abbrev-ref HEAD 2>/dev/null || true)
|
|
65
|
-
git -C "$REPO" worktree remove "$WORKTREE"
|
|
66
|
-
[[ -n "$BRANCH" && "$BRANCH" != "HEAD" ]]
|
|
78
|
+
git -C "$REPO" worktree remove "$WORKTREE"
|
|
79
|
+
if [[ -n "$BRANCH" && "$BRANCH" != "HEAD" ]]; then
|
|
80
|
+
git -C "$REPO" branch -d "$BRANCH" 2>/dev/null || \
|
|
81
|
+
log_to "[claude-worktree-remove] Kept unmerged branch: $BRANCH"
|
|
82
|
+
fi
|
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
set -euo pipefail
|
|
6
6
|
|
|
7
7
|
WORKTREE_PATH="${CODEX_WORKTREE_PATH:-${CURSOR_WORKTREE_PATH:-$(pwd)}}"
|
|
8
|
-
SOURCE_PATH="${CODEX_SOURCE_TREE_PATH:-${ROOT_WORKTREE_PATH:-$HOME/Developer/
|
|
8
|
+
SOURCE_PATH="${CODEX_SOURCE_TREE_PATH:-${ROOT_WORKTREE_PATH:-$HOME/Developer/projects/myapp}}"
|
|
9
9
|
|
|
10
10
|
cd "$WORKTREE_PATH"
|
|
11
11
|
|
|
@@ -99,7 +99,7 @@ These wrappers exist because Claude's `WorktreeCreate` hook has Claude-specific
|
|
|
99
99
|
5. Run `scripts/worktree-up.sh` with `ROOT_WORKTREE_PATH`, `CLAUDE_WORKTREE_NAME`, and `CI=1` set, under a watchdog that hard-kills after `CLAUDE_WORKTREE_SETUP_TIMEOUT` (default 900s).
|
|
100
100
|
6. Echo the worktree path on stdout — and absolutely nothing else.
|
|
101
101
|
|
|
102
|
-
The remove wrapper mirrors steps 3–5 for cleanup, then runs `git worktree remove
|
|
102
|
+
The remove wrapper mirrors steps 3–5 for cleanup, refuses a dirty worktree, then runs `git worktree remove` without force. It uses `git branch -d` and keeps any unmerged branch instead of destroying it.
|
|
103
103
|
|
|
104
104
|
`ROOT_WORKTREE_PATH` is set so the shared `scripts/worktree-up.sh` can locate the source checkout via the existing cascade (`CODEX_SOURCE_TREE_PATH` → `ROOT_WORKTREE_PATH`).
|
|
105
105
|
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
interface:
|
|
2
|
+
display_name: "Fix Pr Comments"
|
|
3
|
+
short_description: "Fetch PR review comments and implement all requested changes"
|
|
4
|
+
icon_small: "./assets/codex-icon.svg"
|
|
5
|
+
icon_large: "./assets/codex-icon.svg"
|
|
6
|
+
brand_color: "#C62EE2"
|
|
7
|
+
default_prompt: "Use $fix-pr-comments to help with this task."
|
|
8
|
+
|
|
9
|
+
policy:
|
|
10
|
+
allow_implicit_invocation: false
|
|
@@ -0,0 +1,17 @@
|
|
|
1
|
+
<!-- @license lucide-static v1.24.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="fix-pr-comments skill icon"
|
|
3
|
+
class="lucide lucide-git-merge"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<circle cx="18" cy="18" r="3" />
|
|
15
|
+
<circle cx="6" cy="6" r="3" />
|
|
16
|
+
<path d="M6 21V9a9 9 0 0 0 9 9" />
|
|
17
|
+
</svg>
|
|
@@ -1,10 +1,31 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: grill-me
|
|
3
|
-
description:
|
|
3
|
+
description: Grill the user in rapid batches to sharpen a plan, decision, or idea. Use when the user asks for Grill Me, wants a rigorous interview, or uses a grill trigger phrase.
|
|
4
|
+
disable-model-invocation: true
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
|
|
7
|
+
# Grill Me
|
|
7
8
|
|
|
8
|
-
|
|
9
|
+
Stress-test the user's thinking until both sides share a precise, actionable understanding.
|
|
9
10
|
|
|
10
|
-
|
|
11
|
+
## Interview loop
|
|
12
|
+
|
|
13
|
+
1. Inspect the available environment first. Resolve discoverable facts through files, tools, documentation, or other in-scope evidence instead of asking the user.
|
|
14
|
+
2. Open the session by recommending voice input: tell the user that answering with the microphone is fastest, and that short answers keyed `1` through `10` are enough.
|
|
15
|
+
3. Ask exactly 10 numbered questions in one batch. Prioritize the highest-leverage unresolved decisions and order them so earlier questions clarify later ones.
|
|
16
|
+
4. For every question, include a concise **Recommended answer** based on current evidence. Make the tradeoff or consequence clear enough for the user to accept, reject, or amend it quickly.
|
|
17
|
+
5. Invite one grouped reply covering `1` through `10`. Accept terse answers, corrections, skipped items, or a blanket acceptance of the recommendations.
|
|
18
|
+
6. After the reply, summarize what is established, call out contradictions or missing dependencies, and investigate any newly discoverable facts.
|
|
19
|
+
7. Ask the next batch of 10 questions. Continue until the important branches of the decision tree are resolved. If fewer than 10 meaningful decisions remain, ask only those remaining; never add filler questions to reach 10.
|
|
20
|
+
8. Present the final shared understanding as a compact decision brief: goal, scope, constraints, chosen approach, rejected alternatives, risks, and acceptance criteria.
|
|
21
|
+
9. Ask for explicit confirmation that the brief is correct before acting on it.
|
|
22
|
+
|
|
23
|
+
## Question quality
|
|
24
|
+
|
|
25
|
+
- Ask decisions only the user can make; investigate facts yourself.
|
|
26
|
+
- Challenge assumptions and compare credible alternatives instead of merely collecting preferences.
|
|
27
|
+
- Keep each question independently answerable and concise enough for a spoken response.
|
|
28
|
+
- Do not repeat settled questions unless new evidence invalidates the earlier answer.
|
|
29
|
+
- Match the user's language.
|
|
30
|
+
|
|
31
|
+
Do not implement the plan until the user confirms the final decision brief.
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
<!-- @license lucide-static v0.294.0 - ISC -->
|
|
2
|
+
<svg role="img" aria-label="grill-me skill icon"
|
|
3
|
+
class="lucide lucide-messages-square"
|
|
4
|
+
xmlns="http://www.w3.org/2000/svg"
|
|
5
|
+
width="128"
|
|
6
|
+
height="128"
|
|
7
|
+
viewBox="0 0 24 24"
|
|
8
|
+
fill="none"
|
|
9
|
+
stroke="#F5F5F5"
|
|
10
|
+
stroke-width="2"
|
|
11
|
+
stroke-linecap="round"
|
|
12
|
+
stroke-linejoin="round"
|
|
13
|
+
>
|
|
14
|
+
<path d="M14 9a2 2 0 0 1-2 2H6l-4 4V4c0-1.1.9-2 2-2h8a2 2 0 0 1 2 2v5Z" />
|
|
15
|
+
<path d="M18 9h2a2 2 0 0 1 2 2v11l-4-4h-6a2 2 0 0 1-2-2v-1" />
|
|
16
|
+
</svg>
|