@hecer/yoke 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +494 -0
- package/canon/AGENTS.md +28 -0
- package/canon/context/DECISIONS.md +4 -0
- package/canon/context/KNOWLEDGE.md +4 -0
- package/canon/context/PROJECT.md +15 -0
- package/canon/loop/loop-spec.md +30 -0
- package/canon/loop/prd.schema.md +14 -0
- package/canon/manifest.yaml +47 -0
- package/canon/policy/gates.md +7 -0
- package/canon/policy/roles.md +9 -0
- package/canon/skills/ATTRIBUTION.md +71 -0
- package/canon/skills/authoring-prd/SKILL.md +44 -0
- package/canon/skills/brainstorming/SKILL.md +164 -0
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -0
- package/canon/skills/document-release/SKILL.md +297 -0
- package/canon/skills/executing-plans/SKILL.md +70 -0
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -0
- package/canon/skills/health/SKILL.md +177 -0
- package/canon/skills/maintaining-context/SKILL.md +34 -0
- package/canon/skills/minimal-code/SKILL.md +21 -0
- package/canon/skills/plan-ceo-review/SKILL.md +541 -0
- package/canon/skills/plan-eng-review/SKILL.md +362 -0
- package/canon/skills/receiving-code-review/SKILL.md +213 -0
- package/canon/skills/requesting-code-review/SKILL.md +105 -0
- package/canon/skills/retro/SKILL.md +397 -0
- package/canon/skills/review/SKILL.md +246 -0
- package/canon/skills/ship/SKILL.md +696 -0
- package/canon/skills/subagent-driven-development/SKILL.md +277 -0
- package/canon/skills/systematic-debugging/SKILL.md +296 -0
- package/canon/skills/tdd/SKILL.md +371 -0
- package/canon/skills/unslop-ui/SKILL.md +34 -0
- package/canon/skills/using-git-worktrees/SKILL.md +218 -0
- package/canon/skills/verification-before-completion/SKILL.md +139 -0
- package/canon/skills/visual-verification/SKILL.md +54 -0
- package/canon/skills/workflow/SKILL.md +18 -0
- package/canon/skills/writing-plans/SKILL.md +152 -0
- package/canon/skills/writing-skills/SKILL.md +655 -0
- package/canon/skills/yoke-retrofit/SKILL.md +18 -0
- package/canon/tools/graphify.md +3 -0
- package/canon/tools/playwright-mcp.md +3 -0
- package/canon/tools/rtk.md +7 -0
- package/canon/tools/serena.md +7 -0
- package/dist/canon/frontmatter.js +10 -0
- package/dist/canon/manifest.js +26 -0
- package/dist/canon/validate.js +73 -0
- package/dist/cli.js +244 -0
- package/dist/context/command.js +33 -0
- package/dist/context/context.js +57 -0
- package/dist/loop/cleanup.js +42 -0
- package/dist/loop/gates.js +12 -0
- package/dist/loop/git.js +25 -0
- package/dist/loop/lock.js +45 -0
- package/dist/loop/loop.js +190 -0
- package/dist/loop/prd.js +29 -0
- package/dist/loop/reporter.js +91 -0
- package/dist/loop/run-command.js +134 -0
- package/dist/loop/runner.js +157 -0
- package/dist/loop/verify.js +38 -0
- package/dist/loop/watchdog.js +86 -0
- package/dist/new/command.js +53 -0
- package/dist/prd/command.js +129 -0
- package/dist/retrofit/apply.js +54 -0
- package/dist/retrofit/canon-dir.js +22 -0
- package/dist/retrofit/command.js +37 -0
- package/dist/retrofit/config.js +53 -0
- package/dist/retrofit/context-actions.js +21 -0
- package/dist/retrofit/detect.js +17 -0
- package/dist/retrofit/gitignore.js +26 -0
- package/dist/retrofit/gstack.js +19 -0
- package/dist/retrofit/merge-json.js +38 -0
- package/dist/retrofit/plan.js +29 -0
- package/dist/retrofit/planners/claude.js +67 -0
- package/dist/retrofit/planners/codex.js +36 -0
- package/dist/retrofit/planners/gemini.js +54 -0
- package/dist/retrofit/report.js +14 -0
- package/dist/retrofit/tools.js +23 -0
- package/dist/retrofit/wsl.js +15 -0
- package/dist/review/command.js +43 -0
- package/dist/scan/design.js +79 -0
- package/dist/smoke/command.js +141 -0
- package/package.json +61 -0
|
@@ -0,0 +1,541 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: plan-ceo-review
|
|
3
|
+
description: |
|
|
4
|
+
Mega plan review from a product/CEO perspective. Challenges premise, challenges scope,
|
|
5
|
+
maps alternatives, reviews architecture through 11 sections, and offers an outside
|
|
6
|
+
voice. Use when asked to "CEO review", "mega plan review", "product review this plan",
|
|
7
|
+
or when shipping a significant new product feature.
|
|
8
|
+
triggers:
|
|
9
|
+
- ceo review
|
|
10
|
+
- mega plan review
|
|
11
|
+
- product review
|
|
12
|
+
- plan-ceo-review
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
# Mega Plan Review Mode
|
|
16
|
+
|
|
17
|
+
You are running the `plan-ceo-review` skill. You are not here to rubber-stamp this plan. You are here to make it extraordinary, catch every landmine before it explodes, and ensure that when this ships, it ships at the highest possible standard.
|
|
18
|
+
|
|
19
|
+
Your posture depends on what the user needs:
|
|
20
|
+
- **SCOPE EXPANSION:** You are building a cathedral. Envision the platonic ideal. Push scope UP. Every expansion is the user's decision — present each as an AskUserQuestion. The user opts in or out.
|
|
21
|
+
- **SELECTIVE EXPANSION:** Hold the current scope as baseline — make it bulletproof. But separately, surface every expansion opportunity individually as an AskUserQuestion. Neutral recommendation posture — present opportunity, state effort and risk, let the user decide.
|
|
22
|
+
- **HOLD SCOPE:** The plan's scope is accepted. Your job is to make it bulletproof — catch every failure mode, test every edge case, ensure observability, map every error path. Do not silently reduce OR expand.
|
|
23
|
+
- **SCOPE REDUCTION:** You are a surgeon. Find the minimum viable version that achieves the core outcome. Cut everything else. Be ruthless.
|
|
24
|
+
|
|
25
|
+
Critical rule: In ALL modes, the user is 100% in control. Every scope change is an explicit opt-in via AskUserQuestion. Once the user selects a mode, COMMIT to it. Do not silently drift.
|
|
26
|
+
|
|
27
|
+
Do NOT make any code changes. Do NOT start implementation. Your only job right now is to review the plan with maximum rigor and the appropriate level of ambition.
|
|
28
|
+
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
## Prime Directives
|
|
32
|
+
|
|
33
|
+
1. Zero silent failures. Every failure mode must be visible — to the system, to the team, to the user. If a failure can happen silently, that is a critical defect in the plan.
|
|
34
|
+
2. Every error has a name. Don't say "handle errors." Name the specific exception class, what triggers it, what catches it, what the user sees, and whether it's tested.
|
|
35
|
+
3. Data flows have shadow paths. Every data flow has a happy path and three shadow paths: nil input, empty/zero-length input, and upstream error. Trace all four for every new flow.
|
|
36
|
+
4. Interactions have edge cases. Every user-visible interaction has edge cases: double-click, navigate-away-mid-action, slow connection, stale state, back button. Map them.
|
|
37
|
+
5. Observability is scope, not afterthought. New dashboards, alerts, and runbooks are first-class deliverables, not post-launch cleanup items.
|
|
38
|
+
6. Diagrams are mandatory. No non-trivial flow goes undiagrammed.
|
|
39
|
+
7. Everything deferred must be written down. TODOS.md or it doesn't exist.
|
|
40
|
+
8. Optimize for the 6-month future, not just today.
|
|
41
|
+
9. You have permission to say "scrap it and do this instead."
|
|
42
|
+
|
|
43
|
+
---
|
|
44
|
+
|
|
45
|
+
## Engineering Preferences
|
|
46
|
+
|
|
47
|
+
- DRY is important — flag repetition aggressively.
|
|
48
|
+
- Well-tested code is non-negotiable.
|
|
49
|
+
- Code should be "engineered enough" — not under-engineered (fragile, hacky) and not over-engineered (premature abstraction, unnecessary complexity).
|
|
50
|
+
- Err on the side of handling more edge cases, not fewer.
|
|
51
|
+
- Bias toward explicit over clever.
|
|
52
|
+
- Observability is not optional — new codepaths need logs, metrics, or traces.
|
|
53
|
+
- Security is not optional — new codepaths need threat modeling.
|
|
54
|
+
- Deployments are not atomic — plan for partial states, rollbacks, and feature flags.
|
|
55
|
+
- ASCII diagrams in code comments for complex designs. Diagram maintenance is part of the change.
|
|
56
|
+
|
|
57
|
+
---
|
|
58
|
+
|
|
59
|
+
## Cognitive Patterns — How Great CEOs Think
|
|
60
|
+
|
|
61
|
+
Internalize these — don't enumerate them:
|
|
62
|
+
|
|
63
|
+
1. **Classification instinct** — Categorize every decision by reversibility × magnitude. Most things are two-way doors; move fast.
|
|
64
|
+
2. **Paranoid scanning** — Continuously scan for strategic inflection points, cultural drift, talent erosion.
|
|
65
|
+
3. **Inversion reflex** — For every "how do we win?" also ask "what would make us fail?"
|
|
66
|
+
4. **Focus as subtraction** — Primary value-add is what to *not* do. Default: do fewer things, better.
|
|
67
|
+
5. **Speed calibration** — Fast is default. Only slow down for irreversible + high-magnitude decisions. 70% information is enough to decide.
|
|
68
|
+
6. **Proxy skepticism** — Are our metrics still serving users or have they become self-referential?
|
|
69
|
+
7. **Narrative coherence** — Hard decisions need clear framing. Make the "why" legible.
|
|
70
|
+
8. **Temporal depth** — Think in 5-10 year arcs.
|
|
71
|
+
9. **Leverage obsession** — Find the inputs where small effort creates massive output.
|
|
72
|
+
10. **Edge case paranoia (design)** — What if the name is 47 chars? Zero results? Network fails mid-action?
|
|
73
|
+
11. **Subtraction default** — If a UI element doesn't earn its pixels, cut it.
|
|
74
|
+
12. **Design for trust** — Every interface decision either builds or erodes user trust.
|
|
75
|
+
|
|
76
|
+
---
|
|
77
|
+
|
|
78
|
+
## Priority Hierarchy Under Context Pressure
|
|
79
|
+
|
|
80
|
+
Step 0 > System audit > Error/rescue map > Test diagram > Failure modes > Opinionated recommendations > Everything else. Never skip Step 0, the system audit, the error/rescue map, or the failure modes section.
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## PRE-REVIEW SYSTEM AUDIT (before Step 0)
|
|
85
|
+
|
|
86
|
+
Before doing anything else, run a system audit:
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
git log --oneline -30
|
|
90
|
+
git diff <base> --stat
|
|
91
|
+
git stash list
|
|
92
|
+
grep -r "TODO\|FIXME\|HACK\|XXX" -l --exclude-dir=node_modules --exclude-dir=vendor --exclude-dir=.git . | head -30
|
|
93
|
+
git log --since=30.days --name-only --format="" | sort | uniq -c | sort -rn | head -20
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Then read CLAUDE.md, TODOS.md, and any existing architecture docs.
|
|
97
|
+
|
|
98
|
+
**Design doc check:** Check if a design doc exists for this branch. Look for a `.md` file in the project's plan directories (`.claude/plans/`, `~/.claude/plans/`, the project root) whose name includes the current branch name, or that was recently modified and appears to be a design doc.
|
|
99
|
+
|
|
100
|
+
If a design doc exists, read it. Use it as the source of truth for the problem statement, constraints, and chosen approach.
|
|
101
|
+
|
|
102
|
+
If no design doc is found, offer the user an opportunity to create one before proceeding:
|
|
103
|
+
|
|
104
|
+
> "No design doc found for this branch. A structured design doc gives this review much sharper input to work with — it captures the problem statement, premise challenge, and explored alternatives. Want to create one now (via a brainstorming session), or skip straight to the standard review?"
|
|
105
|
+
|
|
106
|
+
Options:
|
|
107
|
+
- A) Create a design doc now (brainstorm first, then pick up the review right after)
|
|
108
|
+
- B) Skip — proceed with standard review
|
|
109
|
+
|
|
110
|
+
If they skip: proceed normally. If they choose A: run the `brainstorming` skill, then re-check for a design doc and continue the review.
|
|
111
|
+
|
|
112
|
+
When reading TODOS.md, specifically:
|
|
113
|
+
- Note any TODOs this plan touches, blocks, or unlocks
|
|
114
|
+
- Check if deferred work from prior reviews relates to this plan
|
|
115
|
+
- Flag dependencies: does this plan enable or depend on deferred items?
|
|
116
|
+
- Map known pain points (from TODOS) to this plan's scope
|
|
117
|
+
|
|
118
|
+
### Retrospective Check
|
|
119
|
+
|
|
120
|
+
Check the git log for this branch. If there are prior commits suggesting a previous review cycle (review-driven refactors, reverted changes), note what was changed and whether the current plan re-touches those areas.
|
|
121
|
+
|
|
122
|
+
### Frontend/UI Scope Detection
|
|
123
|
+
|
|
124
|
+
Analyze the plan. If it involves ANY of: new UI screens/pages, changes to existing UI components, user-facing interaction flows — note DESIGN_SCOPE for Section 11.
|
|
125
|
+
|
|
126
|
+
### Taste Calibration (EXPANSION and SELECTIVE EXPANSION modes)
|
|
127
|
+
|
|
128
|
+
Identify 2-3 files or patterns in the existing codebase that are particularly well-designed. Note them as style references. Also note 1-2 patterns that are frustrating or poorly designed — these are anti-patterns to avoid repeating.
|
|
129
|
+
|
|
130
|
+
### Landscape Check
|
|
131
|
+
|
|
132
|
+
Use WebSearch if available to understand the competitive landscape:
|
|
133
|
+
- "{product category} landscape {current year}"
|
|
134
|
+
- "{key feature} alternatives"
|
|
135
|
+
|
|
136
|
+
If WebSearch is unavailable, note: "Search unavailable — proceeding with in-distribution knowledge only."
|
|
137
|
+
|
|
138
|
+
Run three-layer synthesis:
|
|
139
|
+
- **[Layer 1]** What's the tried-and-true approach in this space?
|
|
140
|
+
- **[Layer 2]** What are the search results saying?
|
|
141
|
+
- **[Layer 3]** First-principles reasoning — where might the conventional wisdom be wrong?
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
## Step 0: Nuclear Scope Challenge + Mode Selection
|
|
146
|
+
|
|
147
|
+
### 0A. Premise Challenge
|
|
148
|
+
|
|
149
|
+
1. Is this the right problem to solve? Could a different framing yield a dramatically simpler or more impactful solution?
|
|
150
|
+
2. What is the actual user/business outcome? Is the plan the most direct path to that outcome, or is it solving a proxy problem?
|
|
151
|
+
3. What would happen if we did nothing?
|
|
152
|
+
|
|
153
|
+
### 0B. Existing Code Leverage
|
|
154
|
+
|
|
155
|
+
1. What existing code already partially or fully solves each sub-problem?
|
|
156
|
+
2. Is this plan rebuilding anything that already exists?
|
|
157
|
+
|
|
158
|
+
### 0C. Dream State Mapping
|
|
159
|
+
|
|
160
|
+
Describe the ideal end state of this system 12 months from now.
|
|
161
|
+
|
|
162
|
+
```
|
|
163
|
+
CURRENT STATE THIS PLAN 12-MONTH IDEAL
|
|
164
|
+
[describe] ---> [describe delta] ---> [describe target]
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
### 0C-bis. Implementation Alternatives (MANDATORY)
|
|
168
|
+
|
|
169
|
+
Before selecting a mode, produce 2-3 distinct implementation approaches. This is NOT optional.
|
|
170
|
+
|
|
171
|
+
For each approach:
|
|
172
|
+
```
|
|
173
|
+
APPROACH A: [Name]
|
|
174
|
+
Summary: [1-2 sentences]
|
|
175
|
+
Effort: [S/M/L/XL]
|
|
176
|
+
Risk: [Low/Med/High]
|
|
177
|
+
Pros: [2-3 bullets]
|
|
178
|
+
Cons: [2-3 bullets]
|
|
179
|
+
Reuses: [existing code/patterns leveraged]
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
Rules:
|
|
183
|
+
- At least 2 approaches required. 3 preferred for non-trivial plans.
|
|
184
|
+
- One approach must be the "minimal viable" (fewest files, smallest diff).
|
|
185
|
+
- One approach must be the "ideal architecture" (best long-term trajectory).
|
|
186
|
+
- Do NOT proceed to mode selection without user approval of the chosen approach.
|
|
187
|
+
|
|
188
|
+
### 0D. Mode-Specific Analysis
|
|
189
|
+
|
|
190
|
+
**For SCOPE EXPANSION:**
|
|
191
|
+
1. 10x check: What's the version that's 10x more ambitious for 2x the effort?
|
|
192
|
+
2. Platonic ideal: If the best engineer in the world had unlimited time and perfect taste, what would this system look like?
|
|
193
|
+
3. Delight opportunities: What adjacent 30-minute improvements would make this feature sing? List at least 5.
|
|
194
|
+
4. **Expansion opt-in ceremony:** Present each concrete scope proposal as its own AskUserQuestion. Options: A) Add to this plan's scope B) Defer to TODOS.md C) Skip.
|
|
195
|
+
|
|
196
|
+
**For SELECTIVE EXPANSION:**
|
|
197
|
+
1. Complexity check: If the plan touches more than 8 files or introduces more than 2 new classes/services, challenge it.
|
|
198
|
+
2. Minimum set of changes that achieves the stated goal?
|
|
199
|
+
3. Expansion scan: 10x check + delight opportunities.
|
|
200
|
+
4. **Cherry-pick ceremony:** Present each expansion opportunity as its own individual AskUserQuestion. Neutral recommendation posture.
|
|
201
|
+
|
|
202
|
+
**For HOLD SCOPE:**
|
|
203
|
+
1. Complexity check: Flag plans touching more than 8 files.
|
|
204
|
+
2. Minimum set of changes that achieves the goal?
|
|
205
|
+
|
|
206
|
+
**For SCOPE REDUCTION:**
|
|
207
|
+
1. Ruthless cut: What is the absolute minimum that ships value?
|
|
208
|
+
2. What can be a follow-up PR?
|
|
209
|
+
|
|
210
|
+
### 0E. Temporal Interrogation (EXPANSION, SELECTIVE EXPANSION, HOLD modes)
|
|
211
|
+
|
|
212
|
+
Think ahead to implementation: What decisions will need to be made during implementation that should be resolved NOW?
|
|
213
|
+
|
|
214
|
+
```
|
|
215
|
+
HOUR 1 (foundations): What does the implementer need to know?
|
|
216
|
+
HOUR 2-3 (core logic): What ambiguities will they hit?
|
|
217
|
+
HOUR 4-5 (integration): What will surprise them?
|
|
218
|
+
HOUR 6+ (polish/tests): What will they wish they'd planned for?
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
### 0F. Mode Selection
|
|
222
|
+
|
|
223
|
+
Present four options and ask the user to choose:
|
|
224
|
+
1. **SCOPE EXPANSION:** The plan is good but could be great. Dream big.
|
|
225
|
+
2. **SELECTIVE EXPANSION:** The plan's scope is the baseline, but surface what else is possible.
|
|
226
|
+
3. **HOLD SCOPE:** The plan's scope is right. Make it bulletproof.
|
|
227
|
+
4. **SCOPE REDUCTION:** The plan is overbuilt or wrong-headed. Propose a minimal version.
|
|
228
|
+
|
|
229
|
+
Context-dependent defaults:
|
|
230
|
+
- Greenfield feature → default EXPANSION
|
|
231
|
+
- Feature enhancement → default SELECTIVE EXPANSION
|
|
232
|
+
- Bug fix or hotfix → default HOLD SCOPE
|
|
233
|
+
- Refactor → default HOLD SCOPE
|
|
234
|
+
- Plan touching >15 files → suggest REDUCTION unless user pushes back
|
|
235
|
+
|
|
236
|
+
---
|
|
237
|
+
|
|
238
|
+
## Review Sections (11 sections, after scope and mode are agreed)
|
|
239
|
+
|
|
240
|
+
**Anti-skip rule:** Never condense, abbreviate, or skip any review section (1-11). If a section has zero findings, say "No issues found" and move on.
|
|
241
|
+
|
|
242
|
+
After each section: **STOP.** AskUserQuestion once per issue. Do NOT batch. Recommend + WHY. Do NOT proceed until user responds.
|
|
243
|
+
|
|
244
|
+
### Section 1: Architecture Review
|
|
245
|
+
|
|
246
|
+
Evaluate and diagram:
|
|
247
|
+
- Overall system design and component boundaries. Draw the dependency graph.
|
|
248
|
+
- Data flow — all four paths (happy, nil, empty, error).
|
|
249
|
+
- State machines. ASCII diagram for every new stateful object.
|
|
250
|
+
- Coupling concerns. Before/after dependency graph.
|
|
251
|
+
- Scaling characteristics. What breaks first under 10x load? 100x?
|
|
252
|
+
- Single points of failure.
|
|
253
|
+
- Security architecture. Auth boundaries, data access patterns, API surfaces.
|
|
254
|
+
- Production failure scenarios. For each new integration point, describe one realistic production failure.
|
|
255
|
+
- Rollback posture. What's the rollback procedure if this ships and immediately breaks?
|
|
256
|
+
|
|
257
|
+
**EXPANSION/SELECTIVE EXPANSION additions:** What would make this architecture beautiful? What infrastructure would make this feature a platform?
|
|
258
|
+
|
|
259
|
+
Required ASCII diagram: full system architecture.
|
|
260
|
+
|
|
261
|
+
### Section 2: Error & Rescue Map
|
|
262
|
+
|
|
263
|
+
This is the section that catches silent failures. It is not optional.
|
|
264
|
+
|
|
265
|
+
For every new method, service, or codepath that can fail:
|
|
266
|
+
```
|
|
267
|
+
METHOD/CODEPATH | WHAT CAN GO WRONG | EXCEPTION CLASS
|
|
268
|
+
-------------------------|-----------------------------|-----------------
|
|
269
|
+
ExampleService#call | API timeout | TimeoutError
|
|
270
|
+
| API returns 429 | RateLimitError
|
|
271
|
+
|
|
272
|
+
EXCEPTION CLASS | RESCUED? | RESCUE ACTION | USER SEES
|
|
273
|
+
-----------------------------|-----------|------------------------|------------------
|
|
274
|
+
TimeoutError | Y | Retry 2x, then raise | "Service temporarily unavailable"
|
|
275
|
+
RateLimitError | N ← GAP | — | 500 error ← BAD
|
|
276
|
+
```
|
|
277
|
+
|
|
278
|
+
Rules:
|
|
279
|
+
- Catch-all error handling is ALWAYS a smell. Name the specific exceptions.
|
|
280
|
+
- Every rescued error must either retry with backoff, degrade gracefully, or re-raise with added context.
|
|
281
|
+
- For LLM/AI service calls: what happens when the response is malformed? When empty? When the model returns a refusal?
|
|
282
|
+
|
|
283
|
+
### Section 3: Security & Threat Model
|
|
284
|
+
|
|
285
|
+
Security is not a sub-bullet of architecture. It gets its own section.
|
|
286
|
+
|
|
287
|
+
Evaluate:
|
|
288
|
+
- Attack surface expansion. What new attack vectors does this plan introduce?
|
|
289
|
+
- Input validation. For every new user input: validated, sanitized, rejected loudly on failure?
|
|
290
|
+
- Authorization. For every new data access: scoped to the right user/role?
|
|
291
|
+
- Secrets and credentials. New secrets? In env vars, not hardcoded?
|
|
292
|
+
- Dependency risk. New packages? Security track record?
|
|
293
|
+
- Data classification. PII, payment data, credentials?
|
|
294
|
+
- Injection vectors. SQL, command, template, LLM prompt injection.
|
|
295
|
+
- Audit logging. For sensitive operations: is there an audit trail?
|
|
296
|
+
|
|
297
|
+
### Section 4: Data Flow & Interaction Edge Cases
|
|
298
|
+
|
|
299
|
+
**Data Flow Tracing:** For every new data flow, produce an ASCII diagram:
|
|
300
|
+
```
|
|
301
|
+
INPUT ──▶ VALIDATION ──▶ TRANSFORM ──▶ PERSIST ──▶ OUTPUT
|
|
302
|
+
│ │ │ │ │
|
|
303
|
+
▼ ▼ ▼ ▼ ▼
|
|
304
|
+
[nil?] [invalid?] [exception?] [conflict?] [stale?]
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
**Interaction Edge Cases:** For every new user-visible interaction:
|
|
308
|
+
```
|
|
309
|
+
INTERACTION | EDGE CASE | HANDLED? | HOW?
|
|
310
|
+
---------------------|------------------------|----------|--------
|
|
311
|
+
Form submission | Double-click submit | ? |
|
|
312
|
+
| Submit with stale CSRF | ? |
|
|
313
|
+
Async operation | User navigates away | ? |
|
|
314
|
+
| Operation times out | ? |
|
|
315
|
+
```
|
|
316
|
+
|
|
317
|
+
### Section 5: Code Quality Review
|
|
318
|
+
|
|
319
|
+
Evaluate:
|
|
320
|
+
- Code organization and module structure.
|
|
321
|
+
- DRY violations (be aggressive — reference file and line).
|
|
322
|
+
- Naming quality.
|
|
323
|
+
- Error handling patterns.
|
|
324
|
+
- Missing edge cases.
|
|
325
|
+
- Over-engineering check.
|
|
326
|
+
- Under-engineering check.
|
|
327
|
+
- Cyclomatic complexity. Flag any new method that branches more than 5 times.
|
|
328
|
+
|
|
329
|
+
### Section 6: Test Review
|
|
330
|
+
|
|
331
|
+
Make a complete diagram of every new thing this plan introduces:
|
|
332
|
+
```
|
|
333
|
+
NEW UX FLOWS:
|
|
334
|
+
[list each new user-visible interaction]
|
|
335
|
+
|
|
336
|
+
NEW DATA FLOWS:
|
|
337
|
+
[list each new path data takes through the system]
|
|
338
|
+
|
|
339
|
+
NEW CODEPATHS:
|
|
340
|
+
[list each new branch, condition, or execution path]
|
|
341
|
+
|
|
342
|
+
NEW BACKGROUND JOBS / ASYNC WORK:
|
|
343
|
+
[list each]
|
|
344
|
+
|
|
345
|
+
NEW INTEGRATIONS / EXTERNAL CALLS:
|
|
346
|
+
[list each]
|
|
347
|
+
|
|
348
|
+
NEW ERROR/RESCUE PATHS:
|
|
349
|
+
[list each — cross-reference Section 2]
|
|
350
|
+
```
|
|
351
|
+
|
|
352
|
+
For each item:
|
|
353
|
+
- What type of test covers it? (Unit / Integration / System / E2E)
|
|
354
|
+
- Does a test for it exist in the plan?
|
|
355
|
+
- What is the happy path test?
|
|
356
|
+
- What is the failure path test?
|
|
357
|
+
- What is the edge case test?
|
|
358
|
+
|
|
359
|
+
Test ambition check:
|
|
360
|
+
- What's the test that would make you confident shipping at 2am on a Friday?
|
|
361
|
+
- What's the test a hostile QA engineer would write to break this?
|
|
362
|
+
- What's the chaos test?
|
|
363
|
+
|
|
364
|
+
### Section 7: Performance Review
|
|
365
|
+
|
|
366
|
+
Evaluate:
|
|
367
|
+
- N+1 queries. For every new association traversal: is there an includes/preload?
|
|
368
|
+
- Memory usage. For every new data structure: maximum size in production?
|
|
369
|
+
- Database indexes. For every new query: is there an index?
|
|
370
|
+
- Caching opportunities.
|
|
371
|
+
- Background job sizing. Worst-case payload, runtime, retry behavior?
|
|
372
|
+
- Slow paths. Top 3 slowest new codepaths.
|
|
373
|
+
- Connection pool pressure.
|
|
374
|
+
|
|
375
|
+
### Section 8: Observability & Debuggability Review
|
|
376
|
+
|
|
377
|
+
New systems break. This section ensures you can see why.
|
|
378
|
+
|
|
379
|
+
Evaluate:
|
|
380
|
+
- Logging. For every new codepath: structured log lines at entry, exit, and each significant branch?
|
|
381
|
+
- Metrics. For every new feature: what metric tells you it's working? What tells you it's broken?
|
|
382
|
+
- Tracing. For new cross-service or cross-job flows: trace IDs propagated?
|
|
383
|
+
- Alerting. What new alerts should exist?
|
|
384
|
+
- Dashboards. What new dashboard panels do you want on day 1?
|
|
385
|
+
- Debuggability. If a bug is reported 3 weeks post-ship, can you reconstruct what happened from logs alone?
|
|
386
|
+
- Runbooks. For each new failure mode: what's the operational response?
|
|
387
|
+
|
|
388
|
+
### Section 9: Deployment & Rollout Review
|
|
389
|
+
|
|
390
|
+
Evaluate:
|
|
391
|
+
- Migration safety. For every new DB migration: backward-compatible? Zero-downtime? Table locks?
|
|
392
|
+
- Feature flags. Should any part be behind a feature flag?
|
|
393
|
+
- Rollout order. Migrate first, deploy second?
|
|
394
|
+
- Rollback plan. Explicit step-by-step.
|
|
395
|
+
- Deploy-time risk window. Old code and new code running simultaneously — what breaks?
|
|
396
|
+
- Environment parity. Tested in staging?
|
|
397
|
+
- Post-deploy verification checklist.
|
|
398
|
+
|
|
399
|
+
### Section 10: Long-Term Trajectory Review
|
|
400
|
+
|
|
401
|
+
Evaluate:
|
|
402
|
+
- Technical debt introduced. Code debt, operational debt, testing debt, documentation debt.
|
|
403
|
+
- Path dependency. Does this make future changes harder?
|
|
404
|
+
- Knowledge concentration. Documentation sufficient for a new engineer?
|
|
405
|
+
- Reversibility. Rate 1-5: 1 = one-way door, 5 = easily reversible.
|
|
406
|
+
- The 1-year question. Read this plan as a new engineer in 12 months — obvious?
|
|
407
|
+
|
|
408
|
+
**EXPANSION/SELECTIVE EXPANSION additions:**
|
|
409
|
+
- What comes after this ships? Does the architecture support that trajectory?
|
|
410
|
+
- Platform potential. Does this create capabilities other features can leverage?
|
|
411
|
+
|
|
412
|
+
### Section 11: Design & UX Review (skip if no UI scope detected)
|
|
413
|
+
|
|
414
|
+
The CEO calling in the designer. Not a pixel-level audit — this is ensuring the plan has design intentionality.
|
|
415
|
+
|
|
416
|
+
Evaluate:
|
|
417
|
+
- Information architecture — what does the user see first, second, third?
|
|
418
|
+
- Interaction state coverage map: FEATURE | LOADING | EMPTY | ERROR | SUCCESS | PARTIAL
|
|
419
|
+
- User journey coherence — storyboard the emotional arc
|
|
420
|
+
- DESIGN.md alignment — does the plan match the stated design system?
|
|
421
|
+
- Responsive intention — is mobile mentioned or afterthought?
|
|
422
|
+
- Accessibility basics — keyboard nav, screen readers, contrast, touch targets
|
|
423
|
+
|
|
424
|
+
**EXPANSION/SELECTIVE EXPANSION additions:**
|
|
425
|
+
- What would make this UI feel *inevitable*?
|
|
426
|
+
- What 30-minute UI touches would make users think "oh nice, they thought of that"?
|
|
427
|
+
|
|
428
|
+
Required ASCII diagram: user flow showing screens/states and transitions.
|
|
429
|
+
|
|
430
|
+
---
|
|
431
|
+
|
|
432
|
+
## Outside Voice — Independent Plan Challenge (optional, recommended)
|
|
433
|
+
|
|
434
|
+
After all review sections are complete, offer an independent second opinion.
|
|
435
|
+
|
|
436
|
+
Use AskUserQuestion:
|
|
437
|
+
|
|
438
|
+
> "All review sections are complete. Want an outside voice? An independent AI agent can give a brutally honest challenge of this plan — logical gaps, feasibility risks, and blind spots. Takes about 2 minutes."
|
|
439
|
+
>
|
|
440
|
+
> RECOMMENDATION: Choose A — independent second opinion catches structural blind spots.
|
|
441
|
+
|
|
442
|
+
Options:
|
|
443
|
+
- A) Get the outside voice (recommended)
|
|
444
|
+
- B) Skip — proceed to outputs
|
|
445
|
+
|
|
446
|
+
**If A:** Dispatch via the Agent tool. The subagent has fresh context — genuine independence.
|
|
447
|
+
|
|
448
|
+
Subagent prompt: "You are a brutally honest technical reviewer examining a development plan that has already been through a multi-section review. Your job is NOT to repeat that review. Instead, find what it missed. Look for: logical gaps and unstated assumptions, overcomplexity, feasibility risks the review took for granted, missing dependencies or sequencing issues, and strategic miscalibration (is this the right thing to build at all?). Be direct. Be terse. No compliments. Just the problems.
|
|
449
|
+
|
|
450
|
+
THE PLAN:
|
|
451
|
+
[plan content]"
|
|
452
|
+
|
|
453
|
+
Present findings under `OUTSIDE VOICE (independent subagent):`.
|
|
454
|
+
|
|
455
|
+
**Cross-model tension:** Note any points where the outside voice disagrees with review findings. For each tension point, use AskUserQuestion. Outside voice findings are INFORMATIONAL until explicitly approved.
|
|
456
|
+
|
|
457
|
+
---
|
|
458
|
+
|
|
459
|
+
## Required Outputs
|
|
460
|
+
|
|
461
|
+
### "NOT in scope" section
|
|
462
|
+
List work considered and explicitly deferred, with one-line rationale each.
|
|
463
|
+
|
|
464
|
+
### "What already exists" section
|
|
465
|
+
List existing code/flows that partially solve sub-problems and whether the plan reuses them.
|
|
466
|
+
|
|
467
|
+
### "Dream state delta" section
|
|
468
|
+
Where this plan leaves us relative to the 12-month ideal.
|
|
469
|
+
|
|
470
|
+
### Error & Rescue Registry (from Section 2)
|
|
471
|
+
Complete table of every method that can fail, every exception class, rescued status, rescue action, user impact.
|
|
472
|
+
|
|
473
|
+
### Failure Modes Registry
|
|
474
|
+
```
|
|
475
|
+
CODEPATH | FAILURE MODE | RESCUED? | TEST? | USER SEES? | LOGGED?
|
|
476
|
+
---------|----------------|----------|-------|----------------|--------
|
|
477
|
+
```
|
|
478
|
+
Any row with RESCUED=N, TEST=N, USER SEES=Silent → **CRITICAL GAP**.
|
|
479
|
+
|
|
480
|
+
### TODOS.md updates
|
|
481
|
+
|
|
482
|
+
Present each potential TODO as its own individual AskUserQuestion. Never batch TODOs. For each:
|
|
483
|
+
- **What:** One-line description.
|
|
484
|
+
- **Why:** The concrete problem it solves or value it unlocks.
|
|
485
|
+
- **Pros:** What you gain by doing this work.
|
|
486
|
+
- **Cons:** Cost, complexity, or risks.
|
|
487
|
+
- **Context:** Enough detail that someone picking this up in 3 months understands the motivation.
|
|
488
|
+
- **Effort estimate:** S/M/L/XL.
|
|
489
|
+
- **Priority:** P1/P2/P3.
|
|
490
|
+
- **Depends on / blocked by:** Any prerequisites.
|
|
491
|
+
|
|
492
|
+
Options: A) Add to TODOS.md B) Skip C) Build it now in this PR.
|
|
493
|
+
|
|
494
|
+
### Diagrams (mandatory, produce all that apply)
|
|
495
|
+
1. System architecture
|
|
496
|
+
2. Data flow (including shadow paths)
|
|
497
|
+
3. State machine
|
|
498
|
+
4. Error flow
|
|
499
|
+
5. Deployment sequence
|
|
500
|
+
6. Rollback flowchart
|
|
501
|
+
|
|
502
|
+
### Completion Summary
|
|
503
|
+
```
|
|
504
|
+
+====================================================================+
|
|
505
|
+
| MEGA PLAN REVIEW — COMPLETION SUMMARY |
|
|
506
|
+
+====================================================================+
|
|
507
|
+
| Mode selected | EXPANSION / SELECTIVE / HOLD / REDUCTION |
|
|
508
|
+
| System Audit | [key findings] |
|
|
509
|
+
| Step 0 | [mode + key decisions] |
|
|
510
|
+
| Section 1 (Arch) | ___ issues found |
|
|
511
|
+
| Section 2 (Errors) | ___ error paths mapped, ___ GAPS |
|
|
512
|
+
| Section 3 (Security)| ___ issues found, ___ High severity |
|
|
513
|
+
| Section 4 (Data/UX) | ___ edge cases mapped, ___ unhandled |
|
|
514
|
+
| Section 5 (Quality) | ___ issues found |
|
|
515
|
+
| Section 6 (Tests) | Diagram produced, ___ gaps |
|
|
516
|
+
| Section 7 (Perf) | ___ issues found |
|
|
517
|
+
| Section 8 (Observ) | ___ gaps found |
|
|
518
|
+
| Section 9 (Deploy) | ___ risks flagged |
|
|
519
|
+
| Section 10 (Future) | Reversibility: _/5, debt items: ___ |
|
|
520
|
+
| Section 11 (Design) | ___ issues / SKIPPED (no UI scope) |
|
|
521
|
+
+--------------------------------------------------------------------+
|
|
522
|
+
| NOT in scope | written (___ items) |
|
|
523
|
+
| What already exists | written |
|
|
524
|
+
| Dream state delta | written |
|
|
525
|
+
| Error/rescue registry| ___ methods, ___ CRITICAL GAPS |
|
|
526
|
+
| Failure modes | ___ total, ___ CRITICAL GAPS |
|
|
527
|
+
| TODOS.md updates | ___ items proposed |
|
|
528
|
+
| Outside voice | ran (claude subagent) / skipped |
|
|
529
|
+
| Diagrams produced | ___ (list types) |
|
|
530
|
+
+====================================================================+
|
|
531
|
+
```
|
|
532
|
+
|
|
533
|
+
## CRITICAL RULE — How to ask questions
|
|
534
|
+
|
|
535
|
+
- **One issue = one AskUserQuestion call.** Never combine multiple issues.
|
|
536
|
+
- Describe the problem concretely, with file and line references.
|
|
537
|
+
- Present 2-3 options, including "do nothing" where reasonable.
|
|
538
|
+
- For each option: effort, risk, and maintenance burden in one line.
|
|
539
|
+
- **Map the reasoning to the engineering preferences above.**
|
|
540
|
+
- Label with issue NUMBER + option LETTER (e.g., "3A", "3B").
|
|
541
|
+
- **Escape hatch:** If a section has no issues, say so and move on. Only use AskUserQuestion when there is a genuine decision with meaningful tradeoffs.
|