@kurokeita/add-skill 1.17.2 → 1.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/skills/gh-fix-ci/SKILL.md +13 -5
- package/dist/skills/grill-me/SKILL.md +9 -0
- package/dist/skills/prd-to-tasks/SKILL.md +164 -0
- package/dist/skills/test-driven-development/SKILL.md +70 -345
- package/dist/skills/test-driven-development/deep-modules.md +33 -0
- package/dist/skills/test-driven-development/interface-design.md +35 -0
- package/dist/skills/test-driven-development/mocking.md +68 -0
- package/dist/skills/test-driven-development/refactoring.md +10 -0
- package/dist/skills/test-driven-development/tests.md +61 -0
- package/dist/skills/write-prd/SKILL.md +162 -0
- package/package.json +1 -1
- package/dist/skills/test-driven-development/testing-anti-patterns.md +0 -316
|
@@ -21,7 +21,11 @@ Prereq: authenticate with the standard GitHub CLI once (for example, run `gh aut
|
|
|
21
21
|
|
|
22
22
|
## Quick start
|
|
23
23
|
|
|
24
|
-
-
|
|
24
|
+
- Detect the interpreter first:
|
|
25
|
+
`PYTHON_BIN=$(command -v python || command -v python3 || true)`
|
|
26
|
+
- If `PYTHON_BIN` is empty, skip the bundled script and use the manual `gh` fallback workflow.
|
|
27
|
+
- Otherwise run:
|
|
28
|
+
`"$PYTHON_BIN" "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "<number-or-url>"`
|
|
25
29
|
- Add `--json` if you want machine-friendly output for summarization.
|
|
26
30
|
|
|
27
31
|
## Workflow
|
|
@@ -34,7 +38,11 @@ Prereq: authenticate with the standard GitHub CLI once (for example, run `gh aut
|
|
|
34
38
|
- If the user provides a PR number or URL, use that directly.
|
|
35
39
|
3. Inspect failing checks (GitHub Actions only).
|
|
36
40
|
- Preferred: run the bundled script (handles gh field drift and job-log fallbacks):
|
|
37
|
-
-
|
|
41
|
+
- Resolve the interpreter first:
|
|
42
|
+
`PYTHON_BIN=$(command -v python || command -v python3 || true)`
|
|
43
|
+
- If `PYTHON_BIN` is set, run:
|
|
44
|
+
`"$PYTHON_BIN" "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "<number-or-url>"`
|
|
45
|
+
- If neither `python` nor `python3` is available, skip the bundled script and use the manual fallback below.
|
|
38
46
|
- Add `--json` for machine-friendly output.
|
|
39
47
|
- Manual fallback:
|
|
40
48
|
- `gh pr checks <pr> --json name,state,bucket,link,startedAt,completedAt,workflow`
|
|
@@ -65,6 +73,6 @@ Fetch failing PR checks, pull GitHub Actions logs, and extract a failure snippet
|
|
|
65
73
|
|
|
66
74
|
Usage examples:
|
|
67
75
|
|
|
68
|
-
- `python "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "123"`
|
|
69
|
-
- `python "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "https://github.com/org/repo/pull/123" --json`
|
|
70
|
-
- `python "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --max-lines 200 --context 40`
|
|
76
|
+
- `PYTHON_BIN=$(command -v python || command -v python3 || true); test -n "$PYTHON_BIN" && "$PYTHON_BIN" "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "123"`
|
|
77
|
+
- `PYTHON_BIN=$(command -v python || command -v python3 || true); test -n "$PYTHON_BIN" && "$PYTHON_BIN" "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --pr "https://github.com/org/repo/pull/123" --json`
|
|
78
|
+
- `PYTHON_BIN=$(command -v python || command -v python3 || true); test -n "$PYTHON_BIN" && "$PYTHON_BIN" "<path-to-skill>/scripts/inspect_pr_checks.py" --repo "." --max-lines 200 --context 40`
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grill-me
|
|
3
|
+
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Grill me
|
|
7
|
+
|
|
8
|
+
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
|
|
9
|
+
If a question can be answered by exploring the codebase, explore the codebase instead.
|
|
@@ -0,0 +1,164 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: prd-to-tasks
|
|
3
|
+
description: >
|
|
4
|
+
Break a local PRD into independently grabbable implementation tasks using vertical slices, then save them under
|
|
5
|
+
.prd/{prd-name}/tasks/ in the current workspace. Use when the user wants to convert a PRD to tasks, break a spec
|
|
6
|
+
into actionable work, slice a PRD, or create implementation tasks from a local planning document. Do NOT trigger
|
|
7
|
+
for: writing PRDs, reviewing PRDs without task creation, or remote work item management systems.
|
|
8
|
+
argument-hint: "[prd-name]"
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
# PRD to Tasks
|
|
12
|
+
|
|
13
|
+
Break a local PRD into independently grabbable implementation tasks using vertical slices.
|
|
14
|
+
|
|
15
|
+
Input location:
|
|
16
|
+
|
|
17
|
+
- PRD directory: `.prd/{prd-name}/`
|
|
18
|
+
- Source PRD: `.prd/{prd-name}/prd.md`
|
|
19
|
+
|
|
20
|
+
Output location:
|
|
21
|
+
|
|
22
|
+
- Tasks directory: `.prd/{prd-name}/tasks/`
|
|
23
|
+
- Optional index: `.prd/{prd-name}/tasks/README.md`
|
|
24
|
+
|
|
25
|
+
`{prd-name}` should be a filesystem-safe kebab-case slug. If the user provides a title instead, normalize it to a slug
|
|
26
|
+
and state the slug you used.
|
|
27
|
+
|
|
28
|
+
## Step 1: Locate the PRD
|
|
29
|
+
|
|
30
|
+
If the user provided `$ARGUMENTS`, treat that as the PRD name and normalize it.
|
|
31
|
+
|
|
32
|
+
If not, ask which PRD directory to use.
|
|
33
|
+
|
|
34
|
+
Then:
|
|
35
|
+
|
|
36
|
+
1. Confirm `.prd/{prd-name}/prd.md` exists
|
|
37
|
+
2. Read it fully
|
|
38
|
+
3. Check whether `.prd/{prd-name}/tasks/` already exists
|
|
39
|
+
4. If task files already exist, summarize them and ask whether to overwrite, append, or regenerate selectively
|
|
40
|
+
|
|
41
|
+
If the PRD does not exist, stop and tell the user exactly which path is missing.
|
|
42
|
+
|
|
43
|
+
## Step 2: Explore the codebase
|
|
44
|
+
|
|
45
|
+
If you have not already explored the relevant code in this conversation, do so now:
|
|
46
|
+
|
|
47
|
+
1. Read top-level architecture guidance if present
|
|
48
|
+
2. Search for code related to the PRD concepts
|
|
49
|
+
3. Read nearby guidance files for affected subsystems
|
|
50
|
+
4. Look for similar features and test patterns
|
|
51
|
+
|
|
52
|
+
The PRD tells you *what* to build. The codebase tells you enough about *where* and *how* to create useful tasks.
|
|
53
|
+
|
|
54
|
+
## Step 3: Draft vertical slices
|
|
55
|
+
|
|
56
|
+
Break the PRD into **tracer bullet** tasks.
|
|
57
|
+
|
|
58
|
+
### What is a tracer bullet?
|
|
59
|
+
|
|
60
|
+
A tracer bullet is a thin vertical slice through the whole feature path. It proves that the relevant layers connect
|
|
61
|
+
correctly end to end. It is not a horizontal "build all backend first, then all UI" slice.
|
|
62
|
+
|
|
63
|
+
Prefer slices that are:
|
|
64
|
+
|
|
65
|
+
- Independently verifiable
|
|
66
|
+
- Demoable on their own
|
|
67
|
+
- Thin and end to end
|
|
68
|
+
- Ordered by dependency
|
|
69
|
+
|
|
70
|
+
### HITL vs AFK
|
|
71
|
+
|
|
72
|
+
Each slice should be marked:
|
|
73
|
+
|
|
74
|
+
- `HITL`: needs human input, review, testing on real devices, design signoff, or external coordination
|
|
75
|
+
- `AFK`: can be implemented and validated without waiting on a person
|
|
76
|
+
|
|
77
|
+
### Vertical slice rules
|
|
78
|
+
|
|
79
|
+
- Each slice should deliver a complete behavior, not one technical layer
|
|
80
|
+
- Prefer many thin slices over a few broad ones
|
|
81
|
+
- Use the PRD's module design to inform slicing, but do not create one task per module
|
|
82
|
+
- Order slices so blockers come first
|
|
83
|
+
|
|
84
|
+
## Step 4: Quiz the user
|
|
85
|
+
|
|
86
|
+
Present the proposed slices as a numbered list. For each slice, include:
|
|
87
|
+
|
|
88
|
+
- Title
|
|
89
|
+
- Type: `HITL` or `AFK`
|
|
90
|
+
- Blocked by
|
|
91
|
+
- User stories covered
|
|
92
|
+
|
|
93
|
+
Ask the user whether:
|
|
94
|
+
|
|
95
|
+
- The granularity is right
|
|
96
|
+
- The dependencies are correct
|
|
97
|
+
- Any slices should be merged or split
|
|
98
|
+
- The `HITL` versus `AFK` labels are accurate
|
|
99
|
+
|
|
100
|
+
Iterate until approved.
|
|
101
|
+
|
|
102
|
+
## Step 5: Write task files
|
|
103
|
+
|
|
104
|
+
Only write files after the user approves the breakdown.
|
|
105
|
+
|
|
106
|
+
Create `.prd/{prd-name}/tasks/` if needed.
|
|
107
|
+
|
|
108
|
+
Create one Markdown file per task using zero-padded numeric prefixes:
|
|
109
|
+
|
|
110
|
+
- `.prd/{prd-name}/tasks/01-{task-slug}.md`
|
|
111
|
+
- `.prd/{prd-name}/tasks/02-{task-slug}.md`
|
|
112
|
+
|
|
113
|
+
`{task-slug}` should be kebab-case and derived from the task title.
|
|
114
|
+
|
|
115
|
+
Each task file should use this structure:
|
|
116
|
+
|
|
117
|
+
```md
|
|
118
|
+
# {Task title}
|
|
119
|
+
|
|
120
|
+
- Type: AFK | HITL
|
|
121
|
+
- Status: Proposed
|
|
122
|
+
- Blocked by: None | 01-other-task
|
|
123
|
+
- Source PRD: ../prd.md
|
|
124
|
+
|
|
125
|
+
## What to Build
|
|
126
|
+
|
|
127
|
+
Concise description of the vertical slice. Describe end-to-end behavior, not a layer-by-layer checklist.
|
|
128
|
+
|
|
129
|
+
## Acceptance Criteria
|
|
130
|
+
|
|
131
|
+
- [ ] Criterion 1
|
|
132
|
+
- [ ] Criterion 2
|
|
133
|
+
- [ ] Criterion 3
|
|
134
|
+
|
|
135
|
+
## User Stories Addressed
|
|
136
|
+
|
|
137
|
+
- User story 3
|
|
138
|
+
- User story 7
|
|
139
|
+
|
|
140
|
+
## Notes
|
|
141
|
+
|
|
142
|
+
- Optional implementation hints grounded in the current codebase
|
|
143
|
+
- Reference relevant modules or concepts, but avoid binding the task to fragile file-level details unless necessary
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
Also create or update `.prd/{prd-name}/tasks/README.md` with:
|
|
147
|
+
|
|
148
|
+
- The PRD title
|
|
149
|
+
- The approved slice list in dependency order
|
|
150
|
+
- A short summary of each task
|
|
151
|
+
|
|
152
|
+
## Important constraints
|
|
153
|
+
|
|
154
|
+
- Do not modify `prd.md` unless the user asks
|
|
155
|
+
- Do not delete existing task files unless the user explicitly approves replacement
|
|
156
|
+
- If existing tasks overlap the proposed slices, call that out before writing duplicates
|
|
157
|
+
|
|
158
|
+
## After writing
|
|
159
|
+
|
|
160
|
+
Summarize:
|
|
161
|
+
|
|
162
|
+
- The PRD slug used
|
|
163
|
+
- The task files created or updated
|
|
164
|
+
- Any existing task files that were preserved
|
|
@@ -1,389 +1,114 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
3
|
-
description: Use when
|
|
2
|
+
name: tdd
|
|
3
|
+
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# Test-Driven Development
|
|
6
|
+
# Test-Driven Development
|
|
7
7
|
|
|
8
|
-
##
|
|
8
|
+
## Philosophy
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
|
|
11
11
|
|
|
12
|
-
**
|
|
12
|
+
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
|
|
13
13
|
|
|
14
|
-
**
|
|
14
|
+
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
|
|
15
15
|
|
|
16
|
-
|
|
16
|
+
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
|
|
17
17
|
|
|
18
|
-
|
|
18
|
+
## Anti-Pattern: Horizontal Slices
|
|
19
19
|
|
|
20
|
-
-
|
|
21
|
-
- Bug fixes
|
|
22
|
-
- Refactoring
|
|
23
|
-
- Behavior changes
|
|
20
|
+
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
|
|
24
21
|
|
|
25
|
-
|
|
22
|
+
This produces **crap tests**:
|
|
26
23
|
|
|
27
|
-
-
|
|
28
|
-
-
|
|
29
|
-
-
|
|
24
|
+
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
|
|
25
|
+
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
|
|
26
|
+
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
|
|
27
|
+
- You outrun your headlights, committing to test structure before understanding the implementation
|
|
30
28
|
|
|
31
|
-
|
|
29
|
+
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
|
|
32
30
|
|
|
33
|
-
## The Iron Law
|
|
34
|
-
|
|
35
|
-
```
|
|
36
|
-
NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
|
|
37
31
|
```
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
Implement fresh from tests. Period.
|
|
49
|
-
|
|
50
|
-
## Red-Green-Refactor
|
|
51
|
-
|
|
52
|
-
```dot
|
|
53
|
-
digraph tdd_cycle {
|
|
54
|
-
rankdir=LR;
|
|
55
|
-
red [label="RED\nWrite failing test", shape=box, style=filled, fillcolor="#ffcccc"];
|
|
56
|
-
verify_red [label="Verify fails\ncorrectly", shape=diamond];
|
|
57
|
-
green [label="GREEN\nMinimal code", shape=box, style=filled, fillcolor="#ccffcc"];
|
|
58
|
-
verify_green [label="Verify passes\nAll green", shape=diamond];
|
|
59
|
-
refactor [label="REFACTOR\nClean up", shape=box, style=filled, fillcolor="#ccccff"];
|
|
60
|
-
next [label="Next", shape=ellipse];
|
|
61
|
-
|
|
62
|
-
red -> verify_red;
|
|
63
|
-
verify_red -> green [label="yes"];
|
|
64
|
-
verify_red -> red [label="wrong\nfailure"];
|
|
65
|
-
green -> verify_green;
|
|
66
|
-
verify_green -> refactor [label="yes"];
|
|
67
|
-
verify_green -> green [label="no"];
|
|
68
|
-
refactor -> verify_green [label="stay\ngreen"];
|
|
69
|
-
verify_green -> next;
|
|
70
|
-
next -> red;
|
|
71
|
-
}
|
|
32
|
+
WRONG (horizontal):
|
|
33
|
+
RED: test1, test2, test3, test4, test5
|
|
34
|
+
GREEN: impl1, impl2, impl3, impl4, impl5
|
|
35
|
+
|
|
36
|
+
RIGHT (vertical):
|
|
37
|
+
RED→GREEN: test1→impl1
|
|
38
|
+
RED→GREEN: test2→impl2
|
|
39
|
+
RED→GREEN: test3→impl3
|
|
40
|
+
...
|
|
72
41
|
```
|
|
73
42
|
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
Write one minimal test showing what should happen.
|
|
77
|
-
|
|
78
|
-
<Good>
|
|
79
|
-
```typescript
|
|
80
|
-
test('retries failed operations 3 times', async () => {
|
|
81
|
-
let attempts = 0;
|
|
82
|
-
const operation = () => {
|
|
83
|
-
attempts++;
|
|
84
|
-
if (attempts < 3) throw new Error('fail');
|
|
85
|
-
return 'success';
|
|
86
|
-
};
|
|
43
|
+
## Workflow
|
|
87
44
|
|
|
88
|
-
|
|
45
|
+
### 1. Planning
|
|
89
46
|
|
|
90
|
-
|
|
91
|
-
expect(attempts).toBe(3);
|
|
92
|
-
});
|
|
47
|
+
Before writing any code:
|
|
93
48
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
test('retry works', async () => {
|
|
101
|
-
const mock = jest.fn()
|
|
102
|
-
.mockRejectedValueOnce(new Error())
|
|
103
|
-
.mockRejectedValueOnce(new Error())
|
|
104
|
-
.mockResolvedValueOnce('success');
|
|
105
|
-
await retryOperation(mock);
|
|
106
|
-
expect(mock).toHaveBeenCalledTimes(3);
|
|
107
|
-
});
|
|
108
|
-
```
|
|
109
|
-
|
|
110
|
-
Vague name, tests mock not code
|
|
111
|
-
</Bad>
|
|
49
|
+
- [ ] Confirm with user what interface changes are needed
|
|
50
|
+
- [ ] Confirm with user which behaviors to test (prioritize)
|
|
51
|
+
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
|
|
52
|
+
- [ ] Design interfaces for [testability](interface-design.md)
|
|
53
|
+
- [ ] List the behaviors to test (not implementation steps)
|
|
54
|
+
- [ ] Get user approval on the plan
|
|
112
55
|
|
|
113
|
-
|
|
56
|
+
Ask: "What should the public interface look like? Which behaviors are most important to test?"
|
|
114
57
|
|
|
115
|
-
|
|
116
|
-
- Clear name
|
|
117
|
-
- Real code (no mocks unless unavoidable)
|
|
58
|
+
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
|
|
118
59
|
|
|
119
|
-
###
|
|
60
|
+
### 2. Tracer Bullet
|
|
120
61
|
|
|
121
|
-
|
|
62
|
+
Write ONE test that confirms ONE thing about the system:
|
|
122
63
|
|
|
123
|
-
```bash
|
|
124
|
-
npm test path/to/test.test.ts
|
|
125
64
|
```
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
- Failure message is expected
|
|
131
|
-
- Fails because feature missing (not typos)
|
|
132
|
-
|
|
133
|
-
**Test passes?** You're testing existing behavior. Fix test.
|
|
134
|
-
|
|
135
|
-
**Test errors?** Fix error, re-run until it fails correctly.
|
|
136
|
-
|
|
137
|
-
### GREEN - Minimal Code
|
|
138
|
-
|
|
139
|
-
Write simplest code to pass the test.
|
|
140
|
-
|
|
141
|
-
<Good>
|
|
142
|
-
```typescript
|
|
143
|
-
async function retryOperation<T>(fn: () => Promise<T>): Promise<T> {
|
|
144
|
-
for (let i = 0; i < 3; i++) {
|
|
145
|
-
try {
|
|
146
|
-
return await fn();
|
|
147
|
-
} catch (e) {
|
|
148
|
-
if (i === 2) throw e;
|
|
149
|
-
}
|
|
150
|
-
}
|
|
151
|
-
throw new Error('unreachable');
|
|
152
|
-
}
|
|
153
|
-
```
|
|
154
|
-
Just enough to pass
|
|
155
|
-
</Good>
|
|
156
|
-
|
|
157
|
-
<Bad>
|
|
158
|
-
```typescript
|
|
159
|
-
async function retryOperation<T>(
|
|
160
|
-
fn: () => Promise<T>,
|
|
161
|
-
options?: {
|
|
162
|
-
maxRetries?: number;
|
|
163
|
-
backoff?: 'linear' | 'exponential';
|
|
164
|
-
onRetry?: (attempt: number) => void;
|
|
165
|
-
}
|
|
166
|
-
): Promise<T> {
|
|
167
|
-
// YAGNI
|
|
168
|
-
}
|
|
169
|
-
```
|
|
170
|
-
Over-engineered
|
|
171
|
-
</Bad>
|
|
172
|
-
|
|
173
|
-
Don't add features, refactor other code, or "improve" beyond the test.
|
|
174
|
-
|
|
175
|
-
### Verify GREEN - Watch It Pass
|
|
176
|
-
|
|
177
|
-
**MANDATORY.**
|
|
178
|
-
|
|
179
|
-
```bash
|
|
180
|
-
npm test path/to/test.test.ts
|
|
65
|
+
RED: Write test for first behavior
|
|
66
|
+
↓ RUN TESTS — confirm it fails for the right reason
|
|
67
|
+
GREEN: Write minimal code to pass
|
|
68
|
+
↓ RUN TESTS — confirm it passes
|
|
181
69
|
```
|
|
182
70
|
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
- Test passes
|
|
186
|
-
- Other tests still pass
|
|
187
|
-
- Output pristine (no errors, warnings)
|
|
188
|
-
|
|
189
|
-
**Test fails?** Fix code, not test.
|
|
190
|
-
|
|
191
|
-
**Other tests fail?** Fix now.
|
|
192
|
-
|
|
193
|
-
### REFACTOR - Clean Up
|
|
194
|
-
|
|
195
|
-
After green only:
|
|
196
|
-
|
|
197
|
-
- Remove duplication
|
|
198
|
-
- Improve names
|
|
199
|
-
- Extract helpers
|
|
200
|
-
|
|
201
|
-
Keep tests green. Don't add behavior.
|
|
202
|
-
|
|
203
|
-
### Repeat
|
|
204
|
-
|
|
205
|
-
Next failing test for next feature.
|
|
206
|
-
|
|
207
|
-
## Good Tests
|
|
208
|
-
|
|
209
|
-
| Quality | Good | Bad |
|
|
210
|
-
|---------|------|-----|
|
|
211
|
-
| **Minimal** | One thing. "and" in name? Split it. | `test('validates email and domain and whitespace')` |
|
|
212
|
-
| **Clear** | Name describes behavior | `test('test1')` |
|
|
213
|
-
| **Shows intent** | Demonstrates desired API | Obscures what code should do |
|
|
214
|
-
|
|
215
|
-
## Why Order Matters
|
|
216
|
-
|
|
217
|
-
**"I'll write tests after to verify it works"**
|
|
218
|
-
|
|
219
|
-
Tests written after code pass immediately. Passing immediately proves nothing:
|
|
220
|
-
|
|
221
|
-
- Might test wrong thing
|
|
222
|
-
- Might test implementation, not behavior
|
|
223
|
-
- Might miss edge cases you forgot
|
|
224
|
-
- You never saw it catch the bug
|
|
225
|
-
|
|
226
|
-
Test-first forces you to see the test fail, proving it actually tests something.
|
|
227
|
-
|
|
228
|
-
**"I already manually tested all the edge cases"**
|
|
71
|
+
This is your tracer bullet - proves the path works end-to-end.
|
|
229
72
|
|
|
230
|
-
|
|
73
|
+
**You MUST run the test before writing any implementation.** A test that was never observed to fail might always pass vacuously, be testing the wrong thing, or be broken. The failure message tells you what the test is actually checking.
|
|
231
74
|
|
|
232
|
-
|
|
233
|
-
- Can't re-run when code changes
|
|
234
|
-
- Easy to forget cases under pressure
|
|
235
|
-
- "It worked when I tried it" ≠ comprehensive
|
|
75
|
+
### 3. Incremental Loop
|
|
236
76
|
|
|
237
|
-
|
|
77
|
+
For each remaining behavior:
|
|
238
78
|
|
|
239
|
-
**"Deleting X hours of work is wasteful"**
|
|
240
|
-
|
|
241
|
-
Sunk cost fallacy. The time is already gone. Your choice now:
|
|
242
|
-
|
|
243
|
-
- Delete and rewrite with TDD (X more hours, high confidence)
|
|
244
|
-
- Keep it and add tests after (30 min, low confidence, likely bugs)
|
|
245
|
-
|
|
246
|
-
The "waste" is keeping code you can't trust. Working code without real tests is technical debt.
|
|
247
|
-
|
|
248
|
-
**"TDD is dogmatic, being pragmatic means adapting"**
|
|
249
|
-
|
|
250
|
-
TDD IS pragmatic:
|
|
251
|
-
|
|
252
|
-
- Finds bugs before commit (faster than debugging after)
|
|
253
|
-
- Prevents regressions (tests catch breaks immediately)
|
|
254
|
-
- Documents behavior (tests show how to use code)
|
|
255
|
-
- Enables refactoring (change freely, tests catch breaks)
|
|
256
|
-
|
|
257
|
-
"Pragmatic" shortcuts = debugging in production = slower.
|
|
258
|
-
|
|
259
|
-
**"Tests after achieve the same goals - it's spirit not ritual"**
|
|
260
|
-
|
|
261
|
-
No. Tests-after answer "What does this do?" Tests-first answer "What should this do?"
|
|
262
|
-
|
|
263
|
-
Tests-after are biased by your implementation. You test what you built, not what's required. You verify remembered edge cases, not discovered ones.
|
|
264
|
-
|
|
265
|
-
Tests-first force edge case discovery before implementing. Tests-after verify you remembered everything (you didn't).
|
|
266
|
-
|
|
267
|
-
30 minutes of tests after ≠ TDD. You get coverage, lose proof tests work.
|
|
268
|
-
|
|
269
|
-
## Common Rationalizations
|
|
270
|
-
|
|
271
|
-
| Excuse | Reality |
|
|
272
|
-
|--------|---------|
|
|
273
|
-
| "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
|
|
274
|
-
| "I'll test after" | Tests passing immediately prove nothing. |
|
|
275
|
-
| "Tests after achieve same goals" | Tests-after = "what does this do?" Tests-first = "what should this do?" |
|
|
276
|
-
| "Already manually tested" | Ad-hoc ≠ systematic. No record, can't re-run. |
|
|
277
|
-
| "Deleting X hours is wasteful" | Sunk cost fallacy. Keeping unverified code is technical debt. |
|
|
278
|
-
| "Keep as reference, write tests first" | You'll adapt it. That's testing after. Delete means delete. |
|
|
279
|
-
| "Need to explore first" | Fine. Throw away exploration, start with TDD. |
|
|
280
|
-
| "Test hard = design unclear" | Listen to test. Hard to test = hard to use. |
|
|
281
|
-
| "TDD will slow me down" | TDD faster than debugging. Pragmatic = test-first. |
|
|
282
|
-
| "Manual test faster" | Manual doesn't prove edge cases. You'll re-test every change. |
|
|
283
|
-
| "Existing code has no tests" | You're improving it. Add tests for existing code. |
|
|
284
|
-
|
|
285
|
-
## Red Flags - STOP and Start Over
|
|
286
|
-
|
|
287
|
-
- Code before test
|
|
288
|
-
- Test after implementation
|
|
289
|
-
- Test passes immediately
|
|
290
|
-
- Can't explain why test failed
|
|
291
|
-
- Tests added "later"
|
|
292
|
-
- Rationalizing "just this once"
|
|
293
|
-
- "I already manually tested it"
|
|
294
|
-
- "Tests after achieve the same purpose"
|
|
295
|
-
- "It's about spirit not ritual"
|
|
296
|
-
- "Keep as reference" or "adapt existing code"
|
|
297
|
-
- "Already spent X hours, deleting is wasteful"
|
|
298
|
-
- "TDD is dogmatic, I'm being pragmatic"
|
|
299
|
-
- "This is different because..."
|
|
300
|
-
|
|
301
|
-
**All of these mean: Delete code. Start over with TDD.**
|
|
302
|
-
|
|
303
|
-
## Example: Bug Fix
|
|
304
|
-
|
|
305
|
-
**Bug:** Empty email accepted
|
|
306
|
-
|
|
307
|
-
**RED**
|
|
308
|
-
|
|
309
|
-
```typescript
|
|
310
|
-
test('rejects empty email', async () => {
|
|
311
|
-
const result = await submitForm({ email: '' });
|
|
312
|
-
expect(result.error).toBe('Email required');
|
|
313
|
-
});
|
|
314
|
-
```
|
|
315
|
-
|
|
316
|
-
**Verify RED**
|
|
317
|
-
|
|
318
|
-
```bash
|
|
319
|
-
$ npm test
|
|
320
|
-
FAIL: expected 'Email required', got undefined
|
|
321
79
|
```
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
function submitForm(data: FormData) {
|
|
327
|
-
if (!data.email?.trim()) {
|
|
328
|
-
return { error: 'Email required' };
|
|
329
|
-
}
|
|
330
|
-
// ...
|
|
331
|
-
}
|
|
332
|
-
```
|
|
333
|
-
|
|
334
|
-
**Verify GREEN**
|
|
335
|
-
|
|
336
|
-
```bash
|
|
337
|
-
$ npm test
|
|
338
|
-
PASS
|
|
80
|
+
RED: Write next test
|
|
81
|
+
↓ RUN TESTS — confirm new test fails, existing pass
|
|
82
|
+
GREEN: Minimal code to pass
|
|
83
|
+
↓ RUN TESTS — confirm all pass
|
|
339
84
|
```
|
|
340
85
|
|
|
341
|
-
|
|
342
|
-
Extract validation for multiple fields if needed.
|
|
343
|
-
|
|
344
|
-
## Verification Checklist
|
|
345
|
-
|
|
346
|
-
Before marking work complete:
|
|
347
|
-
|
|
348
|
-
- [ ] Every new function/method has a test
|
|
349
|
-
- [ ] Watched each test fail before implementing
|
|
350
|
-
- [ ] Each test failed for expected reason (feature missing, not typo)
|
|
351
|
-
- [ ] Wrote minimal code to pass each test
|
|
352
|
-
- [ ] All tests pass
|
|
353
|
-
- [ ] Output pristine (no errors, warnings)
|
|
354
|
-
- [ ] Tests use real code (mocks only if unavoidable)
|
|
355
|
-
- [ ] Edge cases and errors covered
|
|
356
|
-
|
|
357
|
-
Can't check all boxes? You skipped TDD. Start over.
|
|
86
|
+
Rules:
|
|
358
87
|
|
|
359
|
-
|
|
88
|
+
- One test at a time
|
|
89
|
+
- Run tests after writing each test (confirm RED) and after each implementation step (confirm GREEN)
|
|
90
|
+
- Only enough code to pass current test
|
|
91
|
+
- Don't anticipate future tests
|
|
92
|
+
- Keep tests focused on observable behavior
|
|
360
93
|
|
|
361
|
-
|
|
362
|
-
|---------|----------|
|
|
363
|
-
| Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. |
|
|
364
|
-
| Test too complicated | Design too complicated. Simplify interface. |
|
|
365
|
-
| Must mock everything | Code too coupled. Use dependency injection. |
|
|
366
|
-
| Test setup huge | Extract helpers. Still complex? Simplify design. |
|
|
94
|
+
### 4. Refactor
|
|
367
95
|
|
|
368
|
-
|
|
96
|
+
After all tests pass, look for [refactor candidates](refactoring.md):
|
|
369
97
|
|
|
370
|
-
|
|
98
|
+
- [ ] Extract duplication
|
|
99
|
+
- [ ] Deepen modules (move complexity behind simple interfaces)
|
|
100
|
+
- [ ] Apply SOLID principles where natural
|
|
101
|
+
- [ ] Consider what new code reveals about existing code
|
|
102
|
+
- [ ] Run tests after each refactor step
|
|
371
103
|
|
|
372
|
-
Never
|
|
104
|
+
**Never refactor while RED.** Get to GREEN first.
|
|
373
105
|
|
|
374
|
-
##
|
|
375
|
-
|
|
376
|
-
When adding mocks or test utilities, read @testing-anti-patterns.md to avoid common pitfalls:
|
|
377
|
-
|
|
378
|
-
- Testing mock behavior instead of real behavior
|
|
379
|
-
- Adding test-only methods to production classes
|
|
380
|
-
- Mocking without understanding dependencies
|
|
381
|
-
|
|
382
|
-
## Final Rule
|
|
106
|
+
## Checklist Per Cycle
|
|
383
107
|
|
|
384
108
|
```
|
|
385
|
-
|
|
386
|
-
|
|
109
|
+
[ ] Test describes behavior, not implementation
|
|
110
|
+
[ ] Test uses public interface only
|
|
111
|
+
[ ] Test would survive internal refactor
|
|
112
|
+
[ ] Code is minimal for this test
|
|
113
|
+
[ ] No speculative features added
|
|
387
114
|
```
|
|
388
|
-
|
|
389
|
-
No exceptions without your human partner's permission.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
# Deep Modules
|
|
2
|
+
|
|
3
|
+
From "A Philosophy of Software Design":
|
|
4
|
+
|
|
5
|
+
**Deep module** = small interface + lots of implementation
|
|
6
|
+
|
|
7
|
+
```
|
|
8
|
+
┌─────────────────────┐
|
|
9
|
+
│ Small Interface │ ← Few methods, simple params
|
|
10
|
+
├─────────────────────┤
|
|
11
|
+
│ │
|
|
12
|
+
│ │
|
|
13
|
+
│ Deep Implementation│ ← Complex logic hidden
|
|
14
|
+
│ │
|
|
15
|
+
│ │
|
|
16
|
+
└─────────────────────┘
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
**Shallow module** = large interface + little implementation (avoid)
|
|
20
|
+
|
|
21
|
+
```
|
|
22
|
+
┌─────────────────────────────────┐
|
|
23
|
+
│ Large Interface │ ← Many methods, complex params
|
|
24
|
+
├─────────────────────────────────┤
|
|
25
|
+
│ Thin Implementation │ ← Just passes through
|
|
26
|
+
└─────────────────────────────────┘
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
When designing interfaces, ask:
|
|
30
|
+
|
|
31
|
+
- Can I reduce the number of methods?
|
|
32
|
+
- Can I simplify the parameters?
|
|
33
|
+
- Can I hide more complexity inside?
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Interface Design for Testability
|
|
2
|
+
|
|
3
|
+
Good interfaces make testing natural:
|
|
4
|
+
|
|
5
|
+
1. **Accept dependencies, don't create them**
|
|
6
|
+
|
|
7
|
+
```typescript
|
|
8
|
+
// Testable
|
|
9
|
+
function processOrder(order, paymentGateway) {}
|
|
10
|
+
|
|
11
|
+
// Hard to test
|
|
12
|
+
function processOrder(order) {
|
|
13
|
+
const gateway = new StripeGateway();
|
|
14
|
+
}
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
2. **Return results, don't produce side effects**
|
|
18
|
+
|
|
19
|
+
```typescript
|
|
20
|
+
// Testable
|
|
21
|
+
function calculateDiscount(cart): Discount {}
|
|
22
|
+
|
|
23
|
+
// Hard to test
|
|
24
|
+
function applyDiscount(cart): void {
|
|
25
|
+
cart.total -= discount;
|
|
26
|
+
}
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
3. **Small surface area**
|
|
30
|
+
- Fewer methods = fewer tests needed
|
|
31
|
+
- Fewer params = simpler test setup
|
|
32
|
+
|
|
33
|
+
4. **Keep test-only helpers out of production APIs**
|
|
34
|
+
|
|
35
|
+
If cleanup or inspection logic exists only for tests, put it in test utilities rather than adding methods to production classes. Production interfaces should reflect real runtime behavior, not test harness needs.
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
# When to Mock
|
|
2
|
+
|
|
3
|
+
Mock at **system boundaries** only:
|
|
4
|
+
|
|
5
|
+
- External APIs (payment, email, etc.)
|
|
6
|
+
- Databases (sometimes - prefer test DB)
|
|
7
|
+
- Time/randomness
|
|
8
|
+
- File system (sometimes)
|
|
9
|
+
|
|
10
|
+
Don't mock:
|
|
11
|
+
|
|
12
|
+
- Your own classes/modules
|
|
13
|
+
- Internal collaborators
|
|
14
|
+
- Anything you control
|
|
15
|
+
|
|
16
|
+
## Designing for Mockability
|
|
17
|
+
|
|
18
|
+
At system boundaries, design interfaces that are easy to mock:
|
|
19
|
+
|
|
20
|
+
**1. Use dependency injection**
|
|
21
|
+
|
|
22
|
+
Pass external dependencies in rather than creating them internally:
|
|
23
|
+
|
|
24
|
+
```typescript
|
|
25
|
+
// Easy to mock
|
|
26
|
+
function processPayment(order, paymentClient) {
|
|
27
|
+
return paymentClient.charge(order.total);
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
// Hard to mock
|
|
31
|
+
function processPayment(order) {
|
|
32
|
+
const client = new StripeClient(process.env.STRIPE_KEY);
|
|
33
|
+
return client.charge(order.total);
|
|
34
|
+
}
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
**2. Prefer SDK-style interfaces over generic fetchers**
|
|
38
|
+
|
|
39
|
+
Create specific functions for each external operation instead of one generic function with conditional logic:
|
|
40
|
+
|
|
41
|
+
```typescript
|
|
42
|
+
// GOOD: Each function is independently mockable
|
|
43
|
+
const api = {
|
|
44
|
+
getUser: (id) => fetch(`/users/${id}`),
|
|
45
|
+
getOrders: (userId) => fetch(`/users/${userId}/orders`),
|
|
46
|
+
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
|
|
47
|
+
};
|
|
48
|
+
|
|
49
|
+
// BAD: Mocking requires conditional logic inside the mock
|
|
50
|
+
const api = {
|
|
51
|
+
fetch: (endpoint, options) => fetch(endpoint, options),
|
|
52
|
+
};
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
The SDK approach means:
|
|
56
|
+
|
|
57
|
+
- Each mock returns one specific shape
|
|
58
|
+
- No conditional logic in test setup
|
|
59
|
+
- Easier to see which endpoints a test exercises
|
|
60
|
+
- Type safety per endpoint
|
|
61
|
+
|
|
62
|
+
## Mock Faithfully
|
|
63
|
+
|
|
64
|
+
When a boundary must be mocked, preserve the parts of reality the test depends on:
|
|
65
|
+
|
|
66
|
+
- Include the full response shape that downstream code relies on, not just the fields used in the immediate assertion
|
|
67
|
+
- Avoid mocking away side effects the behavior under test actually needs
|
|
68
|
+
- If unsure what must remain real, run the test against the real path first and then mock the lowest external boundary
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
# Refactor Candidates
|
|
2
|
+
|
|
3
|
+
After TDD cycle, look for:
|
|
4
|
+
|
|
5
|
+
- **Duplication** → Extract function/class
|
|
6
|
+
- **Long methods** → Break into private helpers (keep tests on public interface)
|
|
7
|
+
- **Shallow modules** → Combine or deepen
|
|
8
|
+
- **Feature envy** → Move logic to where data lives
|
|
9
|
+
- **Primitive obsession** → Introduce value objects
|
|
10
|
+
- **Existing code** the new code reveals as problematic
|
|
@@ -0,0 +1,61 @@
|
|
|
1
|
+
# Good and Bad Tests
|
|
2
|
+
|
|
3
|
+
## Good Tests
|
|
4
|
+
|
|
5
|
+
**Integration-style**: Test through real interfaces, not mocks of internal parts.
|
|
6
|
+
|
|
7
|
+
```typescript
|
|
8
|
+
// GOOD: Tests observable behavior
|
|
9
|
+
test("user can checkout with valid cart", async () => {
|
|
10
|
+
const cart = createCart();
|
|
11
|
+
cart.add(product);
|
|
12
|
+
const result = await checkout(cart, paymentMethod);
|
|
13
|
+
expect(result.status).toBe("confirmed");
|
|
14
|
+
});
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Characteristics:
|
|
18
|
+
|
|
19
|
+
- Tests behavior users/callers care about
|
|
20
|
+
- Uses public API only
|
|
21
|
+
- Survives internal refactors
|
|
22
|
+
- Describes WHAT, not HOW
|
|
23
|
+
- One logical assertion per test
|
|
24
|
+
|
|
25
|
+
## Bad Tests
|
|
26
|
+
|
|
27
|
+
**Implementation-detail tests**: Coupled to internal structure.
|
|
28
|
+
|
|
29
|
+
```typescript
|
|
30
|
+
// BAD: Tests implementation details
|
|
31
|
+
test("checkout calls paymentService.process", async () => {
|
|
32
|
+
const mockPayment = jest.mock(paymentService);
|
|
33
|
+
await checkout(cart, payment);
|
|
34
|
+
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
|
|
35
|
+
});
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
Red flags:
|
|
39
|
+
|
|
40
|
+
- Mocking internal collaborators
|
|
41
|
+
- Testing private methods
|
|
42
|
+
- Asserting on call counts/order
|
|
43
|
+
- Test breaks when refactoring without behavior change
|
|
44
|
+
- Test name describes HOW not WHAT
|
|
45
|
+
- Verifying through external means instead of interface
|
|
46
|
+
|
|
47
|
+
```typescript
|
|
48
|
+
// BAD: Bypasses interface to verify
|
|
49
|
+
test("createUser saves to database", async () => {
|
|
50
|
+
await createUser({ name: "Alice" });
|
|
51
|
+
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
|
|
52
|
+
expect(row).toBeDefined();
|
|
53
|
+
});
|
|
54
|
+
|
|
55
|
+
// GOOD: Verifies through interface
|
|
56
|
+
test("createUser makes user retrievable", async () => {
|
|
57
|
+
const user = await createUser({ name: "Alice" });
|
|
58
|
+
const retrieved = await getUser(user.id);
|
|
59
|
+
expect(retrieved.name).toBe("Alice");
|
|
60
|
+
});
|
|
61
|
+
```
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: write-prd
|
|
3
|
+
description: >
|
|
4
|
+
Create a developer-focused PRD through user interview, codebase exploration, and module design, then save it
|
|
5
|
+
under .prd/{prd-name}/ in the current workspace. Use when the user wants to write a PRD, create a product
|
|
6
|
+
requirements document, plan a new feature, spec out work, or define requirements for implementation. Also
|
|
7
|
+
trigger when the user says things like "let's plan this feature", "I need a spec for...", "write requirements
|
|
8
|
+
for...", or "create a PRD for...". Do NOT trigger for: reviewing existing PRDs, creating child tasks from a PRD,
|
|
9
|
+
or general project management outside the local PRD directory.
|
|
10
|
+
argument-hint: "[prd-name]"
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# Write a PRD
|
|
14
|
+
|
|
15
|
+
Create a developer-focused Product Requirements Document and save it to the local workspace instead of a remote tracker.
|
|
16
|
+
The PRD captures the *what* and *why* of a feature, not line-level implementation details.
|
|
17
|
+
|
|
18
|
+
During the interview phase, explicitly use the `grill-me` skill if it is available in the current workspace. Treat
|
|
19
|
+
that skill as the driver for the questioning loop, while this skill remains responsible for codebase exploration,
|
|
20
|
+
module design, PRD drafting, and filesystem output.
|
|
21
|
+
|
|
22
|
+
Output location:
|
|
23
|
+
|
|
24
|
+
- Directory: `.prd/{prd-name}/`
|
|
25
|
+
- Primary document: `.prd/{prd-name}/prd.md`
|
|
26
|
+
- Optional notes captured during planning: `.prd/{prd-name}/notes.md`
|
|
27
|
+
|
|
28
|
+
`{prd-name}` should be a filesystem-safe kebab-case slug. If the user provides a title with spaces or punctuation,
|
|
29
|
+
normalize it to kebab-case and tell the user which slug you used.
|
|
30
|
+
|
|
31
|
+
You may skip steps if they are clearly unnecessary, but do not rush. The value of this skill is in forcing clarity.
|
|
32
|
+
|
|
33
|
+
## Step 1: Establish the PRD target
|
|
34
|
+
|
|
35
|
+
If the user provided `$ARGUMENTS`, treat that as the intended PRD name and normalize it to a slug.
|
|
36
|
+
|
|
37
|
+
If not, ask for the PRD name before writing files.
|
|
38
|
+
|
|
39
|
+
Before creating files:
|
|
40
|
+
|
|
41
|
+
- Check whether `.prd/{prd-name}/` already exists
|
|
42
|
+
- If it exists, read `prd.md` if present and summarize the current state
|
|
43
|
+
- Ask whether the user wants to overwrite the PRD, revise it in place, or create a new slug
|
|
44
|
+
|
|
45
|
+
## Step 2: Get the problem description
|
|
46
|
+
|
|
47
|
+
Ask the user for a long, detailed description of the problem they want to solve and any candidate solutions they already
|
|
48
|
+
have in mind.
|
|
49
|
+
|
|
50
|
+
If an existing `.prd/{prd-name}/prd.md` exists, summarize it first so the user can correct or extend it instead of repeating
|
|
51
|
+
context.
|
|
52
|
+
|
|
53
|
+
## Step 3: Explore the codebase
|
|
54
|
+
|
|
55
|
+
Before interviewing, ground yourself in the actual codebase:
|
|
56
|
+
|
|
57
|
+
1. Read the top-level `README.md`, `AGENTS.md`, or other architectural guidance if present
|
|
58
|
+
2. Search broadly for relevant terms, domain concepts, and adjacent implementations
|
|
59
|
+
3. Read area-specific guidance files when relevant
|
|
60
|
+
4. Look for similar features, patterns, and existing terminology
|
|
61
|
+
|
|
62
|
+
Do not try to read everything. Use search-driven exploration to understand the current state well enough to ask informed
|
|
63
|
+
questions and write an accurate PRD.
|
|
64
|
+
|
|
65
|
+
## Step 4: Interview relentlessly
|
|
66
|
+
|
|
67
|
+
Invoke the `grill-me` skill for this phase if available, then interview the user until you reach shared understanding.
|
|
68
|
+
Walk down each branch of the design tree and resolve dependencies one by one.
|
|
69
|
+
|
|
70
|
+
Ask one question at a time. Keep going until ambiguity is removed.
|
|
71
|
+
|
|
72
|
+
Ask about:
|
|
73
|
+
|
|
74
|
+
- Edge cases and failure scenarios
|
|
75
|
+
- Existing behavior that must be preserved
|
|
76
|
+
- What success looks like from the user's perspective
|
|
77
|
+
- Who the end users are
|
|
78
|
+
- Constraints from architecture, rollout, compatibility, or operations
|
|
79
|
+
|
|
80
|
+
Continue exploring the codebase while interviewing. If a question can be answered by inspecting code, inspect code instead
|
|
81
|
+
of asking the user.
|
|
82
|
+
|
|
83
|
+
### Scope check
|
|
84
|
+
|
|
85
|
+
If it becomes clear the work is too large for one PRD, tell the user and ask whether to narrow the scope or continue with
|
|
86
|
+
a larger document that will later be split into multiple execution tracks.
|
|
87
|
+
|
|
88
|
+
## Step 5: Sketch module design
|
|
89
|
+
|
|
90
|
+
Sketch the major conceptual modules to build or modify. Look for opportunities to define **deep modules**:
|
|
91
|
+
|
|
92
|
+
- Interfaces simpler than the implementation they hide
|
|
93
|
+
- Testable in isolation
|
|
94
|
+
- Internals can change without rippling outward
|
|
95
|
+
|
|
96
|
+
Check that these modules match the user's expectations and ask which ones need stronger test coverage.
|
|
97
|
+
|
|
98
|
+
## Step 6: Draft the PRD and get approval
|
|
99
|
+
|
|
100
|
+
Write the PRD and show it to the user for approval before saving or overwriting `prd.md`.
|
|
101
|
+
|
|
102
|
+
The PRD should be durable:
|
|
103
|
+
|
|
104
|
+
- Do describe conceptual modules, responsibilities, interfaces, and data flow
|
|
105
|
+
- Do describe architectural patterns and decisions
|
|
106
|
+
- Do not reference specific file paths, class names, or step-by-step implementation procedures
|
|
107
|
+
- Do not include code snippets unless the user explicitly asks for them
|
|
108
|
+
|
|
109
|
+
Use this structure:
|
|
110
|
+
|
|
111
|
+
```md
|
|
112
|
+
# {Human-readable PRD title}
|
|
113
|
+
|
|
114
|
+
## Problem Statement
|
|
115
|
+
|
|
116
|
+
## Solution
|
|
117
|
+
|
|
118
|
+
## User Stories
|
|
119
|
+
1. As a ...
|
|
120
|
+
|
|
121
|
+
## Module Design
|
|
122
|
+
### {Module name}
|
|
123
|
+
- Responsibility:
|
|
124
|
+
- Interface:
|
|
125
|
+
- Status: new | existing
|
|
126
|
+
- Depth: deep | shallow
|
|
127
|
+
|
|
128
|
+
## Implementation Decisions
|
|
129
|
+
|
|
130
|
+
## Testing Decisions
|
|
131
|
+
|
|
132
|
+
## Out of Scope
|
|
133
|
+
|
|
134
|
+
## Open Questions
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Omit `Open Questions` if none remain.
|
|
138
|
+
|
|
139
|
+
## Step 7: Save to the filesystem
|
|
140
|
+
|
|
141
|
+
Only save after the user approves the PRD.
|
|
142
|
+
|
|
143
|
+
Create `.prd/{prd-name}/` if it does not exist.
|
|
144
|
+
|
|
145
|
+
Write:
|
|
146
|
+
|
|
147
|
+
- `.prd/{prd-name}/prd.md`: the approved PRD
|
|
148
|
+
|
|
149
|
+
Optionally write:
|
|
150
|
+
|
|
151
|
+
- `.prd/{prd-name}/notes.md`: short working notes, unresolved investigation points, or source references gathered during
|
|
152
|
+
discovery. Keep this concise and operational, not user-facing.
|
|
153
|
+
|
|
154
|
+
Do not create extra files unless they add real value.
|
|
155
|
+
|
|
156
|
+
## After saving
|
|
157
|
+
|
|
158
|
+
Tell the user:
|
|
159
|
+
|
|
160
|
+
- The slug used
|
|
161
|
+
- The files created or updated
|
|
162
|
+
- Whether this was a new PRD or a revision of an existing directory
|
package/package.json
CHANGED
|
@@ -1,316 +0,0 @@
|
|
|
1
|
-
# Testing Anti-Patterns
|
|
2
|
-
|
|
3
|
-
**Load this reference when:** writing or changing tests, adding mocks, or tempted to add test-only methods to production code.
|
|
4
|
-
|
|
5
|
-
## Overview
|
|
6
|
-
|
|
7
|
-
Tests must verify real behavior, not mock behavior. Mocks are a means to isolate, not the thing being tested.
|
|
8
|
-
|
|
9
|
-
**Core principle:** Test what the code does, not what the mocks do.
|
|
10
|
-
|
|
11
|
-
**Following strict TDD prevents these anti-patterns.**
|
|
12
|
-
|
|
13
|
-
## The Iron Laws
|
|
14
|
-
|
|
15
|
-
```
|
|
16
|
-
1. NEVER test mock behavior
|
|
17
|
-
2. NEVER add test-only methods to production classes
|
|
18
|
-
3. NEVER mock without understanding dependencies
|
|
19
|
-
```
|
|
20
|
-
|
|
21
|
-
## Anti-Pattern 1: Testing Mock Behavior
|
|
22
|
-
|
|
23
|
-
**The violation:**
|
|
24
|
-
|
|
25
|
-
```typescript
|
|
26
|
-
// ❌ BAD: Testing that the mock exists
|
|
27
|
-
test('renders sidebar', () => {
|
|
28
|
-
render(<Page />);
|
|
29
|
-
expect(screen.getByTestId('sidebar-mock')).toBeInTheDocument();
|
|
30
|
-
});
|
|
31
|
-
```
|
|
32
|
-
|
|
33
|
-
**Why this is wrong:**
|
|
34
|
-
|
|
35
|
-
- You're verifying the mock works, not that the component works
|
|
36
|
-
- Test passes when mock is present, fails when it's not
|
|
37
|
-
- Tells you nothing about real behavior
|
|
38
|
-
|
|
39
|
-
**your human partner's correction:** "Are we testing the behavior of a mock?"
|
|
40
|
-
|
|
41
|
-
**The fix:**
|
|
42
|
-
|
|
43
|
-
```typescript
|
|
44
|
-
// ✅ GOOD: Test real component or don't mock it
|
|
45
|
-
test('renders sidebar', () => {
|
|
46
|
-
render(<Page />); // Don't mock sidebar
|
|
47
|
-
expect(screen.getByRole('navigation')).toBeInTheDocument();
|
|
48
|
-
});
|
|
49
|
-
|
|
50
|
-
// OR if sidebar must be mocked for isolation:
|
|
51
|
-
// Don't assert on the mock - test Page's behavior with sidebar present
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
### Gate Function
|
|
55
|
-
|
|
56
|
-
```
|
|
57
|
-
BEFORE asserting on any mock element:
|
|
58
|
-
Ask: "Am I testing real component behavior or just mock existence?"
|
|
59
|
-
|
|
60
|
-
IF testing mock existence:
|
|
61
|
-
STOP - Delete the assertion or unmock the component
|
|
62
|
-
|
|
63
|
-
Test real behavior instead
|
|
64
|
-
```
|
|
65
|
-
|
|
66
|
-
## Anti-Pattern 2: Test-Only Methods in Production
|
|
67
|
-
|
|
68
|
-
**The violation:**
|
|
69
|
-
|
|
70
|
-
```typescript
|
|
71
|
-
// ❌ BAD: destroy() only used in tests
|
|
72
|
-
class Session {
|
|
73
|
-
async destroy() { // Looks like production API!
|
|
74
|
-
await this._workspaceManager?.destroyWorkspace(this.id);
|
|
75
|
-
// ... cleanup
|
|
76
|
-
}
|
|
77
|
-
}
|
|
78
|
-
|
|
79
|
-
// In tests
|
|
80
|
-
afterEach(() => session.destroy());
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
**Why this is wrong:**
|
|
84
|
-
|
|
85
|
-
- Production class polluted with test-only code
|
|
86
|
-
- Dangerous if accidentally called in production
|
|
87
|
-
- Violates YAGNI and separation of concerns
|
|
88
|
-
- Confuses object lifecycle with entity lifecycle
|
|
89
|
-
|
|
90
|
-
**The fix:**
|
|
91
|
-
|
|
92
|
-
```typescript
|
|
93
|
-
// ✅ GOOD: Test utilities handle test cleanup
|
|
94
|
-
// Session has no destroy() - it's stateless in production
|
|
95
|
-
|
|
96
|
-
// In test-utils/
|
|
97
|
-
export async function cleanupSession(session: Session) {
|
|
98
|
-
const workspace = session.getWorkspaceInfo();
|
|
99
|
-
if (workspace) {
|
|
100
|
-
await workspaceManager.destroyWorkspace(workspace.id);
|
|
101
|
-
}
|
|
102
|
-
}
|
|
103
|
-
|
|
104
|
-
// In tests
|
|
105
|
-
afterEach(() => cleanupSession(session));
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
### Gate Function
|
|
109
|
-
|
|
110
|
-
```
|
|
111
|
-
BEFORE adding any method to production class:
|
|
112
|
-
Ask: "Is this only used by tests?"
|
|
113
|
-
|
|
114
|
-
IF yes:
|
|
115
|
-
STOP - Don't add it
|
|
116
|
-
Put it in test utilities instead
|
|
117
|
-
|
|
118
|
-
Ask: "Does this class own this resource's lifecycle?"
|
|
119
|
-
|
|
120
|
-
IF no:
|
|
121
|
-
STOP - Wrong class for this method
|
|
122
|
-
```
|
|
123
|
-
|
|
124
|
-
## Anti-Pattern 3: Mocking Without Understanding
|
|
125
|
-
|
|
126
|
-
**The violation:**
|
|
127
|
-
|
|
128
|
-
```typescript
|
|
129
|
-
// ❌ BAD: Mock breaks test logic
|
|
130
|
-
test('detects duplicate server', () => {
|
|
131
|
-
// Mock prevents config write that test depends on!
|
|
132
|
-
vi.mock('ToolCatalog', () => ({
|
|
133
|
-
discoverAndCacheTools: vi.fn().mockResolvedValue(undefined)
|
|
134
|
-
}));
|
|
135
|
-
|
|
136
|
-
await addServer(config);
|
|
137
|
-
await addServer(config); // Should throw - but won't!
|
|
138
|
-
});
|
|
139
|
-
```
|
|
140
|
-
|
|
141
|
-
**Why this is wrong:**
|
|
142
|
-
|
|
143
|
-
- Mocked method had side effect test depended on (writing config)
|
|
144
|
-
- Over-mocking to "be safe" breaks actual behavior
|
|
145
|
-
- Test passes for wrong reason or fails mysteriously
|
|
146
|
-
|
|
147
|
-
**The fix:**
|
|
148
|
-
|
|
149
|
-
```typescript
|
|
150
|
-
// ✅ GOOD: Mock at correct level
|
|
151
|
-
test('detects duplicate server', () => {
|
|
152
|
-
// Mock the slow part, preserve behavior test needs
|
|
153
|
-
vi.mock('MCPServerManager'); // Just mock slow server startup
|
|
154
|
-
|
|
155
|
-
await addServer(config); // Config written
|
|
156
|
-
await addServer(config); // Duplicate detected ✓
|
|
157
|
-
});
|
|
158
|
-
```
|
|
159
|
-
|
|
160
|
-
### Gate Function
|
|
161
|
-
|
|
162
|
-
```
|
|
163
|
-
BEFORE mocking any method:
|
|
164
|
-
STOP - Don't mock yet
|
|
165
|
-
|
|
166
|
-
1. Ask: "What side effects does the real method have?"
|
|
167
|
-
2. Ask: "Does this test depend on any of those side effects?"
|
|
168
|
-
3. Ask: "Do I fully understand what this test needs?"
|
|
169
|
-
|
|
170
|
-
IF depends on side effects:
|
|
171
|
-
Mock at lower level (the actual slow/external operation)
|
|
172
|
-
OR use test doubles that preserve necessary behavior
|
|
173
|
-
NOT the high-level method the test depends on
|
|
174
|
-
|
|
175
|
-
IF unsure what test depends on:
|
|
176
|
-
Run test with real implementation FIRST
|
|
177
|
-
Observe what actually needs to happen
|
|
178
|
-
THEN add minimal mocking at the right level
|
|
179
|
-
|
|
180
|
-
Red flags:
|
|
181
|
-
- "I'll mock this to be safe"
|
|
182
|
-
- "This might be slow, better mock it"
|
|
183
|
-
- Mocking without understanding the dependency chain
|
|
184
|
-
```
|
|
185
|
-
|
|
186
|
-
## Anti-Pattern 4: Incomplete Mocks
|
|
187
|
-
|
|
188
|
-
**The violation:**
|
|
189
|
-
|
|
190
|
-
```typescript
|
|
191
|
-
// ❌ BAD: Partial mock - only fields you think you need
|
|
192
|
-
const mockResponse = {
|
|
193
|
-
status: 'success',
|
|
194
|
-
data: { userId: '123', name: 'Alice' }
|
|
195
|
-
// Missing: metadata that downstream code uses
|
|
196
|
-
};
|
|
197
|
-
|
|
198
|
-
// Later: breaks when code accesses response.metadata.requestId
|
|
199
|
-
```
|
|
200
|
-
|
|
201
|
-
**Why this is wrong:**
|
|
202
|
-
|
|
203
|
-
- **Partial mocks hide structural assumptions** - You only mocked fields you know about
|
|
204
|
-
- **Downstream code may depend on fields you didn't include** - Silent failures
|
|
205
|
-
- **Tests pass but integration fails** - Mock incomplete, real API complete
|
|
206
|
-
- **False confidence** - Test proves nothing about real behavior
|
|
207
|
-
|
|
208
|
-
**The Iron Rule:** Mock the COMPLETE data structure as it exists in reality, not just fields your immediate test uses.
|
|
209
|
-
|
|
210
|
-
**The fix:**
|
|
211
|
-
|
|
212
|
-
```typescript
|
|
213
|
-
// ✅ GOOD: Mirror real API completeness
|
|
214
|
-
const mockResponse = {
|
|
215
|
-
status: 'success',
|
|
216
|
-
data: { userId: '123', name: 'Alice' },
|
|
217
|
-
metadata: { requestId: 'req-789', timestamp: 1234567890 }
|
|
218
|
-
// All fields real API returns
|
|
219
|
-
};
|
|
220
|
-
```
|
|
221
|
-
|
|
222
|
-
### Gate Function
|
|
223
|
-
|
|
224
|
-
```
|
|
225
|
-
BEFORE creating mock responses:
|
|
226
|
-
Check: "What fields does the real API response contain?"
|
|
227
|
-
|
|
228
|
-
Actions:
|
|
229
|
-
1. Examine actual API response from docs/examples
|
|
230
|
-
2. Include ALL fields system might consume downstream
|
|
231
|
-
3. Verify mock matches real response schema completely
|
|
232
|
-
|
|
233
|
-
Critical:
|
|
234
|
-
If you're creating a mock, you must understand the ENTIRE structure
|
|
235
|
-
Partial mocks fail silently when code depends on omitted fields
|
|
236
|
-
|
|
237
|
-
If uncertain: Include all documented fields
|
|
238
|
-
```
|
|
239
|
-
|
|
240
|
-
## Anti-Pattern 5: Integration Tests as Afterthought
|
|
241
|
-
|
|
242
|
-
**The violation:**
|
|
243
|
-
|
|
244
|
-
```
|
|
245
|
-
✅ Implementation complete
|
|
246
|
-
❌ No tests written
|
|
247
|
-
"Ready for testing"
|
|
248
|
-
```
|
|
249
|
-
|
|
250
|
-
**Why this is wrong:**
|
|
251
|
-
|
|
252
|
-
- Testing is part of implementation, not optional follow-up
|
|
253
|
-
- TDD would have caught this
|
|
254
|
-
- Can't claim complete without tests
|
|
255
|
-
|
|
256
|
-
**The fix:**
|
|
257
|
-
|
|
258
|
-
```
|
|
259
|
-
TDD cycle:
|
|
260
|
-
1. Write failing test
|
|
261
|
-
2. Implement to pass
|
|
262
|
-
3. Refactor
|
|
263
|
-
4. THEN claim complete
|
|
264
|
-
```
|
|
265
|
-
|
|
266
|
-
## When Mocks Become Too Complex
|
|
267
|
-
|
|
268
|
-
**Warning signs:**
|
|
269
|
-
|
|
270
|
-
- Mock setup longer than test logic
|
|
271
|
-
- Mocking everything to make test pass
|
|
272
|
-
- Mocks missing methods real components have
|
|
273
|
-
- Test breaks when mock changes
|
|
274
|
-
|
|
275
|
-
**your human partner's question:** "Do we need to be using a mock here?"
|
|
276
|
-
|
|
277
|
-
**Consider:** Integration tests with real components often simpler than complex mocks
|
|
278
|
-
|
|
279
|
-
## TDD Prevents These Anti-Patterns
|
|
280
|
-
|
|
281
|
-
**Why TDD helps:**
|
|
282
|
-
|
|
283
|
-
1. **Write test first** → Forces you to think about what you're actually testing
|
|
284
|
-
2. **Watch it fail** → Confirms test tests real behavior, not mocks
|
|
285
|
-
3. **Minimal implementation** → No test-only methods creep in
|
|
286
|
-
4. **Real dependencies** → You see what the test actually needs before mocking
|
|
287
|
-
|
|
288
|
-
**If you're testing mock behavior, you violated TDD** - you added mocks without watching test fail against real code first.
|
|
289
|
-
|
|
290
|
-
## Quick Reference
|
|
291
|
-
|
|
292
|
-
| Anti-Pattern | Fix |
|
|
293
|
-
|--------------|-----|
|
|
294
|
-
| Assert on mock elements | Test real component or unmock it |
|
|
295
|
-
| Test-only methods in production | Move to test utilities |
|
|
296
|
-
| Mock without understanding | Understand dependencies first, mock minimally |
|
|
297
|
-
| Incomplete mocks | Mirror real API completely |
|
|
298
|
-
| Tests as afterthought | TDD - tests first |
|
|
299
|
-
| Over-complex mocks | Consider integration tests |
|
|
300
|
-
|
|
301
|
-
## Red Flags
|
|
302
|
-
|
|
303
|
-
- Assertion checks for `*-mock` test IDs
|
|
304
|
-
- Methods only called in test files
|
|
305
|
-
- Mock setup is >50% of test
|
|
306
|
-
- Test fails when you remove mock
|
|
307
|
-
- Can't explain why mock is needed
|
|
308
|
-
- Mocking "just to be safe"
|
|
309
|
-
|
|
310
|
-
## The Bottom Line
|
|
311
|
-
|
|
312
|
-
**Mocks are tools to isolate, not things to test.**
|
|
313
|
-
|
|
314
|
-
If TDD reveals you're testing mock behavior, you've gone wrong.
|
|
315
|
-
|
|
316
|
-
Fix: Test real behavior or question why you're mocking at all.
|