@heihei0299/matt-skills 1.3.1 → 1.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/ask-matt/PHASE-BOUNDARIES.md +55 -0
- package/.agents/skills/ask-matt/SKILL.md +37 -25
- package/.agents/skills/ci-guard/SKILL.md +104 -0
- package/.agents/skills/ci-guard/agents/openai.yaml +5 -0
- package/.agents/skills/code-review/SKILL.md +28 -35
- package/.agents/skills/codebase-design/DEEPENING.md +4 -4
- package/.agents/skills/codebase-design/DESIGN-IT-TWICE.md +10 -10
- package/.agents/skills/codebase-design/SKILL.md +13 -13
- package/.agents/skills/diagnosing-bugs/SKILL.md +34 -30
- package/.agents/skills/diagnosing-bugs/scripts/hitl-loop.template.sh +3 -0
- package/.agents/skills/domain-modeling/ADR-FORMAT.md +11 -11
- package/.agents/skills/domain-modeling/CONTEXT-FORMAT.md +3 -3
- package/.agents/skills/domain-modeling/SKILL.md +10 -10
- package/.agents/skills/grill-me/SKILL.md +1 -1
- package/.agents/skills/grill-with-docs/SKILL.md +1 -1
- package/.agents/skills/grilling/SKILL.md +20 -4
- package/.agents/skills/grilling/agents/openai.yaml +1 -1
- package/.agents/skills/handoff/SKILL.md +1 -1
- package/.agents/skills/improve-codebase-architecture/HTML-REPORT.md +19 -19
- package/.agents/skills/improve-codebase-architecture/SKILL.md +21 -21
- package/.agents/skills/instance-test/SKILL.md +36 -27
- package/.agents/skills/instance-test/agents/openai.yaml +1 -1
- package/.agents/skills/instance-test/references/instances.md +63 -36
- package/.agents/skills/prototype/LOGIC.md +30 -42
- package/.agents/skills/prototype/SKILL.md +7 -7
- package/.agents/skills/prototype/UI.md +23 -23
- package/.agents/skills/research/SKILL.md +1 -1
- package/.agents/skills/resolving-merge-conflicts/SKILL.md +1 -1
- package/.agents/skills/scaffold-functional-test/SKILL.md +77 -0
- package/.agents/skills/scaffold-functional-test/agents/openai.yaml +5 -0
- package/.agents/skills/setup-matt-pocock-skills/SKILL.md +30 -30
- package/.agents/skills/setup-matt-pocock-skills/domain.md +4 -4
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-github.md +5 -5
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-gitlab.md +6 -6
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-local.md +3 -3
- package/.agents/skills/tdd/SKILL.md +9 -7
- package/.agents/skills/teach/GLOSSARY-FORMAT.md +3 -3
- package/.agents/skills/teach/LEARNING-RECORD-FORMAT.md +10 -10
- package/.agents/skills/teach/MISSION-FORMAT.md +4 -4
- package/.agents/skills/teach/RESOURCES-FORMAT.md +2 -2
- package/.agents/skills/teach/SKILL.md +4 -4
- package/.agents/skills/to-questionnaire/SKILL.md +54 -0
- package/.agents/skills/to-questionnaire/agents/openai.yaml +5 -0
- package/.agents/skills/to-spec/SKILL.md +4 -4
- package/.agents/skills/to-tickets/SKILL.md +16 -16
- package/.agents/skills/triage/AGENT-BRIEF.md +9 -9
- package/.agents/skills/triage/OUT-OF-SCOPE.md +15 -15
- package/.agents/skills/triage/SKILL.md +29 -29
- package/.agents/skills/wait-what/SKILL.md +7 -0
- package/.agents/skills/wait-what/agents/openai.yaml +5 -0
- package/.agents/skills/wayfinder/SKILL.md +37 -37
- package/.agents/skills/wizard/SKILL.md +44 -0
- package/.agents/skills/wizard/agents/openai.yaml +3 -0
- package/.agents/skills/wizard/template.sh +204 -0
- package/.agents/skills/writing-for-agents/SKILL-MECHANICS.md +22 -0
- package/.agents/skills/writing-for-agents/SKILL.md +81 -0
- package/.agents/skills/writing-for-agents/agents/openai.yaml +3 -0
- package/README.md +9 -9
- package/bin/cli.js +1 -1
- package/config/proprietary.json +8 -1
- package/package.json +1 -1
- package/scripts/sync-upstream.js +1 -1
- package/template/.opencode/CONTEXT.md +2 -2
- package/template/.opencode/commands/{writing-great-skills.md → writing-for-agents.md} +1 -1
- package/template/.opencode/docs/agents/skill-design.md +3 -3
- package/template/.opencode/skills/ci-guard/SKILL.md +104 -0
- package/template/.opencode/skills/ci-guard/agents/openai.yaml +5 -0
- package/template/.opencode/skills/scaffold-functional-test/SKILL.md +77 -0
- package/template/.opencode/skills/scaffold-functional-test/agents/openai.yaml +5 -0
- package/template/.pi/CONTEXT.md +55 -0
- package/template/.pi/docs/agents/runtime-discipline.md +3 -2
- package/template/.pi/docs/agents/skill-design.md +10 -5
- package/template/.pi/skills/ci-guard/SKILL.md +104 -0
- package/template/.pi/skills/ci-guard/agents/openai.yaml +5 -0
- package/template/.pi/skills/scaffold-functional-test/SKILL.md +77 -0
- package/template/.pi/skills/scaffold-functional-test/agents/openai.yaml +5 -0
- package/template/AGENTS.md +3 -4
- package/.agents/skills/writing-great-skills/GLOSSARY.md +0 -201
- package/.agents/skills/writing-great-skills/SKILL.md +0 -83
- package/.agents/skills/writing-great-skills/agents/openai.yaml +0 -5
- package/template/.opencode/skills/instance-test/SKILL.md +0 -61
- package/template/.opencode/skills/instance-test/agents/openai.yaml +0 -5
- package/template/.opencode/skills/instance-test/references/instances.md +0 -48
- package/template/.pi/skills/instance-test/SKILL.md +0 -61
- package/template/.pi/skills/instance-test/agents/openai.yaml +0 -5
- package/template/.pi/skills/instance-test/references/instances.md +0 -48
|
@@ -9,18 +9,24 @@ A discipline for hard bugs. Skip phases only when explicitly justified.
|
|
|
9
9
|
|
|
10
10
|
When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
|
11
11
|
|
|
12
|
-
##
|
|
12
|
+
## Redact
|
|
13
13
|
|
|
14
|
-
|
|
14
|
+
This skill has you show commands, outputs and captured artifacts. **Redact every secret first**: write `<REDACTED>` in its place. Build loops against env vars, so the credential stays in the environment rather than in what you show. Captured artifacts carry auth headers: quote only the lines that carry the signal.
|
|
15
|
+
|
|
16
|
+
If the redacted output is not enough to diagnose the bug, say so and ask the user.
|
|
17
|
+
|
|
18
|
+
## Phase 1: Build a feedback loop
|
|
19
|
+
|
|
20
|
+
**This is the skill.** Everything else is mechanical. If you have a **tight** pass/fail signal for the bug (one that goes red on _this_ bug), you will find the cause; bisection, hypothesis-testing, and instrumentation all just consume it. If you don't have one, no amount of staring at code will save you.
|
|
15
21
|
|
|
16
22
|
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
|
17
23
|
|
|
18
|
-
### Ways to construct one
|
|
24
|
+
### Ways to construct one, in roughly this order
|
|
19
25
|
|
|
20
|
-
1. **Failing test** at whatever seam reaches the bug
|
|
26
|
+
1. **Failing test** at whatever seam reaches the bug: unit, integration, e2e.
|
|
21
27
|
2. **Curl / HTTP script** against a running dev server.
|
|
22
28
|
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
|
23
|
-
4. **Headless browser script** (Playwright / Puppeteer)
|
|
29
|
+
4. **Headless browser script** (Playwright / Puppeteer) that drives the UI and asserts on DOM/console/network.
|
|
24
30
|
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
|
25
31
|
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
|
26
32
|
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
|
@@ -38,48 +44,48 @@ Treat the loop as a product. Once you have _a_ loop, **tighten** it:
|
|
|
38
44
|
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
|
39
45
|
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
|
40
46
|
|
|
41
|
-
A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight
|
|
47
|
+
A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight, a debugging superpower.
|
|
42
48
|
|
|
43
49
|
### Non-deterministic bugs
|
|
44
50
|
|
|
45
|
-
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not
|
|
51
|
+
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not, so keep raising the rate until it's debuggable.
|
|
46
52
|
|
|
47
53
|
### When you genuinely cannot build a loop
|
|
48
54
|
|
|
49
|
-
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
|
55
|
+
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a redacted captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
|
50
56
|
|
|
51
|
-
### Completion criterion
|
|
57
|
+
### Completion criterion: a tight loop that goes red
|
|
52
58
|
|
|
53
|
-
Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command**
|
|
59
|
+
Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command** (a script path, a test invocation, a curl) that you have **already run at least once** (show the invocation and its output, redacted), and that is:
|
|
54
60
|
|
|
55
|
-
- [ ] **Red-capable
|
|
56
|
-
- [ ] **Deterministic
|
|
57
|
-
- [ ] **Fast
|
|
58
|
-
- [ ] **Agent-runnable
|
|
61
|
+
- [ ] **Red-capable**: it drives the actual bug code path and asserts the **user's exact symptom**, so it can go red on this bug and green once fixed. Not "runs without erroring"; it must be able to _catch this specific bug_.
|
|
62
|
+
- [ ] **Deterministic**: same verdict every run (flaky bugs: a pinned, high reproduction rate, per above).
|
|
63
|
+
- [ ] **Fast**: seconds, not minutes.
|
|
64
|
+
- [ ] **Agent-runnable**: you can run it unattended; a human in the loop only via `scripts/hitl-loop.template.sh`.
|
|
59
65
|
|
|
60
|
-
If you catch yourself reading code to build a theory before this command exists, **stop
|
|
66
|
+
If you catch yourself reading code to build a theory before this command exists, **stop: jumping straight to a hypothesis is the exact failure this skill prevents.** No red-capable command, no Phase 2.
|
|
61
67
|
|
|
62
|
-
## Phase 2
|
|
68
|
+
## Phase 2: Reproduce + minimise
|
|
63
69
|
|
|
64
|
-
Run the loop. Watch it go red
|
|
70
|
+
Run the loop. Watch it go red as the bug appears.
|
|
65
71
|
|
|
66
72
|
Confirm:
|
|
67
73
|
|
|
68
|
-
- [ ] The loop produces the failure mode the **user** described
|
|
74
|
+
- [ ] The loop produces the failure mode the **user** described, not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
|
69
75
|
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
|
70
76
|
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
|
71
77
|
|
|
72
78
|
### Minimise
|
|
73
79
|
|
|
74
|
-
Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut
|
|
80
|
+
Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut, and keep only what's load-bearing for the failure.
|
|
75
81
|
|
|
76
82
|
Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer moving parts left to suspect) and becomes the clean regression test in Phase 5.
|
|
77
83
|
|
|
78
|
-
Done when **every remaining element is load-bearing
|
|
84
|
+
Done when **every remaining element is load-bearing**: removing any one of them makes the loop go green.
|
|
79
85
|
|
|
80
86
|
Do not proceed until you have reproduced **and** minimised.
|
|
81
87
|
|
|
82
|
-
## Phase 3
|
|
88
|
+
## Phase 3: Hypothesise
|
|
83
89
|
|
|
84
90
|
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
|
85
91
|
|
|
@@ -87,11 +93,11 @@ Each hypothesis must be **falsifiable**: state the prediction it makes.
|
|
|
87
93
|
|
|
88
94
|
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
|
89
95
|
|
|
90
|
-
If you cannot state the prediction, the hypothesis is a vibe
|
|
96
|
+
If you cannot state the prediction, the hypothesis is a vibe: discard or sharpen it.
|
|
91
97
|
|
|
92
|
-
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it
|
|
98
|
+
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it; proceed with your ranking if the user is AFK.
|
|
93
99
|
|
|
94
|
-
## Phase 4
|
|
100
|
+
## Phase 4: Instrument
|
|
95
101
|
|
|
96
102
|
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
|
97
103
|
|
|
@@ -105,9 +111,9 @@ Tool preference:
|
|
|
105
111
|
|
|
106
112
|
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
|
107
113
|
|
|
108
|
-
## Phase 5
|
|
114
|
+
## Phase 5: Fix + regression test
|
|
109
115
|
|
|
110
|
-
Write the regression test **before the fix
|
|
116
|
+
Write the regression test **before the fix**, but only if there is a **correct seam** for it.
|
|
111
117
|
|
|
112
118
|
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
|
113
119
|
|
|
@@ -121,7 +127,7 @@ If a correct seam exists:
|
|
|
121
127
|
4. Watch it pass.
|
|
122
128
|
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
|
123
129
|
|
|
124
|
-
## Phase 6
|
|
130
|
+
## Phase 6: Cleanup
|
|
125
131
|
|
|
126
132
|
Required before declaring done:
|
|
127
133
|
|
|
@@ -129,6 +135,4 @@ Required before declaring done:
|
|
|
129
135
|
- [ ] Regression test passes (or absence of seam is documented)
|
|
130
136
|
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
|
131
137
|
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
|
132
|
-
- [ ] The hypothesis that turned out correct is stated in the commit / PR message
|
|
133
|
-
|
|
134
|
-
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
|
138
|
+
- [ ] The hypothesis that turned out correct is stated in the commit / PR message, so the next debugger learns
|
|
@@ -11,6 +11,9 @@
|
|
|
11
11
|
# capture VAR "<question>" → show question, read response into VAR
|
|
12
12
|
#
|
|
13
13
|
# At the end, captured values are printed as KEY=VALUE for the agent to parse.
|
|
14
|
+
#
|
|
15
|
+
# `capture` prints its value back to the terminal, where the agent reads it,
|
|
16
|
+
# so capture observations, and leave signing in to the user as a `step`.
|
|
14
17
|
|
|
15
18
|
set -euo pipefail
|
|
16
19
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
|
|
4
4
|
|
|
5
|
-
Create the `docs/adr/` directory lazily
|
|
5
|
+
Create the `docs/adr/` directory lazily: only when the first ADR is needed.
|
|
6
6
|
|
|
7
7
|
## Template
|
|
8
8
|
|
|
@@ -12,15 +12,15 @@ Create the `docs/adr/` directory lazily — only when the first ADR is needed.
|
|
|
12
12
|
{1-3 sentences: what's the context, what did we decide, and why.}
|
|
13
13
|
```
|
|
14
14
|
|
|
15
|
-
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why
|
|
15
|
+
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why*, not in filling out sections.
|
|
16
16
|
|
|
17
17
|
## Optional sections
|
|
18
18
|
|
|
19
19
|
Only include these when they add genuine value. Most ADRs won't need them.
|
|
20
20
|
|
|
21
|
-
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`)
|
|
22
|
-
- **Considered Options
|
|
23
|
-
- **Consequences
|
|
21
|
+
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`): useful when decisions are revisited
|
|
22
|
+
- **Considered Options**: only when the rejected alternatives are worth remembering
|
|
23
|
+
- **Consequences**: only when non-obvious downstream effects need to be called out
|
|
24
24
|
|
|
25
25
|
## Numbering
|
|
26
26
|
|
|
@@ -30,18 +30,18 @@ Scan `docs/adr/` for the highest existing number and increment by one.
|
|
|
30
30
|
|
|
31
31
|
All three of these must be true:
|
|
32
32
|
|
|
33
|
-
1. **Hard to reverse
|
|
34
|
-
2. **Surprising without context
|
|
35
|
-
3. **The result of a real trade-off
|
|
33
|
+
1. **Hard to reverse**: the cost of changing your mind later is meaningful
|
|
34
|
+
2. **Surprising without context**: a future reader will look at the code and wonder "why on earth did they do it this way?"
|
|
35
|
+
3. **The result of a real trade-off**: there were genuine alternatives and you picked one for specific reasons
|
|
36
36
|
|
|
37
|
-
If a decision is easy to reverse, skip it
|
|
37
|
+
If a decision is easy to reverse, skip it: you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
|
|
38
38
|
|
|
39
39
|
### What qualifies
|
|
40
40
|
|
|
41
41
|
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
|
|
42
42
|
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
|
|
43
|
-
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library
|
|
43
|
+
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library: just the ones that would take a quarter to swap out.
|
|
44
44
|
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
|
|
45
45
|
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
|
|
46
46
|
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
|
|
47
|
-
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it
|
|
47
|
+
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it; otherwise someone will suggest GraphQL again in six months.
|
|
@@ -40,9 +40,9 @@ _Avoid_: Client, buyer, account
|
|
|
40
40
|
|
|
41
41
|
## Contexts
|
|
42
42
|
|
|
43
|
-
- [Ordering](./src/ordering/CONTEXT.md)
|
|
44
|
-
- [Billing](./src/billing/CONTEXT.md)
|
|
45
|
-
- [Fulfillment](./src/fulfillment/CONTEXT.md)
|
|
43
|
+
- [Ordering](./src/ordering/CONTEXT.md): receives and tracks customer orders
|
|
44
|
+
- [Billing](./src/billing/CONTEXT.md): generates invoices and processes payments
|
|
45
|
+
- [Fulfillment](./src/fulfillment/CONTEXT.md): manages warehouse picking and shipping
|
|
46
46
|
|
|
47
47
|
## Relationships
|
|
48
48
|
|
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: domain-modeling
|
|
3
|
-
description: Build and sharpen a project's domain model. Use when
|
|
3
|
+
description: Build and sharpen a project's domain model. Use when discussing codebase terminology, writing or editing a CONTEXT.md, or recording or editing an ADR.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# Domain Modeling
|
|
7
7
|
|
|
8
|
-
Actively build and sharpen the project's domain model as you design. This is the *active* discipline
|
|
8
|
+
Actively build and sharpen the project's domain model as you design. This is the *active* discipline: challenging terms, inventing edge-case scenarios, and writing the glossary and decisions down the moment they crystallise. (Merely *reading* `CONTEXT.md` for vocabulary is not this skill: that's a one-line habit any skill can do. This skill is for when you're changing the model, not just consuming it.)
|
|
9
9
|
|
|
10
10
|
## File structure
|
|
11
11
|
|
|
@@ -37,17 +37,17 @@ If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The ma
|
|
|
37
37
|
│ └── docs/adr/
|
|
38
38
|
```
|
|
39
39
|
|
|
40
|
-
Create files lazily
|
|
40
|
+
Create files lazily: only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
|
|
41
41
|
|
|
42
42
|
## During the session
|
|
43
43
|
|
|
44
44
|
### Challenge against the glossary
|
|
45
45
|
|
|
46
|
-
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y
|
|
46
|
+
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y. Which is it?"
|
|
47
47
|
|
|
48
48
|
### Sharpen fuzzy language
|
|
49
49
|
|
|
50
|
-
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account'
|
|
50
|
+
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account': do you mean the Customer or the User? Those are different things."
|
|
51
51
|
|
|
52
52
|
### Discuss concrete scenarios
|
|
53
53
|
|
|
@@ -55,11 +55,11 @@ When domain relationships are being discussed, stress-test them with specific sc
|
|
|
55
55
|
|
|
56
56
|
### Cross-reference with code
|
|
57
57
|
|
|
58
|
-
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible
|
|
58
|
+
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible. Which is right?"
|
|
59
59
|
|
|
60
60
|
### Update CONTEXT.md inline
|
|
61
61
|
|
|
62
|
-
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up
|
|
62
|
+
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up: capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
|
63
63
|
|
|
64
64
|
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
|
|
65
65
|
|
|
@@ -67,8 +67,8 @@ When a term is resolved, update `CONTEXT.md` right there. Don't batch these up
|
|
|
67
67
|
|
|
68
68
|
Only offer to create an ADR when all three are true:
|
|
69
69
|
|
|
70
|
-
1. **Hard to reverse
|
|
71
|
-
2. **Surprising without context
|
|
72
|
-
3. **The result of a real trade-off
|
|
70
|
+
1. **Hard to reverse**: the cost of changing your mind later is meaningful
|
|
71
|
+
2. **Surprising without context**: a future reader will wonder "why did they do it this way?"
|
|
72
|
+
3. **The result of a real trade-off**: there were genuine alternatives and you picked one for specific reasons
|
|
73
73
|
|
|
74
74
|
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
|
@@ -3,10 +3,26 @@ name: grilling
|
|
|
3
3
|
description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
Interview
|
|
6
|
+
Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**: every decision branches into the decisions that hang off it.
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Work the tree in **rounds**. The **frontier** is every decision whose prerequisites are already settled: the questions you can ask _now_ without guessing at answers you haven't heard yet. Ask the whole frontier in one round: number each question and give your recommended answer. Then wait for the user's answers before the next round.
|
|
9
9
|
|
|
10
|
-
|
|
10
|
+
Format a round like so:
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
```
|
|
13
|
+
❓ **Q1** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
|
|
14
|
+
|
|
15
|
+
➡️ <your recommended answer>
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
❓ **Q2** - **<question title>**: <question body, might be multiple paragraphs, including multiple choices>
|
|
20
|
+
|
|
21
|
+
➡️ <your recommended answer>
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Each round the user answers reshapes the tree: settled decisions push the frontier outward and unblock questions that depended on them. Recompute the frontier and ask the next round. A question whose answer depends on another question still open in this round belongs to a _later_ round, not this one.
|
|
25
|
+
|
|
26
|
+
Finding _facts_ is your job, never the user's. When a frontier question needs a fact from the environment (filesystem, tools, etc.), dispatch a sub-agent to find it; don't ask the user for anything you could look up yourself. Don't block on it: a running exploration is an unsettled prerequisite, so only the questions downstream of it wait for the sub-agent to report; ask the rest of the frontier now. The _decisions_ are the user's: put each to them and wait.
|
|
27
|
+
|
|
28
|
+
The session is done when the frontier is empty: every branch of the design tree visited, nothing left silently assumed. Do not act on it until the user confirms you have reached a shared understanding.
|
|
@@ -7,7 +7,7 @@ disable-model-invocation: true
|
|
|
7
7
|
|
|
8
8
|
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
|
|
9
9
|
|
|
10
|
-
Include a "suggested skills" section in the document, which
|
|
10
|
+
Include a "suggested skills" section in the document, naming which skills the next agent should call the Skill tool for.
|
|
11
11
|
|
|
12
12
|
Do not duplicate content already captured in other artifacts (specs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
|
|
13
13
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# HTML Report Format
|
|
2
2
|
|
|
3
|
-
The architectural review is rendered as a single self-contained HTML file in the OS temp directory. Tailwind and Mermaid both come from CDNs. Mermaid handles graph-shaped diagrams reliably; hand-built divs and inline SVG handle the more editorial visuals (mass diagrams, cross-sections). Mix the two
|
|
3
|
+
The architectural review is rendered as a single self-contained HTML file in the OS temp directory. Tailwind and Mermaid both come from CDNs. Mermaid handles graph-shaped diagrams reliably; hand-built divs and inline SVG handle the more editorial visuals (mass diagrams, cross-sections). Mix the two: don't lean on Mermaid for everything, it'll start to look generic.
|
|
4
4
|
|
|
5
5
|
## Scaffold
|
|
6
6
|
|
|
@@ -9,7 +9,7 @@ The architectural review is rendered as a single self-contained HTML file in the
|
|
|
9
9
|
<html lang="en">
|
|
10
10
|
<head>
|
|
11
11
|
<meta charset="utf-8" />
|
|
12
|
-
<title>Architecture review
|
|
12
|
+
<title>Architecture review for {{repo name}}</title>
|
|
13
13
|
<script src="https://cdn.tailwindcss.com"></script>
|
|
14
14
|
<script type="module">
|
|
15
15
|
import mermaid from "https://cdn.jsdelivr.net/npm/mermaid@11/dist/mermaid.esm.min.mjs";
|
|
@@ -35,7 +35,7 @@ The architectural review is rendered as a single self-contained HTML file in the
|
|
|
35
35
|
|
|
36
36
|
## Header
|
|
37
37
|
|
|
38
|
-
Repo name, date, and a compact legend: solid box = module, dashed line = seam, red arrow = leakage, thick dark box = deep module. No introduction paragraph
|
|
38
|
+
Repo name, date, and a compact legend: solid box = module, dashed line = seam, red arrow = leakage, thick dark box = deep module. No introduction paragraph. Straight into the candidates.
|
|
39
39
|
|
|
40
40
|
## Candidate card
|
|
41
41
|
|
|
@@ -43,20 +43,20 @@ The diagrams carry the weight. Prose is sparse, plain, and uses the glossary ter
|
|
|
43
43
|
|
|
44
44
|
Each candidate is one `<article>`:
|
|
45
45
|
|
|
46
|
-
- **Title
|
|
47
|
-
- **Badge row
|
|
48
|
-
- **Files
|
|
49
|
-
- **Before / After diagram
|
|
50
|
-
- **Problem
|
|
51
|
-
- **Solution
|
|
52
|
-
- **Wins
|
|
53
|
-
- **ADR callout** (if applicable)
|
|
46
|
+
- **Title**: short, names the deepening (e.g. "Collapse the Order intake pipeline").
|
|
47
|
+
- **Badge row**: recommendation strength (`Strong` = emerald, `Worth exploring` = amber, `Speculative` = slate), plus a tag for the dependency category (`in-process`, `local-substitutable`, `ports & adapters`, `mock`).
|
|
48
|
+
- **Files**: monospaced list, `font-mono text-sm`.
|
|
49
|
+
- **Before / After diagram**: the centrepiece. Two columns, side by side. See patterns below.
|
|
50
|
+
- **Problem**: one sentence. What hurts.
|
|
51
|
+
- **Solution**: one sentence. What changes.
|
|
52
|
+
- **Wins**: bullets, ≤6 words each. e.g. "Tests hit one interface", "Pricing logic stops leaking", "Delete 4 shallow wrappers".
|
|
53
|
+
- **ADR callout** (if applicable): one line in an amber-tinted box.
|
|
54
54
|
|
|
55
55
|
No paragraphs of explanation. If the diagram needs a paragraph to be understood, redraw the diagram.
|
|
56
56
|
|
|
57
57
|
## Diagram patterns
|
|
58
58
|
|
|
59
|
-
Pick the pattern that fits the candidate. Mix them. Don't make every diagram look the same
|
|
59
|
+
Pick the pattern that fits the candidate. Mix them. Don't make every diagram look the same. Variety is part of the point.
|
|
60
60
|
|
|
61
61
|
### Mermaid graph (the workhorse for dependencies / call flow)
|
|
62
62
|
|
|
@@ -77,7 +77,7 @@ Use a Mermaid `flowchart` or `graph` when the point is "X calls Y calls Z, and l
|
|
|
77
77
|
|
|
78
78
|
### Hand-built boxes-and-arrows (when Mermaid's layout fights you)
|
|
79
79
|
|
|
80
|
-
Modules as `<div>`s with borders and labels. Arrows as inline SVG `<line>` or `<path>` elements positioned absolutely over a relative container. Reach for this when you want the "after" diagram to feel like one thick-bordered deep module with greyed-out internals
|
|
80
|
+
Modules as `<div>`s with borders and labels. Arrows as inline SVG `<line>` or `<path>` elements positioned absolutely over a relative container. Reach for this when you want the "after" diagram to feel like one thick-bordered deep module with greyed-out internals, since Mermaid won't render that with the right weight.
|
|
81
81
|
|
|
82
82
|
### Cross-section (good for layered shallowness)
|
|
83
83
|
|
|
@@ -85,7 +85,7 @@ Stack horizontal bands (`h-12 border-l-4`) to show layers a call passes through.
|
|
|
85
85
|
|
|
86
86
|
### Mass diagram (good for "interface as wide as implementation")
|
|
87
87
|
|
|
88
|
-
Two rectangles per module
|
|
88
|
+
Two rectangles per module: one for interface surface area, one for implementation. Before: interface rectangle is nearly as tall as the implementation rectangle (shallow). After: interface rectangle is short, implementation rectangle is tall (deep).
|
|
89
89
|
|
|
90
90
|
### Call-graph collapse
|
|
91
91
|
|
|
@@ -96,8 +96,8 @@ Before: a tree of function calls rendered as nested boxes. After: the same tree
|
|
|
96
96
|
- Lean editorial, not corporate-dashboard. Generous whitespace. Serif optional for headings (`font-serif` works well with stone/slate).
|
|
97
97
|
- Colour sparingly: one accent (emerald or indigo) plus red for leakage and amber for warnings.
|
|
98
98
|
- Keep diagrams ~320px tall so before/after sits comfortably side by side without scrolling.
|
|
99
|
-
- Use `text-xs uppercase tracking-wider` for module labels inside diagrams
|
|
100
|
-
- The only scripts are the Tailwind CDN and the Mermaid ESM import. The report is otherwise static
|
|
99
|
+
- Use `text-xs uppercase tracking-wider` for module labels inside diagrams, so they read as schematic, not as UI.
|
|
100
|
+
- The only scripts are the Tailwind CDN and the Mermaid ESM import. The report is otherwise static: no app code, no interactivity beyond Mermaid's own rendering.
|
|
101
101
|
|
|
102
102
|
## Top recommendation section
|
|
103
103
|
|
|
@@ -105,7 +105,7 @@ One larger card. Candidate name, one sentence on why, anchor link to its card. T
|
|
|
105
105
|
|
|
106
106
|
## Tone
|
|
107
107
|
|
|
108
|
-
Plain English, concise
|
|
108
|
+
Plain English, concise, but the architectural nouns and verbs come straight from the `/codebase-design` skill. Concision is not an excuse to drift.
|
|
109
109
|
|
|
110
110
|
**Use exactly:** module, interface, implementation, depth, deep, shallow, seam, adapter, leverage, locality.
|
|
111
111
|
|
|
@@ -113,11 +113,11 @@ Plain English, concise — but the architectural nouns and verbs come straight f
|
|
|
113
113
|
|
|
114
114
|
**Phrasings that fit the style:**
|
|
115
115
|
|
|
116
|
-
- "Order intake module is shallow
|
|
116
|
+
- "Order intake module is shallow: interface nearly matches the implementation."
|
|
117
117
|
- "Pricing leaks across the seam."
|
|
118
118
|
- "Deepen: one interface, one place to test."
|
|
119
119
|
- "Two adapters justify the seam: HTTP in prod, in-memory in tests."
|
|
120
120
|
|
|
121
|
-
**Wins bullets** name the gain in glossary terms: *"locality: bugs concentrate in one module"*, *"leverage: one interface, N call sites"*, *"interface shrinks; implementation absorbs the wrappers"*. Don't write *"easier to maintain"* or *"cleaner code"
|
|
121
|
+
**Wins bullets** name the gain in glossary terms: *"locality: bugs concentrate in one module"*, *"leverage: one interface, N call sites"*, *"interface shrinks; implementation absorbs the wrappers"*. Don't write *"easier to maintain"* or *"cleaner code"*, because those terms aren't in the glossary and don't earn their place.
|
|
122
122
|
|
|
123
123
|
No hedging, no throat-clearing, no "it's worth noting that…". If a sentence could be a bullet, make it a bullet. If a bullet could be cut, cut it. If a term isn't in the `/codebase-design` glossary, reach for one that is before inventing a new one.
|
|
@@ -6,28 +6,28 @@ disable-model-invocation: true
|
|
|
6
6
|
|
|
7
7
|
# Improve Codebase Architecture
|
|
8
8
|
|
|
9
|
-
Surface architectural friction and propose **deepening opportunities
|
|
9
|
+
Surface architectural friction and propose **deepening opportunities**: refactors that turn shallow modules into deep ones. The aim is testability and AI-navigability.
|
|
10
10
|
|
|
11
11
|
This command is _informed_ by the project's domain model and built on a shared design vocabulary:
|
|
12
12
|
|
|
13
|
-
-
|
|
13
|
+
- Call the Skill tool with "codebase-design" for the architecture vocabulary (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and its principles (the deletion test, "the interface is the test surface", "one adapter = hypothetical seam, two = real"). Use these terms exactly in every suggestion, and don't drift into "component," "service," "API," or "boundary."
|
|
14
14
|
- The domain language in `CONTEXT.md` gives names to good seams; ADRs in `docs/adr/` record decisions this command should not re-litigate.
|
|
15
15
|
|
|
16
16
|
## Process
|
|
17
17
|
|
|
18
18
|
### 1. Explore
|
|
19
19
|
|
|
20
|
-
**Scope before you scan
|
|
20
|
+
**Scope before you scan: YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look:
|
|
21
21
|
|
|
22
|
-
- If the user named a direction
|
|
23
|
-
- Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots
|
|
22
|
+
- If the user named a direction (a module, a subsystem, a pain point), take it, and skip the inference below.
|
|
23
|
+
- Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots, the files and areas that keep coming up, and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net.
|
|
24
24
|
|
|
25
25
|
Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first.
|
|
26
26
|
|
|
27
|
-
Then
|
|
27
|
+
Then spawn a sub-agent to walk the codebase. Don't follow rigid heuristics; explore organically and note where you experience friction:
|
|
28
28
|
|
|
29
29
|
- Where does understanding one concept require bouncing between many small modules?
|
|
30
|
-
- Where are modules **shallow
|
|
30
|
+
- Where are modules **shallow**, with an interface nearly as complex as the implementation?
|
|
31
31
|
- Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
|
|
32
32
|
- Where do tightly-coupled modules leak across their seams?
|
|
33
33
|
- Which parts of the codebase are untested, or hard to test through their current interface?
|
|
@@ -36,24 +36,24 @@ Apply the **deletion test** to anything you suspect is shallow: would deleting i
|
|
|
36
36
|
|
|
37
37
|
### 2. Present candidates as an HTML report
|
|
38
38
|
|
|
39
|
-
Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user
|
|
39
|
+
Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user (`xdg-open <path>` on Linux, `open <path>` on macOS, `start <path>` on Windows) and tell them the absolute path.
|
|
40
40
|
|
|
41
|
-
The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals
|
|
41
|
+
The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals: use Mermaid when relationships are graph-shaped (call graphs, dependencies, sequences), and hand-built divs/SVG when you want something more editorial (mass diagrams, cross-sections, collapse animations). Each candidate gets a **before/after visualisation**. Be visual.
|
|
42
42
|
|
|
43
43
|
For each candidate, render a card with:
|
|
44
44
|
|
|
45
|
-
- **Files
|
|
46
|
-
- **Problem
|
|
47
|
-
- **Solution
|
|
48
|
-
- **Benefits
|
|
49
|
-
- **Before / After diagram
|
|
50
|
-
- **Recommendation strength
|
|
45
|
+
- **Files**: which files/modules are involved
|
|
46
|
+
- **Problem**: why the current architecture is causing friction
|
|
47
|
+
- **Solution**: plain English description of what would change
|
|
48
|
+
- **Benefits**: explained in terms of locality and leverage, and how tests would improve
|
|
49
|
+
- **Before / After diagram**: side-by-side, custom-drawn, illustrating the shallowness and the deepening
|
|
50
|
+
- **Recommendation strength**: one of `Strong`, `Worth exploring`, `Speculative`, rendered as a badge
|
|
51
51
|
|
|
52
52
|
End the report with a **Top recommendation** section: which candidate you'd tackle first and why.
|
|
53
53
|
|
|
54
|
-
**Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module"
|
|
54
|
+
**Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module," not "the FooBarHandler," and not "the Order service."
|
|
55
55
|
|
|
56
|
-
**ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007
|
|
56
|
+
**ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007, but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids.
|
|
57
57
|
|
|
58
58
|
See [HTML-REPORT.md](HTML-REPORT.md) for the full HTML scaffold, diagram patterns, and styling guidance.
|
|
59
59
|
|
|
@@ -61,11 +61,11 @@ Do NOT propose interfaces yet. After the file is written, ask the user: "Which o
|
|
|
61
61
|
|
|
62
62
|
### 3. Grilling loop
|
|
63
63
|
|
|
64
|
-
Once the user picks a candidate,
|
|
64
|
+
Once the user picks a candidate, call the Skill tool with "grilling" to walk the decision tree with them: constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
|
|
65
65
|
|
|
66
|
-
Side effects happen inline as decisions crystallize
|
|
66
|
+
Side effects happen inline as decisions crystallize; call the Skill tool with "domain-modeling" to keep the domain model current as you go:
|
|
67
67
|
|
|
68
68
|
- **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md`. Create the file lazily if it doesn't exist.
|
|
69
69
|
- **Sharpening a fuzzy term during the conversation?** Update `CONTEXT.md` right there.
|
|
70
|
-
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing
|
|
71
|
-
- **Want to explore alternative interfaces for the deepened module?**
|
|
70
|
+
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing; skip ephemeral reasons ("not worth it right now") and self-evident ones.
|
|
71
|
+
- **Want to explore alternative interfaces for the deepened module?** Call the Skill tool with "codebase-design" and use its design-it-twice parallel sub-agent pattern.
|