@complexthings/superpowers-agent 9.2.1 → 10.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/skills/claude-handoff/SKILL.md +18 -0
- package/.agents/skills/code-review/SKILL.md +89 -0
- package/.agents/skills/{improve-codebase-architecture → codebase-design}/DEEPENING.md +1 -1
- package/.agents/skills/{improve-codebase-architecture/INTERFACE-DESIGN.md → codebase-design/DESIGN-IT-TWICE.md} +3 -3
- package/.agents/skills/codebase-design/SKILL.md +114 -0
- package/.agents/skills/design-an-interface/SKILL.md +94 -0
- package/.agents/skills/{diagnose → diagnosing-bugs}/SKILL.md +29 -12
- package/.agents/skills/{grill-with-docs → domain-modeling}/CONTEXT-FORMAT.md +1 -4
- package/.agents/skills/domain-modeling/SKILL.md +74 -0
- package/.agents/skills/fable-mode/SKILL.md +95 -0
- package/.agents/skills/git-guardrails-claude-code/SKILL.md +95 -0
- package/.agents/skills/git-guardrails-claude-code/scripts/block-dangerous-git.sh +25 -0
- package/.agents/skills/grill-me/SKILL.md +7 -0
- package/.agents/skills/grill-with-docs/SKILL.md +3 -86
- package/.agents/skills/grilling/SKILL.md +14 -0
- package/.agents/skills/handoff/SKILL.md +2 -1
- package/.agents/skills/i-have-adhd/SKILL.md +120 -0
- package/.agents/skills/implement/SKILL.md +11 -0
- package/.agents/skills/improve-codebase-architecture/HTML-REPORT.md +3 -3
- package/.agents/skills/improve-codebase-architecture/SKILL.md +13 -28
- package/.agents/skills/loop-me/SKILL.md +32 -0
- package/.agents/skills/prototype/SKILL.md +1 -1
- package/.agents/skills/qa/SKILL.md +130 -0
- package/.agents/skills/request-refactor-plan/SKILL.md +68 -0
- package/.agents/skills/research/SKILL.md +12 -0
- package/.agents/skills/resolving-merge-conflicts/SKILL.md +14 -0
- package/.agents/skills/scaffold-exercises/SKILL.md +106 -0
- package/.agents/skills/setup-matt-pocock-skills/SKILL.md +11 -9
- package/.agents/skills/setup-matt-pocock-skills/domain.md +2 -2
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-github.md +23 -0
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-gitlab.md +23 -0
- package/.agents/skills/setup-matt-pocock-skills/issue-tracker-local.md +11 -0
- package/.agents/skills/skill-creator/LICENSE.txt +202 -0
- package/.agents/skills/skill-creator/SKILL.md +485 -0
- package/.agents/skills/skill-creator/agents/analyzer.md +274 -0
- package/.agents/skills/skill-creator/agents/comparator.md +202 -0
- package/.agents/skills/skill-creator/agents/grader.md +223 -0
- package/.agents/skills/skill-creator/assets/eval_review.html +146 -0
- package/.agents/skills/skill-creator/eval-viewer/generate_review.py +471 -0
- package/.agents/skills/skill-creator/eval-viewer/viewer.html +1325 -0
- package/.agents/skills/skill-creator/references/schemas.md +430 -0
- package/.agents/skills/skill-creator/scripts/__init__.py +0 -0
- package/.agents/skills/skill-creator/scripts/__pycache__/__init__.cpython-314.pyc +0 -0
- package/.agents/skills/skill-creator/scripts/__pycache__/run_eval.cpython-314.pyc +0 -0
- package/.agents/skills/skill-creator/scripts/__pycache__/utils.cpython-314.pyc +0 -0
- package/.agents/skills/skill-creator/scripts/aggregate_benchmark.py +401 -0
- package/.agents/skills/skill-creator/scripts/generate_report.py +326 -0
- package/.agents/skills/skill-creator/scripts/improve_description.py +247 -0
- package/.agents/skills/skill-creator/scripts/package_skill.py +136 -0
- package/.agents/skills/skill-creator/scripts/quick_validate.py +103 -0
- package/.agents/skills/skill-creator/scripts/run_eval.py +310 -0
- package/.agents/skills/skill-creator/scripts/run_loop.py +328 -0
- package/.agents/skills/skill-creator/scripts/utils.py +47 -0
- package/.agents/skills/tdd/SKILL.md +17 -90
- package/.agents/skills/tdd/tests.md +16 -0
- package/.agents/skills/teach/GLOSSARY-FORMAT.md +35 -0
- package/.agents/skills/teach/LEARNING-RECORD-FORMAT.md +46 -0
- package/.agents/skills/teach/MISSION-FORMAT.md +31 -0
- package/.agents/skills/teach/RESOURCES-FORMAT.md +32 -0
- package/.agents/skills/teach/SKILL.md +140 -0
- package/.agents/skills/{to-prd → to-spec}/SKILL.md +11 -12
- package/.agents/skills/to-tickets/SKILL.md +114 -0
- package/.agents/skills/triage/AGENT-BRIEF.md +40 -1
- package/.agents/skills/triage/OUT-OF-SCOPE.md +5 -1
- package/.agents/skills/triage/SKILL.md +20 -11
- package/.agents/skills/wayfinder/SKILL.md +127 -0
- package/.agents/skills/writing-great-skills/GLOSSARY.md +201 -0
- package/.agents/skills/writing-great-skills/SKILL.md +83 -0
- package/.agents/superpowers-agent +103 -222
- package/.agents/superpowers-bootstrap.md +3 -3
- package/.agents/templates/AGENTS.md.template +11 -34
- package/.agents/templates/SUPERPOWERS.md.template +4 -4
- package/.github/copilot-instructions.md +23 -99
- package/.github/hooks/rtk-rewrite.json +22 -0
- package/AGENTS.md +7 -6
- package/README.md +53 -174
- package/package.json +2 -2
- package/skills/collaboration/brainstorming/SKILL.md +39 -139
- package/skills/collaboration/brainstorming/skill.json +2 -2
- package/skills/collaboration/leveraging-cli-tools/SKILL.md +70 -71
- package/skills/collaboration/leveraging-cli-tools/references/copilot-instructions.md +30 -0
- package/skills/collaboration/leveraging-cli-tools/scripts/setup-ponytail.sh +185 -0
- package/skills/collaboration/leveraging-cli-tools/scripts/setup-rtk.sh +217 -0
- package/skills/collaboration/leveraging-cli-tools/skill.json +1 -1
- package/skills/meta/create-skill-json/SKILL.md +4 -4
- package/skills/meta/create-skill-json/skill.json +1 -1
- package/skills/meta/create-skill-json/test-scenarios.md +1 -1
- package/skills/setup-skills/SKILL.md +18 -11
- package/skills/setup-skills/skill.json +8 -0
- package/.agents/skills/caveman/SKILL.md +0 -49
- package/.agents/skills/improve-codebase-architecture/LANGUAGE.md +0 -53
- package/.agents/skills/karpathy-guidelines/SKILL.md +0 -75
- package/.agents/skills/review/SKILL.md +0 -78
- package/.agents/skills/tdd/deep-modules.md +0 -33
- package/.agents/skills/tdd/interface-design.md +0 -31
- package/.agents/skills/tdd/refactoring.md +0 -10
- package/.agents/skills/to-issues/SKILL.md +0 -83
- package/.agents/skills/zoom-out/SKILL.md +0 -7
- package/skills/architecture/ABOUT.md +0 -20
- package/skills/architecture/preserving-productive-tensions/SKILL.md +0 -146
- package/skills/architecture/preserving-productive-tensions/skill.json +0 -9
- package/skills/collaboration/brainstorming/spec-document-reviewer-prompt.md +0 -50
- package/skills/collaboration/brainstorming/visual-companion.md +0 -277
- package/skills/collaboration/dispatching-parallel-agents/SKILL.md +0 -174
- package/skills/collaboration/dispatching-parallel-agents/skill.json +0 -9
- package/skills/collaboration/executing-plans/SKILL.md +0 -130
- package/skills/collaboration/executing-plans/skill.json +0 -9
- package/skills/collaboration/finishing-a-development-branch/SKILL.md +0 -261
- package/skills/collaboration/finishing-a-development-branch/skill.json +0 -9
- package/skills/collaboration/leveraging-cli-tools/scripts/slim.py +0 -167
- package/skills/collaboration/receiving-code-review/SKILL.md +0 -233
- package/skills/collaboration/receiving-code-review/skill.json +0 -9
- package/skills/collaboration/requesting-code-review/SKILL.md +0 -110
- package/skills/collaboration/requesting-code-review/code-reviewer.md +0 -146
- package/skills/collaboration/requesting-code-review/skill.json +0 -12
- package/skills/collaboration/subagent-driven-development/SKILL.md +0 -255
- package/skills/collaboration/subagent-driven-development/code-quality-reviewer-prompt.md +0 -26
- package/skills/collaboration/subagent-driven-development/implementer-prompt.md +0 -113
- package/skills/collaboration/subagent-driven-development/skill.json +0 -15
- package/skills/collaboration/subagent-driven-development/spec-reviewer-prompt.md +0 -61
- package/skills/collaboration/using-git-worktrees/SKILL.md +0 -366
- package/skills/collaboration/using-git-worktrees/skill.json +0 -9
- package/skills/collaboration/writing-plans/SKILL.md +0 -121
- package/skills/collaboration/writing-plans/plan-document-reviewer-prompt.md +0 -52
- package/skills/collaboration/writing-plans/skill.json +0 -9
- package/skills/debugging/defense-in-depth/SKILL.md +0 -380
- package/skills/debugging/defense-in-depth/skill.json +0 -9
- package/skills/debugging/root-cause-tracing/SKILL.md +0 -361
- package/skills/debugging/root-cause-tracing/find-polluter.sh +0 -63
- package/skills/debugging/root-cause-tracing/skill.json +0 -12
- package/skills/debugging/systematic-debugging/SKILL.md +0 -299
- package/skills/debugging/systematic-debugging/condition-based-waiting-example.ts +0 -158
- package/skills/debugging/systematic-debugging/condition-based-waiting.md +0 -115
- package/skills/debugging/systematic-debugging/defense-in-depth.md +0 -122
- package/skills/debugging/systematic-debugging/find-polluter.sh +0 -63
- package/skills/debugging/systematic-debugging/root-cause-tracing.md +0 -169
- package/skills/debugging/systematic-debugging/skill.json +0 -9
- package/skills/debugging/systematic-debugging/test-academic.md +0 -14
- package/skills/debugging/systematic-debugging/test-pressure-1.md +0 -58
- package/skills/debugging/systematic-debugging/test-pressure-2.md +0 -68
- package/skills/debugging/systematic-debugging/test-pressure-3.md +0 -69
- package/skills/debugging/verification-before-completion/SKILL.md +0 -143
- package/skills/debugging/verification-before-completion/skill.json +0 -9
- package/skills/finding-skills/SKILL.md +0 -101
- package/skills/finding-skills/skill.json +0 -8
- package/skills/meta/create-agents-md/SKILL.md +0 -182
- package/skills/meta/create-agents-md/skill.json +0 -9
- package/skills/meta/creating-prompts/SKILL.md +0 -349
- package/skills/meta/creating-prompts/examples/do-example.md +0 -65
- package/skills/meta/creating-prompts/examples/plan-example.md +0 -75
- package/skills/meta/creating-prompts/examples/refine-example.md +0 -65
- package/skills/meta/creating-prompts/examples/research-example.md +0 -63
- package/skills/meta/creating-prompts/scripts/get-next-number.sh +0 -27
- package/skills/meta/creating-prompts/skill.json +0 -20
- package/skills/meta/creating-prompts/templates/do-template.md +0 -59
- package/skills/meta/creating-prompts/templates/plan-template.md +0 -58
- package/skills/meta/creating-prompts/templates/refine-template.md +0 -54
- package/skills/meta/creating-prompts/templates/research-template.md +0 -56
- package/skills/meta/using-superpowers/SKILL.md +0 -108
- package/skills/meta/using-superpowers/skill.json +0 -5
- package/skills/meta/writing-prompts/SKILL.md +0 -122
- package/skills/meta/writing-prompts/references/platforms.md +0 -114
- package/skills/meta/writing-prompts/skill.json +0 -9
- package/skills/problem-solving/ABOUT.md +0 -40
- package/skills/problem-solving/collision-zone-thinking/SKILL.md +0 -188
- package/skills/problem-solving/collision-zone-thinking/references/historical-examples.md +0 -393
- package/skills/problem-solving/collision-zone-thinking/skill.json +0 -9
- package/skills/problem-solving/inversion-exercise/SKILL.md +0 -174
- package/skills/problem-solving/inversion-exercise/skill.json +0 -9
- package/skills/problem-solving/meta-pattern-recognition/SKILL.md +0 -116
- package/skills/problem-solving/meta-pattern-recognition/skill.json +0 -9
- package/skills/problem-solving/scale-game/SKILL.md +0 -222
- package/skills/problem-solving/scale-game/skill.json +0 -9
- package/skills/problem-solving/simplification-cascades/SKILL.md +0 -113
- package/skills/problem-solving/simplification-cascades/skill.json +0 -9
- package/skills/problem-solving/when-stuck/SKILL.md +0 -69
- package/skills/problem-solving/when-stuck/skill.json +0 -9
- package/skills/research/ABOUT.md +0 -20
- package/skills/research/tracing-knowledge-lineages/SKILL.md +0 -241
- package/skills/research/tracing-knowledge-lineages/skill.json +0 -9
- package/skills/testing/condition-based-waiting/SKILL.md +0 -359
- package/skills/testing/condition-based-waiting/example.ts +0 -158
- package/skills/testing/condition-based-waiting/skill.json +0 -12
- package/skills/testing/test-driven-development/SKILL.md +0 -434
- package/skills/testing/test-driven-development/skill.json +0 -9
- package/skills/testing/testing-anti-patterns/SKILL.md +0 -298
- package/skills/testing/testing-anti-patterns/skill.json +0 -9
- package/skills/testing/verification-before-completion/SKILL.md +0 -246
- package/skills/testing/verification-before-completion/skill.json +0 -10
- package/skills/using-a-skill/SKILL.md +0 -101
- package/skills/using-a-skill/skill.json +0 -8
- /package/.agents/skills/{diagnose → diagnosing-bugs}/scripts/hitl-loop.template.sh +0 -0
- /package/.agents/skills/{grill-with-docs → domain-modeling}/ADR-FORMAT.md +0 -0
|
@@ -0,0 +1,18 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: claude-handoff
|
|
3
|
+
description: Hand the current conversation off to a fresh background agent that picks up the work immediately.
|
|
4
|
+
argument-hint: "What will the next session be used for?"
|
|
5
|
+
disable-model-invocation: true
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
Write a handoff summary of the current conversation so a fresh agent can continue the work. Instead of saving it, launch a background agent seeded with the summary as its prompt: `claude --bg --name "<descriptive name>" "<handoff summary>"`. It starts in the current working directory and returns immediately; the user manages it with `claude agents`.
|
|
9
|
+
|
|
10
|
+
Always pass `-n`/`--name` with a descriptive name (e.g. `--name "Fix login bug"`) — it sets the display name shown in the job list, session picker, and terminal title.
|
|
11
|
+
|
|
12
|
+
Include a "suggested skills" section in the summary, which suggests skills that the agent should invoke.
|
|
13
|
+
|
|
14
|
+
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
|
|
15
|
+
|
|
16
|
+
Redact any sensitive information, such as API keys, passwords, or personally identifiable information — the summary becomes the agent's prompt.
|
|
17
|
+
|
|
18
|
+
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the summary accordingly.
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: code-review
|
|
3
|
+
description: Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/PRD asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to "review since X".
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
Two-axis review of the diff between `HEAD` and a fixed point the user supplies:
|
|
7
|
+
|
|
8
|
+
- **Standards** — does the code conform to this repo's documented coding standards?
|
|
9
|
+
- **Spec** — does the code faithfully implement the originating issue / PRD / spec?
|
|
10
|
+
|
|
11
|
+
Both axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.
|
|
12
|
+
|
|
13
|
+
The issue tracker should have been provided to you — run `/setup-matt-pocock-skills` if `docs/agents/issue-tracker.md` is missing.
|
|
14
|
+
|
|
15
|
+
## Process
|
|
16
|
+
|
|
17
|
+
### 1. Pin the fixed point
|
|
18
|
+
|
|
19
|
+
Whatever the user said is the fixed point — a commit SHA, branch name, tag, `main`, `HEAD~5`, etc. If they didn't specify one, ask for it.
|
|
20
|
+
|
|
21
|
+
Capture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.
|
|
22
|
+
|
|
23
|
+
Before going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here — not inside two parallel sub-agents.
|
|
24
|
+
|
|
25
|
+
### 2. Identify the spec source
|
|
26
|
+
|
|
27
|
+
Look for the originating spec, in this order:
|
|
28
|
+
|
|
29
|
+
1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.) — fetch via the workflow in `docs/agents/issue-tracker.md`.
|
|
30
|
+
2. A path the user passed as an argument.
|
|
31
|
+
3. A PRD/spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.
|
|
32
|
+
4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report "no spec available".
|
|
33
|
+
|
|
34
|
+
### 3. Identify the standards sources
|
|
35
|
+
|
|
36
|
+
Anything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.
|
|
37
|
+
|
|
38
|
+
On top of whatever the repo documents, the Standards axis always carries the **smell baseline** below — a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:
|
|
39
|
+
|
|
40
|
+
- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.
|
|
41
|
+
- **Always a judgement call.** Each smell is a labelled heuristic ("possible Feature Envy"), never a hard violation — and, like any standard here, skip anything tooling already enforces.
|
|
42
|
+
|
|
43
|
+
Each smell reads *what it is* → *how to fix*; match it against the diff:
|
|
44
|
+
|
|
45
|
+
- **Mysterious Name** — a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.
|
|
46
|
+
- **Duplicated Code** — the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.
|
|
47
|
+
- **Feature Envy** — a method that reaches into another object's data more than its own. → move the method onto the data it envies.
|
|
48
|
+
- **Data Clumps** — the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.
|
|
49
|
+
- **Primitive Obsession** — a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.
|
|
50
|
+
- **Repeated Switches** — the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.
|
|
51
|
+
- **Shotgun Surgery** — one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.
|
|
52
|
+
- **Divergent Change** — one file or module is edited for several unrelated reasons. → split so each module changes for one reason.
|
|
53
|
+
- **Speculative Generality** — abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.
|
|
54
|
+
- **Message Chains** — long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.
|
|
55
|
+
- **Middle Man** — a class or function that mostly just delegates onward. → cut it, call the real target direct.
|
|
56
|
+
- **Refused Bequest** — a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.
|
|
57
|
+
|
|
58
|
+
### 4. Spawn both sub-agents in parallel
|
|
59
|
+
|
|
60
|
+
Send a single message with two `Agent` tool calls. Use the `general-purpose` subagent for both.
|
|
61
|
+
|
|
62
|
+
**Standards sub-agent prompt** — include:
|
|
63
|
+
|
|
64
|
+
- The full diff command and commit list.
|
|
65
|
+
- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full — the sub-agent has no other access to it.
|
|
66
|
+
- The brief: "Report — per file/hunk where relevant — (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls — documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words."
|
|
67
|
+
|
|
68
|
+
**Spec sub-agent prompt** — include:
|
|
69
|
+
|
|
70
|
+
- The diff command and commit list.
|
|
71
|
+
- The path or fetched contents of the spec.
|
|
72
|
+
- The brief: "Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words."
|
|
73
|
+
|
|
74
|
+
If the spec is missing, skip the Spec sub-agent and note this in the final report.
|
|
75
|
+
|
|
76
|
+
### 5. Aggregate
|
|
77
|
+
|
|
78
|
+
Present the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings — the two axes are deliberately separate (see _Why two axes_).
|
|
79
|
+
|
|
80
|
+
End with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes — that's the reranking the separation exists to prevent.
|
|
81
|
+
|
|
82
|
+
## Why two axes
|
|
83
|
+
|
|
84
|
+
A change can pass one axis and fail the other:
|
|
85
|
+
|
|
86
|
+
- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**
|
|
87
|
+
- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**
|
|
88
|
+
|
|
89
|
+
Reporting them separately stops one axis from masking the other.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Deepening
|
|
2
2
|
|
|
3
|
-
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [
|
|
3
|
+
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [SKILL.md](SKILL.md) — **module**, **interface**, **seam**, **adapter**.
|
|
4
4
|
|
|
5
5
|
## Dependency categories
|
|
6
6
|
|
|
@@ -1,8 +1,8 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Design It Twice
|
|
2
2
|
|
|
3
3
|
When the user wants to explore alternative interfaces for a chosen deepening candidate, use this parallel sub-agent pattern. Based on "Design It Twice" (Ousterhout) — your first idea is unlikely to be the best.
|
|
4
4
|
|
|
5
|
-
Uses the vocabulary in [
|
|
5
|
+
Uses the vocabulary in [SKILL.md](SKILL.md) — **module**, **interface**, **seam**, **adapter**, **leverage**.
|
|
6
6
|
|
|
7
7
|
## Process
|
|
8
8
|
|
|
@@ -27,7 +27,7 @@ Prompt each sub-agent with a separate technical brief (file paths, coupling deta
|
|
|
27
27
|
- Agent 3: "Optimise for the most common caller — make the default case trivial."
|
|
28
28
|
- Agent 4 (if applicable): "Design around ports & adapters for cross-seam dependencies."
|
|
29
29
|
|
|
30
|
-
Include both [
|
|
30
|
+
Include both [SKILL.md](SKILL.md) vocabulary and CONTEXT.md vocabulary in the brief so each sub-agent names things consistently with the architecture language and the project's domain language.
|
|
31
31
|
|
|
32
32
|
Each sub-agent outputs:
|
|
33
33
|
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: codebase-design
|
|
3
|
+
description: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Codebase Design
|
|
7
|
+
|
|
8
|
+
Design **deep modules**: a lot of behaviour behind a small interface, placed at a clean seam, testable through that interface. Use this language and these principles wherever code is being designed or restructured. The aim is leverage for callers, locality for maintainers, and testability for everyone.
|
|
9
|
+
|
|
10
|
+
## Glossary
|
|
11
|
+
|
|
12
|
+
Use these terms exactly — don't substitute "component," "service," "API," or "boundary." Consistent language is the whole point.
|
|
13
|
+
|
|
14
|
+
**Module** — anything with an interface and an implementation. Deliberately scale-agnostic: a function, class, package, or tier-spanning slice. _Avoid_: unit, component, service.
|
|
15
|
+
|
|
16
|
+
**Interface** — everything a caller must know to use the module correctly: the type signature, but also invariants, ordering constraints, error modes, required configuration, and performance characteristics. _Avoid_: API, signature (too narrow — they refer only to the type-level surface).
|
|
17
|
+
|
|
18
|
+
**Implementation** — what's inside a module, its body of code. Distinct from **Adapter**: a thing can be a small adapter with a large implementation (a Postgres repo) or a large adapter with a small implementation (an in-memory fake). Reach for "adapter" when the seam is the topic; "implementation" otherwise.
|
|
19
|
+
|
|
20
|
+
**Depth** — leverage at the interface: the amount of behaviour a caller (or test) can exercise per unit of interface they have to learn. A module is **deep** when a large amount of behaviour sits behind a small interface, **shallow** when the interface is nearly as complex as the implementation.
|
|
21
|
+
|
|
22
|
+
**Seam** _(Michael Feathers)_ — a place where you can alter behaviour without editing in that place; the *location* at which a module's interface lives. Where to put the seam is its own design decision, distinct from what goes behind it. _Avoid_: boundary (overloaded with DDD's bounded context).
|
|
23
|
+
|
|
24
|
+
**Adapter** — a concrete thing that satisfies an interface at a seam. Describes *role* (what slot it fills), not substance (what's inside).
|
|
25
|
+
|
|
26
|
+
**Leverage** — what callers get from depth: more capability per unit of interface they learn. One implementation pays back across N call sites and M tests.
|
|
27
|
+
|
|
28
|
+
**Locality** — what maintainers get from depth: change, bugs, knowledge, and verification concentrate in one place rather than spreading across callers. Fix once, fixed everywhere.
|
|
29
|
+
|
|
30
|
+
## Deep vs shallow
|
|
31
|
+
|
|
32
|
+
**Deep module** = small interface + lots of implementation:
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
┌─────────────────────┐
|
|
36
|
+
│ Small Interface │ ← Few methods, simple params
|
|
37
|
+
├─────────────────────┤
|
|
38
|
+
│ │
|
|
39
|
+
│ Deep Implementation│ ← Complex logic hidden
|
|
40
|
+
│ │
|
|
41
|
+
└─────────────────────┘
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**Shallow module** = large interface + little implementation (avoid):
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
┌─────────────────────────────────┐
|
|
48
|
+
│ Large Interface │ ← Many methods, complex params
|
|
49
|
+
├─────────────────────────────────┤
|
|
50
|
+
│ Thin Implementation │ ← Just passes through
|
|
51
|
+
└─────────────────────────────────┘
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
When designing an interface, ask:
|
|
55
|
+
|
|
56
|
+
- Can I reduce the number of methods?
|
|
57
|
+
- Can I simplify the parameters?
|
|
58
|
+
- Can I hide more complexity inside?
|
|
59
|
+
|
|
60
|
+
## Principles
|
|
61
|
+
|
|
62
|
+
- **Depth is a property of the interface, not the implementation.** A deep module can be internally composed of small, mockable, swappable parts — they just aren't part of the interface. A module can have **internal seams** (private to its implementation, used by its own tests) as well as the **external seam** at its interface.
|
|
63
|
+
- **The deletion test.** Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.
|
|
64
|
+
- **The interface is the test surface.** Callers and tests cross the same seam. If you want to test *past* the interface, the module is probably the wrong shape.
|
|
65
|
+
- **One adapter means a hypothetical seam. Two adapters means a real one.** Don't introduce a seam unless something actually varies across it.
|
|
66
|
+
|
|
67
|
+
## Designing for testability
|
|
68
|
+
|
|
69
|
+
Good interfaces make testing natural:
|
|
70
|
+
|
|
71
|
+
1. **Accept dependencies, don't create them.**
|
|
72
|
+
|
|
73
|
+
```typescript
|
|
74
|
+
// Testable
|
|
75
|
+
function processOrder(order, paymentGateway) {}
|
|
76
|
+
|
|
77
|
+
// Hard to test
|
|
78
|
+
function processOrder(order) {
|
|
79
|
+
const gateway = new StripeGateway();
|
|
80
|
+
}
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
2. **Return results, don't produce side effects.**
|
|
84
|
+
|
|
85
|
+
```typescript
|
|
86
|
+
// Testable
|
|
87
|
+
function calculateDiscount(cart): Discount {}
|
|
88
|
+
|
|
89
|
+
// Hard to test
|
|
90
|
+
function applyDiscount(cart): void {
|
|
91
|
+
cart.total -= discount;
|
|
92
|
+
}
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
3. **Small surface area.** Fewer methods = fewer tests needed. Fewer params = simpler test setup.
|
|
96
|
+
|
|
97
|
+
## Relationships
|
|
98
|
+
|
|
99
|
+
- A **Module** has exactly one **Interface** (the surface it presents to callers and tests).
|
|
100
|
+
- **Depth** is a property of a **Module**, measured against its **Interface**.
|
|
101
|
+
- A **Seam** is where a **Module**'s **Interface** lives.
|
|
102
|
+
- An **Adapter** sits at a **Seam** and satisfies the **Interface**.
|
|
103
|
+
- **Depth** produces **Leverage** for callers and **Locality** for maintainers.
|
|
104
|
+
|
|
105
|
+
## Rejected framings
|
|
106
|
+
|
|
107
|
+
- **Depth as ratio of implementation-lines to interface-lines** (Ousterhout): rewards padding the implementation. We use depth-as-leverage instead.
|
|
108
|
+
- **"Interface" as the TypeScript `interface` keyword or a class's public methods**: too narrow — interface here includes every fact a caller must know.
|
|
109
|
+
- **"Boundary"**: overloaded with DDD's bounded context. Say **seam** or **interface**.
|
|
110
|
+
|
|
111
|
+
## Going deeper
|
|
112
|
+
|
|
113
|
+
- **Deepening a cluster given its dependencies** — see [DEEPENING.md](DEEPENING.md): dependency categories, seam discipline, and replace-don't-layer testing.
|
|
114
|
+
- **Exploring alternative interfaces** — see [DESIGN-IT-TWICE.md](DESIGN-IT-TWICE.md): spin up parallel sub-agents to design the interface several radically different ways, then compare on depth, locality, and seam placement.
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: design-an-interface
|
|
3
|
+
description: Generate multiple radically different interface designs for a module using parallel sub-agents. Use when user wants to design an API, explore interface options, compare module shapes, or mentions "design it twice".
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Design an Interface
|
|
7
|
+
|
|
8
|
+
Based on "Design It Twice" from "A Philosophy of Software Design": your first idea is unlikely to be the best. Generate multiple radically different designs, then compare.
|
|
9
|
+
|
|
10
|
+
## Workflow
|
|
11
|
+
|
|
12
|
+
### 1. Gather Requirements
|
|
13
|
+
|
|
14
|
+
Before designing, understand:
|
|
15
|
+
|
|
16
|
+
- [ ] What problem does this module solve?
|
|
17
|
+
- [ ] Who are the callers? (other modules, external users, tests)
|
|
18
|
+
- [ ] What are the key operations?
|
|
19
|
+
- [ ] Any constraints? (performance, compatibility, existing patterns)
|
|
20
|
+
- [ ] What should be hidden inside vs exposed?
|
|
21
|
+
|
|
22
|
+
Ask: "What does this module need to do? Who will use it?"
|
|
23
|
+
|
|
24
|
+
### 2. Generate Designs (Parallel Sub-Agents)
|
|
25
|
+
|
|
26
|
+
Spawn 3+ sub-agents simultaneously using Task tool. Each must produce a **radically different** approach.
|
|
27
|
+
|
|
28
|
+
```
|
|
29
|
+
Prompt template for each sub-agent:
|
|
30
|
+
|
|
31
|
+
Design an interface for: [module description]
|
|
32
|
+
|
|
33
|
+
Requirements: [gathered requirements]
|
|
34
|
+
|
|
35
|
+
Constraints for this design: [assign a different constraint to each agent]
|
|
36
|
+
- Agent 1: "Minimize method count - aim for 1-3 methods max"
|
|
37
|
+
- Agent 2: "Maximize flexibility - support many use cases"
|
|
38
|
+
- Agent 3: "Optimize for the most common case"
|
|
39
|
+
- Agent 4: "Take inspiration from [specific paradigm/library]"
|
|
40
|
+
|
|
41
|
+
Output format:
|
|
42
|
+
1. Interface signature (types/methods)
|
|
43
|
+
2. Usage example (how caller uses it)
|
|
44
|
+
3. What this design hides internally
|
|
45
|
+
4. Trade-offs of this approach
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
### 3. Present Designs
|
|
49
|
+
|
|
50
|
+
Show each design with:
|
|
51
|
+
|
|
52
|
+
1. **Interface signature** - types, methods, params
|
|
53
|
+
2. **Usage examples** - how callers actually use it in practice
|
|
54
|
+
3. **What it hides** - complexity kept internal
|
|
55
|
+
|
|
56
|
+
Present designs sequentially so user can absorb each approach before comparison.
|
|
57
|
+
|
|
58
|
+
### 4. Compare Designs
|
|
59
|
+
|
|
60
|
+
After showing all designs, compare them on:
|
|
61
|
+
|
|
62
|
+
- **Interface simplicity**: fewer methods, simpler params
|
|
63
|
+
- **General-purpose vs specialized**: flexibility vs focus
|
|
64
|
+
- **Implementation efficiency**: does shape allow efficient internals?
|
|
65
|
+
- **Depth**: small interface hiding significant complexity (good) vs large interface with thin implementation (bad)
|
|
66
|
+
- **Ease of correct use** vs **ease of misuse**
|
|
67
|
+
|
|
68
|
+
Discuss trade-offs in prose, not tables. Highlight where designs diverge most.
|
|
69
|
+
|
|
70
|
+
### 5. Synthesize
|
|
71
|
+
|
|
72
|
+
Often the best design combines insights from multiple options. Ask:
|
|
73
|
+
|
|
74
|
+
- "Which design best fits your primary use case?"
|
|
75
|
+
- "Any elements from other designs worth incorporating?"
|
|
76
|
+
|
|
77
|
+
## Evaluation Criteria
|
|
78
|
+
|
|
79
|
+
From "A Philosophy of Software Design":
|
|
80
|
+
|
|
81
|
+
**Interface simplicity**: Fewer methods, simpler params = easier to learn and use correctly.
|
|
82
|
+
|
|
83
|
+
**General-purpose**: Can handle future use cases without changes. But beware over-generalization.
|
|
84
|
+
|
|
85
|
+
**Implementation efficiency**: Does interface shape allow efficient implementation? Or force awkward internals?
|
|
86
|
+
|
|
87
|
+
**Depth**: Small interface hiding significant complexity = deep module (good). Large interface with thin implementation = shallow module (avoid).
|
|
88
|
+
|
|
89
|
+
## Anti-Patterns
|
|
90
|
+
|
|
91
|
+
- Don't let sub-agents produce similar designs - enforce radical difference
|
|
92
|
+
- Don't skip comparison - the value is in contrast
|
|
93
|
+
- Don't implement - this is purely about interface shape
|
|
94
|
+
- Don't evaluate based on implementation effort
|
|
@@ -1,17 +1,17 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
3
|
-
description:
|
|
2
|
+
name: diagnosing-bugs
|
|
3
|
+
description: Diagnosis loop for hard bugs and performance regressions. Use when the user says "diagnose"/"debug this", or reports something broken/throwing/failing/slow.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
#
|
|
6
|
+
# Diagnosing Bugs
|
|
7
7
|
|
|
8
8
|
A discipline for hard bugs. Skip phases only when explicitly justified.
|
|
9
9
|
|
|
10
|
-
When exploring the codebase,
|
|
10
|
+
When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
|
11
11
|
|
|
12
12
|
## Phase 1 — Build a feedback loop
|
|
13
13
|
|
|
14
|
-
**This is the skill.** Everything else is mechanical. If you have a
|
|
14
|
+
**This is the skill.** Everything else is mechanical. If you have a **tight** pass/fail signal for the bug — one that goes red on _this_ bug — you will find the cause; bisection, hypothesis-testing, and instrumentation all just consume it. If you don't have one, no amount of staring at code will save you.
|
|
15
15
|
|
|
16
16
|
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
|
17
17
|
|
|
@@ -30,15 +30,15 @@ Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give
|
|
|
30
30
|
|
|
31
31
|
Build the right feedback loop, and the bug is 90% fixed.
|
|
32
32
|
|
|
33
|
-
###
|
|
33
|
+
### Tighten the loop
|
|
34
34
|
|
|
35
|
-
Treat the loop as a product. Once you have _a_ loop,
|
|
35
|
+
Treat the loop as a product. Once you have _a_ loop, **tighten** it:
|
|
36
36
|
|
|
37
37
|
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
|
38
38
|
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
|
39
39
|
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
|
40
40
|
|
|
41
|
-
A 30-second flaky loop is barely better than no loop
|
|
41
|
+
A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight — a debugging superpower.
|
|
42
42
|
|
|
43
43
|
### Non-deterministic bugs
|
|
44
44
|
|
|
@@ -48,11 +48,20 @@ The goal is not a clean repro but a **higher reproduction rate**. Loop the trigg
|
|
|
48
48
|
|
|
49
49
|
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
|
50
50
|
|
|
51
|
-
|
|
51
|
+
### Completion criterion — a tight loop that goes red
|
|
52
52
|
|
|
53
|
-
|
|
53
|
+
Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command** — a script path, a test invocation, a curl — that you have **already run at least once** (paste the invocation and its output), and that is:
|
|
54
54
|
|
|
55
|
-
|
|
55
|
+
- [ ] **Red-capable** — it drives the actual bug code path and asserts the **user's exact symptom**, so it can go red on this bug and green once fixed. Not "runs without erroring" — it must be able to _catch this specific bug_.
|
|
56
|
+
- [ ] **Deterministic** — same verdict every run (flaky bugs: a pinned, high reproduction rate, per above).
|
|
57
|
+
- [ ] **Fast** — seconds, not minutes.
|
|
58
|
+
- [ ] **Agent-runnable** — you can run it unattended; a human in the loop only via `scripts/hitl-loop.template.sh`.
|
|
59
|
+
|
|
60
|
+
If you catch yourself reading code to build a theory before this command exists, **stop — jumping straight to a hypothesis is the exact failure this skill prevents.** No red-capable command, no Phase 2.
|
|
61
|
+
|
|
62
|
+
## Phase 2 — Reproduce + minimise
|
|
63
|
+
|
|
64
|
+
Run the loop. Watch it go red — the bug appears.
|
|
56
65
|
|
|
57
66
|
Confirm:
|
|
58
67
|
|
|
@@ -60,7 +69,15 @@ Confirm:
|
|
|
60
69
|
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
|
61
70
|
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
|
62
71
|
|
|
63
|
-
|
|
72
|
+
### Minimise
|
|
73
|
+
|
|
74
|
+
Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut — keep only what's load-bearing for the failure.
|
|
75
|
+
|
|
76
|
+
Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer moving parts left to suspect) and becomes the clean regression test in Phase 5.
|
|
77
|
+
|
|
78
|
+
Done when **every remaining element is load-bearing** — removing any one of them makes the loop go green.
|
|
79
|
+
|
|
80
|
+
Do not proceed until you have reproduced **and** minimised.
|
|
64
81
|
|
|
65
82
|
## Phase 3 — Hypothesise
|
|
66
83
|
|
|
@@ -24,13 +24,10 @@ _Avoid_: Client, buyer, account
|
|
|
24
24
|
|
|
25
25
|
## Rules
|
|
26
26
|
|
|
27
|
-
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others
|
|
28
|
-
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
|
|
27
|
+
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others under `_Avoid_`.
|
|
29
28
|
- **Keep definitions tight.** One or two sentences max. Define what it IS, not what it does.
|
|
30
|
-
- **Show relationships.** Use bold term names and express cardinality where obvious.
|
|
31
29
|
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
|
|
32
30
|
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
|
|
33
|
-
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
|
|
34
31
|
|
|
35
32
|
## Single vs multi-context repos
|
|
36
33
|
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: domain-modeling
|
|
3
|
+
description: Build and sharpen a project's domain model. Use when the user wants to pin down domain terminology or a ubiquitous language, record an architectural decision, or when another skill needs to maintain the domain model.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Domain Modeling
|
|
7
|
+
|
|
8
|
+
Actively build and sharpen the project's domain model as you design. This is the *active* discipline — challenging terms, inventing edge-case scenarios, and writing the glossary and decisions down the moment they crystallise. (Merely *reading* `CONTEXT.md` for vocabulary is not this skill — that's a one-line habit any skill can do. This skill is for when you're changing the model, not just consuming it.)
|
|
9
|
+
|
|
10
|
+
## File structure
|
|
11
|
+
|
|
12
|
+
Most repos have a single context:
|
|
13
|
+
|
|
14
|
+
```
|
|
15
|
+
/
|
|
16
|
+
├── CONTEXT.md
|
|
17
|
+
├── docs/
|
|
18
|
+
│ └── adr/
|
|
19
|
+
│ ├── 0001-event-sourced-orders.md
|
|
20
|
+
│ └── 0002-postgres-for-write-model.md
|
|
21
|
+
└── src/
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
/
|
|
28
|
+
├── CONTEXT-MAP.md
|
|
29
|
+
├── docs/
|
|
30
|
+
│ └── adr/ ← system-wide decisions
|
|
31
|
+
├── src/
|
|
32
|
+
│ ├── ordering/
|
|
33
|
+
│ │ ├── CONTEXT.md
|
|
34
|
+
│ │ └── docs/adr/ ← context-specific decisions
|
|
35
|
+
│ └── billing/
|
|
36
|
+
│ ├── CONTEXT.md
|
|
37
|
+
│ └── docs/adr/
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
|
|
41
|
+
|
|
42
|
+
## During the session
|
|
43
|
+
|
|
44
|
+
### Challenge against the glossary
|
|
45
|
+
|
|
46
|
+
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
|
|
47
|
+
|
|
48
|
+
### Sharpen fuzzy language
|
|
49
|
+
|
|
50
|
+
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
|
|
51
|
+
|
|
52
|
+
### Discuss concrete scenarios
|
|
53
|
+
|
|
54
|
+
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
|
|
55
|
+
|
|
56
|
+
### Cross-reference with code
|
|
57
|
+
|
|
58
|
+
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
|
|
59
|
+
|
|
60
|
+
### Update CONTEXT.md inline
|
|
61
|
+
|
|
62
|
+
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
|
63
|
+
|
|
64
|
+
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
|
|
65
|
+
|
|
66
|
+
### Offer ADRs sparingly
|
|
67
|
+
|
|
68
|
+
Only offer to create an ADR when all three are true:
|
|
69
|
+
|
|
70
|
+
1. **Hard to reverse** — the cost of changing your mind later is meaningful
|
|
71
|
+
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
|
|
72
|
+
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
|
73
|
+
|
|
74
|
+
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: fable-mode
|
|
3
|
+
description: Use PROACTIVELY the moment you notice a task has many layers - multiple dependent steps, unknowns that could change the approach, debugging where the first theory might be wrong, or anything that needs verification before handoff. Also use when a task keeps failing or stalling, or when Nate says "fable mode", "think like Fable", "use the Fable skill", "use the Fable method", "work like Fable", "slow down and do this right", or "think this through first". Loads Fable 5's working discipline (the five-gate task loop plus standing habits) so any session, especially one running on Opus 4.8 or Sonnet 5, applies it.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# The Fable Method
|
|
7
|
+
|
|
8
|
+
Fable 5's working discipline, written down so any model can run it. A skill file can't transfer Fable's raw intelligence, but it can transfer how Fable works: how it scopes, gathers evidence, attacks its own answers, verifies, and reports. Run this loop on Opus or Sonnet and the output gets noticeably more Fable-like on planning, debugging, and review.
|
|
9
|
+
|
|
10
|
+
A hard task is anything where the first idea might be wrong: multi-step builds, debugging, research with claims, anything touching data you haven't looked at yet. For a one-file edit or a simple lookup, skip the gates and just do the work.
|
|
11
|
+
|
|
12
|
+
## The loop: five gates, in order
|
|
13
|
+
|
|
14
|
+
Every hard task passes through five gates. A gate must pass before the next one opens. When a task stalls or a result surprises you, name which gate you're at and re-run it.
|
|
15
|
+
|
|
16
|
+
### Gate 1 — Scope before work
|
|
17
|
+
|
|
18
|
+
State what done looks like before touching anything.
|
|
19
|
+
|
|
20
|
+
- Define done in one or two sentences: what artifact exists at the end, what must be true of it, and how you will check that it's true. If you can't write the check, you don't understand the task yet.
|
|
21
|
+
- Check standing rules first (CLAUDE.md, skills, memory). Don't invent an approach the project already has a rule for.
|
|
22
|
+
- Separate known from assumed. Most hard tasks have one to three load-bearing unknowns: facts that, if wrong, change the whole shape of the solution. Name them explicitly.
|
|
23
|
+
- If the request is ambiguous in a way that changes what you'd build, ask one question, aimed at the biggest gap. Otherwise pick the sensible default, say so in one line, and proceed. Ask questions to change outcomes, not to feel safe.
|
|
24
|
+
- Right-size the effort. Match the depth of this process to the stakes of the task. Deep reasoning belongs in planning and review, not in mechanical steps.
|
|
25
|
+
|
|
26
|
+
### Gate 2 — Evidence before reasoning
|
|
27
|
+
|
|
28
|
+
Never design from memory of what a file, API, or dataset "probably" looks like. Open it.
|
|
29
|
+
|
|
30
|
+
- Files and live tool output are sources. Training memory is only a hypothesis generator.
|
|
31
|
+
- Attack the load-bearing unknowns first, with the cheapest probe. A 30-second read of the real data beats an hour of building on a guess.
|
|
32
|
+
- Prefer a thin end-to-end pass over a complete first stage. Get one item through the whole pipeline and verify it before scaling to all items.
|
|
33
|
+
- Keep a live plan for anything with 3+ steps. Slice by dependency, not by category: each step's output feeds the next. The plan is a hypothesis, not a contract.
|
|
34
|
+
|
|
35
|
+
### Gate 3 — Reason adversarially
|
|
36
|
+
|
|
37
|
+
Before committing to an answer, switch roles and try to kill it.
|
|
38
|
+
|
|
39
|
+
- Attack your own emerging answer as a hostile reviewer: what input, state, or reading makes this wrong? Actually test that case; don't just imagine it.
|
|
40
|
+
- Then steelman what survives. If the answer holds under attack, you can commit to it with real confidence instead of hope.
|
|
41
|
+
- Steelman the existing thing before changing it. Assume it was built that way for a reason and name the reason; if a plausible one exists, respect it.
|
|
42
|
+
- When reviewing, finding nothing wrong is a legitimate result. "Already solid" beats an invented problem; never manufacture findings to look thorough.
|
|
43
|
+
- Re-decide after every result. Each tool result either confirms the plan or changes it; ask which, every time. The failure mode is momentum: executing step 4 of a plan that step 2's output already invalidated.
|
|
44
|
+
- Two failed attempts at the same fix means the diagnosis is wrong. Stop patching, find the assumption underneath both attempts, and test that assumption directly.
|
|
45
|
+
|
|
46
|
+
### Gate 4 — Verify before declaring done
|
|
47
|
+
|
|
48
|
+
"It ran" is not verification. Verify at the layer of the claim.
|
|
49
|
+
|
|
50
|
+
- If the claim is "the output is correct," look at the output. If the claim is "the page renders," look at the page. Exit code 0 only proves the layer below the claim.
|
|
51
|
+
- Use evidence you didn't generate. Re-open the file you wrote. Run the code. Screenshot the page and read the screenshot. Diff before against after. Count the things you claimed to count.
|
|
52
|
+
- Re-check against the original request and the standing rules from Gate 1. Did you build what was asked, and did you follow the rules you loaded?
|
|
53
|
+
- Sample the tails, not just the middle: first item, last item, weirdest item. Happy-path spot checks hide the failures that matter.
|
|
54
|
+
- Treat good news as suspect. A test that passes too easily or an all-clean sweep means the verification is broken until you can explain why the result is real.
|
|
55
|
+
- Zero-context test for anything user-facing: would someone with none of this session's context understand it and be able to act on it?
|
|
56
|
+
|
|
57
|
+
### Gate 5 — Report calibrated
|
|
58
|
+
|
|
59
|
+
The report is part of the work, not an afterthought.
|
|
60
|
+
|
|
61
|
+
- Lead with the answer, then the support.
|
|
62
|
+
- Separate verified from assumed, out loud. "I confirmed X by running Y; I'm assuming Z because I couldn't check it."
|
|
63
|
+
- Cite evidence with specifics: file paths, line numbers, the command you ran, the number you saw.
|
|
64
|
+
- Report what you observed, not what you intended. If tests failed, say so with the output. If a step was skipped, say that.
|
|
65
|
+
- Never soften a real problem to be agreeable. Disagreement with concrete reasoning beats compliance. Flag the risk once, concretely, then respect the user's call.
|
|
66
|
+
- Never state as fact what you have not verified this session. Done means the Gate 1 check passed and you watched it pass.
|
|
67
|
+
|
|
68
|
+
## Standing habits (always on, every gate)
|
|
69
|
+
|
|
70
|
+
- Convert relative to absolute: "tomorrow" becomes a date, "the latest version" becomes a version number, "recently" becomes a month.
|
|
71
|
+
- Surface constraints proactively. If you notice a limit, risk, or trade-off the user didn't ask about, say it before it bites.
|
|
72
|
+
- Pick the next action by information per unit cost: the cheapest probe of the biggest remaining unknown beats the largest visible chunk of work.
|
|
73
|
+
- Sort actions by reversibility. Reversible and in scope: just do it. Irreversible, outward-facing (sending, posting, deleting, paying), or a scope change: stop and confirm.
|
|
74
|
+
- Unblock yourself before escalating: read more, search more, try another route. Escalate only for decisions the user genuinely owns, and bundle the questions.
|
|
75
|
+
- Mechanical work repeating 3+ times gets a script, not per-instance reasoning. Reasoning is for judgment; scripts are for repetition.
|
|
76
|
+
- Preserve by default. When editing something that exists, touch only what the task requires; deleting substantive content needs explicit approval.
|
|
77
|
+
|
|
78
|
+
## Smells that mean a gate got skipped
|
|
79
|
+
|
|
80
|
+
- You're building something and haven't opened the real data/file/API response it depends on. (Gate 2)
|
|
81
|
+
- You just said or thought "should work" about anything you can test right now. (Gate 4)
|
|
82
|
+
- You're on attempt three of the same fix. (Gate 3)
|
|
83
|
+
- Your last three actions came from the original plan with no check against intermediate results. (Gate 3)
|
|
84
|
+
- You're about to report done and the evidence is your intention, not an observation. (Gate 4)
|
|
85
|
+
- A result came back surprisingly clean and you moved on without asking why. (Gate 4)
|
|
86
|
+
- You can't say in one sentence what done looks like. (Gate 1)
|
|
87
|
+
|
|
88
|
+
Any one of these: stop, go back to that gate.
|
|
89
|
+
|
|
90
|
+
## Notes
|
|
91
|
+
|
|
92
|
+
- This is a method skill, not a workflow. It changes how you execute the current task; it produces no files of its own.
|
|
93
|
+
- It stacks with task-specific skills (/proveit, /verify, /code-review). Those are the "how to check" tools; this is the discipline of when to reach for them.
|
|
94
|
+
- Don't apply it to trivial work. Forcing all five gates onto a two-minute edit is its own failure mode.
|
|
95
|
+
- If a task keeps failing under this discipline, that's the signal to escalate to a stronger model, not to loosen the process. Keep the discipline either way.
|