@rafinery/cli 0.8.16 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (76) hide show
  1. package/CHANGELOG.md +130 -0
  2. package/bin/rafa.mjs +11 -0
  3. package/blueprint/.claude/agents/atlas.md +21 -3
  4. package/blueprint/.claude/agents/bloom.md +2 -2
  5. package/blueprint/.claude/agents/prism.md +42 -3
  6. package/blueprint/.claude/agents/sage.md +1 -1
  7. package/blueprint/.claude/commands/rafa.md +24 -13
  8. package/blueprint/.claude/rafa/contract.md +23 -0
  9. package/blueprint/.claude/rafa/hooks/brain-commit.mjs +4 -0
  10. package/blueprint/.claude/rafa/hooks/session-start.mjs +66 -0
  11. package/blueprint/.claude/skills/rafa-build/SKILL.md +111 -20
  12. package/blueprint/.claude/skills/rafa-commit/SKILL.md +22 -0
  13. package/blueprint/.claude/skills/rafa-distill/SKILL.md +37 -4
  14. package/blueprint/.claude/skills/rafa-improve/SKILL.md +22 -0
  15. package/blueprint/.claude/skills/rafa-leverage/SKILL.md +6 -1
  16. package/blueprint/.claude/skills/rafa-plan/SKILL.md +36 -1
  17. package/blueprint/.claude/skills/rafa-review/SKILL.md +16 -2
  18. package/blueprint/.claude/skills/rafa-sage/SKILL.md +8 -2
  19. package/blueprint/.claude/skills/rafa-scan/SKILL.md +4 -1
  20. package/lib/brain-repo.mjs +1 -1
  21. package/lib/checkpoint.mjs +26 -15
  22. package/lib/distill.mjs +56 -3
  23. package/lib/distiller/doctrine.mjs +5 -16
  24. package/lib/distiller/schema-ladder.mjs +186 -0
  25. package/lib/doctor.mjs +19 -1
  26. package/lib/facts.mjs +130 -0
  27. package/lib/gate/compile.mjs +12 -0
  28. package/lib/gate/verify-citations.mjs +78 -1
  29. package/lib/hydrate.mjs +45 -0
  30. package/lib/init.mjs +9 -1
  31. package/lib/leverage/engine.mjs +32 -3
  32. package/lib/leverage.mjs +11 -2
  33. package/lib/loop-cache.mjs +67 -0
  34. package/lib/mcp-client.mjs +19 -0
  35. package/lib/push.mjs +18 -4
  36. package/lib/reflex.mjs +5 -1
  37. package/lib/releases.mjs +78 -0
  38. package/lib/review.mjs +32 -6
  39. package/lib/session-facts.mjs +108 -0
  40. package/lib/skill-deps.mjs +377 -0
  41. package/lib/stamp.mjs +24 -0
  42. package/lib/update.mjs +9 -0
  43. package/package.json +3 -2
  44. package/skills-bundle/frontend-design/LICENSE.txt +177 -0
  45. package/skills-bundle/frontend-design/SKILL.md +55 -0
  46. package/{LICENSE → skills-bundle/grill-me/LICENSE} +6 -6
  47. package/skills-bundle/grill-me/SKILL.md +7 -0
  48. package/skills-bundle/grill-me/agents/openai.yaml +5 -0
  49. package/skills-bundle/grilling/SKILL.md +53 -0
  50. package/skills-bundle/improve-codebase-architecture/HTML-REPORT.md +123 -0
  51. package/skills-bundle/improve-codebase-architecture/LICENSE +21 -0
  52. package/skills-bundle/improve-codebase-architecture/SKILL.md +71 -0
  53. package/skills-bundle/improve-codebase-architecture/agents/openai.yaml +5 -0
  54. package/skills-bundle/requesting-code-review/LICENSE +21 -0
  55. package/skills-bundle/requesting-code-review/SKILL.md +95 -0
  56. package/skills-bundle/requesting-code-review/code-reviewer.md +172 -0
  57. package/skills-bundle/skills-manifest.json +72 -0
  58. package/skills-bundle/tdd/LICENSE +21 -0
  59. package/skills-bundle/tdd/SKILL.md +36 -0
  60. package/skills-bundle/tdd/agents/openai.yaml +3 -0
  61. package/skills-bundle/tdd/mocking.md +59 -0
  62. package/skills-bundle/tdd/tests.md +77 -0
  63. package/skills-bundle/vercel-composition-patterns/AGENTS.md +946 -0
  64. package/skills-bundle/vercel-composition-patterns/README.md +60 -0
  65. package/skills-bundle/vercel-composition-patterns/SKILL.md +89 -0
  66. package/skills-bundle/vercel-composition-patterns/metadata.json +11 -0
  67. package/skills-bundle/vercel-composition-patterns/rules/_sections.md +29 -0
  68. package/skills-bundle/vercel-composition-patterns/rules/_template.md +24 -0
  69. package/skills-bundle/vercel-composition-patterns/rules/architecture-avoid-boolean-props.md +100 -0
  70. package/skills-bundle/vercel-composition-patterns/rules/architecture-compound-components.md +112 -0
  71. package/skills-bundle/vercel-composition-patterns/rules/patterns-children-over-render-props.md +87 -0
  72. package/skills-bundle/vercel-composition-patterns/rules/patterns-explicit-variants.md +100 -0
  73. package/skills-bundle/vercel-composition-patterns/rules/react19-no-forwardref.md +42 -0
  74. package/skills-bundle/vercel-composition-patterns/rules/state-context-interface.md +191 -0
  75. package/skills-bundle/vercel-composition-patterns/rules/state-decouple-implementation.md +113 -0
  76. package/skills-bundle/vercel-composition-patterns/rules/state-lift-state.md +125 -0
@@ -0,0 +1,55 @@
1
+ ---
2
+ name: frontend-design
3
+ description: Guidance for distinctive, intentional visual design when building new UI or reshaping an existing one. Helps with aesthetic direction, typography, and making choices that don't read as templated defaults.
4
+ license: Complete terms in LICENSE.txt
5
+ ---
6
+
7
+ # Frontend Design
8
+
9
+ Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else's. This client has already rejected proposals that felt templated, and is paying for a distinctive point of view: make deliberate, opinionated choices about palette, typography, and layout that are specific to this brief, and take one real aesthetic risk you can justify.
10
+
11
+ ## Ground it in the subject
12
+
13
+ If the brief does not pin down what the product or subject is, pin it yourself before designing: name one concrete subject, its audience, and the page's single job, and state your choice. If there's any information in your memory about the human's preferences, context about what they're building, or designs you've made before – use that as a hint. The subject's own world, its materials, instruments, artifacts, and vernacular, is where distinctive choices come from. Build with the brief's real content and subject matter throughout.
14
+
15
+ ## Design principles
16
+
17
+ For web designs, the hero is a thesis. Open with the most characteristic thing in the subject's world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment. Be deliberate with your choice: a big number with a small label, supporting stats, and a gradient accent is the template answer, only use if that's truly the best option.
18
+
19
+ Typography carries the personality of the page. Pair the display and body faces deliberately, not the same families you would reach for on any other project, and set a clear type scale with intentional weights, widths, and spacing. Make the type treatment itself a memorable part of the design, not a neutral delivery vehicle for the content.
20
+
21
+ Structure is information. Structural devices, numbering, eyebrows, dividers, labels, should encode something true about the content, not decorate it. Many generic designs use numbered markers (01 / 02 / 03), but that's only appropriate if the content actually is a sequence - like a real process or a typed timeline where order carries information the reader needs. Question if choices like numbered markers actually make sense before incorporating them.
22
+
23
+ Leverage motion deliberately. Think about where and if animation can serve the subject: a page-load sequence, a scroll-triggered reveal, hover micro-interactions, ambient atmosphere. An orchestrated moment usually lands harder than scattered effects; choose what the direction calls for. However, sometimes less is more, and extra animation contributes to the feeling that the design is AI-generated.
24
+
25
+ Match complexity to the vision. Maximalist directions need elaborate execution; minimal directions need precision in spacing, type, and detail. Elegance is executing the chosen vision well.
26
+
27
+ Consider written content carefully. Often a design brief may not contain real content, and it's up to you to come up with copy. Copy can make a design feel as templated as the design itself. See the below section on writing for more guidance.
28
+
29
+ ## Process: brainstorm, explore, plan, critique, build, critique again
30
+
31
+ For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns. All three are legitimate for some briefs, but they are defaults rather than choices, and they appear regardless of subject. Where the brief pins down a visual direction, follow it exactly — the brief's own words always win, including when it asks for one of these looks. Where it leaves an axis free, don't spend that freedom on one of these defaults. Just like a human designer who's hired, there's often a careful balance between doing what you're good at and taking each project as a chance to experiment and learn.
32
+
33
+ Work in two passes. First, brainstorm a short design plan based on the human's design brief: create a compact token system with color, type, layout, and signature. Color: describe the palette as 4–6 named hex values. Type: the typefaces for 2+ roles (a characterful display face that's used with restraint, a complementary body face, and a utility face for captions or data if needed). Layout: a layout concept, using one-sentence prose descriptions and ASCII wireframes to ideate and compare. Signature: the single unique element this page will be remembered by that embodies the brief in an appropriate way.
34
+
35
+ Then review that plan against the brief before building: if any part of it reads like the generic default you would produce for any similar page (work through a similar prompt to see if you arrive somewhere similar) rather than a choice made for this specific brief — revise that part, say what you changed and why. Only after you've confirmed the relative uniqueness of your design plan should you start to write the code, following the revised plan exactly and deriving every color and type decision from it.
36
+
37
+ When writing the code, be careful of structuring your CSS selector specificities. It's easy to generate CSS classes that cancel each other out (especially with a type-based selector like .section and a element-based selector like .cta). This can happen often with paddings/margins between sections.
38
+
39
+ Try to do a lot of this planning and iteration in your thinking, and only show ideas to the user when you have higher confidence it'll delight them.
40
+
41
+ ## Restraint and self-critique
42
+
43
+ Spend your boldness in one place. Let the signature element be the one memorable thing, keep everything around it quiet and disciplined, and cut any decoration that does not serve the brief. Not taking a risk can be a risk itself! Build to a quality floor without announcing it: responsive down to mobile, visible keyboard focus, reduced motion respected. Critique your own work as you build, taking screenshots if your environment supports it – a picture is worth 1000 tokens. Consider Chanel's advice: before leaving the house, take a look in the mirror and remove one accessory. Human creators have memory and always try to do something new, so if you have a space to quickly jot down notes about what you've tried, it can help you in future passes.
44
+
45
+ ## More on writing in design
46
+
47
+ Words appear in a design for one reason: to make it easier to understand, and therefore easier to use. They are design material, not decoration. Bring the same intentionality to copy that you would bring to spacing and color. Before writing anything, ask what the design needs to say, and how it can best be said to help the person navigate the experience.
48
+
49
+ Write from the end user's side of the screen. Name things by what people control and recognize, never by how the system is built. A person manages notifications, not webhook config. Describe what something does in plain terms rather than selling it. Being specific is always better than being clever.
50
+
51
+ Use active voice as default. A control should say exactly what happens when it's used: "Save changes," not "Submit." An action keeps the same name through the whole flow, so the button that says "Publish" produces a toast that says "Published." The vocabulary of an interface is the signposting for someone navigating the product. Cohesion and consistency are how people learn their way around.
52
+
53
+ Treat failure and emptiness as moments for direction, not mood. Explain what went wrong and how to fix it, in the interface's voice rather than a person's. Errors don't apologize, and they are never vague about what happened. An empty screen is an invitation to act.
54
+
55
+ Keep the register conversational and tuned: plain verbs, sentence case, no filler, with tone matched to the brand and the audience. Let each element do exactly one job. A label labels, an example demonstrates, and nothing quietly does double duty.
@@ -1,6 +1,6 @@
1
- The MIT License
1
+ MIT License
2
2
 
3
- Copyright (c) Atai Barkai
3
+ Copyright (c) 2026 Matt Pocock
4
4
 
5
5
  Permission is hereby granted, free of charge, to any person obtaining a copy
6
6
  of this software and associated documentation files (the "Software"), to deal
@@ -9,13 +9,13 @@ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
9
  copies of the Software, and to permit persons to whom the Software is
10
10
  furnished to do so, subject to the following conditions:
11
11
 
12
- The above copyright notice and this permission notice shall be included in
13
- all copies or substantial portions of the Software.
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
14
 
15
15
  THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
16
  IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
17
  FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
18
  AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
19
  LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
- OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
21
- THE SOFTWARE.
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,7 @@
1
+ ---
2
+ name: grill-me
3
+ description: A relentless interview to sharpen a plan or design.
4
+ disable-model-invocation: true
5
+ ---
6
+
7
+ Run a `/grilling` session.
@@ -0,0 +1,5 @@
1
+ interface:
2
+ display_name: "Grill Me"
3
+ short_description: "Sharpen a plan through interview"
4
+ policy:
5
+ allow_implicit_invocation: false
@@ -0,0 +1,53 @@
1
+ ---
2
+ name: grilling
3
+ description: A relentless, one-question-at-a-time interview that pressure-tests a plan or design until it survives. Invoked by /grill-me, run by rafa's prism plan-gate, or used directly on any draft before committing to it.
4
+ ---
5
+
6
+ # Grilling — the adversarial interview
7
+
8
+ Authored by rafinery (wave 6) to complete the vendored `grill-me` launcher,
9
+ which invokes `/grilling` by name. The same procedure runs inside rafa's prism
10
+ plan-gate; this standalone form is for devs grilling anything by hand.
11
+
12
+ You are not here to improve the plan. You are here to find where it breaks.
13
+ Improvement is the author's job, after you've made the weakness undeniable.
14
+
15
+ ## Rules of the interview
16
+
17
+ 1. **One probe at a time.** Ask a single, concrete question; wait for the
18
+ answer; follow the thread before opening a new one. A barrage lets weak
19
+ answers hide in volume.
20
+ 2. **Attack the plan, never the person.** Every question targets a claim,
21
+ dependency, or omission in the artifact — quote the exact line you're
22
+ probing.
23
+ 3. **No leading the witness.** Ask what breaks it, not "have you considered
24
+ X?" with the fix embedded. The author finds the fix; you find the hole.
25
+ 4. **Concrete beats abstract.** "What happens when the webhook retries twice
26
+ in the same second?" beats "what about concurrency?"
27
+ 5. **Survival ends it.** Stop when the plan survives **three consecutive
28
+ probes** without needing a change — or when a probe forces a change, which
29
+ resets the count. A plan that keeps changing isn't done being grilled.
30
+
31
+ ## The probe ladder (work top to bottom)
32
+
33
+ - **Grounding** — which claims rest on something verified, and which on
34
+ something assumed? Name the assumption; ask for its evidence.
35
+ - **Unstated dependencies** — what must already exist/be true for step N to
36
+ work, and where is that declared? (In rafa plans: a missing `blocked_by`.)
37
+ - **Failure modes** — for each step: what happens when it half-completes, runs
38
+ twice, or runs against stale state? Which failures are silent?
39
+ - **Vacuous acceptance** — can the Done-check/success criterion pass while the
40
+ feature is actually broken? A check that recomputes its expectation the way
41
+ the code does, or that a no-op implementation satisfies, is tautological.
42
+ - **YAGNI** — which parts exist because the intent needs them, and which
43
+ because they seemed thorough? Ask what breaks if a part is deleted.
44
+ - **The deletion test** — for any new module/abstraction: would deleting it
45
+ concentrate complexity somewhere honest, or just smear it around?
46
+
47
+ ## Output
48
+
49
+ End with a verdict, not a summary: **SURVIVED** (three consecutive clean
50
+ probes — say which probes it survived) or **CHANGED** (list each probe that
51
+ forced a change, one line each, quoting the plan line it hit). Never end with
52
+ "looks good" — a grilling that found nothing must say which probes were run
53
+ and survived.
@@ -0,0 +1,123 @@
1
+ # HTML Report Format
2
+
3
+ The architectural review is rendered as a single self-contained HTML file in the OS temp directory. Tailwind and Mermaid both come from CDNs. Mermaid handles graph-shaped diagrams reliably; hand-built divs and inline SVG handle the more editorial visuals (mass diagrams, cross-sections). Mix the two — don't lean on Mermaid for everything, it'll start to look generic.
4
+
5
+ ## Scaffold
6
+
7
+ ```html
8
+ <!doctype html>
9
+ <html lang="en">
10
+ <head>
11
+ <meta charset="utf-8" />
12
+ <title>Architecture review — {{repo name}}</title>
13
+ <script src="https://cdn.tailwindcss.com"></script>
14
+ <script type="module">
15
+ import mermaid from "https://cdn.jsdelivr.net/npm/mermaid@11/dist/mermaid.esm.min.mjs";
16
+ mermaid.initialize({ startOnLoad: true, theme: "neutral", securityLevel: "loose" });
17
+ </script>
18
+ <style>
19
+ /* small custom layer for things Tailwind doesn't cover cleanly:
20
+ dashed seam lines, hand-drawn-feeling arrow heads, etc. */
21
+ .seam { stroke-dasharray: 4 4; }
22
+ .leak { stroke: #dc2626; }
23
+ .deep { background: linear-gradient(135deg, #0f172a, #1e293b); }
24
+ </style>
25
+ </head>
26
+ <body class="bg-stone-50 text-slate-900 font-sans">
27
+ <main class="max-w-5xl mx-auto px-6 py-12 space-y-12">
28
+ <header>...</header>
29
+ <section id="candidates" class="space-y-10">...</section>
30
+ <section id="top-recommendation">...</section>
31
+ </main>
32
+ </body>
33
+ </html>
34
+ ```
35
+
36
+ ## Header
37
+
38
+ Repo name, date, and a compact legend: solid box = module, dashed line = seam, red arrow = leakage, thick dark box = deep module. No introduction paragraph — straight into the candidates.
39
+
40
+ ## Candidate card
41
+
42
+ The diagrams carry the weight. Prose is sparse, plain, and uses the glossary terms (from the `/codebase-design` skill) without ceremony.
43
+
44
+ Each candidate is one `<article>`:
45
+
46
+ - **Title** — short, names the deepening (e.g. "Collapse the Order intake pipeline").
47
+ - **Badge row** — recommendation strength (`Strong` = emerald, `Worth exploring` = amber, `Speculative` = slate), plus a tag for the dependency category (`in-process`, `local-substitutable`, `ports & adapters`, `mock`).
48
+ - **Files** — monospaced list, `font-mono text-sm`.
49
+ - **Before / After diagram** — the centrepiece. Two columns, side by side. See patterns below.
50
+ - **Problem** — one sentence. What hurts.
51
+ - **Solution** — one sentence. What changes.
52
+ - **Wins** — bullets, ≤6 words each. e.g. "Tests hit one interface", "Pricing logic stops leaking", "Delete 4 shallow wrappers".
53
+ - **ADR callout** (if applicable) — one line in an amber-tinted box.
54
+
55
+ No paragraphs of explanation. If the diagram needs a paragraph to be understood, redraw the diagram.
56
+
57
+ ## Diagram patterns
58
+
59
+ Pick the pattern that fits the candidate. Mix them. Don't make every diagram look the same — variety is part of the point.
60
+
61
+ ### Mermaid graph (the workhorse for dependencies / call flow)
62
+
63
+ Use a Mermaid `flowchart` or `graph` when the point is "X calls Y calls Z, and look at the mess." Wrap it in a Tailwind-styled card so it doesn't feel parachuted in. Style with classDef to colour leakage edges red and the deep module dark. Sequence diagrams work well for "before: 6 round-trips; after: 1."
64
+
65
+ ```html
66
+ <div class="rounded-lg border border-slate-200 bg-white p-4">
67
+ <pre class="mermaid">
68
+ flowchart LR
69
+ A[OrderHandler] --> B[OrderValidator]
70
+ B --> C[OrderRepo]
71
+ C -.leak.-> D[PricingClient]
72
+ classDef leak stroke:#dc2626,stroke-width:2px;
73
+ class C,D leak
74
+ </pre>
75
+ </div>
76
+ ```
77
+
78
+ ### Hand-built boxes-and-arrows (when Mermaid's layout fights you)
79
+
80
+ Modules as `<div>`s with borders and labels. Arrows as inline SVG `<line>` or `<path>` elements positioned absolutely over a relative container. Reach for this when you want the "after" diagram to feel like one thick-bordered deep module with greyed-out internals — Mermaid won't render that with the right weight.
81
+
82
+ ### Cross-section (good for layered shallowness)
83
+
84
+ Stack horizontal bands (`h-12 border-l-4`) to show layers a call passes through. Before: 6 thin layers each doing nothing. After: 1 thick band labelled with the consolidated responsibility.
85
+
86
+ ### Mass diagram (good for "interface as wide as implementation")
87
+
88
+ Two rectangles per module — one for interface surface area, one for implementation. Before: interface rectangle is nearly as tall as the implementation rectangle (shallow). After: interface rectangle is short, implementation rectangle is tall (deep).
89
+
90
+ ### Call-graph collapse
91
+
92
+ Before: a tree of function calls rendered as nested boxes. After: the same tree collapsed into one box, with the now-internal calls shown faded inside it.
93
+
94
+ ## Style guidance
95
+
96
+ - Lean editorial, not corporate-dashboard. Generous whitespace. Serif optional for headings (`font-serif` works well with stone/slate).
97
+ - Colour sparingly: one accent (emerald or indigo) plus red for leakage and amber for warnings.
98
+ - Keep diagrams ~320px tall so before/after sits comfortably side by side without scrolling.
99
+ - Use `text-xs uppercase tracking-wider` for module labels inside diagrams — they should read as schematic, not as UI.
100
+ - The only scripts are the Tailwind CDN and the Mermaid ESM import. The report is otherwise static — no app code, no interactivity beyond Mermaid's own rendering.
101
+
102
+ ## Top recommendation section
103
+
104
+ One larger card. Candidate name, one sentence on why, anchor link to its card. That's it.
105
+
106
+ ## Tone
107
+
108
+ Plain English, concise — but the architectural nouns and verbs come straight from the `/codebase-design` skill. Concision is not an excuse to drift.
109
+
110
+ **Use exactly:** module, interface, implementation, depth, deep, shallow, seam, adapter, leverage, locality.
111
+
112
+ **Never substitute:** component, service, unit (for module) · API, signature (for interface) · boundary (for seam) · layer, wrapper (for module, when you mean module).
113
+
114
+ **Phrasings that fit the style:**
115
+
116
+ - "Order intake module is shallow — interface nearly matches the implementation."
117
+ - "Pricing leaks across the seam."
118
+ - "Deepen: one interface, one place to test."
119
+ - "Two adapters justify the seam: HTTP in prod, in-memory in tests."
120
+
121
+ **Wins bullets** name the gain in glossary terms: *"locality: bugs concentrate in one module"*, *"leverage: one interface, N call sites"*, *"interface shrinks; implementation absorbs the wrappers"*. Don't write *"easier to maintain"* or *"cleaner code"* — those terms aren't in the glossary and don't earn their place.
122
+
123
+ No hedging, no throat-clearing, no "it's worth noting that…". If a sentence could be a bullet, make it a bullet. If a bullet could be cut, cut it. If a term isn't in the `/codebase-design` glossary, reach for one that is before inventing a new one.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Matt Pocock
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,71 @@
1
+ ---
2
+ name: improve-codebase-architecture
3
+ description: Scan a codebase for deepening opportunities, present them as a visual HTML report, then grill through whichever one you pick.
4
+ disable-model-invocation: true
5
+ ---
6
+
7
+ # Improve Codebase Architecture
8
+
9
+ Surface architectural friction and propose **deepening opportunities** — refactors that turn shallow modules into deep ones. The aim is testability and AI-navigability.
10
+
11
+ This command is _informed_ by the project's domain model and built on a shared design vocabulary:
12
+
13
+ - Run the `/codebase-design` skill for the architecture vocabulary (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and its principles (the deletion test, "the interface is the test surface", "one adapter = hypothetical seam, two = real"). Use these terms exactly in every suggestion — don't drift into "component," "service," "API," or "boundary."
14
+ - The domain language in `CONTEXT.md` gives names to good seams; ADRs in `docs/adr/` record decisions this command should not re-litigate.
15
+
16
+ ## Process
17
+
18
+ ### 1. Explore
19
+
20
+ **Scope before you scan — YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look:
21
+
22
+ - If the user named a direction — a module, a subsystem, a pain point — take it, and skip the inference below.
23
+ - Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots — the files and areas that keep coming up — and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net.
24
+
25
+ Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first.
26
+
27
+ Then use the Agent tool with `subagent_type=Explore` to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction:
28
+
29
+ - Where does understanding one concept require bouncing between many small modules?
30
+ - Where are modules **shallow** — interface nearly as complex as the implementation?
31
+ - Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
32
+ - Where do tightly-coupled modules leak across their seams?
33
+ - Which parts of the codebase are untested, or hard to test through their current interface?
34
+
35
+ Apply the **deletion test** to anything you suspect is shallow: would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want.
36
+
37
+ ### 2. Present candidates as an HTML report
38
+
39
+ Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user — `xdg-open <path>` on Linux, `open <path>` on macOS, `start <path>` on Windows — and tell them the absolute path.
40
+
41
+ The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals — use Mermaid when relationships are graph-shaped (call graphs, dependencies, sequences), and hand-built divs/SVG when you want something more editorial (mass diagrams, cross-sections, collapse animations). Each candidate gets a **before/after visualisation**. Be visual.
42
+
43
+ For each candidate, render a card with:
44
+
45
+ - **Files** — which files/modules are involved
46
+ - **Problem** — why the current architecture is causing friction
47
+ - **Solution** — plain English description of what would change
48
+ - **Benefits** — explained in terms of locality and leverage, and how tests would improve
49
+ - **Before / After diagram** — side-by-side, custom-drawn, illustrating the shallowness and the deepening
50
+ - **Recommendation strength** — one of `Strong`, `Worth exploring`, `Speculative`, rendered as a badge
51
+
52
+ End the report with a **Top recommendation** section: which candidate you'd tackle first and why.
53
+
54
+ **Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service."
55
+
56
+ **ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007 — but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids.
57
+
58
+ See [HTML-REPORT.md](HTML-REPORT.md) for the full HTML scaffold, diagram patterns, and styling guidance.
59
+
60
+ Do NOT propose interfaces yet. After the file is written, ask the user: "Which of these would you like to explore?"
61
+
62
+ ### 3. Grilling loop
63
+
64
+ Once the user picks a candidate, run the `/grilling` skill to walk the decision tree with them — constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
65
+
66
+ Side effects happen inline as decisions crystallize — run the `/domain-modeling` skill to keep the domain model current as you go:
67
+
68
+ - **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md`. Create the file lazily if it doesn't exist.
69
+ - **Sharpening a fuzzy term during the conversation?** Update `CONTEXT.md` right there.
70
+ - **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing — skip ephemeral reasons ("not worth it right now") and self-evident ones.
71
+ - **Want to explore alternative interfaces for the deepened module?** Run the `/codebase-design` skill and use its design-it-twice parallel sub-agent pattern.
@@ -0,0 +1,5 @@
1
+ interface:
2
+ display_name: "Improve Codebase Architecture"
3
+ short_description: "Find and grill architecture improvements"
4
+ policy:
5
+ allow_implicit_invocation: false
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2025 Jesse Vincent
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,95 @@
1
+ ---
2
+ name: requesting-code-review
3
+ description: Use when completing tasks, implementing major features, or before merging to verify work meets requirements
4
+ ---
5
+
6
+ # Requesting Code Review
7
+
8
+ Dispatch a code reviewer subagent to catch issues before they cascade. The reviewer gets precisely crafted context for evaluation — never your session's history.
9
+
10
+ **Core principle:** Review early, review often.
11
+
12
+ ## When to Request Review
13
+
14
+ **Mandatory:**
15
+ - After each task in subagent-driven development
16
+ - After completing major feature
17
+ - Before merge to main
18
+
19
+ **Optional but valuable:**
20
+ - When stuck (fresh perspective)
21
+ - Before refactoring (baseline check)
22
+ - After fixing complex bug
23
+
24
+ ## How to Request
25
+
26
+ **1. Get git SHAs:**
27
+ ```bash
28
+ BASE_SHA=$(git rev-parse HEAD~1) # or origin/main
29
+ HEAD_SHA=$(git rev-parse HEAD)
30
+ ```
31
+
32
+ **2. Dispatch code reviewer subagent:**
33
+
34
+ Dispatch a `general-purpose` subagent, filling the template at [code-reviewer.md](code-reviewer.md)
35
+
36
+ **Placeholders:**
37
+ - `{DESCRIPTION}` - Brief summary of what you built
38
+ - `{PLAN_OR_REQUIREMENTS}` - What it should do
39
+ - `{BASE_SHA}` - Starting commit
40
+ - `{HEAD_SHA}` - Ending commit
41
+
42
+ **3. Act on feedback:**
43
+ - Fix Critical issues immediately
44
+ - Fix Important issues before proceeding
45
+ - Note Minor issues for later
46
+ - Push back if reviewer is wrong (with reasoning)
47
+
48
+ ## Example
49
+
50
+ ```
51
+ [Just completed Task 2: Add verification function]
52
+
53
+ You: Let me request code review before proceeding.
54
+
55
+ BASE_SHA=$(git log --oneline | grep "Task 1" | head -1 | awk '{print $1}')
56
+ HEAD_SHA=$(git rev-parse HEAD)
57
+
58
+ [Dispatch code reviewer subagent]
59
+ DESCRIPTION: Added verifyIndex() and repairIndex() with 4 issue types
60
+ PLAN_OR_REQUIREMENTS: Task 2 from docs/superpowers/plans/deployment-plan.md
61
+ BASE_SHA: a7981ec
62
+ HEAD_SHA: 3df7661
63
+
64
+ [Subagent returns]:
65
+ Strengths: Clean architecture, real tests
66
+ Issues:
67
+ Important: Missing progress indicators
68
+ Minor: Magic number (100) for reporting interval
69
+ Assessment: Ready to proceed
70
+
71
+ You: [Fix progress indicators]
72
+ [Continue to Task 3]
73
+ ```
74
+
75
+ ## Common Rationalizations
76
+
77
+ | Excuse | Reality |
78
+ |--------|---------|
79
+ | "I'll just review the diff myself instead of dispatching a reviewer" | You're the coordinator — reviewing the diff inline burns the context window you need to keep driving the work. Dispatch a reviewer subagent: the diff and the evaluation live in its context, and only the findings come back to you. |
80
+ | "The reviewer needs my whole session history to understand the change" | Hand it precisely crafted context, never your session's history. That keeps the reviewer on the work product, not your thought process. |
81
+
82
+ ## Red Flags
83
+
84
+ **Never:**
85
+ - Skip review because "it's simple"
86
+ - Ignore Critical issues
87
+ - Proceed with unfixed Important issues
88
+ - Argue with valid technical feedback
89
+
90
+ **If reviewer wrong:**
91
+ - Push back with technical reasoning
92
+ - Show code/tests that prove it works
93
+ - Request clarification
94
+
95
+ See template at: [code-reviewer.md](code-reviewer.md)