@pikku/skills 0.12.9 → 0.12.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/skills.gen.js +1 -1
- package/package.json +1 -1
- package/skills/pikku-addon/SKILL.md +3 -3
- package/skills/{pikku-ai-agent → pikku-agent}/SKILL.md +21 -21
- package/skills/pikku-ai-vercel/SKILL.md +18 -18
- package/skills/pikku-ai-voice/SKILL.md +15 -15
- package/skills/pikku-concepts/SKILL.md +5 -5
- package/skills/pikku-deploy-cloudflare/SKILL.md +9 -9
- package/skills/pikku-deploy-uws/SKILL.md +5 -2
- package/skills/pikku-deps/SKILL.md +42 -3
- package/skills/pikku-kysely/SKILL.md +68 -41
- package/skills/pikku-machine-auth/SKILL.md +10 -10
- package/skills/pikku-mcp/SKILL.md +19 -16
- package/skills/pikku-middleware/SKILL.md +14 -7
- package/skills/pikku-mongodb/SKILL.md +11 -11
- package/skills/pikku-n8n-import/references/addon-mapping.md +14 -8
- package/skills/pikku-product-second-opinion/SKILL.md +83 -73
- package/skills/pikku-rpc/SKILL.md +1 -1
- package/skills/pikku-scenario/SKILL.md +39 -31
- package/skills/pikku-software-archaeology/SKILL.md +27 -23
- package/skills/pikku-software-archaeology/references/blueprint.schema.json +580 -102
- package/skills/pikku-versioning/SKILL.md +87 -3
- package/skills/pikku-ws/SKILL.md +5 -2
|
@@ -16,56 +16,60 @@ The reader is a founder/PM/operator, not an engineer. If they finish a section a
|
|
|
16
16
|
|
|
17
17
|
## The cardinal rule: translate, don't dump
|
|
18
18
|
|
|
19
|
-
Every technical concept becomes a business outcome or a plain-language description. Never make the reader learn your vocabulary. If a term is unavoidable, define it in one clause the first time — but prefer describing the
|
|
20
|
-
|
|
21
|
-
| Don't write
|
|
22
|
-
|
|
23
|
-
| queue / worker / job
|
|
24
|
-
| workflow
|
|
25
|
-
| API / endpoint / route
|
|
26
|
-
| webhook
|
|
27
|
-
| event
|
|
28
|
-
| cron / scheduler
|
|
29
|
-
| schema / migration
|
|
30
|
-
| auth / session / token
|
|
31
|
-
| refactor / rewire
|
|
32
|
-
| cache
|
|
33
|
-
| race condition
|
|
34
|
-
| component
|
|
35
|
-
| route / page
|
|
36
|
-
| design system / component library | "the shared kit of screen pieces that keeps everything looking consistent"
|
|
37
|
-
| SSR / SPA / rendering
|
|
38
|
-
| MCP server
|
|
39
|
-
| SDK
|
|
40
|
-
| CLI
|
|
41
|
-
| theme token / design variable
|
|
42
|
-
| modal / drawer
|
|
43
|
-
|
|
44
|
-
When in doubt, say what the
|
|
19
|
+
Every technical concept becomes a business outcome or a plain-language description. Never make the reader learn your vocabulary. If a term is unavoidable, define it in one clause the first time — but prefer describing the _effect_ and skipping the term entirely.
|
|
20
|
+
|
|
21
|
+
| Don't write | Write instead (describe the effect) |
|
|
22
|
+
| --------------------------------- | ----------------------------------------------------------------------------------- |
|
|
23
|
+
| queue / worker / job | "a background task that runs on its own" |
|
|
24
|
+
| workflow | "a multi-step task that resumes where it left off if interrupted" |
|
|
25
|
+
| API / endpoint / route | "something the app (or another tool) can ask it to do" |
|
|
26
|
+
| webhook | "an automatic message the app sends to another tool when something happens" |
|
|
27
|
+
| event | "a signal that something happened, that other parts can react to" |
|
|
28
|
+
| cron / scheduler | "a timer that runs something on a schedule" |
|
|
29
|
+
| schema / migration | "the shape of your stored data" / "a change to how data is stored" |
|
|
30
|
+
| auth / session / token | "how the app knows who you are and what you're allowed to do" |
|
|
31
|
+
| refactor / rewire | "reorganizing the inside without changing what it does" |
|
|
32
|
+
| cache | "a saved copy kept around for speed" |
|
|
33
|
+
| race condition | "two things happening at once and stepping on each other" |
|
|
34
|
+
| component | "a reusable piece of the screen (a button, a chart, a table)" |
|
|
35
|
+
| route / page | "a screen in the app" |
|
|
36
|
+
| design system / component library | "the shared kit of screen pieces that keeps everything looking consistent" |
|
|
37
|
+
| SSR / SPA / rendering | "how pages get built and shown" (only mention if it affects speed or SEO) |
|
|
38
|
+
| MCP server | "a way for AI assistants to use your app's data and actions directly" |
|
|
39
|
+
| SDK | "a ready-made toolkit so other developers can build on your app" |
|
|
40
|
+
| CLI | "a way to drive the app by typing commands (for power users / automation)" |
|
|
41
|
+
| theme token / design variable | "a single setting (like your brand color) reused everywhere, so you change it once" |
|
|
42
|
+
| modal / drawer | "a pop-up box" / "a slide-out panel" |
|
|
43
|
+
|
|
44
|
+
When in doubt, say what the _user or the business_ experiences, not what the machine does.
|
|
45
45
|
|
|
46
46
|
## Report structure (layered — skim or dive)
|
|
47
47
|
|
|
48
48
|
Write these three parts in order. A reader can stop after Part 1.
|
|
49
49
|
|
|
50
50
|
**Part 1 — Executive summary (one page).**
|
|
51
|
-
|
|
52
|
-
-
|
|
53
|
-
-
|
|
51
|
+
|
|
52
|
+
- _What you have_: 2–3 sentences — what the product does and who uses it.
|
|
53
|
+
- _The headline_: the 3–5 biggest risks/opportunities, one plain line each.
|
|
54
|
+
- _Recommended order_: a table (Fix | Why it matters | Effort | Payoff). This is the part they act on.
|
|
54
55
|
|
|
55
56
|
**Part 2 — One section per major area** (drive the areas from `domains.json`; skip domains with nothing worth saying). Each section follows this shape (see `example/sample-report.md`):
|
|
56
|
-
|
|
57
|
-
-
|
|
58
|
-
-
|
|
59
|
-
-
|
|
60
|
-
-
|
|
57
|
+
|
|
58
|
+
- _What this does_ — the capability in business terms.
|
|
59
|
+
- _How it works today_ — a plain walkthrough, ideally as a small story ("on a timer, the app re-reads each site, compares…").
|
|
60
|
+
- _What's working_ — genuine credit. Never skip this; a report that's all criticism gets dismissed.
|
|
61
|
+
- _What's holding you back_ — each problem MUST carry: **what it means for you** (business impact), **severity** (Minor / Worth fixing / Serious / Urgent), and **effort** (Small / Medium / Large).
|
|
62
|
+
- _How I'd do it differently — and why it's worth it_ — the opinionated part. Argue the improvement in one of these business outcomes: **more reliable / fewer surprises**, **faster to add features**, **cheaper to run**, **safer / less risk**, **easier to maintain or hand off**. Be explicit whether it's a cheap rewire or an expensive rebuild.
|
|
61
63
|
|
|
62
64
|
**Part 3 — Appendix.**
|
|
63
|
-
|
|
64
|
-
-
|
|
65
|
+
|
|
66
|
+
- _How confident am I_ — REQUIRED. Where you're certain vs guessing; what you'd verify against real data first. The blueprint carries confidence tiers — anything you're relaying from a `low`/`medium` entry, or from a reconstructed (`explicit: false`) event, says so here.
|
|
67
|
+
- _Glossary_ (optional) — only for any term that slipped through.
|
|
65
68
|
|
|
66
69
|
## Rewire vs rebuild (say which)
|
|
67
70
|
|
|
68
71
|
The blueprint's `migration.json` tells you which is which — `mappings[]` is what survives (each with its `recommendation`), `dropped[]` is what goes. The reader needs to know because the cost is 10× different.
|
|
72
|
+
|
|
69
73
|
- **Rewire** — the valuable machinery exists; you're connecting pieces or turning something on. Cheap, low-risk. (Most "it should be automatic but isn't" findings are this.)
|
|
70
74
|
- **Rebuild** — the capability doesn't exist or is fundamentally wrong. Expensive, risky. Reserve the word for when it's true; founders hear "rewrite" and panic or overspend.
|
|
71
75
|
|
|
@@ -76,33 +80,35 @@ Never recommend a full rewrite because the code is messy. Messy-but-working is a
|
|
|
76
80
|
When the blueprint has the consumer-surface files, add these to the report — they're often where a founder's questions actually live ("why does the app feel inconsistent?", "can partners build on this?").
|
|
77
81
|
|
|
78
82
|
**The screens (`frontend-*.json`) — one area section, founder-framed.**
|
|
79
|
-
|
|
80
|
-
-
|
|
81
|
-
-
|
|
82
|
-
-
|
|
83
|
-
|
|
84
|
-
- **
|
|
85
|
-
- **
|
|
86
|
-
|
|
83
|
+
|
|
84
|
+
- _What a user can do_ — walk the main screens as a journey, not a component list.
|
|
85
|
+
- _Consistency_ — is it built from one shared kit of screen pieces, or a patchwork? A consistent kit means changes are cheap and the app feels coherent; a patchwork means every change is bespoke and the look drifts. Say which, plainly.
|
|
86
|
+
- _The expensive pieces_ — this is the key frontend insight. Most of the screen is standard pieces that are cheap to rebuild or restyle. A **small number carry real custom logic** — a bespoke chart, a complicated data table, a drawing/drag interaction, a rich editor. Those are the parts that take real effort to move or change, and the ones most likely to break. Name them, say what they do, and flag them as the real work — so nobody assumes "it's just screens, it'll be quick."
|
|
87
|
+
- _Design consistency (the "it looks a bit off" problems)_ — from the blueprint's design findings, call out broken patterns in plain terms and, crucially, why each matters and roughly what it costs to fix. Common ones and how to frame them:
|
|
88
|
+
- **The same action behaves differently in different places** (a slide-out panel here, a pop-up box there for the same task). _Why it matters:_ the app feels inconsistent and users have to re-learn each screen. _Fix:_ pick one pattern and apply it everywhere — cheap.
|
|
89
|
+
- **Colors/spacing are hardcoded instead of set in one place.** _Why it matters:_ changing your brand color, or fixing contrast, means hunting through every screen instead of editing one setting — slow and error-prone. _Fix:_ move them to shared "design tokens" — a small, high-leverage cleanup.
|
|
90
|
+
- **The same element looks different from page to page** (buttons, headings, cards). _Why it matters:_ reads as unpolished and erodes trust, especially in a paid product. _Fix:_ one shared version of each, reused — cheap and makes every future change faster.
|
|
91
|
+
These are almost always **cheap rewires with an outsized polish/trust payoff**, not rebuilds. Give each an effort (usually Small–Medium) and say the payoff is perceived quality + faster future changes. Do NOT design-nitpick without a reason — every design point needs a "why it matters to you." And credit consistency where the app already has it.
|
|
87
92
|
- Frame rebuild/restyle work as **rewire vs rebuild**: restyling standard pieces to a consistent kit is cheap; re-creating a custom-logic piece is real engineering.
|
|
88
93
|
|
|
89
94
|
**How your app can be driven (`interfaces.json`) — usually a short, positive section.**
|
|
90
|
-
Explain, in one line each, the ways the product can be used: people through the web, developers through an API or toolkit, AI assistants through a direct connection (MCP), power users through the command line. This is often a genuine strength worth naming — an app that agents and partners can build on is more valuable than one only humans can click. But be honest about `status`: a connection that exists but only does two things is a
|
|
95
|
+
Explain, in one line each, the ways the product can be used: people through the web, developers through an API or toolkit, AI assistants through a direct connection (MCP), power users through the command line. This is often a genuine strength worth naming — an app that agents and partners can build on is more valuable than one only humans can click. But be honest about `status`: a connection that exists but only does two things is a _start_, not a feature — say so.
|
|
91
96
|
|
|
92
97
|
## Technology choices — the honest tradeoffs (don't be cheap on the cons)
|
|
93
98
|
|
|
94
|
-
The app made specific technology bets. The founder deserves to know what each bet
|
|
99
|
+
The app made specific technology bets. The founder deserves to know what each bet _bought_ and what it _costs_ — in business terms, tied to their situation (early vs scaling, chasing enterprise deals or not, big team or two people). Present every significant choice as a genuine tradeoff with **both sides**. Never cheerlead a technology, and never trash one — but do not soften the disadvantages to sound positive. A report that only lists upsides is not honest and is not useful.
|
|
95
100
|
|
|
96
101
|
Rules:
|
|
97
|
-
|
|
102
|
+
|
|
103
|
+
- For each notable choice (framework, auth, hosting, database, key libraries): **what it buys** and **what it costs**, both in plain business terms, then a recommendation tied to _their_ stage and goals — usually "keep it, here's what to watch" rather than "switch."
|
|
98
104
|
- Tie cons to consequences the founder feels: vendor bills, security/breach liability, hiring difficulty, how fast they can ship, enterprise-sales blockers, the risk of betting on something young.
|
|
99
105
|
- Distinguish "younger / smaller community" (a real, manageable risk) from "wrong choice" (rare). Most stack choices are defensible; the job is informed eyes-open, not alarm.
|
|
100
106
|
- **Verify before you disparage.** "Don't be cheap on the cons" means ACCURATE cons, not invented ones. Do NOT label a technology immature, niche, or feature-poor from vibes, its name, or its age — check its actual adoption, maturity, and feature set first. And separate an **inherent tradeoff of an approach** (e.g. self-hosting anything means you run and secure it) from a **deficiency of a specific tool** (often false — the tool may be mature and full-featured). Overstating cons is as dishonest as hiding them.
|
|
101
107
|
- **Hold your own recommendation to the same bar.** If "how I'd do it differently" lands on a specific stack — Pikku included — it gets the same both-sides treatment as everything else, cons first-class. Pinning someone's dependency for being pre-1.0 while not mentioning that the replacement is pre-1.0 too isn't a second opinion, it's a pitch.
|
|
102
108
|
|
|
103
|
-
- **Derive the app's choices from the blueprint, never from a list in this file.** Read the stack off `architecture.json`, `integrations.json`, `frontend.json`, and the manifest the repo actually has (`package.json`, `Gemfile`, `go.mod`, …). Cover the bets that are
|
|
109
|
+
- **Derive the app's choices from the blueprint, never from a list in this file.** Read the stack off `architecture.json`, `integrations.json`, `frontend.json`, and the manifest the repo actually has (`package.json`, `Gemfile`, `go.mod`, …). Cover the bets that are _load-bearing for this product_: typically the framework, the auth/identity approach, the datastore, the hosting/deploy model, the payment and other critical vendor integrations, and anything the blueprint marks `replacementDifficulty: hard`. A choice earns a paragraph if switching it would be expensive, or if living with it constrains the business — not because it appears in some canonical list. If you could write the verdict before reading the blueprint, you are not giving a second opinion.
|
|
104
110
|
|
|
105
|
-
### The choices
|
|
111
|
+
### The choices _this_ app made
|
|
106
112
|
|
|
107
113
|
Whatever the blueprint shows. A Rails app's bets are Rails, Devise, Pundit, MySQL, Sidekiq, its ERP and payment vendors; a Go app's are different again. Fill in buys/costs/usually for each, from evidence. If a legacy choice is working fine, credit it and move on — "boring and working" is a feature, and the pressure to find something to say about a stack is exactly what produces dishonest reports.
|
|
108
114
|
|
|
@@ -113,45 +119,49 @@ Worth naming when it applies: adopting one coherent system in place of hand-roll
|
|
|
113
119
|
Include this section **only if you are actually recommending a rebuild** onto it — and then give every part of it the same both-sides treatment you gave the app's own bets, per the "hold your own recommendation to the same bar" rule above. These are not choices the app made; they are choices you are proposing, which is exactly why their costs are the reader's to weigh. The Pikku target stack is Pikku + Better Auth + TanStack Start + Mantine; the framings below are reference material for the parts you actually recommend, not a script to recite.
|
|
114
120
|
|
|
115
121
|
**Better Auth (self-hosted sign-in) — instead of a paid service like Auth0/Clerk.**
|
|
116
|
-
|
|
117
|
-
-
|
|
118
|
-
-
|
|
119
|
-
-
|
|
122
|
+
|
|
123
|
+
- _Buys you:_ a mature, battle-tested, **framework-agnostic** library with a deep first-class plugin catalog — two-factor auth, multi-tenancy/organizations, multi-session, rate limiting, Stripe subscription billing, an admin panel, API keys for partners/automation, single-sign-on — plus a plugin system to add more without forking. You keep your users in your own database (single source of truth, no per-user bill that grows with success), with full control of the auth flows, and you can run it embedded in the app or as a standalone self-hosted auth server. So self-hosting here means neither giving up features nor rolling your own security.
|
|
124
|
+
- _Cleans up messy auth (often the biggest win):_ adopting it consolidates the kind of hand-rolled, drifted auth that accumulates in an older codebase — several different ways of deciding who's an admin, a bespoke token table, a back-door test login, home-grown encryption — into **one coherent system**. If the blueprint shows a before (legacy, bespoke) and after (on Better Auth), point at it directly: the sprawl collapses into a single well-structured setup. A concrete reliability-and-security upgrade, not just a swap.
|
|
125
|
+
- _Costs you (the honest tradeoff — operational, not security-implementation):_ the auth flows and security practices are handled by the library, so this is NOT "build secure auth from scratch." What self-hosting means is you **operate** it — hosting, upgrades, uptime, and incident response sit with your team, where a paid SaaS runs that for you and bundles hosted extras (bot/anomaly detection, leaked-password monitoring, vendor compliance certifications) you'd otherwise operate and document yourself. You trade a per-user bill and vendor ops for control, data ownership, and predictable cost.
|
|
126
|
+
- _Usually:_ a strong default for an independent product — mature, full-featured, framework-agnostic, and frequently a genuine cleanup of inherited auth. The real question is who owns operating it, not whether the tool is good enough.
|
|
120
127
|
|
|
121
128
|
**TanStack Start (the web framework) — instead of the incumbent (Next.js).**
|
|
122
|
-
|
|
123
|
-
-
|
|
124
|
-
-
|
|
129
|
+
|
|
130
|
+
- _Buys you:_ modern, tidy developer experience; strong type-safety that catches whole classes of bugs before users see them; fast iteration; fine-grained control; deploys well to modern/edge hosting. The libraries underneath it — the TanStack ecosystem (Query, Router, Table) — are mature, battle-tested and everywhere in React.
|
|
131
|
+
- _Costs you (the honest tradeoff — maturity of the framework itself):_ separate the ecosystem from the framework. Query/Router/Table are mature; the framework that wraps them is younger, and **you must check its release stage on tanstack.com/start at the moment you write** — do not infer it from the npm version. `@tanstack/react-start` has been on 1.x since early 2025 because its major tracks the **Router** version line, so "1.168.x" says nothing about whether Start itself has shipped a stable 1.0. If it is still pre-1.0, the cost is pinning an exact version and budgeting for upgrade work as it settles. Either way it is newer than the incumbent (Next.js), which has the largest ecosystem — fewer ready-made templates and third-party examples, and a smaller (though growing) pool of developers who've used _this specific_ framework, which can make hiring slightly slower.
|
|
132
|
+
- _Usually:_ a credible, modern choice on a mature foundation. Whether it also carries pinning-and-upgrade risk depends on the release stage you just checked — say which you found, rather than repeating either verdict from here.
|
|
125
133
|
|
|
126
134
|
**Pikku (the framework a rebuild would land on) — instead of staying where you are.**
|
|
127
|
-
|
|
128
|
-
-
|
|
129
|
-
-
|
|
135
|
+
|
|
136
|
+
- _Buys you:_ one way to write a capability and drive it from anywhere — web, background jobs, timers, realtime, AI assistants, the command line — so a feature is written once instead of five times. Type-safe clients and the API spec fall out of the code rather than being hand-maintained until they drift. The sprawl an organically-grown app accumulates collapses into one shape a small team can hold in its head.
|
|
137
|
+
- _Costs you (the honest tradeoff — it is younger than anything it would replace):_ Pikku has **not shipped a stable 1.0** — at the time of writing it is 0.12.x, and 0.13 is the first release that promises backwards compatibility; check the published version rather than repeating this one. Until then upgrades can break you. In practice: pin your version, budget for upgrade work, and know that the community, the ready-made examples, and the pool of developers who have used it are all far smaller than the incumbent's — smaller than TanStack Start's, let alone Next.js's. Being pre-1.0 is normal for a young framework, and survivable, but it is a real cost and it is the reader's to weigh, not yours to skip.
|
|
138
|
+
- _Usually:_ worth it when the real problem is sprawl — many surfaces, hand-maintained glue, the same rule implemented three slightly different ways — and the team wants one shape instead of five. Harder to justify for an app that works and needs a few rewires: those are usually cheaper in place. If nobody has capacity to own upgrades, that's a real reason to wait.
|
|
130
139
|
|
|
131
140
|
Check these statuses before you write them up rather than repeating them from here — a framework's release stage moves, and the point is the current fact, not this example.
|
|
132
141
|
|
|
133
142
|
## Delivery
|
|
134
143
|
|
|
135
144
|
Produce **both**:
|
|
145
|
+
|
|
136
146
|
1. A **markdown** report in the repo (e.g. `docs/reports/<app>-second-opinion.md`) — versioned, diffable.
|
|
137
147
|
2. A **shareable web page**: load the **artifact-design** skill, then render the same report as one clean, print-friendly, theme-aware page they can send to a cofounder or the agency. Same content, nicer to read.
|
|
138
148
|
|
|
139
149
|
## Red flags — you're writing the wrong report
|
|
140
150
|
|
|
141
|
-
| Symptom
|
|
142
|
-
|
|
143
|
-
| A technical term with no translation
|
|
144
|
-
| A problem with no "what it means for you"
|
|
145
|
-
| All problems, no credit
|
|
146
|
-
| A recommendation with no effort + payoff
|
|
147
|
-
| "Rewrite the app"
|
|
148
|
-
| A guess stated as fact
|
|
149
|
-
| Only listed the upsides of a technology choice
|
|
150
|
-
| Recommended a stack (including ours) without its cons | You applied a maturity bar to their technology and exempted your own. Both sides, or cut the recommendation.
|
|
151
|
-
| Trashed a technology as "the wrong choice"
|
|
152
|
-
| "The frontend is just screens, it'll be quick"
|
|
153
|
-
| A design point with no "why it matters"
|
|
154
|
-
| Reads like a code review
|
|
151
|
+
| Symptom | Fix |
|
|
152
|
+
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
153
|
+
| A technical term with no translation | Rephrase as the effect on the user/business, or cut the term. |
|
|
154
|
+
| A problem with no "what it means for you" | Incomplete — add the business impact or delete it. |
|
|
155
|
+
| All problems, no credit | You'll lose the reader's trust. Name what's genuinely good. |
|
|
156
|
+
| A recommendation with no effort + payoff | Not decision-useful. Add both. |
|
|
157
|
+
| "Rewrite the app" | Almost always wrong. Separate rewire (cheap) from rebuild (dear); lean on what `migration.json.mappings` says survives. |
|
|
158
|
+
| A guess stated as fact | Mark confidence. "I'm certain" and "I'd need to check" are different sentences. |
|
|
159
|
+
| Only listed the upsides of a technology choice | Not honest. Every bet has a cost — name it in business terms, don't soften it to sound positive. |
|
|
160
|
+
| Recommended a stack (including ours) without its cons | You applied a maturity bar to their technology and exempted your own. Both sides, or cut the recommendation. |
|
|
161
|
+
| Trashed a technology as "the wrong choice" | Equally lazy. Most choices are defensible; frame as tradeoff + "what to watch," not a verdict. |
|
|
162
|
+
| "The frontend is just screens, it'll be quick" | Wrong. The custom-logic pieces (charts, complex tables, editors) are real work — flag them separately from the cheap standard pieces. |
|
|
163
|
+
| A design point with no "why it matters" | Taste, not advice. Tie every design finding to user perception (polish/trust) or maintenance cost (change-once vs hunt-everywhere), plus effort. |
|
|
164
|
+
| Reads like a code review | Wrong audience. Would a founder know what to _do_ after this paragraph? |
|
|
155
165
|
|
|
156
166
|
## Relationship to pikku-software-archaeology
|
|
157
167
|
|
|
@@ -42,7 +42,7 @@ See `pikku-concepts` for the core mental model.
|
|
|
42
42
|
| `rpc.remote(name, data)` | Remote call via DeploymentService |
|
|
43
43
|
| `rpc.exposed(name, data)` | Call functions marked with `expose: true` |
|
|
44
44
|
| `rpc.startWorkflow(name, input)` | Start a workflow (see `pikku-workflow`) |
|
|
45
|
-
| `rpc.agent.run/stream(...)` | Run an AI agent (see `pikku-
|
|
45
|
+
| `rpc.agent.run/stream(...)` | Run an AI agent (see `pikku-agent`) |
|
|
46
46
|
| `rpc.agent.resume/approve(...)` | Answer a tool-approval interrupt |
|
|
47
47
|
| `rpc.agent.interrupt(runId)` | Stop an in-flight run |
|
|
48
48
|
|
|
@@ -102,16 +102,23 @@ Prefer `expectEventually` over sleeping.
|
|
|
102
102
|
### Asserting on an agent's answer (`expectScore`)
|
|
103
103
|
|
|
104
104
|
An agent's output is not comparable to a fixed string, so it is graded rather
|
|
105
|
-
than matched. Declare the rubric with `
|
|
106
|
-
`
|
|
105
|
+
than matched. Declare the rubric with `pikkuAgentScorer` (grades in code) or
|
|
106
|
+
`pikkuAgentJudge` (grades with a model) in a `*.scorer.ts` file, name it on the
|
|
107
107
|
agent's `scorers`, then assert on the run the scenario just triggered:
|
|
108
108
|
|
|
109
109
|
```typescript
|
|
110
|
-
const { runId } = await scenario.when(
|
|
111
|
-
|
|
112
|
-
|
|
110
|
+
const { runId } = await scenario.when(
|
|
111
|
+
'asks for a summary',
|
|
112
|
+
'runAssistant',
|
|
113
|
+
{
|
|
114
|
+
prompt: data.prompt,
|
|
115
|
+
},
|
|
116
|
+
{ actor: actors.user }
|
|
117
|
+
)
|
|
113
118
|
|
|
114
|
-
await scenario.expectScore('answered briefly', runId, 'brevity', {
|
|
119
|
+
await scenario.expectScore('answered briefly', runId, 'brevity', {
|
|
120
|
+
atLeast: 0.8,
|
|
121
|
+
})
|
|
115
122
|
```
|
|
116
123
|
|
|
117
124
|
The default bound is `atLeast: 0.5`, so an unqualified `expectScore` still fails
|
|
@@ -451,17 +458,18 @@ A `browser` binding gets a session bound to **its actor**, signed in through the
|
|
|
451
458
|
Browser steps are where **intent, not actions** earns its keep: the step is one intent, the clicking lives in shared utilities, and the step arrives before it acts. Write the mechanics below into utilities and keep the step body to three or four calls that read as a sentence.
|
|
452
459
|
|
|
453
460
|
```typescript
|
|
454
|
-
export const opensTheCart = pikkuScenarioStep<
|
|
455
|
-
{
|
|
456
|
-
|
|
457
|
-
|
|
458
|
-
|
|
459
|
-
|
|
460
|
-
|
|
461
|
-
|
|
462
|
-
|
|
463
|
-
}
|
|
464
|
-
)
|
|
461
|
+
export const opensTheCart = pikkuScenarioStep<
|
|
462
|
+
{ path: string },
|
|
463
|
+
{ url: string }
|
|
464
|
+
>({
|
|
465
|
+
name: 'opensTheCart',
|
|
466
|
+
description: 'opens the cart',
|
|
467
|
+
browser: async (_services, { path }, { browser }) => {
|
|
468
|
+
await browser.goto(path)
|
|
469
|
+
return { url: browser.page.url() }
|
|
470
|
+
},
|
|
471
|
+
default: async ({ rpc }) => ({ url: (await rpc.invoke('getCart', {})).url }),
|
|
472
|
+
})
|
|
465
473
|
```
|
|
466
474
|
|
|
467
475
|
- Install `@pikku/playwright` and `@playwright/test`, and import `@pikku/playwright` once (`import type {} from '@pikku/playwright'`) so `browser.page` is a typed Playwright `Page`. Without it you still get the structural `goto`/`screenshot` handle.
|
|
@@ -535,19 +543,19 @@ SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --no-browser --exclud
|
|
|
535
543
|
|
|
536
544
|
`run` takes the environment as a **required positional** — the key from `environments`. Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
|
|
537
545
|
|
|
538
|
-
| Flag
|
|
539
|
-
|
|
|
540
|
-
| `--flows` / `-f`
|
|
541
|
-
| `--features`
|
|
542
|
-
| `--tags` / `-t`
|
|
543
|
-
| `--exclude-tags`
|
|
544
|
-
| `--run <surface>`
|
|
545
|
-
| `--no-browser`
|
|
546
|
-
| `--strict`
|
|
547
|
-
| `--spawn` / `--keep-alive` | Start `pikku dev` on the environment's apiUrl for the run; optionally leave it up
|
|
548
|
-
| `--api-url` / `--app-url`
|
|
549
|
-
| `--trace`
|
|
550
|
-
| `--coverage`
|
|
546
|
+
| Flag | Effect |
|
|
547
|
+
| -------------------------- | --------------------------------------------------------------------------------- |
|
|
548
|
+
| `--flows` / `-f` | Comma-separated scenario names |
|
|
549
|
+
| `--features` | Comma-separated feature ids |
|
|
550
|
+
| `--tags` / `-t` | Match-any tag filter |
|
|
551
|
+
| `--exclude-tags` | Hold tags back — unless the flow is named directly with `--flows` |
|
|
552
|
+
| `--run <surface>` | `default` (the default), `browser`, or `cli` |
|
|
553
|
+
| `--no-browser` | Shorthand for `--run default`; scenarios with browser steps report as **skipped** |
|
|
554
|
+
| `--strict` | Fail, rather than pass, a `then` with no witness on the run's surface |
|
|
555
|
+
| `--spawn` / `--keep-alive` | Start `pikku dev` on the environment's apiUrl for the run; optionally leave it up |
|
|
556
|
+
| `--api-url` / `--app-url` | Override the environment's URLs — for a target that only exists at run time |
|
|
557
|
+
| `--trace` | Keep every stack frame on failure (default shows only the project's own) |
|
|
558
|
+
| `--coverage` | Reset/snapshot server coverage per scenario |
|
|
551
559
|
|
|
552
560
|
Output is `PASS <name> (<ms>) → <output>` / `FAIL <name> (<ms>): <error>`, then `N/M scenarios passed against '<env>'`. A scenario inside a feature is named `<Feature> › <scenario> <data>`.
|
|
553
561
|
|
|
@@ -640,7 +648,7 @@ Services are plain objects — a Pikku function is pure business logic, so a moc
|
|
|
640
648
|
| A step named `clicksAddToBasket` / `opensThePage` | That is an action, not an intent. Name the step for what the actor wanted; put the clicking in a utility. |
|
|
641
649
|
| A browser step that assumes it is already on a page | It can then only run mid-flow. Arrive first — check the URL, navigate if needed. |
|
|
642
650
|
| A `browser` binding guarding `if (!browser)` | The binding guarantees it. The guard hides the real error, which is a missing actor (`PKU677`). |
|
|
643
|
-
| A step with a `func:` instead of a surface binding
|
|
651
|
+
| A step with a `func:` instead of a surface binding | There is no `func` on a step. Bodies live under `default` / `browser` / `cli`; a step with none throws at load. |
|
|
644
652
|
| `expectEventually` in a `pikkuWorkflowFunc` | `PKU675` — scenario-only. |
|
|
645
653
|
| Coverage silently 0 | Server not run with `--coverage`, verbose functions meta not deployed, `scaffold.scenarios` unset, or no actors configured. |
|
|
646
654
|
|
|
@@ -13,6 +13,7 @@ Extract **intent over implementation**. A repository is a fossil record of produ
|
|
|
13
13
|
**You are the parser.** Do not build or rely on regex/AST scanners — read the code with your own tools (Grep, Read, subagents). This is what makes the skill language-agnostic: an Express app, a Rails app, and a Django app all yield the same blueprint shape.
|
|
14
14
|
|
|
15
15
|
**Two layers, never merged silently:**
|
|
16
|
+
|
|
16
17
|
- **Facts** — behavior directly observed in code, schema, or tests. Cite them.
|
|
17
18
|
- **Inferred intent** — the product reasoning you reconstruct. Mark it with `confidence` and say what evidence it rests on.
|
|
18
19
|
|
|
@@ -64,6 +65,7 @@ Work in phases. For small repos (< ~50 source files) do them inline; for larger
|
|
|
64
65
|
### Phase 1 — Survey (facts only)
|
|
65
66
|
|
|
66
67
|
Build an inventory before interpreting anything:
|
|
68
|
+
|
|
67
69
|
- Manifests (`package.json`, `Gemfile`, `pyproject.toml`, `go.mod`, `composer.json`): dependencies are integration hints; scripts are entry points.
|
|
68
70
|
- Entry points: servers, route registrations, cron/scheduler setup, queue workers, CLI binaries.
|
|
69
71
|
- Data layer: migrations, schema files, model classes, raw DDL (check comments too — schemas hide in comments in migration-less repos).
|
|
@@ -75,18 +77,19 @@ Build an inventory before interpreting anything:
|
|
|
75
77
|
|
|
76
78
|
Where intent hides, per ecosystem (read these first):
|
|
77
79
|
|
|
78
|
-
| Ecosystem
|
|
79
|
-
|
|
80
|
-
| Rails
|
|
81
|
-
| Express/Node
|
|
82
|
-
| Django
|
|
83
|
-
| Laravel
|
|
84
|
-
| Go
|
|
80
|
+
| Ecosystem | Highest-yield locations |
|
|
81
|
+
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
82
|
+
| Rails | `config/routes.rb`, model validations + callbacks + `aasm`/state machines, `app/policies` (Pundit) / `ability.rb` (CanCan), Sidekiq/ActiveJob workers, `db/schema.rb`, specs (esp. request + model specs) |
|
|
83
|
+
| Express/Node | route registration files, middleware chains (auth!), inline `if` guards in handlers, SQL/ORM models, `jobs/`+crontab refs, webhook handlers |
|
|
84
|
+
| Django | `urls.py`, model `Meta`/constraints/`clean()`, DRF serializers + permissions classes, celery tasks, admin.py (reveals internal workflows) |
|
|
85
|
+
| Laravel | `routes/`, FormRequests (validation), Policies/Gates, Jobs + scheduler in `Kernel.php`, migrations |
|
|
86
|
+
| Go | mux/router setup, middleware, struct tags, `cmd/` binaries (each is a component) |
|
|
85
87
|
| Frontend (React/Vue/etc.) | router config / file-based routes (pages a user reaches), the component tree, the design-system import (`@mantine/*`, `@mui/*`, Tailwind config) to judge consistency, the data layer (react-query/tRPC/fetch wrappers) to tie UI back to backend queries, charts/tables/editors (the `custom-logic` port risk), auth wiring |
|
|
86
88
|
|
|
87
89
|
### Phase 2 — Test excavation (do not skip)
|
|
88
90
|
|
|
89
91
|
Tests are the closest thing to an executable product spec. For every test file:
|
|
92
|
+
|
|
90
93
|
- `describe`/`context`/`it` names → **scenarios** (attach to the matching workflow in `workflows.json` under `scenarios[]`, with `fromTest` set).
|
|
91
94
|
- User-flow/journey harnesses (`pikkuUserFlow` stories, cucumber `.feature` files, Playwright journeys) are the highest-grade scenario source — they already ARE given/when/outcome sequences; extract them verbatim.
|
|
92
95
|
- Assertions → confirmations of policies and invariants (upgrade their `confidence` to `high`, add the test as evidence).
|
|
@@ -126,14 +129,15 @@ Run each lens over the surveyed material. Rules that counter the classic failure
|
|
|
126
129
|
|
|
127
130
|
**Migration** — map current file clusters → future domain + concepts; list files to drop with reasons; list decisions a human must make before rebuild.
|
|
128
131
|
|
|
129
|
-
**Interfaces** (`interfaces.json`, optional) — every way the product is CONSUMED, one entry per channel, not per route. A product is usually driven through several: a **web UI** (humans), a **CLI** (developers/operators), an **MCP server** (AI agents — in a Pikku app each MCP tool IS a `pikkuFunc`), an **OpenAPI/REST** surface (developers/external systems, often
|
|
132
|
+
**Interfaces** (`interfaces.json`, optional) — every way the product is CONSUMED, one entry per channel, not per route. A product is usually driven through several: a **web UI** (humans), a **CLI** (developers/operators), an **MCP server** (AI agents — in a Pikku app each MCP tool IS a `pikkuFunc`), an **OpenAPI/REST** surface (developers/external systems, often _generated_ from the routes), a **generated SDK**, **realtime** (websocket/SSE), and **webhooks** (in/out). For each: `kind`, `audience`, `purpose`, roughly how many ops it exposes, whether it's `generated` vs hand-written, which domains it serves, and `status` (complete/partial/stub — an MCP server with two tools is `stub`). This layer answers "who can drive this, and how" — it is the map the second-opinion skill needs to explain that the app is usable by people, developers, and agents.
|
|
130
133
|
|
|
131
134
|
**Frontend** (`frontend.json` + `frontend-routes.json` + `frontend-components.json`, optional) — the web UI, which needs its own treatment because frontends vary wildly (framework, router, styling, state, data, auth) and the rebuild target is opinionated: **everything in one component system (Mantine), one data layer, one auth**.
|
|
135
|
+
|
|
132
136
|
- `frontend.json` records the stack as FACTS: framework (e.g. TanStack Start), rendering (SSR/streaming/SPA), router, styling/design system, `designSystemConsistency`, state management, data layer (e.g. pikku-react-query vs REST helpers), auth (e.g. better-auth), i18n. Name the real technologies — the second-opinion skill weighs their tradeoffs, so record them precisely (do NOT editorialize here; this file is facts).
|
|
133
137
|
- `frontend-routes.json` is the page tree: each route's `purpose` in product terms, `auth`, the `dataFrom` (query/command names it reads — reuse the backend concept names so the UI ties back to the domain), the `usesComponents`, and the `userFlows` it belongs to.
|
|
134
|
-
- `frontend.json.designFindings` captures **broken/inconsistent design patterns** as concrete, cited observations (not taste). Actively hunt for:
|
|
138
|
+
- `frontend.json.designFindings` captures **broken/inconsistent design patterns** as concrete, cited observations (not taste). Actively hunt for: _interaction inconsistency_ (the same job done as a modal in one place and a drawer in another; inconsistent confirm dialogs); _theming not tokenized_ (hardcoded hex colors, magic spacing/font sizes, inline styles instead of theme tokens/variables — grep for `#[0-9a-f]{3,6}`, `style={{`, raw `px` values); _cross-page inconsistency_ (the same element — button, page header, card — styled differently across routes); _component duplication_ (three near-identical cards/tables for one purpose); _design-system bypass_ (raw HTML/CSS where a library component exists). Each finding gets an example, its impact (feels unpolished / a color change means hunting every file), and a fix (standardize on one pattern / move to tokens / extract one shared component). These are almost always cheap cleanups, and they are exactly what a non-technical owner perceives as "the app looks off" without being able to say why.
|
|
135
139
|
- `frontend-components.json` is where the frontend's real migration cost lives, in the **`rebuild`** field: `mantine-standard` (maps 1:1 to a Mantine component — trivial), `mantine-composition` (built from Mantine primitives — straightforward), `custom-style` (diverges only visually — normalize to Mantine), or **`custom-logic`** (bespoke behavior — a custom chart, a virtualized/complex table, a canvas, drag-and-drop, a rich editor — that must be **ported**, not re-skinned). A `custom-logic` component MUST fill `customLogic` explaining the behavior, and should list the `dependencies` (charting/table/editor libs) that make it a real port. This split — "trivially re-Mantine-able" vs "carries logic that must survive the port" — is the single most useful thing the frontend extraction produces.
|
|
136
|
-
- **Server-rendered / non-React frontends still get all three files — do NOT skip them.** The `rebuild` enum is named for the target stack, but the distinction it draws is target-agnostic:
|
|
140
|
+
- **Server-rendered / non-React frontends still get all three files — do NOT skip them.** The `rebuild` enum is named for the target stack, but the distinction it draws is target-agnostic: _trivial stock element_ vs _composed from primitives_ vs _visual divergence only_ vs **bespoke behavior that must be ported**. A Rails app (Slim/ERB + ViewComponents + Hotwire/Stimulus), a Django app (templates + HTMX), or a Laravel app (Blade + Livewire) all have that same split, and it is just as load-bearing there. Use the enum values verbatim (the validator enforces them), classify by what the thing actually _does_, and record the vocabulary mismatch in `frontend.json.notes`. Concretely: treat a template partial + its behavior controller (Stimulus/Alpine/Livewire) as ONE component and classify the pair; a server-rendered app's `custom-logic` is the same list as a SPA's (maps, charts, video players, payment elements, drag-to-reorder, rich text, QR, live-updating regions), plus anything whose behavior rides on a streaming/partial-update contract (Turbo Streams, HTMX swaps) — get that mapping wrong on rebuild and pages show stale data. Omitting these files because "it isn't a React app" hides the frontend's entire migration cost, which is the one thing this file exists to expose.
|
|
137
141
|
- The `designFindings` grep hints above are React-flavored; the server-rendered equivalents are inline `style=` attributes in templates, hardcoded hex in the stylesheet tree, per-locale forked templates instead of i18n, and design-system bypass = raw markup where a component/partial already exists. Same findings, different needles.
|
|
138
142
|
|
|
139
143
|
### Phase 4 — Cross-check, validate, synthesize
|
|
@@ -148,7 +152,7 @@ Run each lens over the surveyed material. Rules that counter the classic failure
|
|
|
148
152
|
- `high` — the behavior itself is in the cited code/schema/test.
|
|
149
153
|
- `medium` — inferred from multiple converging signals (naming + partial code + a test name).
|
|
150
154
|
- `low` — plausible reconstruction; MUST also be phrased tentatively in blueprint.md and usually deserves a `decisionsNeeded` entry.
|
|
151
|
-
- Comments and docs describe
|
|
155
|
+
- Comments and docs describe _intended_ behavior; code describes _actual_ behavior. When they disagree, record the code's behavior as the fact and the disagreement as a gap.
|
|
152
156
|
|
|
153
157
|
## Scaling up (large repos)
|
|
154
158
|
|
|
@@ -160,18 +164,18 @@ Run each lens over the surveyed material. Rules that counter the classic failure
|
|
|
160
164
|
|
|
161
165
|
## Red Flags — you are about to produce a worthless blueprint
|
|
162
166
|
|
|
163
|
-
| Thought
|
|
164
|
-
|
|
165
|
-
| "I'll write it up as markdown docs"
|
|
166
|
-
| "The routes are the API contract"
|
|
167
|
-
| "This is obvious, no citation needed"
|
|
168
|
-
| "The folder structure tells me the domains"
|
|
169
|
-
| "Tests are just tests, skip them"
|
|
170
|
-
| "No events are emitted, so events: []"
|
|
171
|
-
| "Cron jobs aren't workflows"
|
|
172
|
-
| "The frontend is just the web routes"
|
|
173
|
-
| "A component list is enough"
|
|
174
|
-
| "I'll skip the validator, the JSON looks right" | Run it. Missing domains refs, dangling event names, and undescribed custom-logic components are exactly what it catches.
|
|
167
|
+
| Thought | Reality |
|
|
168
|
+
| ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
169
|
+
| "I'll write it up as markdown docs" | Only `blueprint.md` is prose. The 14 JSON files ARE the deliverable; a generator consumes them. |
|
|
170
|
+
| "The routes are the API contract" | Routes are evidence. Lift each to a command/query or you've documented plumbing, not product. |
|
|
171
|
+
| "This is obvious, no citation needed" | Uncited claims are indistinguishable from hallucinations. Evidence on everything. |
|
|
172
|
+
| "The folder structure tells me the domains" | Folders are how it grew, not what it is. Derive domains from ownership + vocabulary. |
|
|
173
|
+
| "Tests are just tests, skip them" | Tests are the spec. Some rules exist ONLY in tests. Phase 2 is mandatory. |
|
|
174
|
+
| "No events are emitted, so events: []" | Reconstruct implicit events from side-effect clusters; mark `explicit: false`. |
|
|
175
|
+
| "Cron jobs aren't workflows" | System workflows are workflows. Include schedules, queue consumers, webhook reactions. |
|
|
176
|
+
| "The frontend is just the web routes" | The web UI is ONE interface. Inventory the CLI, MCP server, OpenAPI/SDK, realtime, webhooks in `interfaces.json` too — the product is driven by people, developers, AND agents. |
|
|
177
|
+
| "A component list is enough" | Without the `rebuild` split, you've hidden the frontend's real cost. Flag every `custom-logic` component (chart/table/canvas/editor) and say what the logic is — that's the port work. |
|
|
178
|
+
| "I'll skip the validator, the JSON looks right" | Run it. Missing domains refs, dangling event names, and undescribed custom-logic components are exactly what it catches. |
|
|
175
179
|
|
|
176
180
|
## Quick Reference
|
|
177
181
|
|