@pikku/skills 0.12.2 → 0.12.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (68) hide show
  1. package/dist/skills.gen.js +1 -1
  2. package/package.json +1 -1
  3. package/skills/pikku-addon/SKILL.md +56 -29
  4. package/skills/pikku-ai-agent/SKILL.md +197 -105
  5. package/skills/pikku-ai-vercel/SKILL.md +57 -18
  6. package/skills/pikku-ai-voice/SKILL.md +126 -52
  7. package/skills/pikku-audit/SKILL.md +35 -13
  8. package/skills/pikku-aws/SKILL.md +66 -16
  9. package/skills/pikku-backblaze/SKILL.md +44 -11
  10. package/skills/pikku-better-auth/SKILL.md +80 -34
  11. package/skills/pikku-cli/SKILL.md +67 -18
  12. package/skills/pikku-cli/references/complete-example.md +2 -0
  13. package/skills/pikku-concepts/SKILL.md +82 -10
  14. package/skills/pikku-concepts/references/concept-mapping.md +2 -2
  15. package/skills/pikku-config/SKILL.md +134 -52
  16. package/skills/pikku-cron/SKILL.md +13 -6
  17. package/skills/pikku-deploy-azure/SKILL.md +83 -28
  18. package/skills/pikku-deploy-cloudflare/SKILL.md +79 -37
  19. package/skills/pikku-deploy-express/SKILL.md +40 -4
  20. package/skills/pikku-deploy-fastify/SKILL.md +22 -1
  21. package/skills/pikku-deploy-lambda/SKILL.md +99 -19
  22. package/skills/pikku-deploy-nextjs/SKILL.md +49 -5
  23. package/skills/pikku-deploy-uws/SKILL.md +54 -1
  24. package/skills/pikku-deps/SKILL.md +29 -8
  25. package/skills/pikku-emails/SKILL.md +36 -5
  26. package/skills/pikku-fabric/SKILL.md +30 -5
  27. package/skills/pikku-fabric-debug/SKILL.md +5 -1
  28. package/skills/pikku-feature/SKILL.md +12 -7
  29. package/skills/pikku-gateway-slack/SKILL.md +72 -11
  30. package/skills/pikku-http/SKILL.md +18 -5
  31. package/skills/pikku-http/references/http-options.md +10 -5
  32. package/skills/pikku-i18n/SKILL.md +18 -7
  33. package/skills/pikku-info/SKILL.md +18 -8
  34. package/skills/pikku-jose/SKILL.md +35 -6
  35. package/skills/pikku-knowledge/SKILL.md +3 -3
  36. package/skills/pikku-kysely/SKILL.md +78 -15
  37. package/skills/pikku-machine-auth/SKILL.md +36 -1
  38. package/skills/pikku-mcp/SKILL.md +159 -149
  39. package/skills/pikku-middleware/SKILL.md +17 -5
  40. package/skills/pikku-mongodb/SKILL.md +10 -2
  41. package/skills/pikku-n8n-import/SKILL.md +14 -6
  42. package/skills/pikku-permissions/SKILL.md +102 -22
  43. package/skills/pikku-pino/SKILL.md +12 -4
  44. package/skills/pikku-product-second-opinion/SKILL.md +3 -3
  45. package/skills/pikku-queue/SKILL.md +45 -16
  46. package/skills/pikku-react/SKILL.md +41 -14
  47. package/skills/pikku-react-query/SKILL.md +14 -10
  48. package/skills/pikku-realtime/SKILL.md +44 -22
  49. package/skills/pikku-redis/SKILL.md +12 -3
  50. package/skills/pikku-rpc/SKILL.md +23 -12
  51. package/skills/pikku-rtl/SKILL.md +21 -17
  52. package/skills/pikku-scenario/SKILL.md +285 -50
  53. package/skills/pikku-schedule/SKILL.md +39 -6
  54. package/skills/pikku-schema-ajv/SKILL.md +24 -2
  55. package/skills/pikku-schema-cfworker/SKILL.md +22 -2
  56. package/skills/pikku-security/SKILL.md +54 -9
  57. package/skills/pikku-services/SKILL.md +49 -9
  58. package/skills/pikku-services/references/audit-wire-service.md +2 -1
  59. package/skills/pikku-software-archaeology/README.md +16 -6
  60. package/skills/pikku-software-archaeology/references/pikku-mapping.md +30 -30
  61. package/skills/pikku-template-clone/SKILL.md +10 -5
  62. package/skills/pikku-trigger/SKILL.md +50 -6
  63. package/skills/pikku-versioning/SKILL.md +46 -17
  64. package/skills/pikku-websocket/SKILL.md +72 -44
  65. package/skills/pikku-workflow/SKILL.md +35 -1
  66. package/skills/pikku-workflow/references/workflow-reference.md +13 -8
  67. package/skills/pikku-workflows-client/SKILL.md +13 -6
  68. package/skills/pikku-ws/SKILL.md +44 -8
@@ -36,14 +36,24 @@ See `pikku-concepts` for the core mental model.
36
36
 
37
37
  ### RPC Methods (on `wire.rpc`)
38
38
 
39
- Four ways to call functions via RPC:
40
-
41
39
  | Method | Purpose |
42
40
  | -------------------------------- | ----------------------------------------- |
43
41
  | `rpc.invoke(name, data)` | Internal call to any wired function |
44
42
  | `rpc.remote(name, data)` | Remote call via DeploymentService |
45
43
  | `rpc.exposed(name, data)` | Call functions marked with `expose: true` |
46
- | `rpc.startWorkflow(name, input)` | Start a workflow |
44
+ | `rpc.startWorkflow(name, input)` | Start a workflow (see `pikku-workflow`) |
45
+ | `rpc.agent.run/stream(...)` | Run an AI agent (see `pikku-ai-agent`) |
46
+ | `rpc.agent.resume/approve(...)` | Answer a tool-approval interrupt |
47
+ | `rpc.agent.interrupt(runId)` | Stop an in-flight run |
48
+
49
+ `rpc.invoke`, `rpc.remote` and `rpc.startWorkflow` are typed off the generated
50
+ RPC map, so the name and the payload are checked. `rpc.exposed` is deliberately
51
+ `(name: string, data: any) => Promise<any>` — it exists to dispatch a name that
52
+ arrived from outside, which by definition cannot be checked at compile time.
53
+ Reach for `rpc.invoke` whenever the name is known statically.
54
+
55
+ `rpc` also carries `depth` (how deep the current RPC chain is, so runaway
56
+ recursion is visible) and `global`.
47
57
 
48
58
  ### Exposed Functions
49
59
 
@@ -61,17 +71,18 @@ const greet = pikkuSessionlessFunc({
61
71
 
62
72
  ### HTTP RPC Endpoint
63
73
 
64
- Expose all `expose: true` functions over HTTP:
74
+ The `POST /rpc/:rpcName` endpoint that dispatches every `expose: true` function
75
+ is **generated, not hand-written**. Turn it on and let codegen own it:
65
76
 
66
- ```typescript
67
- wireHTTP({
68
- route: '/rpc/:rpcName',
69
- method: 'post',
70
- auth: false,
71
- func: rpcCaller,
72
- })
77
+ ```bash
78
+ pikku enable rpc # sets scaffold.rpc = true (auth required)
79
+ pikku enable rpc --noAuth # sets scaffold.rpc = { auth: false } (public)
73
80
  ```
74
81
 
82
+ This writes `rpc-public.gen.ts` with an `rpcCaller` function and its `wireHTTP`
83
+ call already wired. Do not write that wiring yourself — a hand-rolled copy
84
+ collides with the generated route on the same path.
85
+
75
86
  ## Usage Patterns
76
87
 
77
88
  ### Internal Function Composition
@@ -114,7 +125,7 @@ RPC calls go through Pikku's middleware and permission pipeline. Direct imports
114
125
  After `npx pikku all`:
115
126
 
116
127
  ```typescript
117
- import { pikkuRPC } from '.pikku/pikku-rpc.gen.js'
128
+ import { pikkuRPC } from '#pikku/pikku-rpc.gen.js'
118
129
 
119
130
  pikkuRPC.setServerUrl('http://localhost:4002')
120
131
 
@@ -6,10 +6,11 @@ installGroups: [core]
6
6
 
7
7
  # Pikku RTL (Arabic + English)
8
8
 
9
- This skill sits **on top of** `pikku-i18n`. That skill maps a locale to `t()`
10
- tokens; this one adds the second axis: a locale also has a **direction**.
11
- Arabic is not special-cased — it is just another locale file (`ar.json`,
12
- registered `satisfies typeof en`) plus the document being told it is `rtl`.
9
+ This skill sits **on top of** `pikku-i18n`. That skill compiles a locale's
10
+ messages into typed `m.*()` functions; this one adds the second axis: a locale
11
+ also has a **direction**. Arabic is not special-cased — it is just another
12
+ `messages/ar.json` listed in `project.inlang/settings.json`, plus the document
13
+ being told it is `rtl`.
13
14
 
14
15
  ## The one idea
15
16
 
@@ -21,9 +22,9 @@ things right and Arabic, Hebrew, Farsi and Urdu all work with zero per-component
21
22
 
22
23
  ## Agent Operating Procedure
23
24
 
24
- 1. **Tokens first.** Every visible string is already a `t()` token via
25
- `pikku-i18n`. Arabic copy goes in `i18n/ar.json`, mirroring `en.json`'s keys,
26
- registered with `satisfies typeof en` so a missing key is a compile error.
25
+ 1. **Messages first.** Every visible string is already an `m.*()` message via
26
+ `pikku-i18n`. Arabic copy goes in `messages/ar.json`, mirroring `en.json`'s
27
+ keys with the `{param}` names kept identical.
27
28
  2. **Add the direction helper** to the i18n config (one home for locale→dir):
28
29
  ```ts
29
30
  const RTL_LOCALES = new Set(['ar', 'he', 'fa', 'ur'])
@@ -82,7 +83,7 @@ matching `dir` on `<html>`.
82
83
 
83
84
  ```tsx
84
85
  import { DirectionProvider, MantineProvider } from '@mantine/core'
85
- import i18n, { detectLocale, localeDir } from './i18n/config'
86
+ import { detectLocale, localeDir } from './i18n/config'
86
87
 
87
88
  const locale =
88
89
  typeof window !== 'undefined' ? detectLocale(window.location.pathname) : 'en'
@@ -137,8 +138,9 @@ const html = `<!doctype html>
137
138
  </html>`
138
139
  ```
139
140
 
140
- i18next's active language must match: call `i18n.changeLanguage(locale)` before
141
- `renderToString` so the SSR'd text and `dir` agree.
141
+ Paraglide's active locale must match: set it (via the i18n config's
142
+ `setActiveLocale` / `overwriteGetLocale` bridge) before `renderToString`, so the
143
+ SSR'd text and `dir` agree.
142
144
 
143
145
  ### Next.js app-router (test-harness next-ssr / next-static)
144
146
 
@@ -188,16 +190,17 @@ Prefer logical icon components if your icon set ships them.
188
190
  which is inconsistent. Add an Arabic-capable family (e.g. _Noto Sans Arabic_,
189
191
  _IBM Plex Sans Arabic_) to `font-family` so both scripts look intentional.
190
192
  - **Numerals:** don't hardcode digits. Format numbers/dates with
191
- `Intl.NumberFormat`/`Intl.DateTimeFormat` (or i18next formatters) given the
192
- active locale, so Western vs Arabic-Indic digits follow the locale choice.
193
+ `Intl.NumberFormat`/`Intl.DateTimeFormat` given the active locale, so Western
194
+ vs Arabic-Indic digits follow the locale choice.
193
195
  - **Line height:** Arabic diacritics sit tall — a slightly larger `line-height`
194
196
  on Arabic body text avoids clipping. Keep it locale-scoped, not global.
195
197
 
196
198
  ## Adding Arabic to an existing app — checklist
197
199
 
198
- 1. `i18n/ar.json` mirroring `en.json`; register
199
- `ar: { translation: ar satisfies typeof en }` and add `'ar'` to
200
- `supportedLocales`. (Type-complete or it won't compile — the deploy blocks.)
200
+ 1. `messages/ar.json` mirroring `en.json`; add `"ar"` to `locales` in
201
+ `project.inlang/settings.json` and recompile. Keys missing from `ar.json`
202
+ fall back to the base locale per message rather than failing the build, so
203
+ diff the two files rather than trusting `tsc` to catch a gap here.
201
204
  2. Confirm the `localeDir` helper includes `ar` (it does by default).
202
205
  3. Confirm the root sets `dir` from the locale (recipe above).
203
206
  4. Sweep the app's styles: replace every `left/right`, `ml/mr`, `text-align:
@@ -215,5 +218,6 @@ left` with the flow-relative equivalent; revert any manual `row-reverse`.
215
218
  per-locale `if (rtl)` layout branches. Set `dir` once; let layout follow.
216
219
  - Don't set `dir` on individual components — it belongs on `<html>` so the whole
217
220
  document (and Mantine) agrees.
218
- - Don't translate Arabic copy outside the `t()` token system; an RTL language is
219
- a normal locale, governed by `pikku-i18n`.
221
+ - Don't translate Arabic copy outside the message system; an RTL language is a
222
+ normal locale, governed by `pikku-i18n`. There is no `t()` and no i18next in a
223
+ Pikku frontend — the string comes from `m.some__key()`.
@@ -6,7 +6,8 @@ description: >-
6
6
  over the real transport against a running server — so a flow doubles as an e2e test and a
7
7
  staged/production health check. Covers scenario.do / expectEventually / expectError /
8
8
  expectService, declared steps via pikkuScenarioStep (including browser steps driven by
9
- @pikku/playwright), actors and environments in pikku.config.json, SCENARIO_ACTOR_SECRET, the
9
+ @pikku/playwright) written as intent rather than as clicks, with the actions factored into
10
+ shared browser utilities, actors and environments in pikku.config.json, SCENARIO_ACTOR_SECRET, the
10
11
  `pikku scenario list|run` commands, live function coverage via `pikku dev --coverage`, and
11
12
  plain unit tests for pure function logic. TRIGGER when: user asks about scenarios, testing a
12
13
  Pikku function, test coverage, end-to-end flows, browser/UI e2e, or health checks. DO NOT
@@ -35,7 +36,7 @@ A scenario is a `pikkuScenario` export that drives the app **as real actors over
35
36
  Consequences that matter, and bite if ignored:
36
37
 
37
38
  - **There is no state reset.** A scenario runs against a live server. Scope what you create (unique ids, your own rows) and never assume a clean database.
38
- - **Every effect runs as somebody, or as a declared step.** `scenario.do(...)` without `{ actor }` throws `Scenario tried to run '<rpc>' as an internal step…` — there is no bare internal-RPC step. The other way to do work is `scenario.step/given/when/then`, which runs a `pikkuScenarioStep`; its actor is optional (setup steps have none) unless it declares `browser: true`.
39
+ - **Every effect runs as somebody, or as a declared step.** `scenario.do(...)` without `{ actor }` throws `Scenario tried to run '<rpc>' as an internal step…` — there is no bare internal-RPC step. The other way to do work is `scenario.given/when/then`, which runs a `pikkuScenarioStep`; its actor is optional (setup steps have none) unless it declares `browser: true`.
39
40
  - **Actors must be configured and signed in**, or the scenario cannot run.
40
41
 
41
42
  Scenarios live in `srcDirectories` like any other function — by convention `*.scenario.ts`.
@@ -84,13 +85,14 @@ A scenario takes the same config fields as a workflow (`title`, `description`, `
84
85
 
85
86
  ### The scenario API
86
87
 
87
- | Call | Purpose |
88
- | ------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- |
89
- | `scenario.do(step, rpc, data, { actor })` | Run an RPC as that actor. The step name is what appears in the run output. |
90
- | `scenario.expectEventually(step, rpc, data, predicate, { actor, within, interval })` | Poll until `predicate(out)` passes or `within` elapses. For anything asynchronous — queues, workers, eventual state. |
91
- | `scenario.expectError(step, rpc, data, { actor, matches })` | Assert the call **fails**. For fault injection and negative paths. |
92
- | `scenario.expectService(step, 'service.method', { actor, calledWith })` | Assert a stubbed service was called. Requires the server to run with `--test`. |
93
- | `scenario.step(step, stepName, data, { actor })` | Run a declared `pikkuScenarioStep`. `given`/`when`/`then` are the same call with a keyword in the rendered prose. |
88
+ | Call | Purpose |
89
+ | ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- |
90
+ | `scenario.do(step, rpc, data, { actor })` | Run an RPC as that actor. The step name is what appears in the run output. |
91
+ | `scenario.expectEventually(step, rpc, data, predicate, { actor, within, interval })` | Poll until `predicate(out)` passes or `within` elapses. For anything asynchronous — queues, workers, eventual state. |
92
+ | `scenario.expectError(step, rpc, data, { actor, matches })` | Assert the call **fails**. For fault injection and negative paths. |
93
+ | `scenario.expectService(step, 'service.method', { actor, calledWith })` | Assert a stubbed service was called. Requires the server to run with `--test`. |
94
+ | `scenario.given(stepName, step, data, { actor })` | Run a declared `pikkuScenarioStep` as setup. `when` is the same call; `then` also makes the step's bindings witnesses. |
95
+ | `scenario.runScheduledTask(name)` | Fire a wired scheduler on the target now, rather than waiting for its cron. |
94
96
 
95
97
  `expectEventually` is **scenario-only**. Calling it from a `pikkuWorkflowFunc` is a critical inspector error (`PKU675`) pointing you at `pikkuScenario`.
96
98
 
@@ -110,7 +112,9 @@ export const credentialScenario = pikkuScenario({
110
112
  tags: ['scenario', 'credential'],
111
113
  before: resetsCredentials,
112
114
  after: removesInstalledAddon,
113
- func: async (services, data, { scenario, actors }) => { /* … */ },
115
+ func: async (services, data, { scenario, actors }) => {
116
+ /* … */
117
+ },
114
118
  })
115
119
  ```
116
120
 
@@ -123,7 +127,7 @@ export const credentialScenario = pikkuScenario({
123
127
  | Neither runs when the run is suspended or waiting — teardown only fires at a terminal outcome. |
124
128
  | Hooks are **not** ladder rows. The runner records nothing for them; a failure is labelled by phase. |
125
129
 
126
- A hook reaches the app the same way the body does: through `wire.actors`. If you want cleanup to be *visible* on the ladder, make it an ordinary `scenario.then(...)` instead.
130
+ A hook reaches the app the same way the body does: through `wire.actors`. If you want cleanup to be _visible_ on the ladder, make it an ordinary `scenario.then(...)` instead.
127
131
 
128
132
  Hooks are scenario-only. A `before`/`after` on a `pikkuWorkflowFunc` never runs — a workflow is durable and resumable, so a callback that reran on every replay would have no honest meaning.
129
133
 
@@ -154,18 +158,212 @@ export const credentialFeature = pikkuFeature({
154
158
  })
155
159
  ```
156
160
 
157
- | Rule |
158
- | ---------------------------------------------------------------------------------------------------------------------------------------- |
159
- | The **export identifier is the feature's id**; `name` is the human-readable label. Both must be exported or the build fails. |
160
- | A `{ scenario, data }` entry is gherkin's `Examples:` — one run per entry. `data` is typed against that scenario's input. |
161
- | Feature hooks run **once around the whole group** (`before → a → b → c → after`), _not_ per scenario. `after` runs in a `finally`. |
162
- | There is deliberately **no `Background:`**. Per-scenario setup is the scenario's own `before`, referencing a shared function. |
163
- | A scenario's effective tags are its own **plus** the feature's, so `--tags credential` selects through the feature. |
164
- | A scenario need not belong to a feature — one with no input still runs standalone. |
161
+ | Rule |
162
+ | --------------------------------------------------------------------------------------------------------------------------------------------- |
163
+ | The **export identifier is the feature's id**; `name` is the human-readable label. Both must be exported or the build fails. |
164
+ | A `{ scenario, data }` entry is gherkin's `Examples:` — one run per entry. `data` is typed against that scenario's input. |
165
+ | Feature hooks run **once around the whole group** (`before → a → b → c → after`), _not_ per scenario. `after` runs in a `finally`. |
166
+ | There is deliberately **no `Background:`**. Per-scenario setup is the scenario's own `before`, referencing a shared function. |
167
+ | A scenario's effective tags are its own **plus** the feature's, so `--tags credential` selects through the feature. |
168
+ | A scenario need not belong to a feature — one with no input still runs standalone. |
165
169
  | Membership is resolved by **object identity** at runtime, which is why a loop works and why a scenario built inline in a feature is an error. |
166
170
 
167
171
  The **feature is the run unit**: `--flows` on a scenario whose every feature entry carries `data` errors and names the features containing it, because the feature is what supplies that data. Use `--features` for those. A scenario referenced bare anywhere, or in no feature at all, still runs standalone.
168
172
 
173
+ ### Steps describe intent, not actions
174
+
175
+ A scenario records what someone was **trying to do**, never the keystrokes they used to do it. This is the one decision that determines whether a suite survives its first redesign, and it applies to every step name you write.
176
+
177
+ | Action ladder — wrong | Intent ladder — right |
178
+ | ----------------------------------- | --------------------------------------------------- |
179
+ | `Given opens /shop` | `Given the shopper is browsing the shop` |
180
+ | `When clicks the category filter` | `When the shopper buys the £5 strawberry milkshake` |
181
+ | `And clicks "Drinks"` | `Then it is in their basket` |
182
+ | `And clicks the first product card` | |
183
+ | `And clicks Add to basket` | |
184
+ | `Then sees "1 item"` | |
185
+
186
+ Three things go wrong with the left-hand column, and all three are expensive:
187
+
188
+ - **A layout change rewrites every scenario that touched that screen.** In the right-hand column it rewrites one function.
189
+ - **The report is the deliverable.** `buys the £5 strawberry milkshake` is readable by someone who has never seen the app; `clicks [data-testid=add]` tells them nothing about whether the product works.
190
+ - **An action step cannot arrive on its own.** It assumes the previous click left the browser somewhere, so the scenario only runs front-to-back, as a whole, in one order.
191
+
192
+ So there are three layers, and only two of them are named in the report:
193
+
194
+ | Layer | What it is | On the ladder |
195
+ | -------------------------- | -------------------------------------- | ---------------- |
196
+ | Scenario | The flow, written as intents | yes — the ladder |
197
+ | Step (`pikkuScenarioStep`) | One intent | yes — one row |
198
+ | Utility | An ordinary TS function over `browser` | no |
199
+
200
+ Utilities are **not steps**. They are plain exported functions, they take the browser handle, and they hold the clicking:
201
+
202
+ ```typescript
203
+ // shop.browser.ts — shared actions. Not steps: nothing here is an intent.
204
+ import type { PikkuBrowserWire } from '@pikku/core/workflow'
205
+ import type {} from '@pikku/playwright'
206
+
207
+ /** Arrive on the shop, from wherever the browser happens to be. */
208
+ export const ensureOnShop = async (browser: PikkuBrowserWire) => {
209
+ if (!new URL(browser.page.url()).pathname.startsWith('/shop')) {
210
+ await browser.goto('/shop')
211
+ }
212
+ await browser
213
+ .locate({ testId: 'product-grid' })
214
+ .first()
215
+ .waitFor({ state: 'visible' })
216
+ }
217
+
218
+ export const searchFor = async (browser: PikkuBrowserWire, query: string) => {
219
+ await browser.locate({ testId: 'shop-search' }).first().fill(query)
220
+ await browser.page.keyboard.press('Enter')
221
+ }
222
+
223
+ export const filterByCategory = async (
224
+ browser: PikkuBrowserWire,
225
+ category: string
226
+ ) => {
227
+ await browser.locate({ testId: 'category-filter' }).first().click()
228
+ await browser
229
+ .locate({ testId: 'category-option', where: { 'data-category': category } })
230
+ .first()
231
+ .click()
232
+ }
233
+
234
+ export const addToBasket = async (browser: PikkuBrowserWire, name: string) => {
235
+ const card = browser
236
+ .locate({ testId: 'product-card', containing: name })
237
+ .first()
238
+ await card.waitFor({ state: 'visible' })
239
+ await card.locate('[data-testid=add-to-basket]').click()
240
+ }
241
+ ```
242
+
243
+ The step composes them, and it is the step — one row — that the report shows:
244
+
245
+ ```typescript
246
+ export const buysTheItem = pikkuScenarioStep<
247
+ { name: string },
248
+ { name: string }
249
+ >({
250
+ name: 'buysTheItem',
251
+ description: 'finds one item in the shop and puts it in the basket',
252
+ template: 'buys the {name}',
253
+ // One intent, one implementation per surface an actor can drive it through.
254
+ browser: async (_services, { name }, { browser }) => {
255
+ await ensureOnShop(browser)
256
+ await searchFor(browser, name)
257
+ await addToBasket(browser, name)
258
+ return { name }
259
+ },
260
+ default: async ({ rpc }, { name }) => {
261
+ const item = await rpc.invoke('findItemByName', { name })
262
+ await rpc.invoke('addToBasket', { itemId: item.id })
263
+ return { name }
264
+ },
265
+ })
266
+ ```
267
+
268
+ The bindings are **alternatives**: `pikku scenario run --run browser` clicks through the shop, `--run cli` drives it over the websocket, `--run default` (the fast suite, and the default) takes the server-side path — and all of them report the same sentence.
269
+
270
+ ```typescript
271
+ await scenario.when(
272
+ 'buys a milkshake',
273
+ 'buysTheItem',
274
+ { name: '£5 strawberry milkshake' },
275
+ { actor: actors.shopper }
276
+ )
277
+ // reporter renders: When the shopper buys the £5 strawberry milkshake ✓ 1.2s
278
+ ```
279
+
280
+ **Every intent step begins by arriving.** `ensureOnShop` is not defensive noise — it is what lets a scenario start at any step, run alone, and be reordered without touching it. It checks first and navigates only if needed, so a scenario already on the shop pays nothing. This is about the _browser's_ starting position, not the database: there is still no state reset (see above), and you still scope what you create.
281
+
282
+ **The same utilities, a different intent.** A scenario about filtering has filtering as its subject, so there the filter _is_ the intent — same helper, its own step:
283
+
284
+ ```typescript
285
+ export const filtersTheShop = pikkuScenarioStep<
286
+ { category: string },
287
+ { shown: number }
288
+ >({
289
+ name: 'filtersTheShop',
290
+ description: 'narrows the catalogue to one category',
291
+ template: 'filters the shop by {category}',
292
+ browser: async (_services, { category }, { browser }) => {
293
+ await ensureOnShop(browser)
294
+ await filterByCategory(browser, category)
295
+ return {
296
+ shown: await browser.locate({ testId: 'product-card' }).count(),
297
+ }
298
+ },
299
+ default: async ({ rpc }, { category }) => ({
300
+ shown: (await rpc.invoke('listItems', { categorySlug: category })).length,
301
+ }),
302
+ })
303
+ ```
304
+
305
+ Two scenarios, two intents, one set of utilities. That is the shape to aim for: when a helper is reused by a step whose _subject_ it is, promote it to a step there — never the reverse.
306
+
307
+ **Non-browser steps need none of this.** Without a browser there is no navigation to absorb and no DOM to hide, so an intent maps to one RPC and `scenario.do` names it directly:
308
+
309
+ ```typescript
310
+ const order = await scenario.do(
311
+ 'Shopper checks out',
312
+ 'createOrder',
313
+ { basketId, shippingAddress },
314
+ { actor: actors.shopper }
315
+ )
316
+ ```
317
+
318
+ Reach for a `pikkuScenarioStep` on the non-browser side only when one intent genuinely spans several RPCs, or when the step asserts something the RPC result alone does not say.
319
+
320
+ ### `then` bindings are witnesses, not alternatives
321
+
322
+ This is the one place the surface bindings do **not** behave like a switch, and it is the part worth reading twice.
323
+
324
+ On a `given` or `when`, the bindings are alternatives — clicking Buy and calling `createOrder` are two ways to cause one effect, so exactly one runs.
325
+
326
+ On a `then`, they are not two implementations of one assertion. They are two _different claims_:
327
+
328
+ | binding | what it actually proves |
329
+ | --------- | -------------------------------------------------------------- |
330
+ | `default` | the order row says `paid` — the system of record is right |
331
+ | `browser` | the confirmation panel says paid — the truth reached the human |
332
+
333
+ The gap between them is the bug nobody catches: 200 OK, database correct, user still watching a spinner. So a `then` runs **every** binding it declares and fails if they disagree.
334
+
335
+ ```typescript
336
+ export const seesTheOrderConfirmed = pikkuScenarioStep<
337
+ { orderId: string },
338
+ { status: string }
339
+ >({
340
+ name: 'seesTheOrderConfirmed',
341
+ template: 'sees order {orderId} confirmed',
342
+ // Both run on `--run browser`. Each returns what it observed, and the runner
343
+ // compares them — so this fails when the page disagrees with the database.
344
+ browser: async (_services, { orderId }, { browser }) => ({
345
+ status: await browser
346
+ .locate({ testId: 'order-status', where: { 'data-order': orderId } })
347
+ .getAttribute('data-status'),
348
+ }),
349
+ default: async ({ rpc }, { orderId }) => ({
350
+ status: (await rpc.invoke('getOrder', { orderId })).status,
351
+ }),
352
+ })
353
+ ```
354
+
355
+ Three rules follow, and they are the ones that get broken:
356
+
357
+ - **A browser witness must observe on the page.** One that quietly calls an RPC to check the result is worse than no binding at all — it reports a tick for a surface it never looked at.
358
+ - **Return what you observed, don't just assert.** A witness returning a value lets the runner diff the two. A witness that only throws still works, but it can never disagree with anything, so it proves less. Read structured state with `where` on the test-id selector rather than parsing translated copy.
359
+ - **A step with no binding for the run's surface is counted, not excused.** `--run browser` prints `n/m steps ran on browser` over _every_ step, so an action that quietly fell back to the server lowers the number just as an assertion does. A `then` that fell back is additionally named — `--strict` fails on those, because a sentence saying the actor saw something nobody looked at is a different problem from a shortcut. Not being in the UI _is_ the finding: do not add a browser binding that fakes it.
360
+
361
+ **Always give a `then` a `default` witness.** It is the floor every run can fall back to, and an assertion with no witness the run can execute is fatal (`ScenarioNoWitness`) — not a coverage gap. The distinction is the point: a `then` checked server-side under `--run browser` did happen, it just wasn't seen where the prose claims; one checked nowhere never happened at all, and without the error it would return `undefined` and render as a tick. A browser-only `then` is therefore a step that fails the fast suite, which is rarely what you want.
362
+
363
+ **Every scenario must assert.** A flow of only `given`/`when` is a PKU680 critical — it proves nothing threw. Since coverage counts every step, an assertion-free ladder of browser-bound actions would score a perfect `3/3` while checking nothing, so clicking through the UI and never looking at the result is the cheapest way to fake the number. The rule closes that.
364
+
365
+ Assertions with no possible browser witness are a different thing and should not be written as a `then`: "the audit log recorded it" is a system check, and "the receipt email arrives" is `expectEventually`, which is always out-of-band and always server-side.
366
+
169
367
  ### Declared steps (`pikkuScenarioStep`)
170
368
 
171
369
  `scenario.do` can only name an RPC. A **step** is a named, typed unit of scenario behaviour whose body is an ordinary pikku function — so it can call several RPCs as its actor, assert, or drive a browser.
@@ -181,12 +379,16 @@ export const buysAnApple = pikkuScenarioStep<
181
379
  name: 'buysAnApple',
182
380
  description: 'buys an apple',
183
381
  template: 'buys {qty} apples',
184
- func: async (_services, { qty }, { scenarioStep }) => {
382
+ default: async (_services, { qty }, { scenarioStep }) => {
185
383
  return await requireActor(scenarioStep).invoke('placeOrder', { qty })
186
384
  },
187
385
  })
188
386
  ```
189
387
 
388
+ A step's body always lives under a **surface binding** — `default`, `browser` or
389
+ `cli` — never under a `func`. Declaring none throws at load time: at minimum give
390
+ it a `default`.
391
+
190
392
  ```typescript
191
393
  await scenario.given(
192
394
  'buys an apple',
@@ -201,7 +403,7 @@ Rules that bite:
201
403
 
202
404
  - **The step is referenced by its typed string name, not by importing the const** — exactly like `workflow.do`. The name is the step's `pikkuFuncId` and is checked against the generated step map. A non-literal target is a critical error (`PKU678`).
203
405
  - **Steps are not RPCs.** They are deliberately never network-callable — a browser-driving step must not be.
204
- - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal *is* the assertion.
406
+ - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal _is_ the assertion.
205
407
  - **`actor` and `env` are optional on the wire**, because a pure assertion step needs neither. Narrow them with `requireActor(scenarioStep)` and `requireScenarioEnv(scenarioStep)` from `@pikku/core/workflow` rather than a local guard — both name the step and say what to pass. `env` is `{ apiUrl, appUrl? }` from the environment the run targets, and is how a raw-HTTP step learns the target's URL: a step runs in the CLI process, where there is no `variables` service and `process.env` is not the answer.
206
408
  - **Steps default to `retries: 0`**, unlike ordinary workflow steps. Retrying a failed assertion is wrong; pass `retries` explicitly if a step is genuinely flaky-by-nature.
207
409
  - **Step results are persisted**, so return JSON-serialisable data — never a `Locator` or a client object.
@@ -210,22 +412,24 @@ Rules that bite:
210
412
 
211
413
  ### Browser steps
212
414
 
213
- A step declaring `browser: true` gets `wire.browser` — a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
415
+ Declaring a `browser` binding is the whole switch: inside that binding `wire.browser` is guaranteed present and non-optional, and a step without one never sees a browser at all. There is nothing to null-check.
416
+
417
+ A `browser` binding gets a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
418
+
419
+ Browser steps are where **intent, not actions** earns its keep: the step is one intent, the clicking lives in shared utilities, and the step arrives before it acts. Write the mechanics below into utilities and keep the step body to three or four calls that read as a sentence.
214
420
 
215
421
  ```typescript
216
- export const opensTheCart = pikkuScenarioStep<
217
- { path: string },
218
- { url: string },
219
- true
220
- >({
221
- name: 'opensTheCart',
222
- description: 'opens the cart',
223
- browser: true,
224
- func: async (_services, { path }, { browser }) => {
225
- await browser.goto(path)
226
- return { url: browser.page.url() }
227
- },
228
- })
422
+ export const opensTheCart = pikkuScenarioStep<{ path: string }, { url: string }>(
423
+ {
424
+ name: 'opensTheCart',
425
+ description: 'opens the cart',
426
+ browser: async (_services, { path }, { browser }) => {
427
+ await browser.goto(path)
428
+ return { url: browser.page.url() }
429
+ },
430
+ default: async ({ rpc }) => ({ url: (await rpc.invoke('getCart', {})).url }),
431
+ }
432
+ )
229
433
  ```
230
434
 
231
435
  - Install `@pikku/playwright` and `@playwright/test`, and import `@pikku/playwright` once (`import type {} from '@pikku/playwright'`) so `browser.page` is a typed Playwright `Page`. Without it you still get the structural `goto`/`screenshot` handle.
@@ -242,8 +446,14 @@ Personas, actors and environments live in `pikku.config.json`:
242
446
  "scenarios": {
243
447
  "personas": {
244
448
  "shopper": { "description": "Buys things here", "primary": true },
245
- "support": { "description": "Answers for the shop", "proficiency": "power" },
246
- "reminders": { "description": "The shop chasing abandoned carts", "kind": "system" }
449
+ "support": {
450
+ "description": "Answers for the shop",
451
+ "proficiency": "power"
452
+ },
453
+ "reminders": {
454
+ "description": "The shop chasing abandoned carts",
455
+ "kind": "system"
456
+ }
247
457
  },
248
458
  "actors": {
249
459
  "shopper": {
@@ -288,9 +498,24 @@ SCENARIO_ACTOR_SECRET=… pikku scenario run local
288
498
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --flows orderSupportScenario
289
499
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --features credentialFeature
290
500
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --tags smoke,scenario
501
+ SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --no-browser --exclude-tags ai-live
291
502
  ```
292
503
 
293
- `run` takes the environment as a **required positional** — the key from `scenarios.environments`. `--flows`/`-f` filters by scenario name, `--features` by feature id, `--tags`/`-t` by tag (match-any). Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
504
+ `run` takes the environment as a **required positional** — the key from `environments`. Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
505
+
506
+ | Flag | Effect |
507
+ | ----------------------- | --------------------------------------------------------------------------------------- |
508
+ | `--flows` / `-f` | Comma-separated scenario names |
509
+ | `--features` | Comma-separated feature ids |
510
+ | `--tags` / `-t` | Match-any tag filter |
511
+ | `--exclude-tags` | Hold tags back — unless the flow is named directly with `--flows` |
512
+ | `--run <surface>` | `default` (the default), `browser`, or `cli` |
513
+ | `--no-browser` | Shorthand for `--run default`; scenarios with browser steps report as **skipped** |
514
+ | `--strict` | Fail, rather than pass, a `then` with no witness on the run's surface |
515
+ | `--spawn` / `--keep-alive` | Start `pikku dev` on the environment's apiUrl for the run; optionally leave it up |
516
+ | `--api-url` / `--app-url` | Override the environment's URLs — for a target that only exists at run time |
517
+ | `--trace` | Keep every stack frame on failure (default shows only the project's own) |
518
+ | `--coverage` | Reset/snapshot server coverage per scenario |
294
519
 
295
520
  Output is `PASS <name> (<ms>) → <output>` / `FAIL <name> (<ms>): <error>`, then `N/M scenarios passed against '<env>'`. A scenario inside a feature is named `<Feature> › <scenario> <data>`.
296
521
 
@@ -302,10 +527,16 @@ Coverage is attributed by running scenarios against a server that is collecting
302
527
 
303
528
  Prerequisite in `pikku.config.json`:
304
529
 
305
- ```json
306
- { "scaffold": { "scenarios": "auth" } }
530
+ ```bash
531
+ pikku enable scenarios # sets scaffold.scenarios = true (session required)
532
+ pikku enable scenarios --noAuth # sets scaffold.scenarios = { "auth": false }
307
533
  ```
308
534
 
535
+ `scaffold.scenarios` is a boolean or `{ auth?, path? }`. The legacy string forms
536
+ (`"auth"` / `"no-auth"`) are **rejected by the config loader**, not reinterpreted —
537
+ under a shape where a string could be a path, silently reading one as a flag
538
+ would be worse than failing.
539
+
309
540
  `scaffold.scenarios` generates the coverage and stub RPCs into your project (`pikkuScenarioTakeLiveCoverage`, `pikkuScenarioResetLiveCoverage`, `pikkuScenarioResetStubs`, `pikkuScenarioGetStubCalls`), so scenario runs work against any server. The coverage RPC reads `<outDir>/function/pikku-functions-meta-verbose.gen.json` off disk at request time — codegen always writes it, but it has to be deployed alongside the app or the RPC returns `null`.
310
541
 
311
542
  ```bash
@@ -366,16 +597,20 @@ Services are plain objects — a Pikku function is pure business logic, so a moc
366
597
 
367
598
  ## Red flags
368
599
 
369
- | Smell | Why it's wrong |
370
- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
371
- | `pikku tests …` | Removed in #865. Use `pikku scenario`. |
372
- | `.feature` files / Gherkin for function tests | Scenarios are TypeScript, not Gherkin. The in-process cucumber function world was deleted. |
373
- | `scenario.do(...)` with no `{ actor }` | Throws. Every step runs as somebody. |
374
- | A scenario per function | Scenarios are user flows. One flow covers many functions; that is the point. |
375
- | Assuming a clean database | There is no state reset — it may be a staging server. Scope what you create. |
376
- | `sleep()` before asserting | Use `expectEventually`. |
377
- | `expectEventually` in a `pikkuWorkflowFunc` | `PKU675` — scenario-only. |
378
- | Coverage silently 0 | Server not run with `--coverage`, verbose functions meta not deployed, `scaffold.scenarios` unset, or no actors configured. |
600
+ | Smell | Why it's wrong |
601
+ | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
602
+ | `pikku tests …` | Removed in #865. Use `pikku scenario`. |
603
+ | `.feature` files / Gherkin for function tests | Scenarios are TypeScript, not Gherkin. The in-process cucumber function world was deleted. |
604
+ | `scenario.do(...)` with no `{ actor }` | Throws. Every step runs as somebody. |
605
+ | A scenario per function | Scenarios are user flows. One flow covers many functions; that is the point. |
606
+ | Assuming a clean database | There is no state reset — it may be a staging server. Scope what you create. |
607
+ | `sleep()` before asserting | Use `expectEventually`. |
608
+ | A step named `clicksAddToBasket` / `opensThePage` | That is an action, not an intent. Name the step for what the actor wanted; put the clicking in a utility. |
609
+ | A browser step that assumes it is already on a page | It can then only run mid-flow. Arrive first — check the URL, navigate if needed. |
610
+ | A `browser` binding guarding `if (!browser)` | The binding guarantees it. The guard hides the real error, which is a missing actor (`PKU677`). |
611
+ | A step with a `func:` instead of a surface binding | There is no `func` on a step. Bodies live under `default` / `browser` / `cli`; a step with none throws at load. |
612
+ | `expectEventually` in a `pikkuWorkflowFunc` | `PKU675` — scenario-only. |
613
+ | Coverage silently 0 | Server not run with `--coverage`, verbose functions meta not deployed, `scaffold.scenarios` unset, or no actors configured. |
379
614
 
380
615
  `@pikku/cucumber` is a **browser/e2e** harness (`Actor`, `BrowserWorld`, `PersonaData`, `DbUtils`) — out of scope here.
381
616