@pikku/skills 0.12.35 → 0.12.38

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/README.md +9 -4
  2. package/dist/index.d.ts +7 -4
  3. package/dist/index.js +9 -5
  4. package/dist/skills.gen.d.ts +1 -0
  5. package/dist/skills.gen.js +5 -3
  6. package/dist/snippets.d.ts +26 -0
  7. package/dist/snippets.js +148 -0
  8. package/package.json +2 -2
  9. package/skills/pikku-addon/SKILL.md +70 -24
  10. package/skills/pikku-addon/references/addon-package-manifest.md +9 -4
  11. package/skills/pikku-addon/references/openapi.md +130 -0
  12. package/skills/pikku-agent/references/agents.md +3 -1
  13. package/skills/pikku-auth/references/better-auth.md +33 -2
  14. package/skills/pikku-build/SKILL.md +94 -6
  15. package/skills/pikku-build/references/app.md +82 -12
  16. package/skills/pikku-build/references/design.md +16 -5
  17. package/skills/pikku-build/references/feature.md +23 -96
  18. package/skills/pikku-build/references/openapi.md +119 -0
  19. package/skills/pikku-build/references/platform.md +4 -0
  20. package/skills/pikku-build/references/quick.md +16 -6
  21. package/skills/pikku-changes/SKILL.md +172 -0
  22. package/skills/pikku-concepts/SKILL.md +33 -138
  23. package/skills/pikku-concepts/references/bootstrap.md +58 -0
  24. package/skills/pikku-concepts/references/concept-mapping.md +16 -0
  25. package/skills/pikku-concepts/references/language.md +87 -0
  26. package/skills/pikku-deploy/SKILL.md +1 -1
  27. package/skills/pikku-fabric/SKILL.md +13 -13
  28. package/skills/pikku-guide/SKILL.md +264 -0
  29. package/skills/pikku-kysely/SKILL.md +1 -1
  30. package/skills/pikku-n8n-import/SKILL.md +4 -3
  31. package/skills/pikku-react/references/client.md +12 -0
  32. package/skills/pikku-realtime/SKILL.md +6 -6
  33. package/skills/pikku-report/SKILL.md +143 -0
  34. package/skills/pikku-scenario/SKILL.md +71 -562
  35. package/skills/pikku-scenario/references/browser.md +59 -0
  36. package/skills/pikku-scenario/references/coverage.md +70 -0
  37. package/skills/pikku-scenario/references/personas.md +149 -0
  38. package/skills/pikku-scenario/references/steps.md +366 -0
  39. package/skills/pikku-service-backends/SKILL.md +1 -1
  40. package/skills/pikku-wiring/SKILL.md +1 -1
  41. package/skills/pikku-workflow/SKILL.md +7 -8
@@ -32,10 +32,16 @@ Use this skill as an execution checklist, not reference material.
32
32
 
33
33
  ## Pick the reference
34
34
 
35
- | You are… | Read |
36
- | ---------------------------------------------------------------- | --------------------------- |
37
- | Writing or running scenarios | this skill |
38
- | Running a persona as a model-driven virtual user against a stage | `references/persona-run.md` |
35
+ This skill covers writing and running scenarios end to end. Four topics are one level down, and
36
+ each says when to open it:
37
+
38
+ | Read | For |
39
+ | --------------------------- | ------------------------------------------------------------------------------------------------------------ |
40
+ | `references/steps.md` | Authoring a `pikkuScenarioStep` — intent, witnesses, what a step is given |
41
+ | `references/browser.md` | Browser bindings, and locating by message key in a translated app |
42
+ | `references/personas.md` | Personas vs actors, `definePersonas`, upstream credentials for actors, and the human "Sign in as …" switcher |
43
+ | `references/coverage.md` | Live coverage, filling it, and unit tests for pure logic |
44
+ | `references/persona-run.md` | Running a persona as a model-driven virtual user against a stage |
39
45
 
40
46
  ## What a scenario is
41
47
 
@@ -180,6 +186,12 @@ Hooks are scenario-only. A `before`/`after` on a `pikkuWorkflowFunc` never runs
180
186
 
181
187
  ### Grouping scenarios (`pikkuFeature`)
182
188
 
189
+ A feature is also cited by a page of the user guide: `pikku scenario guide`
190
+ fills each cited block with the feature's recordings and screenshots, and
191
+ nothing else (see **pikku-guide**). A scenario's `title` captions its recording
192
+ and a screenshot's `name` captions the still, so write both as copy a user would
193
+ read. Mark plumbing features `document: false`.
194
+
183
195
  `pikkuFeature` groups scenarios the way gherkin's `Feature:` groups `Scenario:`. Scenarios are referenced by **imported identifier**, so a renamed or deleted scenario is a compile error rather than a silent skip:
184
196
 
185
197
  ```typescript
@@ -217,525 +229,77 @@ export const credentialFeature = pikkuFeature({
217
229
 
218
230
  The **feature is the run unit**: `--flows` on a scenario whose every feature entry carries `data` errors and names the features containing it, because the feature is what supplies that data. Use `--features` for those. A scenario referenced bare anywhere, or in no feature at all, still runs standalone.
219
231
 
220
- ### Steps describe intent, not actions
221
-
222
- A scenario records what someone was **trying to do**, never the keystrokes they used to do it. This is the one decision that determines whether a suite survives its first redesign, and it applies to every step name you write.
223
-
224
- | Action ladder — wrong | Intent ladder — right |
225
- | ----------------------------------- | ----------------------------------------------- |
226
- | `Given opens /shop` | `Given shopper is browsing the shop` |
227
- | `When clicks the category filter` | `When shopper buys the £5 strawberry milkshake` |
228
- | `And clicks "Drinks"` | `Then it is in their basket` |
229
- | `And clicks the first product card` | |
230
- | `And clicks Add to basket` | |
231
- | `Then sees "1 item"` | |
232
-
233
- Three things go wrong with the left-hand column, and all three are expensive:
234
-
235
- - **A layout change rewrites every scenario that touched that screen.** In the right-hand column it rewrites one function.
236
- - **The report is the deliverable.** `buys the £5 strawberry milkshake` is readable by someone who has never seen the app; `clicks [data-testid=add]` tells them nothing about whether the product works.
237
- - **An action step cannot arrive on its own.** It assumes the previous click left the browser somewhere, so the scenario only runs front-to-back, as a whole, in one order.
238
-
239
- So there are three layers, and only two of them are named in the report:
240
-
241
- | Layer | What it is | On the ladder |
242
- | -------------------------- | -------------------------------------- | ---------------- |
243
- | Scenario | The flow, written as intents | yes — the ladder |
244
- | Step (`pikkuScenarioStep`) | One intent | yes — one row |
245
- | Utility | An ordinary TS function over `browser` | no |
246
-
247
- Utilities are **not steps**. They are plain exported functions, they take the browser handle, and they hold the clicking:
248
-
249
- ```typescript
250
- // shop.browser.ts — shared actions. Not steps: nothing here is an intent.
251
- import type { PikkuBrowserWire } from '#pikku/scenarios'
252
- import type {} from '@pikku/playwright'
253
-
254
- /** Arrive on the shop, from wherever the browser happens to be. */
255
- export const ensureOnShop = async (browser: PikkuBrowserWire) => {
256
- if (!new URL(browser.page.url()).pathname.startsWith('/shop')) {
257
- await browser.goto('/shop')
258
- }
259
- await browser
260
- .locate({ testId: 'product-grid' })
261
- .first()
262
- .waitFor({ state: 'visible' })
263
- }
264
-
265
- export const searchFor = async (browser: PikkuBrowserWire, query: string) => {
266
- await browser.locate({ testId: 'shop-search' }).first().fill(query)
267
- await browser.page.keyboard.press('Enter')
268
- }
269
-
270
- export const filterByCategory = async (
271
- browser: PikkuBrowserWire,
272
- category: string
273
- ) => {
274
- await browser.locate({ testId: 'category-filter' }).first().click()
275
- await browser
276
- .locate({ testId: 'category-option', where: { 'data-category': category } })
277
- .first()
278
- .click()
279
- }
280
-
281
- export const addToBasket = async (browser: PikkuBrowserWire, name: string) => {
282
- const card = browser
283
- .locate({ testId: 'product-card', containing: name })
284
- .first()
285
- await card.waitFor({ state: 'visible' })
286
- await card.locate('[data-testid=add-to-basket]').click()
287
- }
288
- ```
289
-
290
- The step composes them, and it is the step — one row — that the report shows:
291
-
292
- ```typescript
293
- export const buysTheItem = pikkuScenarioStep<
294
- { name: string },
295
- { name: string }
296
- >({
297
- name: 'buysTheItem',
298
- description: 'finds one item in the shop and puts it in the basket',
299
- template: 'buys the {name}',
300
- // One intent, one implementation per surface an actor can drive it through.
301
- browser: async (_services, { name }, { browser }) => {
302
- await ensureOnShop(browser)
303
- await searchFor(browser, name)
304
- await addToBasket(browser, name)
305
- return { name }
306
- },
307
- default: async ({ rpc }, { name }) => {
308
- const item = await rpc.invoke('findItemByName', { name })
309
- await rpc.invoke('addToBasket', { itemId: item.id })
310
- return { name }
311
- },
312
- })
313
- ```
314
-
315
- The bindings are **alternatives**: `pikku scenario run --run browser` clicks through the shop, `--run cli` drives it over the websocket, `--run default` (the fast suite, and the default) takes the server-side path — and all of them report the same sentence.
316
-
317
- ```typescript
318
- await scenario.when(
319
- 'buys a milkshake',
320
- 'buysTheItem',
321
- { name: '£5 strawberry milkshake' },
322
- { actor: actors.shopper }
323
- )
324
- // reporter renders: When shopper buys the £5 strawberry milkshake ✓ 1.2s
325
- ```
326
-
327
- **Every intent step begins by arriving.** `ensureOnShop` is not defensive noise — it is what lets a scenario start at any step, run alone, and be reordered without touching it. It checks first and navigates only if needed, so a scenario already on the shop pays nothing. This is about the _browser's_ starting position, not the database: there is still no state reset (see above), and you still scope what you create.
328
-
329
- **The same utilities, a different intent.** A scenario about filtering has filtering as its subject, so there the filter _is_ the intent — same helper, its own step:
330
-
331
- ```typescript
332
- export const filtersTheShop = pikkuScenarioStep<
333
- { category: string },
334
- { shown: number }
335
- >({
336
- name: 'filtersTheShop',
337
- description: 'narrows the catalogue to one category',
338
- template: 'filters the shop by {category}',
339
- browser: async (_services, { category }, { browser }) => {
340
- await ensureOnShop(browser)
341
- await filterByCategory(browser, category)
342
- return {
343
- shown: await browser.locate({ testId: 'product-card' }).count(),
344
- }
345
- },
346
- default: async ({ rpc }, { category }) => ({
347
- shown: (await rpc.invoke('listItems', { categorySlug: category })).length,
348
- }),
349
- })
350
- ```
351
-
352
- Two scenarios, two intents, one set of utilities. That is the shape to aim for: when a helper is reused by a step whose _subject_ it is, promote it to a step there — never the reverse.
353
-
354
- **Non-browser steps need none of this.** Without a browser there is no navigation to absorb and no DOM to hide, so an intent maps to one RPC and `scenario.do` names it directly:
355
-
356
- ```typescript
357
- const order = await scenario.do(
358
- 'Shopper checks out',
359
- 'createOrder',
360
- { basketId, shippingAddress },
361
- { actor: actors.shopper }
362
- )
363
- ```
364
-
365
- Reach for a `pikkuScenarioStep` on the non-browser side only when one intent genuinely spans several RPCs, or when the step asserts something the RPC result alone does not say.
366
-
367
- ### What language the prose is in
368
-
369
- A scenario carries two kinds of text, and they do not share a language.
370
-
371
- **Identifiers are English.** The exported const (`buysAnApple`,
372
- `credentialFeature`), the step's `name` — which is its `pikkuFuncId`, the typed
373
- string the generated step map is keyed by — the file name, and every helper in
374
- `*.browser.ts`. These bind to generated code and to `pikku scenario list`; they
375
- are English in every project regardless of who the product is for or what
376
- language the team speaks. There is no setting that changes this.
377
-
378
- **Prose follows `metaLocale` in `pikku.config.json`** (default `en`). That is a
379
- step's `description` and `template`, a feature's `name` and `description`, a
380
- scenario's `title`, and the positional step names passed to
381
- `scenario.given/when/then`. Read the field before you write any of them.
382
-
383
- This split is the same one the feature table already states — _the export
384
- identifier is the feature's id; `name` is the human-readable label_ — applied to
385
- language. The report is the deliverable, and it is read by the team; the
386
- identifier is an API, and it is read by the toolchain.
387
-
388
- ```typescript
389
- // pikku.config.json: { "metaLocale": "de" }
390
- export const buysAnApple = pikkuScenarioStep<
391
- { qty: number },
392
- { orderId: string }
393
- >({
394
- name: 'buysAnApple', // identifier — English, always
395
- description: 'kauft einen Apfel', // prose — follows locale
396
- template: 'kauft {qty} Äpfel', // prose — follows locale
397
- actor: true,
398
- default: async (_services, { qty }, { actor }) =>
399
- await actor.invoke('placeOrder', { qty }),
400
- })
401
- ```
402
-
403
- Note what does **not** change: `placeOrder` is still `placeOrder`, and the file
404
- is still `apple.scenario.ts`.
405
-
406
- A product with a non-English UI is not on its own a reason to set `metaLocale` — that
407
- is the app's language, not the team's. Ask, or leave it `en`.
408
-
409
- **Where a non-`en` `metaLocale` still shows English, today.** The reporter composes a
410
- sentence as `<Keyword> <actor> <template>` (`composeStepProse`), and the keyword is
411
- an English literal. The Console translates the Given/When/Then keywords into its own
412
- UI language; the CLI reporter does not, so `metaLocale: "de"` gives you German step
413
- prose inside an English frame — `Given shopper kauft 1 Äpfel`. Write templates that read
414
- acceptably in that frame rather than trying to defeat it. A second gap: where a
415
- function or scenario declares no `title`, the Console falls back to splitting the
416
- **identifier** into an English-looking label (`toEnglishName`), so under a
417
- non-`en` `metaLocale` meta is worth authoring rather than leaving to the fallback.
418
-
419
- ### `then` bindings are witnesses, not alternatives
420
-
421
- This is the one place the surface bindings do **not** behave like a switch, and it is the part worth reading twice.
422
-
423
- On a `given` or `when`, the bindings are alternatives — clicking Buy and calling `createOrder` are two ways to cause one effect, so exactly one runs.
424
-
425
- On a `then`, they are not two implementations of one assertion. They are two _different claims_:
426
-
427
- | binding | what it actually proves |
428
- | --------- | -------------------------------------------------------------- |
429
- | `default` | the order row says `paid` — the system of record is right |
430
- | `browser` | the confirmation panel says paid — the truth reached the human |
431
-
432
- The gap between them is the bug nobody catches: 200 OK, database correct, user still watching a spinner. So a `then` runs **every** binding it declares and fails if they disagree.
433
-
434
- ```typescript
435
- export const seesTheOrderConfirmed = pikkuScenarioStep<
436
- { orderId: string },
437
- { status: string }
438
- >({
439
- name: 'seesTheOrderConfirmed',
440
- // Both bindings run as the persona, so the step declares one and the runner
441
- // injects `wire.actor` — non-optional in every binding.
442
- actor: true,
443
- template: 'sees order {orderId} confirmed',
444
- // Both run on `--run browser`. Each returns what it observed, and the runner
445
- // compares them — so this fails when the page disagrees with the database.
446
- browser: async (_services, { orderId }, { browser }) => ({
447
- status: await browser
448
- .locate({ testId: 'order-status', where: { 'data-order': orderId } })
449
- .getAttribute('data-status'),
450
- }),
451
- // Through the actor, not through a `rpc` service — see "What a step is given".
452
- default: async (_services, { orderId }, { actor }) => ({
453
- status: (await actor.invoke('getOrder', { orderId })).status,
454
- }),
455
- })
456
- ```
457
-
458
- Three rules follow, and they are the ones that get broken:
459
-
460
- - **A browser witness must observe on the page.** One that quietly calls an RPC to check the result is worse than no binding at all — it reports a tick for a surface it never looked at.
461
- - **Return what you observed, don't just assert.** A witness returning a value lets the runner diff the two. A witness that only throws still works, but it can never disagree with anything, so it proves less. Read structured state with `where` on the test-id selector rather than parsing translated copy.
462
- - **A step with no binding for the run's surface is counted, not excused.** `--run browser` prints `n/m steps ran on browser` over _every_ step, so an action that quietly fell back to the server lowers the number just as an assertion does. A `then` that fell back is additionally named — `--strict` fails on those, because a sentence saying the actor saw something nobody looked at is a different problem from a shortcut. Not being in the UI _is_ the finding: do not add a browser binding that fakes it.
463
-
464
- **Always give a `then` a `default` witness.** It is the floor every run can fall back to, and an assertion with no witness the run can execute is fatal (`ScenarioNoWitness`) — not a coverage gap. The distinction is the point: a `then` checked server-side under `--run browser` did happen, it just wasn't seen where the prose claims; one checked nowhere never happened at all, and without the error it would return `undefined` and render as a tick. A browser-only `then` is therefore a step that fails the fast suite, which is rarely what you want.
465
-
466
- **Every scenario must assert.** A flow of only `given`/`when` is a PKU680 critical — it proves nothing threw. Since coverage counts every step, an assertion-free ladder of browser-bound actions would score a perfect `3/3` while checking nothing, so clicking through the UI and never looking at the result is the cheapest way to fake the number. The rule closes that.
467
-
468
- Assertions with no possible browser witness are a different thing and should not be written as a `then`: "the audit log recorded it" is a system check, and "the receipt email arrives" is `expectEventually`, which is always out-of-band and always server-side.
469
-
470
- ### Declared steps (`pikkuScenarioStep`)
471
-
472
- `scenario.do` can only name an RPC. A **step** is a named, typed unit of scenario behaviour whose body is an ordinary pikku function — so it can call several RPCs as its actor, assert, or drive a browser.
473
-
474
- ```typescript
475
- import { pikkuScenarioStep } from '#pikku/scenarios'
476
-
477
- export const buysAnApple = pikkuScenarioStep<
478
- { qty: number },
479
- { orderId: string }
480
- >({
481
- name: 'buysAnApple',
482
- description: 'buys an apple',
483
- template: 'buys {qty} apples',
484
- actor: true,
485
- default: async (_services, { qty }, { actor }) => {
486
- return await actor.invoke('placeOrder', { qty })
487
- },
488
- })
489
- ```
490
-
491
- A step's body always lives under a **surface binding** — `default`, `browser` or
492
- `cli` — never under a `func`. Declaring none throws at load time: at minimum give
493
- it a `default`.
494
-
495
- ```typescript
496
- await scenario.given(
497
- 'buys an apple',
498
- 'buysAnApple',
499
- { qty: 1 },
500
- { actor: actors.shopper }
501
- )
502
- // reporter renders: Given shopper buys 1 apples ✓ 412ms
503
- ```
504
-
505
- Rules that bite:
506
-
507
- - **The step is referenced by its typed string name, not by importing the const** — exactly like `workflow.do`. The name is the step's `pikkuFuncId` and is checked against the generated step map. A non-literal target is a critical error (`PKU678`).
508
- - **Steps are not RPCs.** They are deliberately never network-callable — a browser-driving step must not be.
509
- - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal _is_ the assertion.
510
- - **A step that runs as somebody declares `actor: true`**, and the runner injects `wire.actor` — non-optional inside every binding, with no guard to write and nothing to unwrap. A `browser` binding implies it, because a window is opened as somebody. Leave it off for a step with no persona to be: an assertion over what an earlier step returned, or one that posts credentials precisely because it must not reuse an actor's session. Dispatching a step that declared it without `{ actor: actors.x }` fails before the body runs (`ScenarioActorRequired`); a step that did not declare it has no `actor` on its wire at all.
511
- - **`env` is optional on the wire**, because most steps need nothing from it. Narrow it with `requireScenarioEnv(scenarioStep)` from `#pikku/scenario` rather than a local guard — it names the step and says what to pass. `env` is `{ apiUrl, appUrl? }` from the environment the run targets, and is how a raw-HTTP step learns the target's URL: a step runs in the CLI process, where there is no `variables` service and `process.env` is not the answer.
512
- - **Steps default to `retries: 0`**, unlike ordinary workflow steps. Retrying a failed assertion is wrong; pass `retries` explicitly if a step is genuinely flaky-by-nature.
513
- - **Step results are persisted**, so return JSON-serialisable data — never a `Locator` or a client object.
514
- - **`description` documents the step; `template` is what the report renders.** `template`'s `{placeholders}` are filled from the input the step was called with, so one step reads differently for each call — `sees {state} addon {packageName}` reports as "sees available addon @pikku/addon-stripe". Reflect every input field in the template, and type the values so they read as words (`state?: 'installed' | 'available'`, not `installed?: boolean`). A placeholder with no value renders as nothing and the whitespace collapses.
515
- - **Never write the actor into the prose.** The reporter renders the actor as the sentence's subject, so a step authored as `` `'sam' creates the client` `` run as `{ actor: actors.sam }` reads "Given sam 'sam' creates the client" — and the hardcoded name desyncs the moment the call site changes actor. Write a bare third-person predicate (`creates the {name} client`) and let the actor supply the subject. Prose that opens with its own actor's key — quoted, capitalised or possessive — is `PKU681`; naming someone **else** mid-sentence ("sends nadia an invite") is ordinary prose and is left alone, as is an actor keyed after a role noun used as a noun ("creates the admin client" as `actors.admin`).
516
- - Prose precedence is `options.description` → the step's `template` → the step's own `description` → the positional step name. Repeated names get `#1`, `#2` ordinals, so a `for` loop over a data set is how you write a Scenario Outline. A loop-generated step name is not statically known, so it is matched back to its declaration by step function instead — which works as long as that function's call sites agree on their phase, actor and prose. Two call sites that disagree make the loop step report under its bare runtime name.
517
-
518
- #### What a step is given
519
-
520
- A step has the signature of an ordinary pikku function, which makes it look as
521
- though it runs where the application runs. It does not — **it runs in the CLI
522
- process**, and the services object is built there, by hand:
232
+ ## Steps
523
233
 
524
- ```typescript
525
- { logger, workflowService, workflowRunService, agentRunner? }
526
- ```
527
-
528
- That is the whole list. There is no `kysely`, no `variables`, no `secrets`, and
529
- none of the project's own singleton or wire services. A step that destructures
530
- one gets `undefined` and fails on first use — `Cannot read properties of
531
- undefined (reading 'selectFrom')` — which reads like a broken container and is
532
- not.
533
-
534
- `rpc` is the trap worth naming, because it is present and it throws. It is a
535
- `guardRpc` whose every member refuses:
536
-
537
- > Scenario tried to run 'getOrder' as an internal step. Every workflow.do in a
538
- > scenario must carry { actor: actors.x } so it executes against 'local'.
539
-
540
- The same guard covers `rpc.agent.run/stream/resume/approve/interrupt` and
541
- `startWorkflow`.
234
+ `scenario.do` can only name an RPC. A **step** — `pikkuScenarioStep` — is a named, typed unit of
235
+ scenario behaviour whose body is an ordinary pikku function, so it can call several RPCs as its
236
+ actor, assert, or drive a browser. It is referenced by its typed string name (checked against the
237
+ generated step map), its body lives under a surface binding (`default` / `browser` / `cli`, never a
238
+ `func`), and one row of the report is one step.
542
239
 
543
- This is the design, not a gap: **everything a step touches of the application
544
- goes over the wire as somebody.** A test that could reach into the database
545
- would be testing a different program from the one a person uses. So there are
546
- exactly three ways in, and they are all through the actor:
240
+ **Read `references/steps.md` before writing one.** It carries the four things that decide whether a
241
+ suite survives its first redesign:
547
242
 
548
- - `actor.invoke(name, data)` — typed over the exposed RPC map, carrying the
549
- actor's session. Declare `actor: true` and destructure it off the wire.
550
- - `.invokeRaw(name, data, { headers })` — same call, reporting
551
- `{ status, ok, body }`, for when the refusal is the assertion.
552
- - a plain `fetch` against `requireScenarioEnv(scenarioStep).apiUrl`, for
553
- anything not an RPC — a websocket, a file upload, a webhook.
243
+ - **What a step is given** — a step runs in the CLI process, not where the app runs. There is no
244
+ `kysely`, no `secrets`, and `rpc` is present but throws on purpose. Everything reaches the app
245
+ through the actor.
246
+ - **Steps describe intent, not actions** — `buys the £5 strawberry milkshake`, never
247
+ `clicks [data-testid=add]`. The clicking lives in plain utilities the step composes.
248
+ - **`then` bindings are witnesses, not alternatives** — a `then` runs _every_ binding it declares
249
+ and fails when they disagree, because "database right, user still watching a spinner" is the bug
250
+ nobody catches.
251
+ - **What language the prose is in** — identifiers are English in every project; `description`,
252
+ `template` and step names follow `metaLocale`.
554
253
 
555
- Two consequences follow, and both shape how steps get written:
254
+ Browser bindings, and locating by message key rather than rendered copy, are in
255
+ `references/browser.md`.
556
256
 
557
- - **A step cannot observe anything the app does not publish.** If a test needs a
558
- fact the client never sees, the fix is to emit it on the stream or expose it
559
- as an RPC — which usually improves the product, since a client debugging the
560
- same problem needed it too.
561
- - **`agentRunner` is conditional.** It is built only when the project declares
562
- agents, and `createDevAgentRunner` needs a base URL _and_ a key together
563
- (`OPENAI_BASE_URL` + `OPENAI_API_KEY`, or the LiteLLM pair). With a key alone
564
- it returns nothing and `agentRunner` is `undefined`, so `actor.converse`
565
- fails before the persona says anything. A suite that would rather own its own
566
- model can pass an `llm` to `runConversation` instead of relying on this one.
257
+ **Every scenario must assert.** A flow of only `given`/`when` is a PKU680 critical — it proves
258
+ nothing threw. Coverage counts every step, so an assertion-free ladder of browser actions would
259
+ score a perfect run while checking nothing.
567
260
 
568
- ### Browser steps
569
-
570
- Declaring a `browser` binding is the whole switch: inside that binding `wire.browser` is guaranteed present and non-optional, and a step without one never sees a browser at all. There is nothing to null-check.
571
-
572
- A `browser` binding gets a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
573
-
574
- Browser steps are where **intent, not actions** earns its keep: the step is one intent, the clicking lives in shared utilities, and the step arrives before it acts. Write the mechanics below into utilities and keep the step body to three or four calls that read as a sentence.
575
-
576
- ```typescript
577
- export const opensTheCart = pikkuScenarioStep<
578
- { path: string },
579
- { url: string }
580
- >({
581
- name: 'opensTheCart',
582
- description: 'opens the cart',
583
- browser: async (_services, { path }, { browser }) => {
584
- await browser.goto(path)
585
- return { url: browser.page.url() }
586
- },
587
- default: async (_services, _data, { actor }) => ({
588
- url: (await actor.invoke('getCart', {})).url,
589
- }),
590
- })
591
- ```
592
-
593
- - Install `@pikku/playwright` and `@playwright/test`, and import `@pikku/playwright` once (`import type {} from '@pikku/playwright'`) so `browser.page` is a typed Playwright `Page`. Without it you still get the structural `goto`/`screenshot` handle.
594
- - The environment needs an `appUrl` beside its `apiUrl`. `pikku scenario run` fails fast before running anything if a browser scenario has no `appUrl` or the driver is not installed.
595
- - `pikku scenario run <env> --no-browser` **skips** scenarios containing browser steps and reports them as skipped — it does not fail them. That is how a machine with no browser stays green.
596
- - Playwright auto-waits; do not wrap `page.click` in `expectEventually`.
597
-
598
- #### Locate by message key, never by rendered copy
599
-
600
- If the app is translated, **no step may contain a user-visible string.** `getByLabel('Full Name')` passes only while the browser happens to render the base locale, and any copy edit turns it into a selector timeout that points at the wizard rather than at the rename that caused it — the test looks broken where it is merely stale.
261
+ ## Configuration
601
262
 
602
- The message catalogue already holds the string under a key. Read it from there. Type the lookup off the catalogue JSON so a renamed or misspelled key is a **compile** error rather than a run-time timeout:
263
+ Personas are declared in TypeScript. One `definePersonas({ … })` call for the
264
+ whole project — codegen builds the `PersonaId` union from it, materialises one
265
+ scenario actor per person, and seeds a user row each:
603
266
 
604
- ```typescript
605
- // tests/scenarios/i18n.ts
606
- import type messages from '../../../../apps/web/messages/en.json'
267
+ ```ts snippet:definePersonas
607
268
 
608
- export type MessageKey = keyof typeof messages
609
-
610
- export const t = (key: MessageKey, locale = baseLocale): string => {
611
- /* … */
612
- }
613
269
  ```
614
270
 
615
- ```typescript
616
- await page.getByLabel(t('jobs_apply_fullname')).fill(identity.name)
617
- await page
618
- .getByRole('button', { name: t('jobs_apply_submit'), exact: true })
619
- .click()
620
- ```
621
-
622
- - Type off `messages/<baseLocale>.json`, **not** the generated Paraglide output — `i18n/paraglide/` is build output, so typing against it makes the tests unbuildable until the app has been built. The JSON is the tracked source.
623
- - Fall back to the base locale for a key a locale has not translated. That is what Paraglide does at run time, so a helper that throws instead would disagree with the screen the test is looking at.
624
- - This is not only about locators. A copy literal passed to a **project helper** (`pick('Where would you like to work?', …)`) reaches the DOM the same way, and so does a pane name quoted back in a failure message. `pikku fabric validate` scans every string in a `*.steps.ts` / `*.scenario.ts` against the base catalogue and errors on any verbatim match, wherever it sits — except comments, and the `name` / `description` / `template` declared directly on a `pikkuFeature`, `pikkuScenario` or `pikkuScenarioStep`, which are Console meta written in the project's `locale` rather than app copy.
625
- - A regex locator (`{ name: /^Next$/i }`) hides the literal but not the problem. `{ name: t('key'), exact: true }` is both stricter and locale-correct.
626
- - Strings the catalogue does not own — a test id, a fixture filename, a seeded value — stay literal. The catalogue is the test for whether something is copy.
627
-
628
- ## Configuration
629
-
630
- Personas, actors and environments live in `pikku.config.json`:
271
+ `pikku.config.json` carries the settings around them — nothing about a person:
631
272
 
632
273
  ```json
633
274
  {
634
275
  "scenarios": {
635
- "personas": {
636
- "shopper": { "description": "Buys things here", "primary": true },
637
- "support": {
638
- "description": "Answers for the shop",
639
- "proficiency": "power"
640
- },
641
- "reminders": {
642
- "description": "The shop chasing abandoned carts",
643
- "kind": "system"
644
- }
645
- },
646
- "actors": {
647
- "shopper": {
648
- "email": "shopper@actors.local",
649
- "name": "Shopper",
650
- "jobTitle": "First-time buyer",
651
- "personality": "Impatient shopper who abandons slow checkouts"
652
- },
653
- "shopperB": { "persona": "shopper", "email": "shopper-b@actors.local" }
654
- },
655
- "environments": {
656
- "local": {
657
- "apiUrl": "http://localhost:4077",
658
- "signInPath": "/api/auth/sign-in/actor"
659
- }
276
+ "emailDomain": "actors.example.com",
277
+ "browserDriver": "@pikku/playwright",
278
+ "model": "claude-sonnet-4-5"
279
+ },
280
+ "environments": {
281
+ "local": {
282
+ "apiUrl": "http://localhost:4077",
283
+ "signInPath": "/api/auth/sign-in/actor"
660
284
  }
661
285
  }
662
286
  }
663
287
  ```
664
288
 
665
- ### Personas and actors
666
-
667
- A **persona** is a kind of person; an **actor** is one body that signs in as one. Above, `support` is declared only as a persona — its actor is materialised (`support@actors.local`), so `actors.support` works without an `actors` entry. Write an actor by hand only when you need something the materialised one wouldn't have:
668
-
669
- - a **real email or personality** for it, like `shopper`;
670
- - a **second body of the same persona**, like `shopperB` — which is what tenant isolation, peer sharing, and "another member's row" scenarios are made of. Two actors of one persona must be two different users, so **two actors sharing an email is an error**.
671
-
672
- A persona holds only what is true of that kind of person for the app's whole lifetime — `description`, `primary` (whose experience the product is), `kind`, `proficiency`. What someone is trying to get done, and the circumstances they are doing it in, belong to the **scenario**, not to them.
289
+ - `scenarios.emailDomain` is the mail domain actor addresses are built on.
290
+ - `scenarios.browserDriver` is the package driving `browser` bindings.
291
+ - `scenarios.model` is the model a persona thinks with (`actor.converse`, `pikku virtual-user run`).
292
+ - `environments.<name>` are the targets a run can point at. The key is the required positional of `pikku scenario run`; `apiUrl` is required and `signInPath`/`rpcPath`/`sessionPath` have defaults.
673
293
 
674
- `kind: "system"` is the app acting on its own — a schedule, a cleanup, a send. It gets **no actor**: there is nobody to sign in. Give it one by hand only if it genuinely has a service account.
675
-
676
- #### Declaring personas in TypeScript
677
-
678
- `definePersonas({ … })` is the code form of the block above, and there may be
679
- **one call in the whole codebase** — one place to read the set from, one place
680
- to add to it. A second anywhere, including in the same file, is a critical.
681
- Generated files are exempt and never claim the slot.
682
-
683
- > [!WARNING]
684
- > The declaration is **read from source, never evaluated** — the CLI writes it
685
- > to JSON that a deployed stage carries without the app. So every value has to
686
- > be statically knowable, and a value that is not comes out as `undefined`
687
- > rather than as an error. Only `name` is checked, so a computed `personality`,
688
- > `jobTitle` or `description` is dropped in silence and the persona runs with a
689
- > blank temperament.
690
-
691
- What that admits and what it does not:
692
-
693
- ```typescript
694
- personality: 'Wound up and short with it.' // read
695
- personality: `Wound up and short with it.
696
- Says what she wants in a few blunt words.` // read — no ${} in it
697
- personality: 'Wound up. ' + 'Short with it.' // dropped, silently
698
- personality: TEMPERAMENTS.impatient // dropped, silently
699
- ```
700
-
701
- A no-substitution template literal is a string literal as far as the reader is
702
- concerned, so it is the way to write a long personality across several lines —
703
- not a concatenation, and not a `prettier-ignore`d single line. Its newlines and
704
- leading indentation are kept verbatim and reach the model that way, which is
705
- harmless but worth knowing before you align it to the surrounding code.
706
-
707
- One more thing worth knowing before writing a rich persona: **`actor.converse`
708
- builds its prompt from `name`, `jobTitle`, `personality` and the scenario's
709
- `task` only.** Fields like `disposition`, `goals` and `roles` are read and
710
- stored, and the console shows them, but they do not reach the conversing
711
- persona's instructions. Anything that must shape how someone talks belongs in
712
- `personality` or in the task.
713
-
714
- An actor with no `persona` is its own persona, so a project that never declares any keeps working unchanged.
294
+ **`references/personas.md`** covers the parts that bite: one entry is one
295
+ person and one actor, emails are derived and never written, `runnable: false`
296
+ for someone only acted upon, `definePersonas` being read from source rather
297
+ than evaluated, and the same actor list powering a human "Sign in as …"
298
+ switcher.
715
299
 
716
300
  - `environments.<name>.apiUrl` is required. `signInPath` defaults to `/auth/sign-in/actor`, `rpcPath` to `/rpc`.
717
301
  - **`SCENARIO_ACTOR_SECRET` is an environment variable and never goes in `pikku.config.json`.** It signs actors in. `pikku scenario run` throws without it; a server auto-building actors warns and runs without them.
718
302
 
719
- ### The same actors sign a human in
720
-
721
- Declared actors are not only for automated runs. `signInPath` is Better Auth's
722
- `actor` plugin (see `pikku-auth`, a separate install), which any caller can post to — so the
723
- frontend gets a one-click "Sign in as …" switcher over the **same** list, and an
724
- app can be reviewed as each kind of user without anyone knowing a seed password.
725
-
726
- The sandbox dev server bakes both halves into the frontend from the declared
727
- personas: `VITE_DEV_ACTORS` (the JSON actor list) and `VITE_DEV_ACTOR_SECRETS`
728
- (`{ email: credential }`, one per persona — `SCENARIO_ACTOR_SECRET` itself never
729
- goes in a bundle; see **pikku-auth**). Neither var is set in a production
730
- build, so the control renders nothing there — but gate the reads on your
731
- bundler's dev flag anyway (`import.meta.env.DEV ? … : undefined`) so no
732
- credential reaches a production bundle in the first place.
733
-
734
- Do not hand-roll the switcher: `useDevActors()` (`pikku-react`, a separate install) is the logic and
735
- `<DevActorSwitcher />` from `@pikku/mantine/dev` is a ready rendering of it.
736
- `pikku fabric validate` **requires** any frontend with a login screen to ship
737
- one — without it a reviewer is locked out of their own sandbox.
738
-
739
303
  ## Running
740
304
 
741
305
  ```bash
@@ -744,7 +308,7 @@ SCENARIO_ACTOR_SECRET=… pikku scenario run local
744
308
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --flows orderSupportScenario
745
309
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --features credentialFeature
746
310
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --tags smoke,scenario
747
- SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --no-browser --exclude-tags ai-live
311
+ SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --exclude-tags ai-live
748
312
  ```
749
313
 
750
314
  `run` takes the environment as a **required positional** — the key from `environments`. Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
@@ -756,7 +320,8 @@ SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --no-browser --exclud
756
320
  | `--tags` / `-t` | Match-any tag filter |
757
321
  | `--exclude-tags` | Hold tags back — unless the flow is named directly with `--flows` |
758
322
  | `--run <surface>` | `default` (the default), `browser`, or `cli` |
759
- | `--no-browser` | Shorthand for `--run default`; scenarios with browser steps report as **skipped** |
323
+ | `--screenshots` | Write asked-for screenshots under `.pikku/scenario-runs/<run>/<scenario>` |
324
+ | `--video <mode>` | `failed` (the default), `all`, or `off` |
760
325
  | `--strict` | Fail, rather than pass, a `then` with no witness on the run's surface |
761
326
  | `--spawn` / `--keep-alive` | Start `pikku dev` on the environment's apiUrl for the run; optionally leave it up |
762
327
  | `--api-url` / `--app-url` | Override the environment's URLs — for a target that only exists at run time |
@@ -769,74 +334,18 @@ Output is `PASS <name> (<ms>) → <output>` / `FAIL <name> (<ms>): <error>`, the
769
334
 
770
335
  ## Coverage
771
336
 
772
- Coverage is attributed by running scenarios against a server that is collecting it. It is **not** derived from unit tests.
773
-
774
- Prerequisite in `pikku.config.json`:
775
-
776
- ```bash
777
- pikku enable scenarios # sets scaffold.scenarios = true
778
- ```
779
-
780
- `scaffold.scenarios` is a boolean or `{ path? }` — whether the surface exists
781
- and where it is written. A bare string is **rejected by the config loader**, not
782
- reinterpreted: under a shape where a string could be a path, silently reading
783
- one as a flag would be worse than failing.
784
-
785
- `scaffold.scenarios` generates the coverage and stub RPCs into your project (`pikkuScenarioTakeLiveCoverage`, `pikkuScenarioResetLiveCoverage`, `pikkuScenarioResetStubs`, `pikkuScenarioGetStubCalls`), so scenario runs work against any server. The coverage RPC reads `<outDir>/function/pikku-functions-meta-verbose.gen.json` off disk at request time — codegen always writes it, but it has to be deployed alongside the app or the RPC returns `null`.
337
+ Coverage is attributed by running scenarios against a server that is collecting it — it is **not**
338
+ derived from unit tests:
786
339
 
787
340
  ```bash
341
+ pikku enable scenarios # sets scaffold.scenarios = true
788
342
  pikku dev --coverage # V8 precise coverage, in-process
789
- pikku dev --coverage --test # also enable stubs (needed for expectService)
790
343
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --coverage
791
344
  ```
792
345
 
793
- The run resets coverage before each scenario and snapshots after, writing **`<outDir>/coverage/scenario-coverage.json`**:
794
-
795
- ```jsonc
796
- {
797
- "generatedAt": "…",
798
- "environment": "local",
799
- "scenarios": {
800
- "<name>": {/* FunctionCoverageReport */},
801
- },
802
- }
803
- ```
804
-
805
- Coverage is best-effort: it disables itself with a warning if the server is not collecting or the first actor cannot invoke, and it needs at least one configured actor. If you get no coverage, check those first.
806
-
807
- **There is no AI-prompt output.** The old `--ai-out` flag died with `pikku tests`; nothing replaced it. To find what needs work, read `scenario-coverage.json` yourself and cross-reference `pikku meta functions list` for input/output schemas.
808
-
809
- ### Filling coverage
810
-
811
- 1. `pikku scenario run <env> --coverage`, then read `<outDir>/coverage/scenario-coverage.json` to see what is unexercised.
812
- 2. `pikku meta functions list` for those functions' schemas.
813
- 3. Write a `pikkuScenario` that reaches them **through a real user flow** with an actor — not a scenario per function. Scenarios are flows; coverage is a consequence.
814
- 4. Re-run to confirm.
815
-
816
- ## Unit tests for pure logic
817
-
818
- Scenarios are the repo-idiomatic way to test functions, and the only thing that contributes to live coverage. For pure logic with heavy branching, a plain unit test calling `func` directly is still valid and cheap:
819
-
820
- ```typescript
821
- import { describe, test } from 'node:test'
822
- import assert from 'node:assert'
823
-
824
- describe('createTodo', () => {
825
- test('creates a todo', async () => {
826
- const services = {
827
- todoStore: { add: async (title: string) => ({ id: '1', title }) },
828
- }
829
- const result = await createTodo.func(services as any, { title: 'Buy milk' })
830
- assert.equal(result.title, 'Buy milk')
831
- })
832
- })
833
- ```
834
-
835
- ```bash
836
- node --import tsx --test src/**/*.test.ts
837
- ```
838
-
839
- Services are plain objects — a Pikku function is pure business logic, so a mock is just the shape the function destructures. Build real services via the `pikkuServices` / `pikkuWireServices` factories when a test needs them.
346
+ That writes `<outDir>/coverage/scenario-coverage.json`. **`references/coverage.md`** has the
347
+ `scaffold.scenarios` shape, why coverage silently reads zero, the four-step loop for filling a gap,
348
+ and where a plain unit test is still the right tool.
840
349
 
841
350
  ## Red flags
842
351