@pikku/skills 0.12.40 → 0.12.42

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,292 @@
1
+ # Scenarios — writing journeys that stay proven
2
+
3
+ A scenario is a user journey run as one of your personas, over the real transport, with that
4
+ persona's session. It is the only kind of test worth writing here, because a passing one proves the
5
+ app works the way a signed-in person experiences it.
6
+
7
+ The traps below all cost a real milestone a red run, and most of them are invisible on the run that
8
+ introduces them. Read the section that matches what you are writing.
9
+
10
+ 1. [The shape of a scenario](#the-shape-of-a-scenario)
11
+ 2. [Extraction: what survives, and what silently does not](#extraction-what-survives-and-what-silently-does-not)
12
+ 3. [What to assert](#what-to-assert)
13
+ 4. [There is no state reset](#there-is-no-state-reset)
14
+ 5. [Steps rot as the app grows](#steps-rot-as-the-app-grows)
15
+ 6. [Browser scenarios](#browser-scenarios)
16
+ 7. [Running them](#running-them)
17
+ 8. [Coverage — which functions have actually been run](#coverage--which-functions-have-actually-been-run)
18
+
19
+ ---
20
+
21
+ ## The shape of a scenario
22
+
23
+ Three ship in `packages/functions/test/scenarios/` — keep them green — and every milestone's gherkin
24
+ block from §5 becomes one more.
25
+
26
+ ```typescript
27
+ import { pikkuScenario } from '#pikku/scenarios'
28
+
29
+ export const tenantReportsAFaultScenario = pikkuScenario<void, { id: string }>({
30
+ title: 'A tenant reports a fault and the owner sees it',
31
+ description: 'The report lands on the owning landlord’s queue, and nobody else’s',
32
+ tags: ['scenario', 'maintenance'],
33
+ func: async (_services, _data, { scenario, actors }) => {
34
+ const report = await scenario.do(
35
+ 'reports a broken boiler',
36
+ 'createMaintenanceReport',
37
+ { summary: 'No hot water' },
38
+ { actor: actors.chidi },
39
+ )
40
+ await scenario.then(
41
+ 'appears on the owner’s queue',
42
+ 'reportShowsOnQueue',
43
+ { id: report.id },
44
+ { actor: actors.amina },
45
+ )
46
+ await scenario.then(
47
+ 'is invisible to the other owner',
48
+ 'reportIsNotVisible',
49
+ { id: report.id },
50
+ { actor: actors.bilal },
51
+ )
52
+ return { id: report.id }
53
+ },
54
+ })
55
+ ```
56
+
57
+ **`do` takes an RPC name; `given`/`when`/`then` take a declared step.** A step is a
58
+ `pikkuScenarioStep` that says what a person is doing and holds one implementation per surface
59
+ (server-side by default, plus a `browser` one that drives the page). Reaching for an RPC name in a
60
+ `then` will not resolve.
61
+
62
+ **`SCENARIO_ACTOR_SECRET` must be in `.env`** (app.md §6, before the first run). Without it
63
+ `/api/auth/sign-in/actor` is disabled — every scenario then fails at sign-in, before its first step,
64
+ for a reason that reads like an auth bug. `pikku scenario run` reads it from the environment, so
65
+ source `.env` first (`set -a && . ./.env && set +a`) when you run outside `bun run dev`.
66
+
67
+ ---
68
+
69
+ ## Extraction: what survives, and what silently does not
70
+
71
+ A scenario body is extracted as a DSL workflow, so it is not ordinary TypeScript. The failures here
72
+ are the expensive kind: two of the three produce a GREEN suite that proves the wrong thing.
73
+
74
+ **Every scenario must assert, and `return await scenario.then(...)` does not count.** A ladder of
75
+ `given`/`when` with no `then` is a PKU680 critical — it fails `pikku all`, so it stops codegen
76
+ rather than a test. Coverage counts every step, so without that rule an assertion-free ladder of
77
+ clicks would score a perfect run while checking nothing. The extractor reads the body statically and
78
+ does not see a `then` in `return` position: seven refusal scenarios here, each ending
79
+ `return await scenario.then('is refused …', …)`, were all reported as never asserting. Bind it —
80
+ `const asserted = await scenario.then(...)`, then `return asserted` on the next line — which is also
81
+ how the value stays inspectable when the step's output is what the scenario returns.
82
+
83
+ **Only `const`/`let`, `if`/`else`, `switch`, `for..of`, `return`, `throw` and workflow calls
84
+ survive.** A counting `for` is refused by PKU679, and so is a `for..of` whose iterable is written
85
+ inline — it must be a named identifier or a field (`data.items`). Bind the seat numbers, the ids,
86
+ the rows to a `const` above the loop and iterate that. One milestone here lost a codegen round to
87
+ each half of that rule, because the first half does not imply the second.
88
+
89
+ **Write each scenario's setup out in full rather than sharing a local helper.** A helper holding
90
+ setup steps does not fail extraction — the steps are recorded and the suite goes green — but the
91
+ extractor cannot bind an `actor` that arrives as a function parameter, so the transcript credits
92
+ every setup step to whoever the last literal binding named. Seven permission scenarios here read
93
+ "henrik sets up her company Salon Nordlicht" in a suite whose whole point was that the company is
94
+ finja's. A scenario body is a recorded document, and the duplication is the price of it saying who
95
+ did what.
96
+
97
+ ---
98
+
99
+ ## What to assert
100
+
101
+ **Write the refusals, and assert WHY.** One persona reaching for another's row has to be rejected,
102
+ and that rejection is a scenario — it is how you prove access control instead of asserting it. But
103
+ only if the step reads the reason: "not ok" is also what a malformed request returns, so a refusal
104
+ step that stops at the status code passes on a call the function never even ran. One here posted
105
+ straight to `/rpc/<name>` with the function's input as the body; that route validates an ENVELOPE
106
+ (`{ rpcName, data }`), so the input read as a bag of unknown properties and came back 422. Asserting
107
+ the refusal MENTIONED the rule — a company, an owner, a scope — turned a green false positive into a
108
+ one-line fix.
109
+
110
+ **Assert totals as deltas.** A screen that counts or sums every row — a revenue tile, a queue count,
111
+ a dashboard — is reporting the whole history of a database nobody resets. Read the summary before
112
+ the journey, read it after, and assert what the journey moved. An absolute ("the failed count is
113
+ zero") is a claim about every run that came before. And when a tile turns out to be one no journey
114
+ in the app can move at all, that is a finding about the app, not an assertion to force: assert it
115
+ held still, say why in the step's doc comment, and tell the user.
116
+
117
+ **Green twice is not the same as unchanged twice — count the rows.** A save that appends where it
118
+ should replace passes every assertion while doubling a table. One here re-sent a product's variants
119
+ without their ids, so the addon read each as new: 1, 2, 4, 8, and by the nineteenth save 262,144
120
+ rows, every run green until the request crossed a body-size limit and surfaced as a `413` that read
121
+ like an infrastructure fault. The assertion nobody writes is the count, and it is one SQL query.
122
+
123
+ **One screen's extra field does not belong on the shared output schema.** A detail page almost
124
+ always wants one column the list does not — when the licence was handed over, who last touched the
125
+ row. Extending the shared `XDetail` that four functions already return bumps the contract of all
126
+ four, for a field three of them never render. Extend at the new function's own output instead —
127
+ `XDetail.extend({ assignedAt })`, named for the screen that asked. Say in the new type's doc comment
128
+ WHY it is not on the base, or the next build merges them back.
129
+
130
+ ---
131
+
132
+ ## There is no state reset
133
+
134
+ A scenario runs against a live server: scope what you create to your own rows and unique ids, and
135
+ never assume a clean database. A scenario that leaves durable state changed has to put it back — the
136
+ one that archives a product unarchives it, the one that cancels a plan restarts it — because its own
137
+ second run starts where its first one stopped.
138
+
139
+ **Run the suite twice and require the second run green.** A suite that only passes on a fresh
140
+ database is a suite that passes once. Two corollaries, both of which cost a milestone a red run:
141
+
142
+ - **Name nothing a setup step might already own.** `setsUpHerCompany` returns the company that actor
143
+ already has rather than renaming it, so a later step passed the literal name it had asked for and
144
+ was told no such company exists. Read the name, slug or id back off the step's output and pass
145
+ THAT.
146
+ - **Put rows somewhere the other scenarios are not.** An import seeded at the same coordinates as
147
+ another scenario's salons accumulated one row per run until it crowded that scenario's own salon
148
+ out of a nearest-N list. Anything a scenario asserts by proximity, recency or a top-N cut is
149
+ asserting against every row every previous run left behind.
150
+
151
+ **Moving the clock forward runs every rule between here and there.** A step that time-travels so a
152
+ scheduled job will fire is not asking for that job — it is asking for all of them, in order. One
153
+ swept to 2030 to reach a four-week chase and found the pile it was about to assert on empty, because
154
+ an unreturned pen reactivates a membership sixty days after cancellation and the sweep had walked
155
+ straight past that window. Travel to the day the rule under test fires and no further, computed off
156
+ a date the scenario read back (`daysAfter(periodEnd, 1)`), never to a round far-future date — and
157
+ when a sweep surprises you, the next rule in the calendar is the first place to look.
158
+
159
+ ---
160
+
161
+ ## Steps rot as the app grows
162
+
163
+ A step is shared, so it is the one thing in the suite that a milestone which never mentions it can
164
+ break. Whatever a step selects by, ask what ELSE will match it after the suite has run a hundred
165
+ times.
166
+
167
+ **Drive a new step from both sides in the milestone that adds it.** A step is only as proven as its
168
+ best-exercised branch: one here read the wrong field off a raw invocation (`attempt.data`; the
169
+ payload is `attempt.body`), so its found-case could never pass — invisible for as long as every
170
+ caller asked for ABSENCE.
171
+
172
+ **When you add a writer of a row an existing step selects by recency, give that step an explicit
173
+ filter in the same change.** "The newest X" stops meaning "the one this scenario just made" the
174
+ moment a second function produces X: a renewal job that raised invoices broke three invoice
175
+ scenarios that had been green for months.
176
+
177
+ **A selector on IDENTITY alone rots the same way once rows gain a lifecycle.** One reused the first
178
+ licence assigned to an email regardless of its state, which was correct until a new milestone's
179
+ refusal scenarios left that actor holding cancelled ones — it then handed back a dead licence, read
180
+ as success, and failed a scenario two milestones older at a step that needed a live one.
181
+
182
+ **A step's input is a recorded contract, and it does not take a version.** Widening one — an extra
183
+ optional field so a step can name a particular row — trips PKU861 exactly like a function's does,
184
+ but the fix that works for a function does not work here: adding `version: 2` to a
185
+ `pikkuScenarioStep` makes the runner unable to find the step at all, and every scenario using it
186
+ fails with `Function not found`. Add a NEW step beside the old one. That is the better answer
187
+ anyway, because a step that has grown an optional field is usually two questions wearing one name —
188
+ "does she have an invoice like this" and "what became of the invoice I am holding" — and the
189
+ scenarios read better once they say which one they are asking.
190
+
191
+ ---
192
+
193
+ ## Browser scenarios
194
+
195
+ **A click returns before its effect lands — assert the effect, then navigate.** A browser step's
196
+ click resolves when the button was pressed, not when the mutation it fired came back. So a step that
197
+ presses "Add to cart" and the next one that opens `/app/cart` are in a race with the `onSuccess`
198
+ that writes the cart token to `localStorage`, and the loser arrives at an empty basket. It passed on
199
+ the first run here and failed on the second, which is the worst way to find out. Put an assertion on
200
+ the confirmation between them — the button's own "Added", the toast, the row that appeared — so the
201
+ navigation waits on the write instead of on luck. This is also why a browser scenario that is only
202
+ clicks reads better than it tests: every `when` that writes wants a `then` before the next page.
203
+
204
+ **A control the browser cannot NAME is a control it cannot drive.** A testid is derived from a
205
+ message key at build time, which has two consequences that only show up when a scenario tries to
206
+ press something:
207
+
208
+ - A label chosen at runtime — `label={isCancel ? m.a() : m.b()}` — derives no key at all, so the
209
+ field is unreachable and the failure reads as a missing element rather than a conditional. Write
210
+ the two controls out separately, each with its own static call.
211
+ - Every row of a list carries the SAME keys, so `pause` on a list of twelve licenses is twelve
212
+ matches and a strict-mode violation. Scoping by text does not save it either, because a button's
213
+ own text is "Pause" and not the row's. Give the row an address of its own —
214
+ `data-testid={id.slice(0, 8)}` on the card, rendered beside the title so a person can read it too
215
+ — and address the control `within` it.
216
+
217
+ Both are screen defects before they are test defects: a field whose label changes identity under it,
218
+ and a list whose rows are indistinguishable to anyone on the phone to support.
219
+
220
+ **A route nested under an existing screen is unreachable until its parent renders an `Outlet`.** A
221
+ milestone added `/app/academy/$slug` under an `/app/academy` that already had a component of its
222
+ own; the parent swallowed the child, so the editor's URL rendered the list — every link, every route
223
+ file and every type check looked right, and the only thing that noticed was a browser scenario that
224
+ OPENED the child path and found the parent's controls on screen. When a milestone deepens a path an
225
+ earlier one already owns, split the parent into a layout (`Outlet`) and an `index` route in the same
226
+ change, and make one scenario open the child by path rather than reach it by clicking.
227
+
228
+ **Never hard-code the target's origin.** A raw-HTTP step — a webhook, a check-in a scanner posts —
229
+ runs in the CLI process, not on the server, so it has to be told where to post. That is
230
+ `wire.scenarioStep.env.apiUrl`, which carries the environment and whatever `--api-url` overrode it.
231
+ A literal `http://localhost:3000` does not merely break on another port: it silently posts into
232
+ whatever else is listening there, so the scenario goes green having never touched this app at all.
233
+
234
+ ---
235
+
236
+ ## Running them
237
+
238
+ ```sh
239
+ bunx --bun pikku scenario run local --spawn # server-side, the fast path
240
+ bunx --bun pikku scenario run local --spawn --run browser # the same journeys, driven as a human
241
+ ```
242
+
243
+ In a multi-app project that one run covers both frontends: each persona carries
244
+ its own `app` and `@pikku/playwright` resolves the base url from the
245
+ environment's `appUrls` map, so there is no second environment to run.
246
+
247
+ `--spawn` starts and stops the server for the run; drop it if `bun run dev` is already up. The
248
+ browser pass needs the environment's `appUrl` and a browser driver installed — without them the run
249
+ fails fast rather than half-running.
250
+
251
+ **Run the whole suite, not the milestone's own scenarios.** The milestone's scenarios are the ones
252
+ you wrote to pass; the regression lives in someone else's. Tightening what "archived" means is a
253
+ one-function change that reads as local and quietly breaks the milestone-01 scenario nobody re-ran.
254
+
255
+ **Restart the server after adding a function.** Hot reload does not register a new RPC and does not
256
+ re-run `afterStart`, so a fresh function answers 404 and anything provisioned at boot is missing —
257
+ failures that read like a wiring bug and are nothing but a stale process.
258
+
259
+ ---
260
+
261
+ ## Coverage — which functions have actually been run
262
+
263
+ Green scenarios tell you the journeys you wrote still work. They say nothing about the code you
264
+ never wrote a journey for, and that gap is invisible without measuring it:
265
+
266
+ ```sh
267
+ bunx --bun pikku dev --coverage # server, instrumented
268
+ bunx --bun pikku scenario run local --coverage # against that server
269
+ ```
270
+
271
+ That writes `coverage/scenario-coverage.json` — which functions each journey exercised. **A function
272
+ no scenario touches has never been run by anything but you, by hand, once.** It compiles, it
273
+ typechecks, `pikku all` is happy, and nobody has proven it does what it says.
274
+
275
+ Run it **as each milestone closes**, not once at the end. Coverage read per milestone is a short
276
+ list you can act on — the milestone you just built either covered its own functions or it did not.
277
+ Read for the first time after ten milestones it is a wall of red that nobody triages, and the honest
278
+ response to a wall of red is to ignore it.
279
+
280
+ Every gap is one of three things, and naming which is the point of looking:
281
+
282
+ - **A missing scenario** — the function matters and no journey reaches it. Write the journey.
283
+ Refusal paths dominate this category, because it is the case you are least likely to have clicked
284
+ through by hand.
285
+ - **A function that should not exist** — nothing reaches it because nothing needs it. Delete it. An
286
+ unused exposed function is also reachable over `POST /rpc/:rpcName`, so this is a security finding,
287
+ not only dead weight.
288
+ - **Genuinely deferred** — real, not yet reachable from the UI. Say so in the milestone note that
289
+ will cover it, so the gap is a decision rather than a hole.
290
+
291
+ Report the number when you hand the milestone over. A number nobody says out loud is a number nobody
292
+ acts on.
@@ -132,7 +132,7 @@ exactly where it gets shipped past. Run the browser pass, and run it **for every
132
132
  environment in `pikkufabric.config.json`**, not just the first:
133
133
 
134
134
  ```sh
135
- bunx --bun pikku scenario run local-admin --spawn --run browser
135
+ bunx --bun pikku scenario run local --spawn --run browser
136
136
  ```
137
137
 
138
138
  `bun run build` is what type-checks each frontend (each app's `tsc` script runs
@@ -115,7 +115,7 @@ flags; the "Read" column is the skill that teaches the thing, where one does.
115
115
  | `doc` | The installed API surface | this skill |
116
116
  | `meta` / `info` | What the project declares, machine- and human-readable | `pikku-meta` |
117
117
  | `validate` | Every check that applies — app structure, an addon's published file set | `pikku-build`, `pikku-addon` |
118
- | `versions` / `semver` | Contract hashes, breaking-change detection, the release semver | `pikku-meta` |
118
+ | `versions` / `release` | Contract hashes, breaking-change detection, versioned releases | `pikku-meta` |
119
119
  | `audit` / `update` | Advisories; which `@pikku/*` can move and what peers that needs | `pikku-meta` |
120
120
  | `scopes` / `roles` | Declared authorization scopes; roles from `defineSystemRole` | `pikku-auth` |
121
121
  | `knowledge` | The knowledge base — what this app is, in its users' language | `pikku-knowledge` |
@@ -376,6 +376,12 @@ returns 0, which tells you nothing about whether it worked:
376
376
  | 3 | the deployment is blocked and nothing the CLI can do will unblock it |
377
377
  | 4 | the wait hit `--timeout` with the deployment still in flight |
378
378
 
379
+ On a failure or timeout, `apply` prints the tail of the builder's own log (and
380
+ carries `buildLog` / `imageBuildLog` on the `--json` result). Read it before
381
+ touching code: `fabric logs` serves the running stage, not the build. When the
382
+ builder recorded nothing, the CLI says so — that is usually fabric-side, so run
383
+ `pikku fabric smoke` before assuming the project is broken.
384
+
379
385
  Fabric parks every deploy at a gate after the plan phase (`status: suspended`).
380
386
  Why it parked is the whole story, and it is `statusReason`, not `status`:
381
387
 
@@ -433,6 +439,21 @@ the one the branch tracks; a stale `origin` left over from scaffolding blocks
433
439
  the deploy with "local HEAD … ≠ remote …" even though your code is pushed.
434
440
  `git branch --set-upstream-to=<remote>/main main` before deploying.
435
441
 
442
+ ### The first user on a deployed stage
443
+
444
+ When sign-up is off, the first account has to come from outside the app.
445
+ `pikku fabric user add <email>` is the CLI form of the console's Add user:
446
+
447
+ ```bash
448
+ pikku fabric user add ada@example.com --name Ada # prompts; blank generates one
449
+ pikku fabric user add ada@example.com --password '<pw>' -b staging
450
+ ```
451
+
452
+ It mints a short-lived operator token for the stage and calls the stage's own
453
+ `admin:createUser`, so the stage must wire `@pikku/addon-admin` as `admin` —
454
+ a 404 is refused by name and nothing is created. A generated password is printed
455
+ once; one you passed is never echoed. With no TTY, pass `--password` or pipe it.
456
+
436
457
  ## Versioning
437
458
 
438
459
  Functions with `expose: true` are versioned via `versions.pikku.json`. When you change a function's input or output schema, you must bump its version number — otherwise `pikku all` will report a breaking change and callers' generated clients become stale.