@pikku/skills 0.12.34 → 0.12.37

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/README.md +9 -4
  2. package/dist/index.d.ts +7 -4
  3. package/dist/index.js +9 -5
  4. package/dist/skills.gen.d.ts +1 -0
  5. package/dist/skills.gen.js +5 -3
  6. package/dist/snippets.d.ts +26 -0
  7. package/dist/snippets.js +148 -0
  8. package/package.json +2 -2
  9. package/skills/pikku-addon/SKILL.md +41 -30
  10. package/skills/pikku-addon/references/addon-package-manifest.md +9 -4
  11. package/skills/pikku-addon/references/openapi.md +99 -0
  12. package/skills/pikku-agent/references/agents.md +3 -1
  13. package/skills/pikku-architect/SKILL.md +12 -0
  14. package/skills/pikku-auth/references/better-auth.md +16 -0
  15. package/skills/pikku-build/SKILL.md +29 -0
  16. package/skills/pikku-build/references/app.md +52 -4
  17. package/skills/pikku-build/references/feature.md +23 -96
  18. package/skills/pikku-build/references/quick.md +12 -3
  19. package/skills/pikku-changes/SKILL.md +172 -0
  20. package/skills/pikku-concepts/SKILL.md +33 -138
  21. package/skills/pikku-concepts/references/bootstrap.md +58 -0
  22. package/skills/pikku-concepts/references/concept-mapping.md +16 -0
  23. package/skills/pikku-concepts/references/language.md +87 -0
  24. package/skills/pikku-deploy/SKILL.md +1 -1
  25. package/skills/pikku-fabric/SKILL.md +26 -13
  26. package/skills/pikku-guide/SKILL.md +264 -0
  27. package/skills/pikku-knowledge/SKILL.md +10 -0
  28. package/skills/pikku-kysely/SKILL.md +1 -1
  29. package/skills/pikku-mantine/SKILL.md +80 -0
  30. package/skills/pikku-n8n-import/SKILL.md +4 -3
  31. package/skills/pikku-react/references/client.md +12 -0
  32. package/skills/pikku-realtime/SKILL.md +6 -6
  33. package/skills/pikku-report/SKILL.md +143 -0
  34. package/skills/pikku-scenario/SKILL.md +71 -563
  35. package/skills/pikku-scenario/references/browser.md +59 -0
  36. package/skills/pikku-scenario/references/coverage.md +70 -0
  37. package/skills/pikku-scenario/references/personas.md +87 -0
  38. package/skills/pikku-scenario/references/steps.md +366 -0
  39. package/skills/pikku-service-backends/SKILL.md +1 -1
  40. package/skills/pikku-wiring/SKILL.md +1 -1
  41. package/skills/pikku-wiring/references/http.md +8 -0
  42. package/skills/pikku-wiring/references/mcp.md +59 -0
  43. package/skills/pikku-workflow/SKILL.md +7 -8
@@ -0,0 +1,59 @@
1
+ # Browser steps
2
+
3
+ Declaring a `browser` binding is the whole switch: inside that binding `wire.browser` is guaranteed present and non-optional, and a step without one never sees a browser at all. There is nothing to null-check.
4
+
5
+ A `browser` binding gets a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
6
+
7
+ Browser steps are where **intent, not actions** earns its keep: the step is one intent, the clicking lives in shared utilities, and the step arrives before it acts. Write the mechanics below into utilities and keep the step body to three or four calls that read as a sentence.
8
+
9
+ ```typescript
10
+ export const opensTheCart = pikkuScenarioStep<
11
+ { path: string },
12
+ { url: string }
13
+ >({
14
+ name: 'opensTheCart',
15
+ description: 'opens the cart',
16
+ browser: async (_services, { path }, { browser }) => {
17
+ await browser.goto(path)
18
+ return { url: browser.page.url() }
19
+ },
20
+ default: async (_services, _data, { actor }) => ({
21
+ url: (await actor.invoke('getCart', {})).url,
22
+ }),
23
+ })
24
+ ```
25
+
26
+ - Install `@pikku/playwright` and `@playwright/test`, and import `@pikku/playwright` once (`import type {} from '@pikku/playwright'`) so `browser.page` is a typed Playwright `Page`. Without it you still get the structural `goto`/`screenshot` handle.
27
+ - The environment needs an `appUrl` beside its `apiUrl`. `pikku scenario run` fails fast before running anything if a browser scenario has no `appUrl` or the driver is not installed.
28
+ - A browser step only launches a browser under `--run browser`; under the default surface it takes its default path instead. A scenario with no binding for the run's surface and no default is reported as **could not run** and fails the run — hold it back with `--exclude-tags`, not with the expectation of a silent skip.
29
+ - Playwright auto-waits; do not wrap `page.click` in `expectEventually`.
30
+
31
+ ## Locate by message key, never by rendered copy
32
+
33
+ If the app is translated, **no step may contain a user-visible string.** `getByLabel('Full Name')` passes only while the browser happens to render the base locale, and any copy edit turns it into a selector timeout that points at the wizard rather than at the rename that caused it — the test looks broken where it is merely stale.
34
+
35
+ The message catalogue already holds the string under a key. Read it from there. Type the lookup off the catalogue JSON so a renamed or misspelled key is a **compile** error rather than a run-time timeout:
36
+
37
+ ```typescript
38
+ // tests/scenarios/i18n.ts
39
+ import type messages from '../../../../apps/web/messages/en.json'
40
+
41
+ export type MessageKey = keyof typeof messages
42
+
43
+ export const t = (key: MessageKey, locale = baseLocale): string => {
44
+ /* … */
45
+ }
46
+ ```
47
+
48
+ ```typescript
49
+ await page.getByLabel(t('jobs_apply_fullname')).fill(identity.name)
50
+ await page
51
+ .getByRole('button', { name: t('jobs_apply_submit'), exact: true })
52
+ .click()
53
+ ```
54
+
55
+ - Type off `messages/<baseLocale>.json`, **not** the generated Paraglide output — `i18n/paraglide/` is build output, so typing against it makes the tests unbuildable until the app has been built. The JSON is the tracked source.
56
+ - Fall back to the base locale for a key a locale has not translated. That is what Paraglide does at run time, so a helper that throws instead would disagree with the screen the test is looking at.
57
+ - This is not only about locators. A copy literal passed to a **project helper** (`pick('Where would you like to work?', …)`) reaches the DOM the same way, and so does a pane name quoted back in a failure message. `pikku fabric validate` scans every string in a `*.steps.ts` / `*.scenario.ts` against the base catalogue and errors on any verbatim match, wherever it sits — except comments, and the `name` / `description` / `template` declared directly on a `pikkuFeature`, `pikkuScenario` or `pikkuScenarioStep`, which are Console meta written in the project's `locale` rather than app copy.
58
+ - A regex locator (`{ name: /^Next$/i }`) hides the literal but not the problem. `{ name: t('key'), exact: true }` is both stricter and locale-correct.
59
+ - Strings the catalogue does not own — a test id, a fixture filename, a seeded value — stay literal. The catalogue is the test for whether something is copy.
@@ -0,0 +1,70 @@
1
+ # Coverage
2
+
3
+ Coverage is attributed by running scenarios against a server that is collecting it. It is **not** derived from unit tests.
4
+
5
+ Prerequisite in `pikku.config.json`:
6
+
7
+ ```bash
8
+ pikku enable scenarios # sets scaffold.scenarios = true
9
+ ```
10
+
11
+ `scaffold.scenarios` is a boolean or `{ path? }` — whether the surface exists
12
+ and where it is written. A bare string is **rejected by the config loader**, not
13
+ reinterpreted: under a shape where a string could be a path, silently reading
14
+ one as a flag would be worse than failing.
15
+
16
+ `scaffold.scenarios` generates the coverage and stub RPCs into your project (`pikkuScenarioTakeLiveCoverage`, `pikkuScenarioResetLiveCoverage`, `pikkuScenarioResetStubs`, `pikkuScenarioGetStubCalls`), so scenario runs work against any server. The coverage RPC reads `<outDir>/function/pikku-functions-meta-verbose.gen.json` off disk at request time — codegen always writes it, but it has to be deployed alongside the app or the RPC returns `null`.
17
+
18
+ ```bash
19
+ pikku dev --coverage # V8 precise coverage, in-process
20
+ pikku dev --coverage --test # also enable stubs (needed for expectService)
21
+ SCENARIO_ACTOR_SECRET=… pikku scenario run local --coverage
22
+ ```
23
+
24
+ The run resets coverage before each scenario and snapshots after, writing **`<outDir>/coverage/scenario-coverage.json`**:
25
+
26
+ ```jsonc
27
+ {
28
+ "generatedAt": "…",
29
+ "environment": "local",
30
+ "scenarios": {
31
+ "<name>": {/* FunctionCoverageReport */},
32
+ },
33
+ }
34
+ ```
35
+
36
+ Coverage is best-effort: it disables itself with a warning if the server is not collecting or the first actor cannot invoke, and it needs at least one configured actor. If you get no coverage, check those first.
37
+
38
+ **There is no AI-prompt output.** The old `--ai-out` flag died with `pikku tests`; nothing replaced it. To find what needs work, read `scenario-coverage.json` yourself and cross-reference `pikku meta functions list` for input/output schemas.
39
+
40
+ ## Filling coverage
41
+
42
+ 1. `pikku scenario run <env> --coverage`, then read `<outDir>/coverage/scenario-coverage.json` to see what is unexercised.
43
+ 2. `pikku meta functions list` for those functions' schemas.
44
+ 3. Write a `pikkuScenario` that reaches them **through a real user flow** with an actor — not a scenario per function. Scenarios are flows; coverage is a consequence.
45
+ 4. Re-run to confirm.
46
+
47
+ ## Unit tests for pure logic
48
+
49
+ Scenarios are the repo-idiomatic way to test functions, and the only thing that contributes to live coverage. For pure logic with heavy branching, a plain unit test calling `func` directly is still valid and cheap:
50
+
51
+ ```typescript
52
+ import { describe, test } from 'node:test'
53
+ import assert from 'node:assert'
54
+
55
+ describe('createTodo', () => {
56
+ test('creates a todo', async () => {
57
+ const services = {
58
+ todoStore: { add: async (title: string) => ({ id: '1', title }) },
59
+ }
60
+ const result = await createTodo.func(services as any, { title: 'Buy milk' })
61
+ assert.equal(result.title, 'Buy milk')
62
+ })
63
+ })
64
+ ```
65
+
66
+ ```bash
67
+ node --import tsx --test src/**/*.test.ts
68
+ ```
69
+
70
+ Services are plain objects — a Pikku function is pure business logic, so a mock is just the shape the function destructures. Build real services via the `pikkuServices` / `pikkuWireServices` factories when a test needs them.
@@ -0,0 +1,87 @@
1
+ # Personas and actors
2
+
3
+ A **persona** is a person your product is for; an **actor** is one body that signs in as them. Every entry materialises exactly one actor, so `actors.<id>` exists for each declared persona and there is no second way to declare a login.
4
+
5
+ - Two people of the same kind are two entries, not one persona with two logins — "you see yours, not theirs" is only testable with two customers.
6
+ - Never write an email address: each is derived from the persona id and `scenarios.emailDomain`, and a hand-written one signs in as somebody who was never created.
7
+ - `roles` is typechecked against `defineSystemRole`; an undeclared role is a build error.
8
+ - A person who is only ever acted *upon* — the account an admin bans — sets `runnable: false`: declared and seeded, never signed in, because a run as them would race the scenario that acts on them.
9
+ - A persona holds only what is true of that kind of person for the app's whole lifetime (`name`, `jobTitle`, `description`, `personality`, `roles`, `goals`, `disposition`). What someone is trying to get done, and the circumstances they are doing it in, belong to the **scenario**, not to them.
10
+
11
+ ## Declaring personas in TypeScript
12
+
13
+ There may be **one `definePersonas` call in the whole codebase** — one place to
14
+ read the set from, one place to add to it. A second anywhere, including in the
15
+ same file, is a critical. Generated files are exempt and never claim the slot.
16
+
17
+ > [!WARNING]
18
+ > The declaration is **read from source, never evaluated** — the CLI writes it
19
+ > to JSON that a deployed stage carries without the app. So every value has to
20
+ > be statically knowable, and a value that is not comes out as `undefined`
21
+ > rather than as an error. Only `name` is checked, so a computed `personality`,
22
+ > `jobTitle` or `description` is dropped in silence and the persona runs with a
23
+ > blank temperament.
24
+
25
+ What that admits and what it does not:
26
+
27
+ ```typescript
28
+ personality: 'Wound up and short with it.' // read
29
+ personality: `Wound up and short with it.
30
+ Says what she wants in a few blunt words.` // read — no ${} in it
31
+ personality: 'Wound up. ' + 'Short with it.' // dropped, silently
32
+ personality: TEMPERAMENTS.impatient // dropped, silently
33
+ ```
34
+
35
+ A no-substitution template literal is a string literal as far as the reader is
36
+ concerned, so it is the way to write a long personality across several lines —
37
+ not a concatenation, and not a `prettier-ignore`d single line. Its newlines and
38
+ leading indentation are kept verbatim and reach the model that way, which is
39
+ harmless but worth knowing before you align it to the surrounding code.
40
+
41
+ One more thing worth knowing before writing a rich persona: **`actor.converse`
42
+ builds its prompt from `name`, `jobTitle`, `personality` and the scenario's
43
+ `task` only.** Fields like `disposition`, `goals` and `roles` are read and
44
+ stored, and the console shows them, but they do not reach the conversing
45
+ persona's instructions. Anything that must shape how someone talks belongs in
46
+ `personality` or in the task.
47
+
48
+ A project that never declares a persona keeps working: a scenario that names no actor needs none.
49
+
50
+ - `environments.<name>.apiUrl` is required. `signInPath` defaults to `/auth/sign-in/actor`, `rpcPath` to `/rpc`.
51
+ - **`SCENARIO_ACTOR_SECRET` is an environment variable and never goes in `pikku.config.json`.** It signs actors in. `pikku scenario run` throws without it; a server auto-building actors warns and runs without them.
52
+
53
+ ## The same actors sign a human in
54
+
55
+ Declared actors are not only for automated runs. `signInPath` is Better Auth's
56
+ `actor` plugin (see `pikku-auth`, a separate install), which any caller can post to — so the
57
+ frontend gets a one-click "Sign in as …" switcher over the **same** list, and an
58
+ app can be reviewed as each kind of user without anyone knowing a seed password.
59
+
60
+ The sandbox dev server bakes both halves into the frontend from the declared
61
+ personas: `VITE_DEV_ACTORS` (the JSON actor list) and `VITE_DEV_ACTOR_SECRETS`
62
+ (`{ email: credential }`, one per persona — `SCENARIO_ACTOR_SECRET` itself never
63
+ goes in a bundle; see **pikku-auth**). Neither var is set in a production
64
+ build, so the control renders nothing there — but gate the reads on your
65
+ bundler's dev flag anyway (`import.meta.env.DEV ? … : undefined`) so no
66
+ credential reaches a production bundle in the first place.
67
+
68
+ Do not hand-roll the switcher: `useDevActors()` (`pikku-react`, a separate install) is the logic and
69
+ `<DevActorSwitcher />` from `@pikku/mantine/dev` is a ready rendering of it.
70
+ `pikku fabric validate` **requires** any frontend with a login screen to ship
71
+ one — without it a reviewer is locked out of their own sandbox.
72
+ When the switcher is missing, it is one of three things, and none of them
73
+ errors:
74
+
75
+ - **The frontend was not started by the dev script.** The two `VITE_DEV_*` vars
76
+ are computed by `bun run dev` and read by vite once, at boot. A bare `vite dev`
77
+ — including one restarted by hand — has an empty list and renders nothing.
78
+ - **`SCENARIO_ACTOR_SECRET` is not in `.env`.** No root secret, no per-persona
79
+ credentials, and the switcher filters out every actor it cannot sign in.
80
+ - **It is not mounted on the page you are looking at.** The template mounts it
81
+ on the login screen. A public homepage that replaces the `/` → `/app`
82
+ redirect needs its own `<DevActorSwitcher />` in the public layout.
83
+
84
+ When the switcher is there but signing in fails with `401 Invalid actor
85
+ secret`, check which server answered before checking the secret: a frontend
86
+ whose dev proxy (`VITE_API_PROXY`, default `http://localhost:3000`) points at
87
+ another project's API sends the sign-in there.
@@ -0,0 +1,366 @@
1
+ # Writing steps
2
+
3
+ A `pikkuScenarioStep` is the unit a scenario's ladder is made of. Everything here is about
4
+ authoring one well, because a step is the thing that survives — or does not survive — the app's
5
+ first redesign.
6
+
7
+ 1. [What a step is](#what-a-step-is)
8
+ 2. [What a step is given](#what-a-step-is-given)
9
+ 3. [Steps describe intent, not actions](#steps-describe-intent-not-actions)
10
+ 4. [`then` bindings are witnesses, not alternatives](#then-bindings-are-witnesses-not-alternatives)
11
+ 5. [What language the prose is in](#what-language-the-prose-is-in)
12
+
13
+ Browser bindings have their own file: **`references/browser.md`**.
14
+
15
+ ---
16
+
17
+ ## What a step is
18
+
19
+ `scenario.do` can only name an RPC. A **step** is a named, typed unit of scenario behaviour whose body is an ordinary pikku function — so it can call several RPCs as its actor, assert, or drive a browser.
20
+
21
+ ```typescript
22
+ import { pikkuScenarioStep } from '#pikku/scenarios'
23
+
24
+ export const buysAnApple = pikkuScenarioStep<
25
+ { qty: number },
26
+ { orderId: string }
27
+ >({
28
+ name: 'buysAnApple',
29
+ description: 'buys an apple',
30
+ template: 'buys {qty} apples',
31
+ actor: true,
32
+ default: async (_services, { qty }, { actor }) => {
33
+ return await actor.invoke('placeOrder', { qty })
34
+ },
35
+ })
36
+ ```
37
+
38
+ A step's body always lives under a **surface binding** — `default`, `browser` or
39
+ `cli` — never under a `func`. Declaring none throws at load time: at minimum give
40
+ it a `default`.
41
+
42
+ ```typescript
43
+ await scenario.given(
44
+ 'buys an apple',
45
+ 'buysAnApple',
46
+ { qty: 1 },
47
+ { actor: actors.shopper }
48
+ )
49
+ // reporter renders: Given shopper buys 1 apples ✓ 412ms
50
+ ```
51
+
52
+ Rules that bite:
53
+
54
+ - **The step is referenced by its typed string name, not by importing the const** — exactly like `workflow.do`. The name is the step's `pikkuFuncId` and is checked against the generated step map. A non-literal target is a critical error (`PKU678`).
55
+ - **Steps are not RPCs.** They are deliberately never network-callable — a browser-driving step must not be.
56
+ - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal _is_ the assertion.
57
+ - **A step that runs as somebody declares `actor: true`**, and the runner injects `wire.actor` — non-optional inside every binding, with no guard to write and nothing to unwrap. A `browser` binding implies it, because a window is opened as somebody. Leave it off for a step with no persona to be: an assertion over what an earlier step returned, or one that posts credentials precisely because it must not reuse an actor's session. Dispatching a step that declared it without `{ actor: actors.x }` fails before the body runs (`ScenarioActorRequired`); a step that did not declare it has no `actor` on its wire at all.
58
+ - **`env` is optional on the wire**, because most steps need nothing from it. Narrow it with `requireScenarioEnv(scenarioStep)` from `#pikku/scenario` rather than a local guard — it names the step and says what to pass. `env` is `{ apiUrl, appUrl? }` from the environment the run targets, and is how a raw-HTTP step learns the target's URL: a step runs in the CLI process, where there is no `variables` service and `process.env` is not the answer.
59
+ - **Steps default to `retries: 0`**, unlike ordinary workflow steps. Retrying a failed assertion is wrong; pass `retries` explicitly if a step is genuinely flaky-by-nature.
60
+ - **Step results are persisted**, so return JSON-serialisable data — never a `Locator` or a client object.
61
+ - **`description` documents the step; `template` is what the report renders.** `template`'s `{placeholders}` are filled from the input the step was called with, so one step reads differently for each call — `sees {state} addon {packageName}` reports as "sees available addon @pikku/addon-stripe". Reflect every input field in the template, and type the values so they read as words (`state?: 'installed' | 'available'`, not `installed?: boolean`). A placeholder with no value renders as nothing and the whitespace collapses.
62
+ - **Never write the actor into the prose.** The reporter renders the actor as the sentence's subject, so a step authored as `` `'sam' creates the client` `` run as `{ actor: actors.sam }` reads "Given sam 'sam' creates the client" — and the hardcoded name desyncs the moment the call site changes actor. Write a bare third-person predicate (`creates the {name} client`) and let the actor supply the subject. Prose that opens with its own actor's key — quoted, capitalised or possessive — is `PKU681`; naming someone **else** mid-sentence ("sends nadia an invite") is ordinary prose and is left alone, as is an actor keyed after a role noun used as a noun ("creates the admin client" as `actors.admin`).
63
+ - Prose precedence is `options.description` → the step's `template` → the step's own `description` → the positional step name. Repeated names get `#1`, `#2` ordinals, so a `for` loop over a data set is how you write a Scenario Outline. A loop-generated step name is not statically known, so it is matched back to its declaration by step function instead — which works as long as that function's call sites agree on their phase, actor and prose. Two call sites that disagree make the loop step report under its bare runtime name.
64
+
65
+ ## What a step is given
66
+
67
+ A step has the signature of an ordinary pikku function, which makes it look as
68
+ though it runs where the application runs. It does not — **it runs in the CLI
69
+ process**, and the services object is built there, by hand:
70
+
71
+ ```typescript
72
+ { logger, workflowService, workflowRunService, agentRunner? }
73
+ ```
74
+
75
+ That is the whole list. There is no `kysely`, no `variables`, no `secrets`, and
76
+ none of the project's own singleton or wire services. A step that destructures
77
+ one gets `undefined` and fails on first use — `Cannot read properties of
78
+ undefined (reading 'selectFrom')` — which reads like a broken container and is
79
+ not.
80
+
81
+ `rpc` is the trap worth naming, because it is present and it throws. It is a
82
+ `guardRpc` whose every member refuses:
83
+
84
+ > Scenario tried to run 'getOrder' as an internal step. Every workflow.do in a
85
+ > scenario must carry { actor: actors.x } so it executes against 'local'.
86
+
87
+ The same guard covers `rpc.agent.run/stream/resume/approve/interrupt` and
88
+ `startWorkflow`.
89
+
90
+ This is the design, not a gap: **everything a step touches of the application
91
+ goes over the wire as somebody.** A test that could reach into the database
92
+ would be testing a different program from the one a person uses. So there are
93
+ exactly three ways in, and they are all through the actor:
94
+
95
+ - `actor.invoke(name, data)` — typed over the exposed RPC map, carrying the
96
+ actor's session. Declare `actor: true` and destructure it off the wire.
97
+ - `.invokeRaw(name, data, { headers })` — same call, reporting
98
+ `{ status, ok, body }`, for when the refusal is the assertion.
99
+ - a plain `fetch` against `requireScenarioEnv(scenarioStep).apiUrl`, for
100
+ anything not an RPC — a websocket, a file upload, a webhook.
101
+
102
+ Two consequences follow, and both shape how steps get written:
103
+
104
+ - **A step cannot observe anything the app does not publish.** If a test needs a
105
+ fact the client never sees, the fix is to emit it on the stream or expose it
106
+ as an RPC — which usually improves the product, since a client debugging the
107
+ same problem needed it too.
108
+ - **`agentRunner` is conditional.** It is built only when the project declares
109
+ agents, and `createDevAgentRunner` needs a base URL _and_ a key together
110
+ (`OPENAI_BASE_URL` + `OPENAI_API_KEY`, or the LiteLLM pair). With a key alone
111
+ it returns nothing and `agentRunner` is `undefined`, so `actor.converse`
112
+ fails before the persona says anything. A suite that would rather own its own
113
+ model can pass an `llm` to `runConversation` instead of relying on this one.
114
+ ---
115
+
116
+ ## Steps describe intent, not actions
117
+
118
+ A scenario records what someone was **trying to do**, never the keystrokes they used to do it. This is the one decision that determines whether a suite survives its first redesign, and it applies to every step name you write.
119
+
120
+ | Action ladder — wrong | Intent ladder — right |
121
+ | ----------------------------------- | ----------------------------------------------- |
122
+ | `Given opens /shop` | `Given shopper is browsing the shop` |
123
+ | `When clicks the category filter` | `When shopper buys the £5 strawberry milkshake` |
124
+ | `And clicks "Drinks"` | `Then it is in their basket` |
125
+ | `And clicks the first product card` | |
126
+ | `And clicks Add to basket` | |
127
+ | `Then sees "1 item"` | |
128
+
129
+ Three things go wrong with the left-hand column, and all three are expensive:
130
+
131
+ - **A layout change rewrites every scenario that touched that screen.** In the right-hand column it rewrites one function.
132
+ - **The report is the deliverable.** `buys the £5 strawberry milkshake` is readable by someone who has never seen the app; `clicks [data-testid=add]` tells them nothing about whether the product works.
133
+ - **An action step cannot arrive on its own.** It assumes the previous click left the browser somewhere, so the scenario only runs front-to-back, as a whole, in one order.
134
+
135
+ So there are three layers, and only two of them are named in the report:
136
+
137
+ | Layer | What it is | On the ladder |
138
+ | -------------------------- | -------------------------------------- | ---------------- |
139
+ | Scenario | The flow, written as intents | yes — the ladder |
140
+ | Step (`pikkuScenarioStep`) | One intent | yes — one row |
141
+ | Utility | An ordinary TS function over `browser` | no |
142
+
143
+ Utilities are **not steps**. They are plain exported functions, they take the browser handle, and they hold the clicking:
144
+
145
+ ```typescript
146
+ // shop.browser.ts — shared actions. Not steps: nothing here is an intent.
147
+ import type { PikkuBrowserWire } from '#pikku/scenarios'
148
+ import type {} from '@pikku/playwright'
149
+
150
+ /** Arrive on the shop, from wherever the browser happens to be. */
151
+ export const ensureOnShop = async (browser: PikkuBrowserWire) => {
152
+ if (!new URL(browser.page.url()).pathname.startsWith('/shop')) {
153
+ await browser.goto('/shop')
154
+ }
155
+ await browser
156
+ .locate({ testId: 'product-grid' })
157
+ .first()
158
+ .waitFor({ state: 'visible' })
159
+ }
160
+
161
+ export const searchFor = async (browser: PikkuBrowserWire, query: string) => {
162
+ await browser.locate({ testId: 'shop-search' }).first().fill(query)
163
+ await browser.page.keyboard.press('Enter')
164
+ }
165
+
166
+ export const filterByCategory = async (
167
+ browser: PikkuBrowserWire,
168
+ category: string
169
+ ) => {
170
+ await browser.locate({ testId: 'category-filter' }).first().click()
171
+ await browser
172
+ .locate({ testId: 'category-option', where: { 'data-category': category } })
173
+ .first()
174
+ .click()
175
+ }
176
+
177
+ export const addToBasket = async (browser: PikkuBrowserWire, name: string) => {
178
+ const card = browser
179
+ .locate({ testId: 'product-card', containing: name })
180
+ .first()
181
+ await card.waitFor({ state: 'visible' })
182
+ await card.locate('[data-testid=add-to-basket]').click()
183
+ }
184
+ ```
185
+
186
+ The step composes them, and it is the step — one row — that the report shows:
187
+
188
+ ```typescript
189
+ export const buysTheItem = pikkuScenarioStep<
190
+ { name: string },
191
+ { name: string }
192
+ >({
193
+ name: 'buysTheItem',
194
+ description: 'finds one item in the shop and puts it in the basket',
195
+ template: 'buys the {name}',
196
+ // One intent, one implementation per surface an actor can drive it through.
197
+ browser: async (_services, { name }, { browser }) => {
198
+ await ensureOnShop(browser)
199
+ await searchFor(browser, name)
200
+ await addToBasket(browser, name)
201
+ return { name }
202
+ },
203
+ default: async ({ rpc }, { name }) => {
204
+ const item = await rpc.invoke('findItemByName', { name })
205
+ await rpc.invoke('addToBasket', { itemId: item.id })
206
+ return { name }
207
+ },
208
+ })
209
+ ```
210
+
211
+ The bindings are **alternatives**: `pikku scenario run --run browser` clicks through the shop, `--run cli` drives it over the websocket, `--run default` (the fast suite, and the default) takes the server-side path — and all of them report the same sentence.
212
+
213
+ ```typescript
214
+ await scenario.when(
215
+ 'buys a milkshake',
216
+ 'buysTheItem',
217
+ { name: '£5 strawberry milkshake' },
218
+ { actor: actors.shopper }
219
+ )
220
+ // reporter renders: When shopper buys the £5 strawberry milkshake ✓ 1.2s
221
+ ```
222
+
223
+ **Every intent step begins by arriving.** `ensureOnShop` is not defensive noise — it is what lets a scenario start at any step, run alone, and be reordered without touching it. It checks first and navigates only if needed, so a scenario already on the shop pays nothing. This is about the _browser's_ starting position, not the database: there is still no state reset (see above), and you still scope what you create.
224
+
225
+ **The same utilities, a different intent.** A scenario about filtering has filtering as its subject, so there the filter _is_ the intent — same helper, its own step:
226
+
227
+ ```typescript
228
+ export const filtersTheShop = pikkuScenarioStep<
229
+ { category: string },
230
+ { shown: number }
231
+ >({
232
+ name: 'filtersTheShop',
233
+ description: 'narrows the catalogue to one category',
234
+ template: 'filters the shop by {category}',
235
+ browser: async (_services, { category }, { browser }) => {
236
+ await ensureOnShop(browser)
237
+ await filterByCategory(browser, category)
238
+ return {
239
+ shown: await browser.locate({ testId: 'product-card' }).count(),
240
+ }
241
+ },
242
+ default: async ({ rpc }, { category }) => ({
243
+ shown: (await rpc.invoke('listItems', { categorySlug: category })).length,
244
+ }),
245
+ })
246
+ ```
247
+
248
+ Two scenarios, two intents, one set of utilities. That is the shape to aim for: when a helper is reused by a step whose _subject_ it is, promote it to a step there — never the reverse.
249
+
250
+ **Non-browser steps need none of this.** Without a browser there is no navigation to absorb and no DOM to hide, so an intent maps to one RPC and `scenario.do` names it directly:
251
+
252
+ ```typescript
253
+ const order = await scenario.do(
254
+ 'Shopper checks out',
255
+ 'createOrder',
256
+ { basketId, shippingAddress },
257
+ { actor: actors.shopper }
258
+ )
259
+ ```
260
+
261
+ Reach for a `pikkuScenarioStep` on the non-browser side only when one intent genuinely spans several RPCs, or when the step asserts something the RPC result alone does not say.
262
+ ---
263
+
264
+ ## `then` bindings are witnesses, not alternatives
265
+
266
+ This is the one place the surface bindings do **not** behave like a switch, and it is the part worth reading twice.
267
+
268
+ On a `given` or `when`, the bindings are alternatives — clicking Buy and calling `createOrder` are two ways to cause one effect, so exactly one runs.
269
+
270
+ On a `then`, they are not two implementations of one assertion. They are two _different claims_:
271
+
272
+ | binding | what it actually proves |
273
+ | --------- | -------------------------------------------------------------- |
274
+ | `default` | the order row says `paid` — the system of record is right |
275
+ | `browser` | the confirmation panel says paid — the truth reached the human |
276
+
277
+ The gap between them is the bug nobody catches: 200 OK, database correct, user still watching a spinner. So a `then` runs **every** binding it declares and fails if they disagree.
278
+
279
+ ```typescript
280
+ export const seesTheOrderConfirmed = pikkuScenarioStep<
281
+ { orderId: string },
282
+ { status: string }
283
+ >({
284
+ name: 'seesTheOrderConfirmed',
285
+ // Both bindings run as the persona, so the step declares one and the runner
286
+ // injects `wire.actor` — non-optional in every binding.
287
+ actor: true,
288
+ template: 'sees order {orderId} confirmed',
289
+ // Both run on `--run browser`. Each returns what it observed, and the runner
290
+ // compares them — so this fails when the page disagrees with the database.
291
+ browser: async (_services, { orderId }, { browser }) => ({
292
+ status: await browser
293
+ .locate({ testId: 'order-status', where: { 'data-order': orderId } })
294
+ .getAttribute('data-status'),
295
+ }),
296
+ // Through the actor, not through a `rpc` service — see "What a step is given".
297
+ default: async (_services, { orderId }, { actor }) => ({
298
+ status: (await actor.invoke('getOrder', { orderId })).status,
299
+ }),
300
+ })
301
+ ```
302
+
303
+ Three rules follow, and they are the ones that get broken:
304
+
305
+ - **A browser witness must observe on the page.** One that quietly calls an RPC to check the result is worse than no binding at all — it reports a tick for a surface it never looked at.
306
+ - **Return what you observed, don't just assert.** A witness returning a value lets the runner diff the two. A witness that only throws still works, but it can never disagree with anything, so it proves less. Read structured state with `where` on the test-id selector rather than parsing translated copy.
307
+ - **A step with no binding for the run's surface is counted, not excused.** `--run browser` prints `n/m steps ran on browser` over _every_ step, so an action that quietly fell back to the server lowers the number just as an assertion does. A `then` that fell back is additionally named — `--strict` fails on those, because a sentence saying the actor saw something nobody looked at is a different problem from a shortcut. Not being in the UI _is_ the finding: do not add a browser binding that fakes it.
308
+
309
+ **Always give a `then` a `default` witness.** It is the floor every run can fall back to, and an assertion with no witness the run can execute is fatal (`ScenarioNoWitness`) — not a coverage gap. The distinction is the point: a `then` checked server-side under `--run browser` did happen, it just wasn't seen where the prose claims; one checked nowhere never happened at all, and without the error it would return `undefined` and render as a tick. A browser-only `then` is therefore a step that fails the fast suite, which is rarely what you want.
310
+
311
+ **Every scenario must assert.** A flow of only `given`/`when` is a PKU680 critical — it proves nothing threw. Since coverage counts every step, an assertion-free ladder of browser-bound actions would score a perfect `3/3` while checking nothing, so clicking through the UI and never looking at the result is the cheapest way to fake the number. The rule closes that.
312
+
313
+ Assertions with no possible browser witness are a different thing and should not be written as a `then`: "the audit log recorded it" is a system check, and "the receipt email arrives" is `expectEventually`, which is always out-of-band and always server-side.
314
+ ---
315
+
316
+ ## What language the prose is in
317
+
318
+ A scenario carries two kinds of text, and they do not share a language.
319
+
320
+ **Identifiers are English.** The exported const (`buysAnApple`,
321
+ `credentialFeature`), the step's `name` — which is its `pikkuFuncId`, the typed
322
+ string the generated step map is keyed by — the file name, and every helper in
323
+ `*.browser.ts`. These bind to generated code and to `pikku scenario list`; they
324
+ are English in every project regardless of who the product is for or what
325
+ language the team speaks. There is no setting that changes this.
326
+
327
+ **Prose follows `metaLocale` in `pikku.config.json`** (default `en`). That is a
328
+ step's `description` and `template`, a feature's `name` and `description`, a
329
+ scenario's `title`, and the positional step names passed to
330
+ `scenario.given/when/then`. Read the field before you write any of them.
331
+
332
+ This split is the same one the feature table already states — _the export
333
+ identifier is the feature's id; `name` is the human-readable label_ — applied to
334
+ language. The report is the deliverable, and it is read by the team; the
335
+ identifier is an API, and it is read by the toolchain.
336
+
337
+ ```typescript
338
+ // pikku.config.json: { "metaLocale": "de" }
339
+ export const buysAnApple = pikkuScenarioStep<
340
+ { qty: number },
341
+ { orderId: string }
342
+ >({
343
+ name: 'buysAnApple', // identifier — English, always
344
+ description: 'kauft einen Apfel', // prose — follows locale
345
+ template: 'kauft {qty} Äpfel', // prose — follows locale
346
+ actor: true,
347
+ default: async (_services, { qty }, { actor }) =>
348
+ await actor.invoke('placeOrder', { qty }),
349
+ })
350
+ ```
351
+
352
+ Note what does **not** change: `placeOrder` is still `placeOrder`, and the file
353
+ is still `apple.scenario.ts`.
354
+
355
+ A product with a non-English UI is not on its own a reason to set `metaLocale` — that
356
+ is the app's language, not the team's. Ask, or leave it `en`.
357
+
358
+ **Where a non-`en` `metaLocale` still shows English, today.** The reporter composes a
359
+ sentence as `<Keyword> <actor> <template>` (`composeStepProse`), and the keyword is
360
+ an English literal. The Console translates the Given/When/Then keywords into its own
361
+ UI language; the CLI reporter does not, so `metaLocale: "de"` gives you German step
362
+ prose inside an English frame — `Given shopper kauft 1 Äpfel`. Write templates that read
363
+ acceptably in that frame rather than trying to defeat it. A second gap: where a
364
+ function or scenario declares no `title`, the Console falls back to splitting the
365
+ **identifier** into an English-looking label (`toEnglishName`), so under a
366
+ non-`en` `metaLocale` meta is worth authoring rather than leaving to the fallback.
@@ -23,7 +23,7 @@ Use this skill as an execution checklist, not reference material.
23
23
  1. Discover before editing. Run the relevant `pikku meta ... --json` command and inspect only the focused output you need.
24
24
  2. Identify the source files that own the behavior. Do not start by reading generated output, `.pikku`, `node_modules`, vendored packages, or broad build artifacts.
25
25
  3. Make the smallest source change that satisfies the task. Keep generated files generated, and avoid hand-editing SDKs, schema output, or typegen.
26
- 4. Validate with the narrowest relevant command first, then run `pikku-verify` or `pikku all` when functions, wirings, schemas, or generated clients may have changed.
26
+ 4. Validate with the narrowest relevant command first, then run `pikku all` when functions, wirings, schemas, or generated clients may have changed.
27
27
  5. If validation fails, fix the source cause and rerun validation. Do not paper over generated errors by editing generated files.
28
28
 
29
29
  Constructor shapes and method signatures come from `pikku doc` — run
@@ -23,7 +23,7 @@ Use this skill as an execution checklist, not reference material.
23
23
  1. Discover before editing. Run the relevant `pikku meta ... --json` command and inspect only the focused output you need.
24
24
  2. Identify the source files that own the behavior. Do not start by reading generated output, `.pikku`, `node_modules`, vendored packages, or broad build artifacts.
25
25
  3. Make the smallest source change that satisfies the task. Keep generated files generated, and avoid hand-editing SDKs, schema output, or typegen.
26
- 4. Validate with the narrowest relevant command first, then run `pikku-verify` or `pikku all` when functions, wirings, schemas, or generated clients may have changed.
26
+ 4. Validate with the narrowest relevant command first, then run `pikku all` when functions, wirings, schemas, or generated clients may have changed.
27
27
  5. If validation fails, fix the source cause and rerun validation. Do not paper over generated errors by editing generated files.
28
28
 
29
29
  ## Before you start
@@ -147,6 +147,14 @@ it is optional because the same function can be reached over plain HTTP or RPC,
147
147
  where there is no stream to send on. The `if (channel)` guard is what lets one
148
148
  function serve both; the return value is the non-streaming answer.
149
149
 
150
+ A function that throws once the stream is open cannot answer with a status code,
151
+ so the runner ends the stream with `{ type: 'error', errorText }` then
152
+ `{ type: 'done' }`. A route whose client parses a different event protocol says
153
+ so with `streamProtocol`, and the failure is written in that one instead —
154
+ `streamProtocol: 'agui'` ends the stream with a single AG-UI `RUN_ERROR` and
155
+ nothing after it. The default is `'pikku'`; the generated agent stream routes
156
+ set `'agui'`.
157
+
150
158
  ### Generated Fetch Client
151
159
 
152
160
  After `npx pikku all`, a type-safe client is generated: