@pikku/skills 0.12.4 → 0.12.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/dist/skills.gen.js +1 -1
  2. package/package.json +1 -1
  3. package/skills/pikku-addon/SKILL.md +74 -33
  4. package/skills/pikku-ai-agent/SKILL.md +197 -105
  5. package/skills/pikku-ai-vercel/SKILL.md +57 -18
  6. package/skills/pikku-ai-voice/SKILL.md +126 -52
  7. package/skills/pikku-audit/SKILL.md +35 -13
  8. package/skills/pikku-aws/SKILL.md +66 -16
  9. package/skills/pikku-backblaze/SKILL.md +44 -11
  10. package/skills/pikku-better-auth/SKILL.md +45 -10
  11. package/skills/pikku-cli/SKILL.md +67 -18
  12. package/skills/pikku-cli/references/complete-example.md +2 -0
  13. package/skills/pikku-concepts/SKILL.md +75 -10
  14. package/skills/pikku-config/SKILL.md +56 -14
  15. package/skills/pikku-cron/SKILL.md +13 -6
  16. package/skills/pikku-deploy-azure/SKILL.md +83 -28
  17. package/skills/pikku-deploy-cloudflare/SKILL.md +79 -37
  18. package/skills/pikku-deploy-express/SKILL.md +40 -4
  19. package/skills/pikku-deploy-fastify/SKILL.md +22 -1
  20. package/skills/pikku-deploy-lambda/SKILL.md +99 -19
  21. package/skills/pikku-deploy-nextjs/SKILL.md +49 -5
  22. package/skills/pikku-deploy-uws/SKILL.md +54 -1
  23. package/skills/pikku-deps/SKILL.md +29 -8
  24. package/skills/pikku-emails/SKILL.md +36 -5
  25. package/skills/pikku-fabric/SKILL.md +27 -2
  26. package/skills/pikku-fabric-debug/SKILL.md +5 -1
  27. package/skills/pikku-feature/SKILL.md +6 -1
  28. package/skills/pikku-gateway-slack/SKILL.md +72 -11
  29. package/skills/pikku-http/SKILL.md +18 -5
  30. package/skills/pikku-http/references/http-options.md +10 -5
  31. package/skills/pikku-i18n/SKILL.md +18 -7
  32. package/skills/pikku-info/SKILL.md +18 -8
  33. package/skills/pikku-jose/SKILL.md +35 -6
  34. package/skills/pikku-knowledge/SKILL.md +50 -7
  35. package/skills/pikku-kysely/SKILL.md +78 -15
  36. package/skills/pikku-machine-auth/SKILL.md +36 -1
  37. package/skills/pikku-mcp/SKILL.md +159 -149
  38. package/skills/pikku-middleware/SKILL.md +17 -5
  39. package/skills/pikku-mongodb/SKILL.md +10 -2
  40. package/skills/pikku-n8n-import/SKILL.md +14 -6
  41. package/skills/pikku-permissions/SKILL.md +102 -22
  42. package/skills/pikku-pino/SKILL.md +12 -4
  43. package/skills/pikku-product-second-opinion/SKILL.md +3 -3
  44. package/skills/pikku-queue/SKILL.md +45 -16
  45. package/skills/pikku-react/SKILL.md +41 -14
  46. package/skills/pikku-react-query/SKILL.md +14 -10
  47. package/skills/pikku-realtime/SKILL.md +44 -22
  48. package/skills/pikku-redis/SKILL.md +12 -3
  49. package/skills/pikku-rpc/SKILL.md +23 -12
  50. package/skills/pikku-rtl/SKILL.md +21 -17
  51. package/skills/pikku-scenario/SKILL.md +141 -76
  52. package/skills/pikku-schedule/SKILL.md +39 -6
  53. package/skills/pikku-schema-ajv/SKILL.md +24 -2
  54. package/skills/pikku-schema-cfworker/SKILL.md +22 -2
  55. package/skills/pikku-security/SKILL.md +54 -9
  56. package/skills/pikku-services/SKILL.md +49 -9
  57. package/skills/pikku-services/references/audit-wire-service.md +2 -1
  58. package/skills/pikku-template-clone/SKILL.md +10 -5
  59. package/skills/pikku-trigger/SKILL.md +50 -6
  60. package/skills/pikku-versioning/SKILL.md +46 -17
  61. package/skills/pikku-websocket/SKILL.md +72 -44
  62. package/skills/pikku-workflow/SKILL.md +123 -11
  63. package/skills/pikku-workflow/references/workflow-reference.md +13 -8
  64. package/skills/pikku-workflows-client/SKILL.md +13 -6
  65. package/skills/pikku-ws/SKILL.md +44 -8
@@ -36,14 +36,24 @@ See `pikku-concepts` for the core mental model.
36
36
 
37
37
  ### RPC Methods (on `wire.rpc`)
38
38
 
39
- Four ways to call functions via RPC:
40
-
41
39
  | Method | Purpose |
42
40
  | -------------------------------- | ----------------------------------------- |
43
41
  | `rpc.invoke(name, data)` | Internal call to any wired function |
44
42
  | `rpc.remote(name, data)` | Remote call via DeploymentService |
45
43
  | `rpc.exposed(name, data)` | Call functions marked with `expose: true` |
46
- | `rpc.startWorkflow(name, input)` | Start a workflow |
44
+ | `rpc.startWorkflow(name, input)` | Start a workflow (see `pikku-workflow`) |
45
+ | `rpc.agent.run/stream(...)` | Run an AI agent (see `pikku-ai-agent`) |
46
+ | `rpc.agent.resume/approve(...)` | Answer a tool-approval interrupt |
47
+ | `rpc.agent.interrupt(runId)` | Stop an in-flight run |
48
+
49
+ `rpc.invoke`, `rpc.remote` and `rpc.startWorkflow` are typed off the generated
50
+ RPC map, so the name and the payload are checked. `rpc.exposed` is deliberately
51
+ `(name: string, data: any) => Promise<any>` — it exists to dispatch a name that
52
+ arrived from outside, which by definition cannot be checked at compile time.
53
+ Reach for `rpc.invoke` whenever the name is known statically.
54
+
55
+ `rpc` also carries `depth` (how deep the current RPC chain is, so runaway
56
+ recursion is visible) and `global`.
47
57
 
48
58
  ### Exposed Functions
49
59
 
@@ -61,17 +71,18 @@ const greet = pikkuSessionlessFunc({
61
71
 
62
72
  ### HTTP RPC Endpoint
63
73
 
64
- Expose all `expose: true` functions over HTTP:
74
+ The `POST /rpc/:rpcName` endpoint that dispatches every `expose: true` function
75
+ is **generated, not hand-written**. Turn it on and let codegen own it:
65
76
 
66
- ```typescript
67
- wireHTTP({
68
- route: '/rpc/:rpcName',
69
- method: 'post',
70
- auth: false,
71
- func: rpcCaller,
72
- })
77
+ ```bash
78
+ pikku enable rpc # sets scaffold.rpc = true (auth required)
79
+ pikku enable rpc --noAuth # sets scaffold.rpc = { auth: false } (public)
73
80
  ```
74
81
 
82
+ This writes `rpc-public.gen.ts` with an `rpcCaller` function and its `wireHTTP`
83
+ call already wired. Do not write that wiring yourself — a hand-rolled copy
84
+ collides with the generated route on the same path.
85
+
75
86
  ## Usage Patterns
76
87
 
77
88
  ### Internal Function Composition
@@ -114,7 +125,7 @@ RPC calls go through Pikku's middleware and permission pipeline. Direct imports
114
125
  After `npx pikku all`:
115
126
 
116
127
  ```typescript
117
- import { pikkuRPC } from '.pikku/pikku-rpc.gen.js'
128
+ import { pikkuRPC } from '#pikku/pikku-rpc.gen.js'
118
129
 
119
130
  pikkuRPC.setServerUrl('http://localhost:4002')
120
131
 
@@ -6,10 +6,11 @@ installGroups: [core]
6
6
 
7
7
  # Pikku RTL (Arabic + English)
8
8
 
9
- This skill sits **on top of** `pikku-i18n`. That skill maps a locale to `t()`
10
- tokens; this one adds the second axis: a locale also has a **direction**.
11
- Arabic is not special-cased — it is just another locale file (`ar.json`,
12
- registered `satisfies typeof en`) plus the document being told it is `rtl`.
9
+ This skill sits **on top of** `pikku-i18n`. That skill compiles a locale's
10
+ messages into typed `m.*()` functions; this one adds the second axis: a locale
11
+ also has a **direction**. Arabic is not special-cased — it is just another
12
+ `messages/ar.json` listed in `project.inlang/settings.json`, plus the document
13
+ being told it is `rtl`.
13
14
 
14
15
  ## The one idea
15
16
 
@@ -21,9 +22,9 @@ things right and Arabic, Hebrew, Farsi and Urdu all work with zero per-component
21
22
 
22
23
  ## Agent Operating Procedure
23
24
 
24
- 1. **Tokens first.** Every visible string is already a `t()` token via
25
- `pikku-i18n`. Arabic copy goes in `i18n/ar.json`, mirroring `en.json`'s keys,
26
- registered with `satisfies typeof en` so a missing key is a compile error.
25
+ 1. **Messages first.** Every visible string is already an `m.*()` message via
26
+ `pikku-i18n`. Arabic copy goes in `messages/ar.json`, mirroring `en.json`'s
27
+ keys with the `{param}` names kept identical.
27
28
  2. **Add the direction helper** to the i18n config (one home for locale→dir):
28
29
  ```ts
29
30
  const RTL_LOCALES = new Set(['ar', 'he', 'fa', 'ur'])
@@ -82,7 +83,7 @@ matching `dir` on `<html>`.
82
83
 
83
84
  ```tsx
84
85
  import { DirectionProvider, MantineProvider } from '@mantine/core'
85
- import i18n, { detectLocale, localeDir } from './i18n/config'
86
+ import { detectLocale, localeDir } from './i18n/config'
86
87
 
87
88
  const locale =
88
89
  typeof window !== 'undefined' ? detectLocale(window.location.pathname) : 'en'
@@ -137,8 +138,9 @@ const html = `<!doctype html>
137
138
  </html>`
138
139
  ```
139
140
 
140
- i18next's active language must match: call `i18n.changeLanguage(locale)` before
141
- `renderToString` so the SSR'd text and `dir` agree.
141
+ Paraglide's active locale must match: set it (via the i18n config's
142
+ `setActiveLocale` / `overwriteGetLocale` bridge) before `renderToString`, so the
143
+ SSR'd text and `dir` agree.
142
144
 
143
145
  ### Next.js app-router (test-harness next-ssr / next-static)
144
146
 
@@ -188,16 +190,17 @@ Prefer logical icon components if your icon set ships them.
188
190
  which is inconsistent. Add an Arabic-capable family (e.g. _Noto Sans Arabic_,
189
191
  _IBM Plex Sans Arabic_) to `font-family` so both scripts look intentional.
190
192
  - **Numerals:** don't hardcode digits. Format numbers/dates with
191
- `Intl.NumberFormat`/`Intl.DateTimeFormat` (or i18next formatters) given the
192
- active locale, so Western vs Arabic-Indic digits follow the locale choice.
193
+ `Intl.NumberFormat`/`Intl.DateTimeFormat` given the active locale, so Western
194
+ vs Arabic-Indic digits follow the locale choice.
193
195
  - **Line height:** Arabic diacritics sit tall — a slightly larger `line-height`
194
196
  on Arabic body text avoids clipping. Keep it locale-scoped, not global.
195
197
 
196
198
  ## Adding Arabic to an existing app — checklist
197
199
 
198
- 1. `i18n/ar.json` mirroring `en.json`; register
199
- `ar: { translation: ar satisfies typeof en }` and add `'ar'` to
200
- `supportedLocales`. (Type-complete or it won't compile — the deploy blocks.)
200
+ 1. `messages/ar.json` mirroring `en.json`; add `"ar"` to `locales` in
201
+ `project.inlang/settings.json` and recompile. Keys missing from `ar.json`
202
+ fall back to the base locale per message rather than failing the build, so
203
+ diff the two files rather than trusting `tsc` to catch a gap here.
201
204
  2. Confirm the `localeDir` helper includes `ar` (it does by default).
202
205
  3. Confirm the root sets `dir` from the locale (recipe above).
203
206
  4. Sweep the app's styles: replace every `left/right`, `ml/mr`, `text-align:
@@ -215,5 +218,6 @@ left` with the flow-relative equivalent; revert any manual `row-reverse`.
215
218
  per-locale `if (rtl)` layout branches. Set `dir` once; let layout follow.
216
219
  - Don't set `dir` on individual components — it belongs on `<html>` so the whole
217
220
  document (and Mantine) agrees.
218
- - Don't translate Arabic copy outside the `t()` token system; an RTL language is
219
- a normal locale, governed by `pikku-i18n`.
221
+ - Don't translate Arabic copy outside the message system; an RTL language is a
222
+ normal locale, governed by `pikku-i18n`. There is no `t()` and no i18next in a
223
+ Pikku frontend — the string comes from `m.some__key()`.
@@ -5,7 +5,7 @@ description: >-
5
5
  test coverage. A scenario (pikkuScenario) drives the app the way users do — steps run as actors
6
6
  over the real transport against a running server — so a flow doubles as an e2e test and a
7
7
  staged/production health check. Covers scenario.do / expectEventually / expectError /
8
- expectService, declared steps via pikkuScenarioStep (including browser steps driven by
8
+ expectService / expectScore, declared steps via pikkuScenarioStep (including browser steps driven by
9
9
  @pikku/playwright) written as intent rather than as clicks, with the actions factored into
10
10
  shared browser utilities, actors and environments in pikku.config.json, SCENARIO_ACTOR_SECRET, the
11
11
  `pikku scenario list|run` commands, live function coverage via `pikku dev --coverage`, and
@@ -36,7 +36,7 @@ A scenario is a `pikkuScenario` export that drives the app **as real actors over
36
36
  Consequences that matter, and bite if ignored:
37
37
 
38
38
  - **There is no state reset.** A scenario runs against a live server. Scope what you create (unique ids, your own rows) and never assume a clean database.
39
- - **Every effect runs as somebody, or as a declared step.** `scenario.do(...)` without `{ actor }` throws `Scenario tried to run '<rpc>' as an internal step…` — there is no bare internal-RPC step. The other way to do work is `scenario.step/given/when/then`, which runs a `pikkuScenarioStep`; its actor is optional (setup steps have none) unless it declares `browser: true`.
39
+ - **Every effect runs as somebody, or as a declared step.** `scenario.do(...)` without `{ actor }` throws `Scenario tried to run '<rpc>' as an internal step…` — there is no bare internal-RPC step. The other way to do work is `scenario.given/when/then`, which runs a `pikkuScenarioStep`; its actor is optional (setup steps have none) unless it declares `browser: true`.
40
40
  - **Actors must be configured and signed in**, or the scenario cannot run.
41
41
 
42
42
  Scenarios live in `srcDirectories` like any other function — by convention `*.scenario.ts`.
@@ -85,18 +85,51 @@ A scenario takes the same config fields as a workflow (`title`, `description`, `
85
85
 
86
86
  ### The scenario API
87
87
 
88
- | Call | Purpose |
89
- | ------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- |
90
- | `scenario.do(step, rpc, data, { actor })` | Run an RPC as that actor. The step name is what appears in the run output. |
91
- | `scenario.expectEventually(step, rpc, data, predicate, { actor, within, interval })` | Poll until `predicate(out)` passes or `within` elapses. For anything asynchronous — queues, workers, eventual state. |
92
- | `scenario.expectError(step, rpc, data, { actor, matches })` | Assert the call **fails**. For fault injection and negative paths. |
93
- | `scenario.expectService(step, 'service.method', { actor, calledWith })` | Assert a stubbed service was called. Requires the server to run with `--test`. |
94
- | `scenario.step(step, stepName, data, { actor })` | Run a declared `pikkuScenarioStep`. `given`/`when`/`then` are the same call with a keyword in the rendered prose. |
88
+ | Call | Purpose |
89
+ | ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- |
90
+ | `scenario.do(step, rpc, data, { actor })` | Run an RPC as that actor. The step name is what appears in the run output. |
91
+ | `scenario.expectEventually(step, rpc, data, predicate, { actor, within, interval })` | Poll until `predicate(out)` passes or `within` elapses. For anything asynchronous — queues, workers, eventual state. |
92
+ | `scenario.expectError(step, rpc, data, { actor, matches })` | Assert the call **fails**. For fault injection and negative paths. |
93
+ | `scenario.expectService(step, 'service.method', { actor, calledWith })` | Assert a stubbed service was called. Requires the server to run with `--test`. |
94
+ | `scenario.expectScore(step, runId, scorer, { atLeast, atMost, reference })` | Grade a finished agent run with a declared scorer and assert the score. See below. |
95
+ | `scenario.given(stepName, step, data, { actor })` | Run a declared `pikkuScenarioStep` as setup. `when` is the same call; `then` also makes the step's bindings witnesses. |
96
+ | `scenario.runScheduledTask(name)` | Fire a wired scheduler on the target now, rather than waiting for its cron. |
95
97
 
96
98
  `expectEventually` is **scenario-only**. Calling it from a `pikkuWorkflowFunc` is a critical inspector error (`PKU675`) pointing you at `pikkuScenario`.
97
99
 
98
100
  Prefer `expectEventually` over sleeping.
99
101
 
102
+ ### Asserting on an agent's answer (`expectScore`)
103
+
104
+ An agent's output is not comparable to a fixed string, so it is graded rather
105
+ than matched. Declare the rubric with `pikkuAIScorer` (grades in code) or
106
+ `pikkuAIJudge` (grades with a model) in a `*.scorer.ts` file, name it on the
107
+ agent's `scorers`, then assert on the run the scenario just triggered:
108
+
109
+ ```typescript
110
+ const { runId } = await scenario.when('asks for a summary', 'runAssistant', {
111
+ prompt: data.prompt,
112
+ }, { actor: actors.user })
113
+
114
+ await scenario.expectScore('answered briefly', runId, 'brevity', { atLeast: 0.8 })
115
+ ```
116
+
117
+ The default bound is `atLeast: 0.5`, so an unqualified `expectScore` still fails
118
+ a run the scorer graded zero. `atMost` is for a rubric where high is the failure
119
+ (sycophancy, verbosity). `reference` supplies the answer key a
120
+ `requiresReference` judge grades against — live traffic has none, so such a
121
+ judge is only ever reachable from a scenario.
122
+
123
+ Grading goes through the `pikkuScenarioGradeRun` instrumentation RPC on the
124
+ server under test, which grades from the snapshot the runtime kept at the end of
125
+ the run. Two consequences: the run must have happened on **that** server and be
126
+ recent, and the grade is returned to the scenario rather than recorded — a
127
+ test's score never lands among the production figures. Sampling is ignored, so a
128
+ scorer set to grade 1% of live traffic still grades every scenario run.
129
+
130
+ Tag any scenario whose scorer is a judge `ai-live`: it costs a model call, and
131
+ the default suite excludes it.
132
+
100
133
  ### Setup and teardown (`before` / `after`)
101
134
 
102
135
  A scenario config takes `before` and `after`. Both have the **same signature as `func`** — `(services, data, wire)` — with the return value discarded:
@@ -111,7 +144,9 @@ export const credentialScenario = pikkuScenario({
111
144
  tags: ['scenario', 'credential'],
112
145
  before: resetsCredentials,
113
146
  after: removesInstalledAddon,
114
- func: async (services, data, { scenario, actors }) => { /* … */ },
147
+ func: async (services, data, { scenario, actors }) => {
148
+ /* … */
149
+ },
115
150
  })
116
151
  ```
117
152
 
@@ -124,7 +159,7 @@ export const credentialScenario = pikkuScenario({
124
159
  | Neither runs when the run is suspended or waiting — teardown only fires at a terminal outcome. |
125
160
  | Hooks are **not** ladder rows. The runner records nothing for them; a failure is labelled by phase. |
126
161
 
127
- A hook reaches the app the same way the body does: through `wire.actors`. If you want cleanup to be *visible* on the ladder, make it an ordinary `scenario.then(...)` instead.
162
+ A hook reaches the app the same way the body does: through `wire.actors`. If you want cleanup to be _visible_ on the ladder, make it an ordinary `scenario.then(...)` instead.
128
163
 
129
164
  Hooks are scenario-only. A `before`/`after` on a `pikkuWorkflowFunc` never runs — a workflow is durable and resumable, so a callback that reran on every replay would have no honest meaning.
130
165
 
@@ -155,14 +190,14 @@ export const credentialFeature = pikkuFeature({
155
190
  })
156
191
  ```
157
192
 
158
- | Rule |
159
- | ---------------------------------------------------------------------------------------------------------------------------------------- |
160
- | The **export identifier is the feature's id**; `name` is the human-readable label. Both must be exported or the build fails. |
161
- | A `{ scenario, data }` entry is gherkin's `Examples:` — one run per entry. `data` is typed against that scenario's input. |
162
- | Feature hooks run **once around the whole group** (`before → a → b → c → after`), _not_ per scenario. `after` runs in a `finally`. |
163
- | There is deliberately **no `Background:`**. Per-scenario setup is the scenario's own `before`, referencing a shared function. |
164
- | A scenario's effective tags are its own **plus** the feature's, so `--tags credential` selects through the feature. |
165
- | A scenario need not belong to a feature — one with no input still runs standalone. |
193
+ | Rule |
194
+ | --------------------------------------------------------------------------------------------------------------------------------------------- |
195
+ | The **export identifier is the feature's id**; `name` is the human-readable label. Both must be exported or the build fails. |
196
+ | A `{ scenario, data }` entry is gherkin's `Examples:` — one run per entry. `data` is typed against that scenario's input. |
197
+ | Feature hooks run **once around the whole group** (`before → a → b → c → after`), _not_ per scenario. `after` runs in a `finally`. |
198
+ | There is deliberately **no `Background:`**. Per-scenario setup is the scenario's own `before`, referencing a shared function. |
199
+ | A scenario's effective tags are its own **plus** the feature's, so `--tags credential` selects through the feature. |
200
+ | A scenario need not belong to a feature — one with no input still runs standalone. |
166
201
  | Membership is resolved by **object identity** at runtime, which is why a loop works and why a scenario built inline in a feature is an error. |
167
202
 
168
203
  The **feature is the run unit**: `--flows` on a scenario whose every feature entry carries `data` errors and names the features containing it, because the feature is what supplies that data. Use `--features` for those. A scenario referenced bare anywhere, or in no feature at all, still runs standalone.
@@ -171,14 +206,14 @@ The **feature is the run unit**: `--flows` on a scenario whose every feature ent
171
206
 
172
207
  A scenario records what someone was **trying to do**, never the keystrokes they used to do it. This is the one decision that determines whether a suite survives its first redesign, and it applies to every step name you write.
173
208
 
174
- | Action ladder — wrong | Intent ladder — right |
175
- | -------------------------------------- | ---------------------------------------------- |
176
- | `Given opens /shop` | `Given the shopper is browsing the shop` |
177
- | `When clicks the category filter` | `When the shopper buys the £5 strawberry milkshake` |
178
- | `And clicks "Drinks"` | `Then it is in their basket` |
179
- | `And clicks the first product card` | |
180
- | `And clicks Add to basket` | |
181
- | `Then sees "1 item"` | |
209
+ | Action ladder — wrong | Intent ladder — right |
210
+ | ----------------------------------- | --------------------------------------------------- |
211
+ | `Given opens /shop` | `Given the shopper is browsing the shop` |
212
+ | `When clicks the category filter` | `When the shopper buys the £5 strawberry milkshake` |
213
+ | `And clicks "Drinks"` | `Then it is in their basket` |
214
+ | `And clicks the first product card` | |
215
+ | `And clicks Add to basket` | |
216
+ | `Then sees "1 item"` | |
182
217
 
183
218
  Three things go wrong with the left-hand column, and all three are expensive:
184
219
 
@@ -188,11 +223,11 @@ Three things go wrong with the left-hand column, and all three are expensive:
188
223
 
189
224
  So there are three layers, and only two of them are named in the report:
190
225
 
191
- | Layer | What it is | On the ladder |
192
- | ------------------------------ | --------------------------------------------------- | ------------- |
193
- | Scenario | The flow, written as intents | yes — the ladder |
194
- | Step (`pikkuScenarioStep`) | One intent | yes — one row |
195
- | Utility | An ordinary TS function over `browser` | no |
226
+ | Layer | What it is | On the ladder |
227
+ | -------------------------- | -------------------------------------- | ---------------- |
228
+ | Scenario | The flow, written as intents | yes — the ladder |
229
+ | Step (`pikkuScenarioStep`) | One intent | yes — one row |
230
+ | Utility | An ordinary TS function over `browser` | no |
196
231
 
197
232
  Utilities are **not steps**. They are plain exported functions, they take the browser handle, and they hold the clicking:
198
233
 
@@ -262,7 +297,7 @@ export const buysTheItem = pikkuScenarioStep<
262
297
  })
263
298
  ```
264
299
 
265
- The two bindings are **alternatives**: `pikku scenario run --run browser` clicks through the shop, `--run default` (the fast suite) takes the server-side path, and both report the same sentence.
300
+ The bindings are **alternatives**: `pikku scenario run --run browser` clicks through the shop, `--run cli` drives it over the websocket, `--run default` (the fast suite, and the default) takes the server-side path — and all of them report the same sentence.
266
301
 
267
302
  ```typescript
268
303
  await scenario.when(
@@ -274,9 +309,9 @@ await scenario.when(
274
309
  // reporter renders: When the shopper buys the £5 strawberry milkshake ✓ 1.2s
275
310
  ```
276
311
 
277
- **Every intent step begins by arriving.** `ensureOnShop` is not defensive noise — it is what lets a scenario start at any step, run alone, and be reordered without touching it. It checks first and navigates only if needed, so a scenario already on the shop pays nothing. This is about the *browser's* starting position, not the database: there is still no state reset (see above), and you still scope what you create.
312
+ **Every intent step begins by arriving.** `ensureOnShop` is not defensive noise — it is what lets a scenario start at any step, run alone, and be reordered without touching it. It checks first and navigates only if needed, so a scenario already on the shop pays nothing. This is about the _browser's_ starting position, not the database: there is still no state reset (see above), and you still scope what you create.
278
313
 
279
- **The same utilities, a different intent.** A scenario about filtering has filtering as its subject, so there the filter *is* the intent — same helper, its own step:
314
+ **The same utilities, a different intent.** A scenario about filtering has filtering as its subject, so there the filter _is_ the intent — same helper, its own step:
280
315
 
281
316
  ```typescript
282
317
  export const filtersTheShop = pikkuScenarioStep<
@@ -299,7 +334,7 @@ export const filtersTheShop = pikkuScenarioStep<
299
334
  })
300
335
  ```
301
336
 
302
- Two scenarios, two intents, one set of utilities. That is the shape to aim for: when a helper is reused by a step whose *subject* it is, promote it to a step there — never the reverse.
337
+ Two scenarios, two intents, one set of utilities. That is the shape to aim for: when a helper is reused by a step whose _subject_ it is, promote it to a step there — never the reverse.
303
338
 
304
339
  **Non-browser steps need none of this.** Without a browser there is no navigation to absorb and no DOM to hide, so an intent maps to one RPC and `scenario.do` names it directly:
305
340
 
@@ -320,11 +355,11 @@ This is the one place the surface bindings do **not** behave like a switch, and
320
355
 
321
356
  On a `given` or `when`, the bindings are alternatives — clicking Buy and calling `createOrder` are two ways to cause one effect, so exactly one runs.
322
357
 
323
- On a `then`, they are not two implementations of one assertion. They are two *different claims*:
358
+ On a `then`, they are not two implementations of one assertion. They are two _different claims_:
324
359
 
325
- | binding | what it actually proves |
326
- | ------- | ----------------------- |
327
- | `default` | the order row says `paid` — the system of record is right |
360
+ | binding | what it actually proves |
361
+ | --------- | -------------------------------------------------------------- |
362
+ | `default` | the order row says `paid` — the system of record is right |
328
363
  | `browser` | the confirmation panel says paid — the truth reached the human |
329
364
 
330
365
  The gap between them is the bug nobody catches: 200 OK, database correct, user still watching a spinner. So a `then` runs **every** binding it declares and fails if they disagree.
@@ -353,7 +388,7 @@ Three rules follow, and they are the ones that get broken:
353
388
 
354
389
  - **A browser witness must observe on the page.** One that quietly calls an RPC to check the result is worse than no binding at all — it reports a tick for a surface it never looked at.
355
390
  - **Return what you observed, don't just assert.** A witness returning a value lets the runner diff the two. A witness that only throws still works, but it can never disagree with anything, so it proves less. Read structured state with `where` on the test-id selector rather than parsing translated copy.
356
- - **A step with no binding for the run's surface is counted, not excused.** `--run browser` prints `n/m steps ran on browser` over *every* step, so an action that quietly fell back to the server lowers the number just as an assertion does. A `then` that fell back is additionally named — `--strict` fails on those, because a sentence saying the actor saw something nobody looked at is a different problem from a shortcut. Not being in the UI *is* the finding: do not add a browser binding that fakes it.
391
+ - **A step with no binding for the run's surface is counted, not excused.** `--run browser` prints `n/m steps ran on browser` over _every_ step, so an action that quietly fell back to the server lowers the number just as an assertion does. A `then` that fell back is additionally named — `--strict` fails on those, because a sentence saying the actor saw something nobody looked at is a different problem from a shortcut. Not being in the UI _is_ the finding: do not add a browser binding that fakes it.
357
392
 
358
393
  **Always give a `then` a `default` witness.** It is the floor every run can fall back to, and an assertion with no witness the run can execute is fatal (`ScenarioNoWitness`) — not a coverage gap. The distinction is the point: a `then` checked server-side under `--run browser` did happen, it just wasn't seen where the prose claims; one checked nowhere never happened at all, and without the error it would return `undefined` and render as a tick. A browser-only `then` is therefore a step that fails the fast suite, which is rarely what you want.
359
394
 
@@ -376,12 +411,16 @@ export const buysAnApple = pikkuScenarioStep<
376
411
  name: 'buysAnApple',
377
412
  description: 'buys an apple',
378
413
  template: 'buys {qty} apples',
379
- func: async (_services, { qty }, { scenarioStep }) => {
414
+ default: async (_services, { qty }, { scenarioStep }) => {
380
415
  return await requireActor(scenarioStep).invoke('placeOrder', { qty })
381
416
  },
382
417
  })
383
418
  ```
384
419
 
420
+ A step's body always lives under a **surface binding** — `default`, `browser` or
421
+ `cli` — never under a `func`. Declaring none throws at load time: at minimum give
422
+ it a `default`.
423
+
385
424
  ```typescript
386
425
  await scenario.given(
387
426
  'buys an apple',
@@ -396,7 +435,7 @@ Rules that bite:
396
435
 
397
436
  - **The step is referenced by its typed string name, not by importing the const** — exactly like `workflow.do`. The name is the step's `pikkuFuncId` and is checked against the generated step map. A non-literal target is a critical error (`PKU678`).
398
437
  - **Steps are not RPCs.** They are deliberately never network-callable — a browser-driving step must not be.
399
- - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal *is* the assertion.
438
+ - **`actor.invoke` is typed over the exposed RPC map**, so the name and the payload are checked and the result comes back narrowed — no cast. `actor.invokeRaw(name, data, { headers })` is the same call reporting `{ status, ok, body }` instead of throwing; use it whenever the refusal _is_ the assertion.
400
439
  - **`actor` and `env` are optional on the wire**, because a pure assertion step needs neither. Narrow them with `requireActor(scenarioStep)` and `requireScenarioEnv(scenarioStep)` from `@pikku/core/workflow` rather than a local guard — both name the step and say what to pass. `env` is `{ apiUrl, appUrl? }` from the environment the run targets, and is how a raw-HTTP step learns the target's URL: a step runs in the CLI process, where there is no `variables` service and `process.env` is not the answer.
401
440
  - **Steps default to `retries: 0`**, unlike ordinary workflow steps. Retrying a failed assertion is wrong; pass `retries` explicitly if a step is genuinely flaky-by-nature.
402
441
  - **Step results are persisted**, so return JSON-serialisable data — never a `Locator` or a client object.
@@ -405,26 +444,24 @@ Rules that bite:
405
444
 
406
445
  ### Browser steps
407
446
 
408
- `browser` is a boolean on the step, and it is the whole switch: `browser: true` and a browser is guaranteed present on the wire, `browser: false` (the default) and there is none. There is no third state and nothing to null-check — `wire.browser` is optional in the type only for the steps that did not ask for one.
447
+ Declaring a `browser` binding is the whole switch: inside that binding `wire.browser` is guaranteed present and non-optional, and a step without one never sees a browser at all. There is nothing to null-check.
409
448
 
410
- A step declaring `browser: true` gets `wire.browser` — a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
449
+ A `browser` binding gets a session bound to **its actor**, signed in through the same `signInPath` + `SCENARIO_ACTOR_SECRET` path the HTTP actors use, so the browser and the RPC calls are one identity. Calling such a step without an actor is a critical error (`PKU677`).
411
450
 
412
451
  Browser steps are where **intent, not actions** earns its keep: the step is one intent, the clicking lives in shared utilities, and the step arrives before it acts. Write the mechanics below into utilities and keep the step body to three or four calls that read as a sentence.
413
452
 
414
453
  ```typescript
415
- export const opensTheCart = pikkuScenarioStep<
416
- { path: string },
417
- { url: string },
418
- true
419
- >({
420
- name: 'opensTheCart',
421
- description: 'opens the cart',
422
- browser: true,
423
- func: async (_services, { path }, { browser }) => {
424
- await browser.goto(path)
425
- return { url: browser.page.url() }
426
- },
427
- })
454
+ export const opensTheCart = pikkuScenarioStep<{ path: string }, { url: string }>(
455
+ {
456
+ name: 'opensTheCart',
457
+ description: 'opens the cart',
458
+ browser: async (_services, { path }, { browser }) => {
459
+ await browser.goto(path)
460
+ return { url: browser.page.url() }
461
+ },
462
+ default: async ({ rpc }) => ({ url: (await rpc.invoke('getCart', {})).url }),
463
+ }
464
+ )
428
465
  ```
429
466
 
430
467
  - Install `@pikku/playwright` and `@playwright/test`, and import `@pikku/playwright` once (`import type {} from '@pikku/playwright'`) so `browser.page` is a typed Playwright `Page`. Without it you still get the structural `goto`/`screenshot` handle.
@@ -441,8 +478,14 @@ Personas, actors and environments live in `pikku.config.json`:
441
478
  "scenarios": {
442
479
  "personas": {
443
480
  "shopper": { "description": "Buys things here", "primary": true },
444
- "support": { "description": "Answers for the shop", "proficiency": "power" },
445
- "reminders": { "description": "The shop chasing abandoned carts", "kind": "system" }
481
+ "support": {
482
+ "description": "Answers for the shop",
483
+ "proficiency": "power"
484
+ },
485
+ "reminders": {
486
+ "description": "The shop chasing abandoned carts",
487
+ "kind": "system"
488
+ }
446
489
  },
447
490
  "actors": {
448
491
  "shopper": {
@@ -487,9 +530,24 @@ SCENARIO_ACTOR_SECRET=… pikku scenario run local
487
530
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --flows orderSupportScenario
488
531
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --features credentialFeature
489
532
  SCENARIO_ACTOR_SECRET=… pikku scenario run local --tags smoke,scenario
533
+ SCENARIO_ACTOR_SECRET=… pikku scenario run local --spawn --no-browser --exclude-tags ai-live
490
534
  ```
491
535
 
492
- `run` takes the environment as a **required positional** — the key from `environments`. `--flows`/`-f` filters by scenario name, `--features` by feature id, `--tags`/`-t` by tag (match-any). Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
536
+ `run` takes the environment as a **required positional** — the key from `environments`. Every filter narrows the same plan, so narrowing a feature to two of its five scenarios still runs the feature's hooks exactly once around those two.
537
+
538
+ | Flag | Effect |
539
+ | ----------------------- | --------------------------------------------------------------------------------------- |
540
+ | `--flows` / `-f` | Comma-separated scenario names |
541
+ | `--features` | Comma-separated feature ids |
542
+ | `--tags` / `-t` | Match-any tag filter |
543
+ | `--exclude-tags` | Hold tags back — unless the flow is named directly with `--flows` |
544
+ | `--run <surface>` | `default` (the default), `browser`, or `cli` |
545
+ | `--no-browser` | Shorthand for `--run default`; scenarios with browser steps report as **skipped** |
546
+ | `--strict` | Fail, rather than pass, a `then` with no witness on the run's surface |
547
+ | `--spawn` / `--keep-alive` | Start `pikku dev` on the environment's apiUrl for the run; optionally leave it up |
548
+ | `--api-url` / `--app-url` | Override the environment's URLs — for a target that only exists at run time |
549
+ | `--trace` | Keep every stack frame on failure (default shows only the project's own) |
550
+ | `--coverage` | Reset/snapshot server coverage per scenario |
493
551
 
494
552
  Output is `PASS <name> (<ms>) → <output>` / `FAIL <name> (<ms>): <error>`, then `N/M scenarios passed against '<env>'`. A scenario inside a feature is named `<Feature> › <scenario> <data>`.
495
553
 
@@ -501,10 +559,16 @@ Coverage is attributed by running scenarios against a server that is collecting
501
559
 
502
560
  Prerequisite in `pikku.config.json`:
503
561
 
504
- ```json
505
- { "scaffold": { "scenarios": "auth" } }
562
+ ```bash
563
+ pikku enable scenarios # sets scaffold.scenarios = true (session required)
564
+ pikku enable scenarios --noAuth # sets scaffold.scenarios = { "auth": false }
506
565
  ```
507
566
 
567
+ `scaffold.scenarios` is a boolean or `{ auth?, path? }`. The legacy string forms
568
+ (`"auth"` / `"no-auth"`) are **rejected by the config loader**, not reinterpreted —
569
+ under a shape where a string could be a path, silently reading one as a flag
570
+ would be worse than failing.
571
+
508
572
  `scaffold.scenarios` generates the coverage and stub RPCs into your project (`pikkuScenarioTakeLiveCoverage`, `pikkuScenarioResetLiveCoverage`, `pikkuScenarioResetStubs`, `pikkuScenarioGetStubCalls`), so scenario runs work against any server. The coverage RPC reads `<outDir>/function/pikku-functions-meta-verbose.gen.json` off disk at request time — codegen always writes it, but it has to be deployed alongside the app or the RPC returns `null`.
509
573
 
510
574
  ```bash
@@ -565,19 +629,20 @@ Services are plain objects — a Pikku function is pure business logic, so a moc
565
629
 
566
630
  ## Red flags
567
631
 
568
- | Smell | Why it's wrong |
569
- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
570
- | `pikku tests …` | Removed in #865. Use `pikku scenario`. |
571
- | `.feature` files / Gherkin for function tests | Scenarios are TypeScript, not Gherkin. The in-process cucumber function world was deleted. |
572
- | `scenario.do(...)` with no `{ actor }` | Throws. Every step runs as somebody. |
573
- | A scenario per function | Scenarios are user flows. One flow covers many functions; that is the point. |
574
- | Assuming a clean database | There is no state reset — it may be a staging server. Scope what you create. |
575
- | `sleep()` before asserting | Use `expectEventually`. |
576
- | A step named `clicksAddToBasket` / `opensThePage` | That is an action, not an intent. Name the step for what the actor wanted; put the clicking in a utility. |
577
- | A browser step that assumes it is already on a page | It can then only run mid-flow. Arrive first — check the URL, navigate if needed. |
578
- | A `browser: true` step guarding `if (!browser)` | `browser: true` guarantees it. The guard hides the real error, which is a missing actor (`PKU677`). |
579
- | `expectEventually` in a `pikkuWorkflowFunc` | `PKU675` — scenario-only. |
580
- | Coverage silently 0 | Server not run with `--coverage`, verbose functions meta not deployed, `scaffold.scenarios` unset, or no actors configured. |
632
+ | Smell | Why it's wrong |
633
+ | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- |
634
+ | `pikku tests …` | Removed in #865. Use `pikku scenario`. |
635
+ | `.feature` files / Gherkin for function tests | Scenarios are TypeScript, not Gherkin. The in-process cucumber function world was deleted. |
636
+ | `scenario.do(...)` with no `{ actor }` | Throws. Every step runs as somebody. |
637
+ | A scenario per function | Scenarios are user flows. One flow covers many functions; that is the point. |
638
+ | Assuming a clean database | There is no state reset — it may be a staging server. Scope what you create. |
639
+ | `sleep()` before asserting | Use `expectEventually`. |
640
+ | A step named `clicksAddToBasket` / `opensThePage` | That is an action, not an intent. Name the step for what the actor wanted; put the clicking in a utility. |
641
+ | A browser step that assumes it is already on a page | It can then only run mid-flow. Arrive first — check the URL, navigate if needed. |
642
+ | A `browser` binding guarding `if (!browser)` | The binding guarantees it. The guard hides the real error, which is a missing actor (`PKU677`). |
643
+ | A step with a `func:` instead of a surface binding | There is no `func` on a step. Bodies live under `default` / `browser` / `cli`; a step with none throws at load. |
644
+ | `expectEventually` in a `pikkuWorkflowFunc` | `PKU675` — scenario-only. |
645
+ | Coverage silently 0 | Server not run with `--coverage`, verbose functions meta not deployed, `scaffold.scenarios` unset, or no actors configured. |
581
646
 
582
647
  `@pikku/cucumber` is a **browser/e2e** harness (`Actor`, `BrowserWorld`, `PersonaData`, `DbUtils`) — out of scope here.
583
648
 
@@ -36,22 +36,55 @@ yarn add @pikku/schedule
36
36
  ```typescript
37
37
  import { InMemorySchedulerService } from '@pikku/schedule'
38
38
 
39
- const scheduler = new InMemorySchedulerService()
39
+ const schedulerService = new InMemorySchedulerService()
40
+ await schedulerService.start() // registers a CronJob per wired scheduled task
40
41
  ```
41
42
 
42
- Implements the scheduler service interface. Schedules are held in memory — they do not survive process restarts. Suitable for development and single-instance deployments.
43
+ It implements core's `SchedulerService` on two mechanisms: `cron` for the
44
+ recurring tasks you declared with `wireScheduler` (see `pikku-cron`), and
45
+ `setTimeout` for one-off delayed RPCs. Both live in process memory, so nothing
46
+ survives a restart and nothing is shared between instances — fine for
47
+ development and a single-instance deployment, wrong for anything else.
48
+
49
+ `PikkuTaskScheduler` is a deprecated alias for the same class.
50
+
51
+ ### Scheduling a one-off RPC
52
+
53
+ ```typescript
54
+ const taskId = await schedulerService.scheduleRPC('5m', 'sendReminder', data, session)
55
+ await schedulerService.getTask(taskId) // { rpcName, scheduledFor, status, … } | null
56
+ await schedulerService.getAllTasks() // pending one-offs only
57
+ await schedulerService.unschedule(taskId) // true when it was still pending
58
+ ```
59
+
60
+ The delay is milliseconds or a duration string (`'30s'`, `'5m'`, `'2h'`). This is
61
+ also the mechanism a workflow's delayed steps use, which is why a workflow that
62
+ sleeps needs a `schedulerService` registered.
43
63
 
44
64
  ## Usage Patterns
45
65
 
46
66
  ### Basic Setup
47
67
 
68
+ The scheduler is a singleton service under the name **`schedulerService`**, and
69
+ it is started in your server bootstrap — declaring it without calling `start()`
70
+ registers no cron jobs, so nothing ever fires:
71
+
48
72
  ```typescript
73
+ // start.ts
49
74
  import { InMemorySchedulerService } from '@pikku/schedule'
50
75
 
51
- const createSingletonServices = pikkuServices(async (config) => {
52
- const scheduler = new InMemorySchedulerService()
53
- return { config, scheduler }
76
+ const schedulerService = new InMemorySchedulerService()
77
+ const singletonServices = await createSingletonServices(config, {
78
+ schedulerService,
54
79
  })
80
+
81
+ await appServer.start()
82
+ await schedulerService.start()
55
83
  ```
56
84
 
57
- For distributed or persistent scheduling, use BullMQ (`BullSchedulerService`) or PgBoss (`PgBossSchedulerService`) from the queue packages instead. See `pikku-queue` for details.
85
+ Call `close()` on shutdown — it stops every cron job and clears pending timers.
86
+
87
+ For distributed or persistent scheduling, take the scheduler service off the
88
+ queue factory instead (`bullFactory.getSchedulerService()`,
89
+ `pgBossFactory.getSchedulerService()`) and register it under the same name. See
90
+ `pikku-queue`.