@alvera-ai/platform-sdk 0.17.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/.agent/AGENTS.md +82 -144
  2. package/.agent/account_management.md +2 -2
  3. package/.agent/action_logs.md +4 -4
  4. package/.agent/ai_agents.md +28 -21
  5. package/.agent/ai_sandbox.md +49 -39
  6. package/.agent/connected_apps.md +3 -3
  7. package/.agent/cookbook/_fixtures/README.md +1 -1
  8. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_generic_table.liquid +1 -1
  9. package/.agent/cookbook/_fixtures/organic-marketing/_lead_submissions_foundation_legal_entity.liquid +80 -0
  10. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_mdm.liquid +2 -1
  11. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_generic_table.liquid +57 -0
  12. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_legal_entity.liquid +30 -0
  13. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_mdm.liquid +44 -0
  14. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_generic_table.liquid +57 -0
  15. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_legal_entity.liquid +36 -0
  16. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_mdm.liquid +41 -0
  17. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_generic_table.liquid +70 -0
  18. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_legal_entity.liquid +52 -0
  19. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_mdm.liquid +42 -0
  20. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_generic_table.liquid +38 -0
  21. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_legal_entity.liquid +48 -0
  22. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_mdm.liquid +49 -0
  23. package/.agent/cookbook/organic-marketing.md +2801 -0
  24. package/.agent/cookbook/payments-compliance.md +2180 -0
  25. package/.agent/cookbook/primary-care.md +2175 -0
  26. package/.agent/cookbook/subscription-saas.md +2403 -0
  27. package/.agent/data_activation_clients.md +65 -52
  28. package/.agent/datalakes.md +338 -171
  29. package/.agent/errors.md +3 -3
  30. package/.agent/generic_tables.md +151 -62
  31. package/.agent/interoperability_contracts.md +57 -22
  32. package/.agent/mdm.md +136 -153
  33. package/.agent/messages.md +36 -34
  34. package/.agent/mock-services.md +1 -1
  35. package/.agent/mutations.md +2 -2
  36. package/.agent/templates.md +14 -13
  37. package/.agent/tools.md +63 -21
  38. package/.agent/type_naming.md +13 -13
  39. package/.agent/workflows.md +99 -53
  40. package/README.md +2 -2
  41. package/dist/bin/platform-sdk.mjs +33 -47
  42. package/dist/bin/platform-sdk.mjs.map +1 -1
  43. package/dist/index.d.mts +565 -379
  44. package/dist/index.d.mts.map +1 -1
  45. package/dist/index.mjs +494 -59
  46. package/dist/index.mjs.map +1 -1
  47. package/package.json +4 -3
  48. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_legal_entity.liquid +0 -88
  49. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_appointment.liquid +0 -47
  50. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_mdm.liquid +0 -24
  51. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_patient.liquid +0 -38
  52. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_compliance_screening.liquid +0 -59
  53. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_mdm.liquid +0 -36
  54. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_mdm.liquid +0 -30
  55. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_payment_account.liquid +0 -55
  56. package/.agent/cookbook/_fixtures/subscription/_customers_subscription_mdm.liquid +0 -20
  57. package/.agent/cookbook/_setup/foundation.md +0 -359
  58. package/.agent/cookbook/_setup/healthcare.md +0 -361
  59. package/.agent/cookbook/_setup/payments.md +0 -365
  60. package/.agent/cookbook/_setup/subscription.md +0 -364
  61. package/.agent/cookbook/action-status-updaters.md +0 -278
  62. package/.agent/cookbook/ai-agent-invoke.md +0 -279
  63. package/.agent/cookbook/appointment-review-sms-workflow.md +0 -801
  64. package/.agent/cookbook/birthday-greeting-sms-trigger.md +0 -696
  65. package/.agent/cookbook/bulk-ingest.md +0 -302
  66. package/.agent/cookbook/contact-us-triage-with-llm.md +0 -663
  67. package/.agent/cookbook/dunning-sms-for-delinquent.md +0 -659
  68. package/.agent/cookbook/generic-tables.md +0 -244
  69. package/.agent/cookbook/invite-team.md +0 -200
  70. package/.agent/cookbook/kyc-notification-on-account-activation.md +0 -661
  71. package/.agent/cookbook/marketing-campaign-send.md +0 -1044
  72. package/.agent/cookbook/paginated-restapi-poller.md +0 -383
  73. package/.agent/cookbook/rest-fetch.md +0 -273
  74. package/.agent/cookbook/sanctions-screening-with-agent-review.md +0 -773
  75. package/.agent/cookbook/score-leads-with-llm-categorization.md +0 -665
  76. package/.agent/cookbook/system-templates.md +0 -165
  77. package/.agent/cookbook/talk-to-data.md +0 -178
  78. package/.agent/cookbook/triage-prospects-by-priority.md +0 -571
  79. package/.agent/cookbook/welcome-sms-for-customers.md +0 -647
  80. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/memorandum-of-association-01.png +0 -0
  81. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/sample_two_page.pdf +0 -0
  82. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/_customers_subscription_customer.liquid +0 -0
  83. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/stripe_customers_batch1.csv +0 -0
@@ -0,0 +1,2403 @@
1
+ ---
2
+ title: "Subscription SaaS: one customer spine, three workflows, three ways in"
3
+ summary: "The whole subscription-billing surface in one walk. Stand up a tenant and a raw datalake, declare one customers table, and run three workflows over it — a welcome SMS gated on reachability, a dunning reminder gated on reachability and KYC, and an LLM agent that bands accounts by priority and fans out one SMS per band. Then fill the same table three more ways — a bulk CSV upload, a pull-based REST fetch, and finally a teammate invited onto the tenant."
4
+ use_case: subscription-saas
5
+ slug: subscription-saas
6
+ vitest_source:
7
+ - integration-tests/tests/subscription-saas/standard-workflow.test.ts
8
+ - integration-tests/tests/subscription-saas/dunning-sms-workflow.test.ts
9
+ - integration-tests/tests/subscription-saas/agent-lead-triage.test.ts
10
+ - integration-tests/tests/subscription-saas/run-dac-bulk.test.ts
11
+ - integration-tests/tests/subscription-saas/run-dac-fetch.test.ts
12
+ - integration-tests/tests/subscription-saas/invite-team.test.ts
13
+ - integration-tests/tests/workspace/bootstrap.test.ts
14
+ status: draft
15
+ ---
16
+
17
+ # Problem
18
+
19
+ A subscription business talks to its customers on a schedule set by the
20
+ billing system. A contract is signed and a welcome message goes out with a
21
+ link to the self-serve portal. An invoice goes unpaid and a reminder goes
22
+ out with a link to pay. An account needs chasing and *how* it is chased
23
+ depends on what kind of account it is — a white-glove note to an enterprise,
24
+ a standard nudge to a self-serve individual.
25
+
26
+ Three jobs, one audience. The classic implementation is three services that
27
+ each query a customers table and each carry their own copy of the
28
+ reachability rule, the KYC rule and the account-tier rule. This walk builds
29
+ all three on **one customers table**, because the customer is the same
30
+ customer in all three and the rules belong to the platform rather than to
31
+ three codebases.
32
+
33
+ The last three scenarios are about how rows get *in*. A row can arrive
34
+ one at a time through an inline ingest, four thousand at a time through a
35
+ file upload, or be pulled from someone else's API on demand — and all three
36
+ land in the same table through the same contract pair. The walk ends by
37
+ inviting a second person onto the tenant, because a billing team is rarely
38
+ one person.
39
+
40
+ **This surface stays raw.** No tokenized or redacted copy is provisioned.
41
+ Every read-back here is a question about this walk's own rows, and the
42
+ priority agent classifies on account type rather than on anyone's name. The
43
+ PII columns are still declared `tokenize` — that declaration is the policy,
44
+ and it goes live the moment someone provisions the pair. If you want a model
45
+ held away from identities today, that is what `payments-compliance` §007
46
+ shows.
47
+
48
+ # Composition
49
+
50
+ | Resource | Why it exists here |
51
+ |-----------------------------------|-------------------------------------------|
52
+ | Tenant + raw datalake | The billing team's own world |
53
+ | One customers generic table | The spine all six scenarios write to |
54
+ | SMS tool, LLM tool | Shared by every workflow below |
55
+ | Contract pair, on two clients | Identity and table row, deliberately split|
56
+ | Three workflows | Welcome, dunning, priority triage |
57
+ | Two connected apps | The portal link, and the pay link |
58
+ | Manual-upload, bulk and REST paths| Three ways into the same table |
59
+ | An invitation | A second person on the tenant |
60
+
61
+ # Walkthrough
62
+
63
+ This cookbook is **self-reliant**: it stands up everything it uses and
64
+ depends on no other file. It is also **idempotent** — every creating step
65
+ looks first and creates only what is missing, so running it twice costs what
66
+ running it once cost. That is not a nicety. Nothing in the platform reclaims
67
+ an abandoned datalake, and a walk that mints a fresh tenant per run leaves
68
+ every previous run's lakes behind forever.
69
+
70
+ ## 001 — sign in as root, and make sure the admin user exists
71
+
72
+ Authenticate as the platform's root admin (`admin@dev.local` /
73
+ `devpassword` in local dev) via the tenantless bootstrap login — keyless by
74
+ structural necessity, since no tenant exists yet to scope a key to, and a
75
+ dev/test-only surface.
76
+
77
+ Then make sure the admin user this walk runs as exists. **The email is
78
+ stable, not per-run**, which is what makes the step idempotent — and it means
79
+ the second run finds the user already signed up. A duplicate signup is
80
+ refused, so the refusal is caught and inspected: if it says the email is
81
+ taken, that is the idempotent path and the walk continues. Any other failure
82
+ is re-raised, because swallowing it would turn a real auth problem into a
83
+ confusing failure three steps later.
84
+
85
+ ```typescript
86
+ ctx.rootSession = await createBootstrapSession({
87
+ baseUrl: process.env.ALVERA_BASE_URL!,
88
+ email: process.env.ALVERA_ROOT_EMAIL!,
89
+ password: process.env.ALVERA_ROOT_PASSWORD!,
90
+ })
91
+ ctx.rootApi = createIsolatedPlatformApi({
92
+ baseUrl: process.env.ALVERA_BASE_URL!,
93
+ sessionToken: ctx.rootSession.sessionToken,
94
+ apiKey: '',
95
+ })
96
+
97
+ ctx.billingEmail = 'cookbook-subscription-saas@dev.local'
98
+ ctx.billingPassword = 'CookbookPass1!'
99
+
100
+ try {
101
+ const signUpResp = await ctx.rootApi.admin.signUp({
102
+ email: ctx.billingEmail,
103
+ password: ctx.billingPassword,
104
+ first_name: 'Cookbook',
105
+ last_name: 'Billing',
106
+ })
107
+ await ctx.rootApi.admin.confirmUser(signUpResp.data.id!)
108
+ } catch (err) {
109
+ // Already provisioned by a previous run. Confirm that is what happened
110
+ // rather than assuming it — a genuine signup failure must not be read as
111
+ // "already there".
112
+ const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
113
+ if (!/taken|already|exist/i.test(detail)) throw err
114
+ }
115
+ ```
116
+
117
+ ## 002 — the find-or-create helper every later step uses
118
+
119
+ Idempotence is one question asked over and over: *is this already here?*
120
+ Rather than answer it a dozen different ways, the walk defines it once.
121
+
122
+ `ensure` takes a label, a lookup and a create. It runs the lookup, returns
123
+ what it finds, and only creates when the lookup comes back empty. The label
124
+ is not decoration — when a run reuses something you did not expect it to, the
125
+ log line naming it is how you find out.
126
+
127
+ Alongside it, `firstNamed` — the lookup half, written once. **A list endpoint
128
+ pages at twenty and is not newest-first**, so a `.find()` over the default
129
+ page silently stops finding things as soon as a lake has a few of them, and
130
+ the failure looks like the resource was never created. Passing
131
+ `page_size: 100` is the cheap fix, and this walk creates enough tools,
132
+ contracts and clients on one lake to need it.
133
+
134
+ Both live on `ctx` rather than as bare functions because each numbered step
135
+ compiles into its own `it()` block, so a plain `function` here would not be
136
+ in scope for the steps that call it.
137
+
138
+ ```typescript
139
+ ctx.ensure = async <T>(
140
+ label: string,
141
+ find: () => Promise<T | undefined>,
142
+ create: () => Promise<T>,
143
+ ): Promise<T> => {
144
+ const existing = await find()
145
+ if (existing !== undefined) {
146
+ console.log(` ↻ reusing ${label}`)
147
+ return existing
148
+ }
149
+ console.log(` + creating ${label}`)
150
+ return await create()
151
+ }
152
+
153
+ // Look one resource up by name across the WHOLE listing, not page one.
154
+ ctx.firstNamed = <T extends { name?: string }>(
155
+ rows: readonly T[] | undefined,
156
+ name: string,
157
+ ): T | undefined => (rows ?? []).find((r) => r.name === name)
158
+ ```
159
+
160
+ ## 003 — the tenant
161
+
162
+ The billing admin signs in without a tenant scope — they may not belong to
163
+ one yet — and the walk finds or creates the tenant by its stable name. The
164
+ server derives the slug; capture it, because every later call is addressed by
165
+ it.
166
+
167
+ ```typescript
168
+ ctx.billingTenantlessSession = await createBootstrapSession({
169
+ baseUrl: process.env.ALVERA_BASE_URL!,
170
+ email: ctx.billingEmail,
171
+ password: ctx.billingPassword,
172
+ })
173
+ ctx.billingTenantlessApi = createIsolatedPlatformApi({
174
+ baseUrl: process.env.ALVERA_BASE_URL!,
175
+ sessionToken: ctx.billingTenantlessSession.sessionToken,
176
+ apiKey: '',
177
+ })
178
+
179
+ const TENANT_NAME = 'Cookbook Subscription SaaS'
180
+
181
+ const tenant = await ctx.ensure(
182
+ `tenant ${TENANT_NAME}`,
183
+ async () => {
184
+ const { data } = await ctx.billingTenantlessApi.tenants.list()
185
+ return (data.data ?? []).find((t: { name?: string }) => t.name === TENANT_NAME)
186
+ },
187
+ async () => {
188
+ const { data } = await ctx.billingTenantlessApi.tenants.create({ name: TENANT_NAME })
189
+ return data
190
+ },
191
+ )
192
+ tenantSlug = tenant.slug!
193
+ ```
194
+
195
+ ## 004 — the tenant-scoped client
196
+
197
+ A tenant-scoped login requires `X-API-Key`, so a key has to exist before the
198
+ admin can sign in against the tenant. Mint one through the platform-admin
199
+ side door with the root bearer; in the web console this is *Settings → API
200
+ Keys*.
201
+
202
+ `data_access_mode: 'raw'` is the ceiling, and on this surface it is the only
203
+ tier there is — no derived lake is provisioned, so raw is what every read
204
+ answers from.
205
+
206
+ **This is the one step in the walk that is not idempotent**, and it is worth
207
+ knowing why rather than discovering it. There is no endpoint that lists a
208
+ tenant's API keys, so there is nothing to look the existing one up with —
209
+ `ensure` has no lookup to run. Each run therefore mints another key. A key
210
+ row is cheap where a datalake is not, so the walk accepts it; if you are
211
+ counting rows in a shared environment, this is the one to count.
212
+
213
+ ```typescript
214
+ const { data: mintedKey } = await ctx.rootApi.admin.createTenantApiKey(tenantSlug, {
215
+ name: 'Cookbook Subscription SaaS Key',
216
+ data_access_mode: 'raw',
217
+ })
218
+ ctx.tenantApiKey = mintedKey.api_key
219
+
220
+ const billingTenantSession = await createSession({
221
+ baseUrl: process.env.ALVERA_BASE_URL!,
222
+ email: ctx.billingEmail,
223
+ password: ctx.billingPassword,
224
+ tenantSlug,
225
+ apiKey: ctx.tenantApiKey,
226
+ })
227
+ api = createIsolatedPlatformApi({
228
+ baseUrl: process.env.ALVERA_BASE_URL!,
229
+ sessionToken: billingTenantSession.sessionToken,
230
+ apiKey: ctx.tenantApiKey,
231
+ })
232
+ ```
233
+
234
+ ## 005 — the datalake
235
+
236
+ One datalake, created `raw`, and on this surface that is the whole story —
237
+ there is no derived pair to provision and nothing to wait for beyond the
238
+ migration in §006.
239
+
240
+ The lookup still filters on `type === 'raw'`. That costs nothing here and
241
+ keeps the step correct if anyone ever turns tokenization on: from that moment
242
+ `datalakes.list` returns three rows, two of which are copies that must never
243
+ be mistaken for the primary.
244
+
245
+ Local-dev defaults match the seeded `dev.exs` setup — `postgres` on
246
+ `localhost:5432`, database `alvera_dev_foundation`, LocalStack S3 on
247
+ `localhost:4566` — and the schema name is stable, because a fresh schema per
248
+ run is the same leak as a fresh tenant per run.
249
+
250
+ ```typescript
251
+ const DB = { host: 'localhost', port: 5432, user: 'postgres', pass: 'postgres', name: 'alvera_dev_foundation' }
252
+ const DB_SCHEMA = 'cookbook_subscription_saas'
253
+ const LAKE_NAME = 'Cookbook Subscription SaaS Datalake'
254
+
255
+ const S3 = {
256
+ cloud_storage_type: 'aws' as const,
257
+ region: 'us-east-1',
258
+ access_key_id: 'test',
259
+ secret_access_key: 'test',
260
+ endpoint: 'http://localhost:4566',
261
+ }
262
+
263
+ const datalake = await ctx.ensure(
264
+ `datalake ${LAKE_NAME}`,
265
+ async () => {
266
+ const { data } = await api.datalakes.list(tenantSlug)
267
+ return (data.data ?? []).find(
268
+ (l: { name?: string; type?: string }) => l.name === LAKE_NAME && l.type === 'raw',
269
+ )
270
+ },
271
+ async () => {
272
+ const { data } = await api.datalakes.create(tenantSlug, {
273
+ name: LAKE_NAME,
274
+ description: 'Subscription SaaS datalake provisioned by the cookbook doctest.',
275
+ timezone: 'America/New_York',
276
+ pool_size: 3,
277
+ type: 'raw',
278
+
279
+ db_writer_host: DB.host,
280
+ db_writer_port: DB.port,
281
+ db_writer_name: DB.name,
282
+ db_writer_schema: DB_SCHEMA,
283
+ db_writer_auth_method: 'password',
284
+ db_writer_user: DB.user,
285
+ db_writer_pass: DB.pass,
286
+ db_writer_enable_ssl: false,
287
+ db_reader_host: DB.host,
288
+ db_reader_port: DB.port,
289
+ db_reader_name: DB.name,
290
+ db_reader_schema: DB_SCHEMA,
291
+ db_reader_auth_method: 'password',
292
+ db_reader_user: DB.user,
293
+ db_reader_pass: DB.pass,
294
+ db_reader_enable_ssl: false,
295
+
296
+ cloud_storage: { ...S3, bucket: 'alvera-platform-dev', base_path: 'cookbook/subscription-saas' },
297
+ })
298
+ return data
299
+ },
300
+ )
301
+ datalakeSlug = datalake.slug!
302
+ ctx.datalakeId = datalake.id!
303
+ ```
304
+
305
+ ## 006 — run the migrations, and wait for ready
306
+
307
+ `datalakes.create` persists the row at `status: 'new'`; it does not run the
308
+ schema DDL. Migration is triggered separately so the operator decides when
309
+ the potentially-slow part happens. `datalakes.migrate` enqueues the job and
310
+ returns immediately with `status: 'enqueued'`; the poll after it is what
311
+ waits for the worker.
312
+
313
+ Migrating is safe to repeat, which is what lets this step stay unguarded on a
314
+ second run. The wait exits the moment the lake reports `ready`, so the
315
+ five-minute ceiling is only ever paid in failure.
316
+
317
+ ```typescript
318
+ const migrateResp = await api.datalakes.migrate(tenantSlug, datalakeSlug)
319
+ if (migrateResp.data.status !== 'enqueued') {
320
+ throw new Error(`datalake migration not enqueued (status: ${migrateResp.data.status})`)
321
+ }
322
+
323
+ const READY_TIMEOUT_MS = 5 * 60_000
324
+ const readyDeadline = Date.now() + READY_TIMEOUT_MS
325
+ let datalakeStatus: string | undefined
326
+ while (Date.now() < readyDeadline) {
327
+ const { data } = await api.datalakes.get(tenantSlug, ctx.datalakeId)
328
+ datalakeStatus = data.status
329
+ if (datalakeStatus === 'ready') break
330
+ await new Promise((r) => setTimeout(r, 5_000))
331
+ }
332
+ if (datalakeStatus !== 'ready') {
333
+ throw new Error(`datalake did not reach :ready within ${READY_TIMEOUT_MS}ms (last: ${datalakeStatus})`)
334
+ }
335
+ ```
336
+
337
+ ## 007 — three helpers the scenarios below share
338
+
339
+ **`waitForFiredRun`.** `workflows.run` only *schedules* a run. It returns
340
+ immediately with a `workflow_run_id`, and the `workflow_run_log_id` and
341
+ `batch_id` a scenario needs are written later, when the run actually fires.
342
+ Two traps live in that gap:
343
+
344
+ - **Poll until `workflow_run_log_id` is a string — not until `status` leaves
345
+ `'scheduled'`.** Those are different moments; the run reaches `processing`
346
+ first and writes the log id a beat later. A predicate on status alone
347
+ releases you to read a `null`, and because `typeof null === 'object'` the
348
+ symptom is a baffling *"expected string, got object"* rather than an
349
+ obvious nil.
350
+ - **Raise on `failed` carrying `failure_reason`** rather than polling to the
351
+ deadline. A scenario blocked on a run that will never fire should say why
352
+ on the first read, not thirty seconds later behind a generic timeout.
353
+
354
+ **`waitForBatches`.** Ingestion is async: `ingest` returns a `batch_id` the
355
+ instant the rows are accepted, and the per-row jobs drain afterwards. This
356
+ polls a client's activation logs until every named batch has actually
357
+ written.
358
+
359
+ **Gate on `dataset_updated`, never on `rows_ingested`.** They answer
360
+ different questions — received versus written — and a batch whose row is
361
+ refused reports `rows_ingested: 1, dataset_updated: 0, status: 'partial'`
362
+ with the reason in `error`. A gate on `rows_ingested` calls that green, the
363
+ next step then runs against a table with nothing in it, and the failure
364
+ surfaces three steps later as *"expected 2 execution logs, got 0"* — which
365
+ reads like a broken workflow rather than a row that never landed.
366
+
367
+ **`deployGenericTable`.** Creating a generic table does not deploy it; it
368
+ rests at `status: 'new'` until a migration runs. The wait that means anything
369
+ asks for the table the same way the next step is about to, and retries until
370
+ it stops erroring — polling the table's own `status: 'deployed'` goes true
371
+ earlier and proves less.
372
+
373
+ ```typescript
374
+ ctx.waitForFiredRun = async (
375
+ runDatalakeSlug: string,
376
+ runId: string,
377
+ timeoutMs = 120_000,
378
+ ): Promise<{ workflowRunLogId: string; batchId: string | null }> => {
379
+ const deadline = Date.now() + timeoutMs
380
+ let lastStatus: string | undefined
381
+ while (Date.now() < deadline) {
382
+ const { data } = await api.workflowRuns.get(tenantSlug, runDatalakeSlug, runId)
383
+ lastStatus = data.status
384
+ if (data.status === 'failed') {
385
+ throw new Error(`workflow run ${runId} failed: ${data.failure_reason ?? 'no failure_reason given'}`)
386
+ }
387
+ if (typeof data.workflow_run_log_id === 'string') {
388
+ return { workflowRunLogId: data.workflow_run_log_id, batchId: data.batch_id ?? null }
389
+ }
390
+ await new Promise((r) => setTimeout(r, 1_000))
391
+ }
392
+ throw new Error(`workflow run ${runId} never fired within ${timeoutMs}ms (last status: ${lastStatus})`)
393
+ }
394
+
395
+ ctx.waitForBatches = async (
396
+ dacSlug: string,
397
+ batchIds: readonly string[],
398
+ timeoutMs = 90_000,
399
+ ): Promise<void> => {
400
+ const targets = new Set(batchIds)
401
+ const deadline = Date.now() + timeoutMs
402
+ let greenCount = 0
403
+ while (Date.now() < deadline) {
404
+ const { data } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
405
+ const green = new Set<string>()
406
+ for (const row of (data.data ?? []) as Array<Record<string, unknown>>) {
407
+ const b = row.batch_id
408
+ if (typeof b !== 'string' || !targets.has(b)) continue
409
+ if (row.status === 'partial' || row.status === 'failed') {
410
+ throw new Error(
411
+ `batch ${b} on ${dacSlug} did not persist its rows (status: ${row.status}): ` +
412
+ `${String(row.error ?? 'no reason given')}`,
413
+ )
414
+ }
415
+ if (typeof row.dataset_updated !== 'number' || row.dataset_updated < 1) continue
416
+ const files = row.output_files
417
+ if (!Array.isArray(files) || files.length === 0) continue
418
+ green.add(b)
419
+ }
420
+ greenCount = green.size
421
+ if (greenCount === targets.size) return
422
+ await new Promise((r) => setTimeout(r, 1_000))
423
+ }
424
+ throw new Error(`only ${greenCount}/${targets.size} batches on ${dacSlug} persisted within ${timeoutMs}ms`)
425
+ }
426
+
427
+ ctx.deployGenericTable = async (tableName: string, timeoutMs = 120_000): Promise<void> => {
428
+ await api.datalakes.migrate(tenantSlug, datalakeSlug)
429
+ const deadline = Date.now() + timeoutMs
430
+ let lastError: unknown
431
+ while (Date.now() < deadline) {
432
+ try {
433
+ await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
434
+ sql: `SELECT 1 FROM ${tableName} LIMIT 1`,
435
+ mode: 'raw',
436
+ })
437
+ return
438
+ } catch (err) {
439
+ lastError = err
440
+ await new Promise((r) => setTimeout(r, 1_000))
441
+ }
442
+ }
443
+ throw new Error(
444
+ `generic table ${tableName} was not readable within ${timeoutMs}ms ` +
445
+ `(last error: ${lastError instanceof Error ? lastError.message : String(lastError)})`,
446
+ )
447
+ }
448
+ ```
449
+
450
+ ## 008 — the SMS tool
451
+
452
+ One SMS tool dispatches every outbound message in this walk — a welcome, a
453
+ dunning reminder and three priority bands. **The per-message body lives on
454
+ the workflow action, never on the tool**, which is what lets one tool serve
455
+ all of them. `intent: 'sms'` tags it for workflow-action use.
456
+
457
+ Local dev points at LocalStack's SNS on `:4566`, so nothing leaves the
458
+ machine.
459
+
460
+ ```typescript
461
+ const smsTool = await ctx.ensure(
462
+ 'SMS tool',
463
+ async () => {
464
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
465
+ return ctx.firstNamed(data.data, 'Cookbook Billing SMS Tool')
466
+ },
467
+ async () => {
468
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
469
+ name: 'Cookbook Billing SMS Tool',
470
+ description: 'SNS-backed SMS dispatcher for every billing scenario, wired to LocalStack.',
471
+ intent: 'sms',
472
+ status: 'active',
473
+ datalake_id: ctx.datalakeId,
474
+ body: {
475
+ tool_body_type: 'sns',
476
+ auth_method: 'access_key',
477
+ region: 'us-east-1',
478
+ phone_number: '+15551234567',
479
+ endpoint_url: 'http://localhost:4566',
480
+ access_key_id: 'test',
481
+ secret_access_key: 'test',
482
+ },
483
+ })
484
+ return data
485
+ },
486
+ )
487
+ toolId = smsTool.id!
488
+ ctx.smsToolId = smsTool.id!
489
+ ```
490
+
491
+ ## 009 — the customers table, which all three workflows share
492
+
493
+ A customer is not one of the datasets the platform ships. The platform models
494
+ exactly one identity — the legal entity — and everything else a build needs
495
+ is a generic table you declare, plus the legal entity the rows resolve to.
496
+ There is no `customer` dataset to reach for.
497
+
498
+ **One table serves all three workflows**, and the column list is the union of
499
+ what they gate on rather than three near-identical tables. `phone` is the
500
+ welcome workflow's reachability gate. `tax_id` is the dunning workflow's KYC
501
+ gate. `customer_type` is what the priority agent bands on. `status` is what
502
+ makes a customer newly contracted.
503
+
504
+ The PII columns are declared `tokenize` and the rest `none`. **That
505
+ declaration is the masking policy, and it is correct whether or not a
506
+ tokenized lake exists today** — this walk provisions none, so nothing masks
507
+ here, and the declaration goes live the moment someone calls
508
+ `provisionTokenization`. Declaring it now is how the policy survives that
509
+ call rather than being remembered afterwards.
510
+
511
+ Do **not** declare a `legal_entity_id` column. The platform stamps it from
512
+ the MDM resolve in §012, and the name is reserved — declaring it does not
513
+ shadow the stamp, it 422s the create, so you lose the table until you drop it.
514
+
515
+ ```typescript
516
+ const customers = await ctx.ensure(
517
+ 'customers table',
518
+ async () => {
519
+ const { data } = await api.genericTables.list(tenantSlug, datalakeSlug, { page_size: 100 })
520
+ return (data.data ?? []).find((t: { title?: string }) => t.title === 'Cookbook Customers')
521
+ },
522
+ async () => {
523
+ const { data } = await api.genericTables.create(tenantSlug, datalakeSlug, {
524
+ title: 'Cookbook Customers',
525
+ description: 'Stripe customers the welcome, dunning and triage workflows all run on.',
526
+ columns: [
527
+ { name: 'customer_number', title: 'Customer Number', type: 'string', description: 'Stripe customer id — unique per customer', is_unique: true, privacy_requirement: 'none' },
528
+ { name: 'customer_type', title: 'Customer Type', type: 'string', description: 'individual / enterprise — the legal-entity branch, and what the priority agent bands on', is_unique: false, privacy_requirement: 'none' },
529
+ { name: 'status', title: 'Status', type: 'string', description: 'Commercial state — contracted is what earns a welcome', is_unique: false, privacy_requirement: 'none' },
530
+ { name: 'name', title: 'Name', type: 'string', description: 'Customer display name', is_unique: false, privacy_requirement: 'tokenize' },
531
+ { name: 'email', title: 'Email', type: 'string', description: 'Billing email', is_unique: false, privacy_requirement: 'tokenize' },
532
+ { name: 'phone', title: 'Phone', type: 'string', description: 'Billing phone — the SMS destination and the reachability gate', is_unique: false, privacy_requirement: 'tokenize' },
533
+ { name: 'tax_id', title: 'Tax ID', type: 'string', description: 'Tax identifier — its presence is the KYC gate', is_unique: false, privacy_requirement: 'tokenize' },
534
+ { name: 'currency', title: 'Currency', type: 'string', description: 'Settlement currency', is_unique: false, privacy_requirement: 'none' },
535
+ { name: 'delinquent', title: 'Delinquent', type: 'boolean', description: 'Whether the latest invoice is past due', is_unique: false, privacy_requirement: 'none' },
536
+ { name: 'address_city', title: 'Address City', type: 'string', description: 'Billing city', is_unique: false, privacy_requirement: 'none' },
537
+ { name: 'address_country', title: 'Address Country', type: 'string', description: 'Billing country', is_unique: false, privacy_requirement: 'none' },
538
+ ],
539
+ })
540
+ return data
541
+ },
542
+ )
543
+ genericTableId = customers.id!
544
+ ctx.customersTableId = customers.id!
545
+ ctx.customersTableName = customers.name!
546
+
547
+ // Creating the table records it; it is not deployed until a migration runs.
548
+ await ctx.deployGenericTable(ctx.customersTableName)
549
+ ```
550
+
551
+ ## 010 — the Stripe data source and the manual-upload tool
552
+
553
+ A workflow runs on rows, and rows arrive through the data-activation chain.
554
+ The chain's first link is a **data source** — a registration of where the
555
+ rows originate. Its `uri` is the system-of-record address, and it matters
556
+ more than it looks: the client injects it into every ingested row as
557
+ `source_uri`, and both contracts below render it onto the identifier. It is
558
+ half of the key that collapses two contracts onto one legal entity.
559
+
560
+ The **manual-upload tool** is the minimal data-exchange tool. It needs no
561
+ endpoint and no credentials, because the rows arrive in the ingest call's own
562
+ body rather than being fetched. `intent: 'data_exchange'` is what
563
+ distinguishes it from the SMS tool above.
564
+
565
+ ```typescript
566
+ const dataSource = await ctx.ensure(
567
+ 'Stripe data source',
568
+ async () => {
569
+ const { data } = await api.dataSources.list(tenantSlug, datalakeSlug, { page_size: 100 })
570
+ return ctx.firstNamed(data.data, 'Cookbook Stripe Source')
571
+ },
572
+ async () => {
573
+ const { data } = await api.dataSources.create(tenantSlug, datalakeSlug, {
574
+ name: 'Cookbook Stripe Source',
575
+ uri: 'stripe.example.com/cookbook-subscription-saas',
576
+ description: 'Stripe billing system — origin of every customer row in this walk.',
577
+ status: 'active',
578
+ is_default: false,
579
+ })
580
+ return data
581
+ },
582
+ )
583
+ dataSourceId = dataSource.id!
584
+ ctx.dataSourceId = dataSource.id!
585
+
586
+ const manualUploadTool = await ctx.ensure(
587
+ 'manual-upload tool',
588
+ async () => {
589
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
590
+ return ctx.firstNamed(data.data, 'Cookbook Manual Upload Tool')
591
+ },
592
+ async () => {
593
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
594
+ name: 'Cookbook Manual Upload Tool',
595
+ description: 'Manual-upload data-exchange tool — backs the inline and bulk ingest paths.',
596
+ intent: 'data_exchange',
597
+ status: 'active',
598
+ datalake_id: ctx.datalakeId,
599
+ data_source_id: dataSourceId,
600
+ body: { tool_body_type: 'manual_upload' },
601
+ })
602
+ return data
603
+ },
604
+ )
605
+ ctx.manualUploadToolId = manualUploadTool.id!
606
+ ```
607
+
608
+ ## 011 — the contract pair
609
+
610
+ One inbound Stripe row becomes two things, so it takes two contracts.
611
+
612
+ **Contract A** (`resource_type: 'legal_entity'`) writes the identity — the
613
+ name, the type, the identification triple. Its `mdm_input_config` is
614
+ `{ type: 'null' }`, because a template that already *is* the subject has
615
+ nothing to resolve.
616
+
617
+ **Contract B** (`resource_type: 'generic_table'`, pinned to §009's table)
618
+ writes the table row, and its `mdm_input_config` emits the same
619
+ `(uri, customer_number)` pair. That is what collapses both onto one legal
620
+ entity per customer, and what earns the row the `legal_entity_id` stamp every
621
+ connected-app link is minted against.
622
+
623
+ **The identifier shape is where this goes wrong, and there are three of them
624
+ on this platform.** `Platform.MDMInput` — what a `mdm_input_config` renders —
625
+ takes `uri` / `type` / `value`. A legal-entity identification, which is
626
+ Contract A's own body, takes `id_type` / `id_number` / `uri`. And
627
+ `POST /mdm/verify` takes `system` / `id_type` / `value`. The overlap is
628
+ uneven: `system` is a genuine alias on the MDM side — it is cast and
629
+ normalised onto `uri` — but `id_type` and `id_number` are not cast at all.
630
+ Unknown keys are dropped rather than refused, so borrowing the legal-entity
631
+ names for the MDM template resolves nothing and fails as a bare
632
+ `mdm_dispatcher_error` naming no field.
633
+
634
+ And the reason the two look interchangeable is that **one becomes the
635
+ other**. `MDMInput.to_identification/1` takes `(uri, type, value)` and
636
+ returns `(uri, id_type, id_number)` — the rename happens inside the platform,
637
+ on the way through. So the identifier reads back under names it will not
638
+ accept on the way in, and that asymmetry is invisible from the stored shape
639
+ alone.
640
+
641
+ The templates are loaded from the vendored fixtures rather than inlined; they
642
+ are long, and a reader who wants them can open the file.
643
+
644
+ ```typescript
645
+ const { readFileSync } = await import('node:fs')
646
+ const { join } = await import('node:path')
647
+ const read = (f: string) =>
648
+ readFileSync(join(process.env.COOKBOOK_FIXTURES_DIR!, 'subscription-saas', f), 'utf8')
649
+
650
+ const leContract = await ctx.ensure(
651
+ 'legal-entity contract',
652
+ async () => {
653
+ const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
654
+ return ctx.firstNamed(data.data, 'Cookbook Customer LE Contract')
655
+ },
656
+ async () => {
657
+ const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
658
+ name: 'Cookbook Customer LE Contract',
659
+ description: 'Stripe customer → LegalEntity (individual or business).',
660
+ resource_type: 'legal_entity',
661
+ type: 'identity',
662
+ generic_table_id: null,
663
+ template_config: { type: 'custom', body: read('_customers_subscription_legal_entity.liquid') },
664
+ mdm_input_config: { type: 'null' },
665
+ })
666
+ return data
667
+ },
668
+ )
669
+ ctx.leContractId = leContract.id!
670
+
671
+ const gtContract = await ctx.ensure(
672
+ 'generic-table contract',
673
+ async () => {
674
+ const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
675
+ return ctx.firstNamed(data.data, 'Cookbook Customer GT Contract')
676
+ },
677
+ async () => {
678
+ const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
679
+ name: 'Cookbook Customer GT Contract',
680
+ description: 'Stripe customer → customers table row, stamped with its subject.',
681
+ resource_type: 'generic_table',
682
+ type: 'identity',
683
+ generic_table_id: ctx.customersTableId,
684
+ template_config: { type: 'custom', body: read('_customers_subscription_generic_table.liquid') },
685
+ mdm_input_config: { type: 'custom', body: read('_customers_subscription_mdm.liquid') },
686
+ })
687
+ return data
688
+ },
689
+ )
690
+ interopContractId = gtContract.id!
691
+ ctx.gtContractId = gtContract.id!
692
+ ```
693
+
694
+ ## 012 — two clients, because the pair must not share one
695
+
696
+ Here is the trap, and it is the most expensive one in this file.
697
+
698
+ The obvious move is to bind both contracts to a single client and ingest
699
+ once. **Do not.** A client fans out per `(row, contract)` at enqueue time,
700
+ and the job key is staggered by *row* — so a pair bound to one client runs
701
+ both contracts on the same row **at the same instant**. Both resolve the same
702
+ subject, both find-or-create it, both miss the other's uncommitted write, and
703
+ the unique index on `(uri, id_type, id_number)` refuses the loser. The legal
704
+ entity is deliberately not part of that key, which is what makes the refusal
705
+ correct rather than a bug: two writers, one subject, one winner.
706
+
707
+ The stagger is the platform's designed guard, and it only helps when the two
708
+ contracts are on different *rows*. So split them across two clients and
709
+ sequence the ingest: identity first, wait, then the table row. Contract B's
710
+ resolve then finds the subject Contract A already committed, and stamps it.
711
+
712
+ Both clients share the tool and the data source. Only the contract list
713
+ differs.
714
+
715
+ ```typescript
716
+ const identityDac = await ctx.ensure(
717
+ 'identity activation client',
718
+ async () => {
719
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
720
+ return ctx.firstNamed(data.data, 'Cookbook Customer Identity DAC')
721
+ },
722
+ async () => {
723
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
724
+ name: 'Cookbook Customer Identity DAC',
725
+ description: 'Writes the legal entity for each Stripe customer. Runs FIRST.',
726
+ tool_id: ctx.manualUploadToolId,
727
+ data_source_id: dataSourceId,
728
+ tool_call: { tool_call_type: 'manual_upload' },
729
+ interoperability_contract_ids: [ctx.leContractId],
730
+ })
731
+ return data
732
+ },
733
+ )
734
+ ctx.identityDacSlug = identityDac.slug!
735
+
736
+ const dataDac = await ctx.ensure(
737
+ 'data activation client',
738
+ async () => {
739
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
740
+ return ctx.firstNamed(data.data, 'Cookbook Customer Data DAC')
741
+ },
742
+ async () => {
743
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
744
+ name: 'Cookbook Customer Data DAC',
745
+ description: 'Writes the customers table row and stamps it with the subject. Runs SECOND.',
746
+ tool_id: ctx.manualUploadToolId,
747
+ data_source_id: dataSourceId,
748
+ tool_call: { tool_call_type: 'manual_upload' },
749
+ interoperability_contract_ids: [ctx.gtContractId],
750
+ })
751
+ return data
752
+ },
753
+ )
754
+ dacId = dataDac.id!
755
+ ctx.dataDacSlug = dataDac.slug!
756
+ ```
757
+
758
+ ## 013 — four customers, identities first
759
+
760
+ Four rows, chosen so that each of the three workflows below can be scoped to
761
+ a pair that splits cleanly:
762
+
763
+ | customer | phone | tax_id | type | proves |
764
+ |-------------------|-------|--------|------------|------------------------------|
765
+ | `CUS-SS-REACH` | yes | yes | individual | passes all three filters |
766
+ | `CUS-SS-NOPHONE` | — | yes | individual | welcome filters it out |
767
+ | `CUS-SS-NOKYC` | yes | — | individual | dunning filters it out |
768
+ | `CUS-SS-ENTERPRISE`| yes | yes | enterprise | the agent bands it high |
769
+
770
+ The customer numbers are **stable, not per-run**. That is what makes the walk
771
+ idempotent — a second run upserts the same four rows on `customer_number`
772
+ rather than adding four more — and it is also what makes the outbound
773
+ messages idempotent, since the action keys are built from these ids.
774
+
775
+ The ingest is **staged**, per §012. Every identity row goes through the
776
+ identity client and is waited on to completion; only then do the table rows
777
+ go through the data client.
778
+
779
+ ```typescript
780
+ ctx.reachNumber = 'CUS-SS-REACH'
781
+ ctx.noPhoneNumber = 'CUS-SS-NOPHONE'
782
+ ctx.noKycNumber = 'CUS-SS-NOKYC'
783
+ ctx.enterpriseNumber = 'CUS-SS-ENTERPRISE'
784
+
785
+ ctx.customerRows = [
786
+ {
787
+ customer_number: ctx.reachNumber,
788
+ customer_type: 'individual',
789
+ status: 'contracted',
790
+ name: 'Olivia Hartmann',
791
+ email: 'olivia.hartmann@example.com',
792
+ phone: '+12025553101',
793
+ tax_id: '311-22-7890',
794
+ currency: 'USD',
795
+ delinquent: 'true',
796
+ address_city: 'Boston',
797
+ address_country: 'US',
798
+ },
799
+ {
800
+ customer_number: ctx.noPhoneNumber,
801
+ customer_type: 'individual',
802
+ status: 'contracted',
803
+ name: 'Noah Pemberton',
804
+ email: 'noah.pemberton@example.com',
805
+ tax_id: '412-88-1200',
806
+ currency: 'USD',
807
+ delinquent: 'false',
808
+ address_city: 'Denver',
809
+ address_country: 'US',
810
+ // phone deliberately absent — the welcome filter must reject this row
811
+ },
812
+ {
813
+ customer_number: ctx.noKycNumber,
814
+ customer_type: 'individual',
815
+ status: 'contracted',
816
+ name: 'Priya Raghunathan',
817
+ email: 'priya.raghunathan@example.com',
818
+ phone: '+12025553103',
819
+ currency: 'USD',
820
+ delinquent: 'true',
821
+ address_city: 'Austin',
822
+ address_country: 'US',
823
+ // tax_id deliberately absent — the dunning filter must reject this row
824
+ },
825
+ {
826
+ customer_number: ctx.enterpriseNumber,
827
+ customer_type: 'enterprise',
828
+ status: 'contracted',
829
+ name: 'Pinnacle Financial Group',
830
+ email: 'ap@pinnaclefg.example.com',
831
+ phone: '+12125559100',
832
+ tax_id: '84-2156789',
833
+ currency: 'USD',
834
+ delinquent: 'true',
835
+ address_city: 'New York',
836
+ address_country: 'US',
837
+ },
838
+ ]
839
+
840
+ // Stage one — the identities.
841
+ const identityIngests = await Promise.all(
842
+ ctx.customerRows.map((row: Record<string, unknown>) =>
843
+ api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.identityDacSlug, { data: row }),
844
+ ),
845
+ )
846
+ await ctx.waitForBatches(
847
+ ctx.identityDacSlug,
848
+ identityIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!),
849
+ )
850
+
851
+ // Stage two — the table rows, now that every subject exists.
852
+ const dataIngests = await Promise.all(
853
+ ctx.customerRows.map((row: Record<string, unknown>) =>
854
+ api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.dataDacSlug, { data: row }),
855
+ ),
856
+ )
857
+ ctx.customerBatchIds = dataIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!)
858
+ await ctx.waitForBatches(ctx.dataDacSlug, ctx.customerBatchIds)
859
+ ```
860
+
861
+ ## 014 — the stamp is the proof that resolution ran
862
+
863
+ The activation log said the rows were written. This asks the table itself,
864
+ which is the independent answer, and it asks for the one column no contract
865
+ writes: `legal_entity_id`.
866
+
867
+ A row present with a null stamp is the failure this step exists to catch. It
868
+ means Contract B's template landed but its `mdm_input_config` resolved
869
+ nothing — the wrong identifier shape, or a subject that was not there yet —
870
+ and every link minted from that row afterwards would point at nobody.
871
+
872
+ ```typescript
873
+ const stampDeadline = Date.now() + 90_000
874
+ let stamped: unknown[][] = []
875
+ while (Date.now() < stampDeadline && stamped.length < 4) {
876
+ const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
877
+ sql: `SELECT customer_number, legal_entity_id
878
+ FROM ${ctx.customersTableName}
879
+ WHERE customer_number IN (
880
+ '${ctx.reachNumber}', '${ctx.noPhoneNumber}',
881
+ '${ctx.noKycNumber}', '${ctx.enterpriseNumber}'
882
+ )
883
+ ORDER BY customer_number`,
884
+ mode: 'raw',
885
+ })
886
+ if (typeof result !== 'string') stamped = result.data
887
+ if (stamped.length < 4) await new Promise((r) => setTimeout(r, 1_000))
888
+ }
889
+ if (stamped.length !== 4) {
890
+ throw new Error(`expected 4 customer rows, found ${stamped.length} within 90s`)
891
+ }
892
+ const unstamped = stamped.filter((row) => !row[1]).map((row) => String(row[0]))
893
+ if (unstamped.length > 0) {
894
+ throw new Error(
895
+ `these rows landed with no legal_entity_id stamp — the MDM resolve did not run: ${unstamped.join(', ')}`,
896
+ )
897
+ }
898
+ ```
899
+
900
+ ## 015 — the billing-portal connected app
901
+
902
+ The welcome SMS carries a deep link to the self-serve billing portal. A
903
+ connected app is a thin registration of that page's URL and hosting mode; the
904
+ page itself lives outside the platform (`mode: 'self_hosted'`). Capture the
905
+ server-derived slug — resolving the link in §019 is addressed by it.
906
+
907
+ ```typescript
908
+ const portalApp = await ctx.ensure(
909
+ 'billing-portal connected app',
910
+ async () => {
911
+ const { data } = await api.connectedApps.list(tenantSlug, datalakeSlug, { page_size: 100 })
912
+ return ctx.firstNamed(data.data, 'Cookbook Billing Portal')
913
+ },
914
+ async () => {
915
+ const { data } = await api.connectedApps.create(tenantSlug, datalakeSlug, {
916
+ name: 'Cookbook Billing Portal',
917
+ description: 'Self-serve billing portal linked from the outbound welcome SMS.',
918
+ mode: 'self_hosted',
919
+ urls: [{ url: 'https://billing.example.local', is_primary: true, label: 'production' }],
920
+ })
921
+ return data
922
+ },
923
+ )
924
+ connectedAppId = portalApp.id!
925
+ ctx.portalAppId = portalApp.id!
926
+ ctx.portalAppSlug = portalApp.slug!
927
+ ```
928
+
929
+ ## 016 — the Welcome SMS workflow
930
+
931
+ Standard shape — a filter, a decision, one action — with three things worth
932
+ reading closely.
933
+
934
+ `dataset_type: 'generic_table'` with `generic_table_id` makes §009's table
935
+ the event source. The row is injected under the `event_dataset` key, so the
936
+ filter reads `event_dataset.phone` directly.
937
+
938
+ `skip_mdm_resolution: false` is set explicitly, and must be. A generic-table
939
+ row is never its own subject — the subject is the legal entity §011's pair
940
+ stamped onto it. With resolution on, the platform loads that subject and
941
+ mints a page token against it, which is what makes
942
+ `{{ connected_app_form_url }}` a *per-customer* link. Skip it and the token
943
+ is never minted and the link renders empty.
944
+
945
+ The `context_datasets` entry queries the `message` dataset for a welcome
946
+ already sent to this legal entity in the last six months — the platform's
947
+ built-in guard against welcoming someone twice, expressed as a context query
948
+ rather than as application code. On a fresh lake it returns empty, which is
949
+ harmless; on the second run of this walk it is one of two guards that stop a
950
+ duplicate message, the other being the idempotency key on the action.
951
+
952
+ ```typescript
953
+ const WELCOME_KEY = 'send_welcome_sms'
954
+
955
+ const welcomeWorkflow = await ctx.ensure(
956
+ 'Welcome SMS workflow',
957
+ async () => {
958
+ const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
959
+ return ctx.firstNamed(data.data, 'Cookbook Welcome SMS Workflow')
960
+ },
961
+ async () => {
962
+ const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
963
+ name: 'Cookbook Welcome SMS Workflow',
964
+ description: 'Sends a welcome SMS to newly contracted customers with a self-serve billing link.',
965
+ dataset_type: 'generic_table',
966
+ generic_table_id: ctx.customersTableId,
967
+ status: 'live',
968
+ tags: ['lifecycle', 'welcome'],
969
+ skip_mdm_resolution: false,
970
+ filter_config: {
971
+ type: 'custom',
972
+ body:
973
+ '{% if event_dataset.status == "contracted" %}' +
974
+ '{% if event_dataset.phone and event_dataset.phone != "" %}true{% endif %}' +
975
+ '{% endif %}',
976
+ },
977
+ decision_config: {
978
+ type: 'custom',
979
+ body: `["${WELCOME_KEY}"]`,
980
+ output_schema: { type: 'array', items: { type: 'string' } },
981
+ },
982
+ context_datasets: [
983
+ {
984
+ dataset_type: 'message',
985
+ where_clause:
986
+ `m.legal_entity_id = '{{ legal_entity_id }}' AND m.decision_key = '${WELCOME_KEY}' ` +
987
+ "AND m.sent_at > NOW() - INTERVAL '6 months'",
988
+ limit: 1,
989
+ position: 0,
990
+ },
991
+ ],
992
+ actions: [
993
+ {
994
+ action_type: 'sms',
995
+ tool_id: ctx.smsToolId,
996
+ decision_key: WELCOME_KEY,
997
+ position: 0,
998
+ trigger_template: 'now',
999
+ idempotency_template: `{{ event_dataset.customer_number }}-${WELCOME_KEY}`,
1000
+ connected_app_id: ctx.portalAppId,
1001
+ connected_app_route: '/portal/welcome',
1002
+ connected_app_metadata_template: '{"customer_number":"{{ event_dataset.customer_number }}"}',
1003
+ tool_call: {
1004
+ tool_call_type: 'sms_request',
1005
+ to: { type: 'custom', body: '{{ event_dataset.phone }}' },
1006
+ body: {
1007
+ type: 'custom',
1008
+ body:
1009
+ 'Hi {{ event_dataset.name }}, welcome to Alvera Billing. ' +
1010
+ 'Self-serve billing portal: {{ connected_app_form_url }}',
1011
+ },
1012
+ sms_type: 'transactional',
1013
+ },
1014
+ },
1015
+ ],
1016
+ })
1017
+ return data
1018
+ },
1019
+ )
1020
+ workflowId = welcomeWorkflow.id!
1021
+ ctx.welcomeWorkflowId = welcomeWorkflow.id!
1022
+ ctx.welcomeWorkflowSlug = welcomeWorkflow.slug!
1023
+
1024
+ if (welcomeWorkflow.skip_mdm_resolution !== false) {
1025
+ throw new Error('a generic-table workflow that mints a per-customer link must keep MDM resolution ON')
1026
+ }
1027
+ ```
1028
+
1029
+ ## 017 — run it across the reachable and unreachable pair
1030
+
1031
+ `manual_override: false` is what makes this a real test — the filter is
1032
+ evaluated, so the `event_dataset.phone` gate genuinely routes each row. With
1033
+ `true` the filter is bypassed and both rows would pass, which proves nothing.
1034
+
1035
+ The where-clause scopes the run to exactly two of the four rows. A
1036
+ generic-table run selects on the table's own columns, so it reads
1037
+ `customer_number` with no alias to prefix.
1038
+
1039
+ A `:partial` batch status is the expected outcome here, not a warning: one
1040
+ row passed and one was filtered, which is precisely the split this scenario
1041
+ is about.
1042
+
1043
+ ```typescript
1044
+ const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, {
1045
+ sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.noPhoneNumber}')`,
1046
+ mode: 'live',
1047
+ manual_override: false,
1048
+ })
1049
+ const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
1050
+ ctx.welcomeRunLogId = fired.workflowRunLogId
1051
+ ctx.welcomeRunBatchId = fired.batchId!
1052
+
1053
+ const deadline = Date.now() + 120_000
1054
+ let status: string | null = null
1055
+ while (Date.now() < deadline) {
1056
+ const { data: log } = await api.workflows.batchLogs.refresh(
1057
+ tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, ctx.welcomeRunLogId,
1058
+ )
1059
+ status = log.status ?? null
1060
+ if (status && status !== 'pending') break
1061
+ await new Promise((r) => setTimeout(r, 2_000))
1062
+ }
1063
+ if (status === 'failed') throw new Error('welcome run reached :failed')
1064
+ if (!status || status === 'pending') {
1065
+ throw new Error('welcome run did not leave :pending within 120s')
1066
+ }
1067
+ ```
1068
+
1069
+ ## 018 — the filter routed, and that is asserted at the log level
1070
+
1071
+ Each row produced a Workflow Execution Log. The reachable customer passed, so
1072
+ its log is `:executing` or `:completed`. The customer with no phone failed the
1073
+ filter, so its log is `:filtered` — a distinct terminal status, and a
1074
+ *correct* outcome rather than an error. The platform looked at the row, the
1075
+ gate rendered empty, and it recorded that the row was intentionally skipped.
1076
+
1077
+ A run where both passed, or both were filtered, would mean the filter is not
1078
+ actually reading `event_dataset.phone`.
1079
+
1080
+ **This assertion is at the execution-log level on purpose, and that is what
1081
+ makes it survive a rerun.** The action inside the passing row's log will
1082
+ refuse to fire a second time — its idempotency key already fired — but the
1083
+ row still passes the filter, so the one-passed-one-filtered split is the same
1084
+ on every run. An assertion written against the action instead would go red on
1085
+ run two for a reason that is not a defect.
1086
+
1087
+ ```typescript
1088
+ const deadline = Date.now() + 90_000
1089
+ let ourLogs: Array<Record<string, unknown>> = []
1090
+ let byStatus: Record<string, number> = {}
1091
+ while (Date.now() < deadline) {
1092
+ const { data: wfLogs } = await api.workflows.workflowLogs.list(
1093
+ tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, { page_size: 100 },
1094
+ )
1095
+ ourLogs = (wfLogs.data ?? [])
1096
+ .map((w) => w as Record<string, unknown>)
1097
+ .filter((w) => w.batch_id === ctx.welcomeRunBatchId)
1098
+ byStatus = {}
1099
+ for (const w of ourLogs) {
1100
+ const st = (w.status as string | undefined) ?? 'unknown'
1101
+ byStatus[st] = (byStatus[st] ?? 0) + 1
1102
+ }
1103
+ const settled = (byStatus.completed ?? 0) + (byStatus.executing ?? 0)
1104
+ if (ourLogs.length === 2 && settled === 1 && (byStatus.filtered ?? 0) === 1) break
1105
+ await new Promise((r) => setTimeout(r, 2_000))
1106
+ }
1107
+ if (ourLogs.length !== 2) {
1108
+ throw new Error(`expected 2 execution logs for the run, got ${ourLogs.length}`)
1109
+ }
1110
+ const passed = (byStatus.executing ?? 0) + (byStatus.completed ?? 0)
1111
+ if (passed !== 1) {
1112
+ throw new Error(`expected 1 row past the filter, got ${passed} — distribution ${JSON.stringify(byStatus)}`)
1113
+ }
1114
+ if ((byStatus.filtered ?? 0) !== 1) {
1115
+ throw new Error(
1116
+ `expected the no-phone customer to be :filtered — distribution ${JSON.stringify(byStatus)}`,
1117
+ )
1118
+ }
1119
+ ```
1120
+
1121
+ ## 019 — the welcome SMS exists, and it carries a working link
1122
+
1123
+ §018 proved the routing. This proves something was actually *sent*, and it is
1124
+ the assertion that survives every rerun.
1125
+
1126
+ **Search on the idempotency key, never on `workflow_id`.** The key is
1127
+ `<customer number>-send_welcome_sms` and the customer number is stable, so it
1128
+ names the same message forever. A workflow id is stable only as long as the
1129
+ workflow resource is: recreate it against a lake that still holds its rows
1130
+ and the id changes while the key does not — the action then correctly refuses
1131
+ to fire, and a search for the new id finds nothing at all.
1132
+
1133
+ Then resolve the `/t/<token>` the platform baked into the body when it minted
1134
+ the page token, and post the tracking a connected-app frontend would post
1135
+ when the customer opens the portal. That closes the loop: the message went
1136
+ out, the link in it resolves to this customer's page, and the open was
1137
+ recorded.
1138
+
1139
+ ```typescript
1140
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
1141
+ search_query: `m.idempotency_key = '${ctx.reachNumber}-send_welcome_sms'`,
1142
+ })
1143
+ if (search.status !== 'completed') {
1144
+ throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
1145
+ }
1146
+
1147
+ const msgDeadline = Date.now() + 90_000
1148
+ let welcomeBody: string | undefined
1149
+ while (Date.now() < msgDeadline && !welcomeBody) {
1150
+ const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
1151
+ userSearchId: search.id!,
1152
+ dataAccessMode: 'raw',
1153
+ })
1154
+ welcomeBody = ((data.data ?? []) as Array<Record<string, unknown>>)
1155
+ .map((m) => String(m.body ?? ''))
1156
+ .find((body) => body.includes('/t/') && body.includes('welcome to Alvera Billing'))
1157
+ if (!welcomeBody) await new Promise((r) => setTimeout(r, 2_000))
1158
+ }
1159
+ if (!welcomeBody) {
1160
+ throw new Error(
1161
+ `no rendered welcome SMS for ${ctx.reachNumber} within 90s — ` +
1162
+ 'the action never fired, or it fired under a different idempotency key',
1163
+ )
1164
+ }
1165
+ const tokenMatch = welcomeBody.match(/\/t\/([A-Za-z0-9_-]+)/)
1166
+ if (!tokenMatch) throw new Error(`no /t/<token> in the rendered body: ${welcomeBody}`)
1167
+
1168
+ const { data: resolved } = await api.connectedApps.resolvePage(
1169
+ tenantSlug, datalakeSlug, ctx.portalAppSlug,
1170
+ { short_path: tokenMatch[1]!, user_agent: 'cookbook-doctest/subscription-saas' },
1171
+ )
1172
+ if (resolved.route_path !== '/portal/welcome') {
1173
+ throw new Error(`resolvePage route_path mismatch: ${resolved.route_path}`)
1174
+ }
1175
+ if (!(resolved.message?.body ?? '').includes('welcome to Alvera Billing')) {
1176
+ throw new Error('the resolved page carries the wrong message')
1177
+ }
1178
+
1179
+ const now = new Date().toISOString()
1180
+ const { data: tracked } = await api.connectedApps.updateMessageTracking(
1181
+ tenantSlug, datalakeSlug, ctx.portalAppSlug,
1182
+ { short_path: tokenMatch[1]!, opened_at: now, form_submitted_at: now },
1183
+ )
1184
+ if (!tracked.message?.opened_at || !tracked.message?.form_submitted_at) {
1185
+ throw new Error('message tracking did not persist opened_at + form_submitted_at')
1186
+ }
1187
+ ```
1188
+
1189
+ ## 020 — the pay-portal connected app
1190
+
1191
+ The dunning SMS carries a different link to a different page: not the
1192
+ welcome portal but the one where an overdue invoice gets paid. A second
1193
+ connected app, registered the same way.
1194
+
1195
+ Two apps rather than two routes on one app is a deliberate choice worth
1196
+ naming. A connected app is the unit a page token is minted against, so
1197
+ separating them means a token issued for a payment page cannot be replayed
1198
+ against the welcome page, and revoking one does not take the other down.
1199
+
1200
+ ```typescript
1201
+ const payApp = await ctx.ensure(
1202
+ 'pay-portal connected app',
1203
+ async () => {
1204
+ const { data } = await api.connectedApps.list(tenantSlug, datalakeSlug, { page_size: 100 })
1205
+ return ctx.firstNamed(data.data, 'Cookbook Payment Portal')
1206
+ },
1207
+ async () => {
1208
+ const { data } = await api.connectedApps.create(tenantSlug, datalakeSlug, {
1209
+ name: 'Cookbook Payment Portal',
1210
+ description: 'Self-serve payment page linked from the outbound dunning SMS.',
1211
+ mode: 'self_hosted',
1212
+ urls: [{ url: 'https://pay.example.local', is_primary: true, label: 'production' }],
1213
+ })
1214
+ return data
1215
+ },
1216
+ )
1217
+ ctx.payAppId = payApp.id!
1218
+ ctx.payAppSlug = payApp.slug!
1219
+ ```
1220
+
1221
+ ## 021 — the Dunning SMS workflow: two gates, not one
1222
+
1223
+ Same table, same shape, one more gate. The welcome workflow asks whether the
1224
+ platform *can* reach this customer. The dunning workflow asks that and
1225
+ whether it is *allowed* to: a customer whose KYC is incomplete — no tax
1226
+ identifier on file — is not chased, however overdue the invoice.
1227
+
1228
+ Both gates are one Liquid filter. That is the whole point of the primitive:
1229
+ the reachability rule and the KYC rule live next to each other on the
1230
+ workflow, where a compliance reviewer can read them, rather than in two
1231
+ branches of a service nobody reviews.
1232
+
1233
+ The recency guard is thirty days rather than the welcome's six months,
1234
+ because that is the shape of the business rule — one reminder per billing
1235
+ cycle, not one reminder per lifetime.
1236
+
1237
+ ```typescript
1238
+ const DUNNING_KEY = 'send_dunning_sms'
1239
+
1240
+ const dunningWorkflow = await ctx.ensure(
1241
+ 'Dunning SMS workflow',
1242
+ async () => {
1243
+ const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
1244
+ return ctx.firstNamed(data.data, 'Cookbook Dunning SMS Workflow')
1245
+ },
1246
+ async () => {
1247
+ const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
1248
+ name: 'Cookbook Dunning SMS Workflow',
1249
+ description: 'Sends a payment reminder to reachable, KYC-complete customers with a self-serve pay link.',
1250
+ dataset_type: 'generic_table',
1251
+ generic_table_id: ctx.customersTableId,
1252
+ status: 'live',
1253
+ tags: ['billing', 'dunning'],
1254
+ skip_mdm_resolution: false,
1255
+ filter_config: {
1256
+ type: 'custom',
1257
+ body:
1258
+ '{% if event_dataset.phone and event_dataset.phone != "" %}' +
1259
+ '{% if event_dataset.tax_id and event_dataset.tax_id != "" %}true{% endif %}' +
1260
+ '{% endif %}',
1261
+ },
1262
+ decision_config: {
1263
+ type: 'custom',
1264
+ body: `["${DUNNING_KEY}"]`,
1265
+ output_schema: { type: 'array', items: { type: 'string' } },
1266
+ },
1267
+ context_datasets: [
1268
+ {
1269
+ dataset_type: 'message',
1270
+ where_clause:
1271
+ `m.legal_entity_id = '{{ legal_entity_id }}' AND m.decision_key = '${DUNNING_KEY}' ` +
1272
+ "AND m.sent_at > NOW() - INTERVAL '30 days'",
1273
+ limit: 1,
1274
+ position: 0,
1275
+ },
1276
+ ],
1277
+ actions: [
1278
+ {
1279
+ action_type: 'sms',
1280
+ tool_id: ctx.smsToolId,
1281
+ decision_key: DUNNING_KEY,
1282
+ position: 0,
1283
+ trigger_template: 'now',
1284
+ idempotency_template: `{{ event_dataset.customer_number }}-${DUNNING_KEY}`,
1285
+ connected_app_id: ctx.payAppId,
1286
+ connected_app_route: '/portal/pay',
1287
+ connected_app_metadata_template: '{"customer_number":"{{ event_dataset.customer_number }}"}',
1288
+ tool_call: {
1289
+ tool_call_type: 'sms_request',
1290
+ to: { type: 'custom', body: '{{ event_dataset.phone }}' },
1291
+ body: {
1292
+ type: 'custom',
1293
+ body:
1294
+ 'Hi {{ event_dataset.name }}, an invoice on account ' +
1295
+ '{{ event_dataset.customer_number }} needs your attention. ' +
1296
+ 'Pay now: {{ connected_app_form_url }}',
1297
+ },
1298
+ sms_type: 'transactional',
1299
+ },
1300
+ },
1301
+ ],
1302
+ })
1303
+ return data
1304
+ },
1305
+ )
1306
+ ctx.dunningWorkflowSlug = dunningWorkflow.slug!
1307
+ ```
1308
+
1309
+ ## 022 — run it, and prove the KYC gate is the one that bit
1310
+
1311
+ The scope is the reachable customer and the one missing a tax identifier.
1312
+ **Both have a phone**, which is what makes this run prove something the
1313
+ welcome run could not: the row that gets filtered here is filtered by the
1314
+ second gate, not the first. Had the scope reused the no-phone customer, a
1315
+ filter that ignored `tax_id` entirely would still have produced the same
1316
+ one-passed-one-filtered split and looked green.
1317
+
1318
+ That is the general lesson about filter tests. A pass/fail split only proves
1319
+ the gate you meant to test if the two rows differ in exactly that field.
1320
+
1321
+ ```typescript
1322
+ const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, {
1323
+ sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.noKycNumber}')`,
1324
+ mode: 'live',
1325
+ manual_override: false,
1326
+ })
1327
+ const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
1328
+ ctx.dunningRunBatchId = fired.batchId!
1329
+
1330
+ const refreshDeadline = Date.now() + 120_000
1331
+ let status: string | null = null
1332
+ while (Date.now() < refreshDeadline) {
1333
+ const { data: log } = await api.workflows.batchLogs.refresh(
1334
+ tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, fired.workflowRunLogId,
1335
+ )
1336
+ status = log.status ?? null
1337
+ if (status && status !== 'pending') break
1338
+ await new Promise((r) => setTimeout(r, 2_000))
1339
+ }
1340
+ if (status === 'failed') throw new Error('dunning run reached :failed')
1341
+ if (!status || status === 'pending') throw new Error('dunning run did not leave :pending within 120s')
1342
+
1343
+ const welDeadline = Date.now() + 90_000
1344
+ let ourLogs: Array<Record<string, unknown>> = []
1345
+ let byStatus: Record<string, number> = {}
1346
+ while (Date.now() < welDeadline) {
1347
+ const { data: wfLogs } = await api.workflows.workflowLogs.list(
1348
+ tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, { page_size: 100 },
1349
+ )
1350
+ ourLogs = (wfLogs.data ?? [])
1351
+ .map((w) => w as Record<string, unknown>)
1352
+ .filter((w) => w.batch_id === ctx.dunningRunBatchId)
1353
+ byStatus = {}
1354
+ for (const w of ourLogs) {
1355
+ const st = (w.status as string | undefined) ?? 'unknown'
1356
+ byStatus[st] = (byStatus[st] ?? 0) + 1
1357
+ }
1358
+ const settled = (byStatus.completed ?? 0) + (byStatus.executing ?? 0)
1359
+ if (ourLogs.length === 2 && settled === 1 && (byStatus.filtered ?? 0) === 1) break
1360
+ await new Promise((r) => setTimeout(r, 2_000))
1361
+ }
1362
+ if (ourLogs.length !== 2) {
1363
+ throw new Error(`expected 2 execution logs for the dunning run, got ${ourLogs.length}`)
1364
+ }
1365
+ if ((byStatus.executing ?? 0) + (byStatus.completed ?? 0) !== 1) {
1366
+ throw new Error(`expected 1 row past the KYC gate — distribution ${JSON.stringify(byStatus)}`)
1367
+ }
1368
+ if ((byStatus.filtered ?? 0) !== 1) {
1369
+ throw new Error(
1370
+ `expected the customer with no tax_id to be :filtered — distribution ${JSON.stringify(byStatus)}`,
1371
+ )
1372
+ }
1373
+ ```
1374
+
1375
+ ## 023 — the reminder went out, and it is a different message on a different app
1376
+
1377
+ Same read-back idiom as §019, and worth doing twice for one reason: it proves
1378
+ the two workflows are genuinely independent. The same customer received two
1379
+ messages, under two idempotency keys, carrying two links that resolve against
1380
+ two connected apps and two routes.
1381
+
1382
+ The key here is `<customer number>-send_dunning_sms`, and it names a row that
1383
+ persists in the lake long after the run that made it. That is what a
1384
+ read-back should ask for — the durable record, not the run.
1385
+
1386
+ ```typescript
1387
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
1388
+ search_query: `m.idempotency_key = '${ctx.reachNumber}-send_dunning_sms'`,
1389
+ })
1390
+ if (search.status !== 'completed') {
1391
+ throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
1392
+ }
1393
+
1394
+ const msgDeadline = Date.now() + 90_000
1395
+ let dunningBody: string | undefined
1396
+ while (Date.now() < msgDeadline && !dunningBody) {
1397
+ const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
1398
+ userSearchId: search.id!,
1399
+ dataAccessMode: 'raw',
1400
+ })
1401
+ dunningBody = ((data.data ?? []) as Array<Record<string, unknown>>)
1402
+ .map((m) => String(m.body ?? ''))
1403
+ .find((body) => body.includes('/t/') && body.includes('needs your attention'))
1404
+ if (!dunningBody) await new Promise((r) => setTimeout(r, 2_000))
1405
+ }
1406
+ if (!dunningBody) {
1407
+ throw new Error(`no rendered dunning SMS for ${ctx.reachNumber} within 90s`)
1408
+ }
1409
+ if (!dunningBody.includes(ctx.reachNumber)) {
1410
+ throw new Error(`the dunning body does not name the overdue account: ${dunningBody}`)
1411
+ }
1412
+
1413
+ const token = dunningBody.match(/\/t\/([A-Za-z0-9_-]+)/)
1414
+ if (!token) throw new Error(`no /t/<token> in the dunning body: ${dunningBody}`)
1415
+
1416
+ const { data: resolved } = await api.connectedApps.resolvePage(
1417
+ tenantSlug, datalakeSlug, ctx.payAppSlug,
1418
+ { short_path: token[1]!, user_agent: 'cookbook-doctest/subscription-saas' },
1419
+ )
1420
+ if (resolved.route_path !== '/portal/pay') {
1421
+ throw new Error(`the dunning link resolved to the wrong route: ${resolved.route_path}`)
1422
+ }
1423
+ ```
1424
+
1425
+ ## 024 — the LLM tool
1426
+
1427
+ The welcome and dunning workflows fire the *same* action for everything that
1428
+ passes their filter. The third scenario is the opposite shape: which message
1429
+ goes out depends on what kind of account it is, and that is a judgement over
1430
+ the row rather than a rule about it.
1431
+
1432
+ That judgement needs a model, and a model needs a tool. This one is a
1433
+ **provider adapter**, and the two halves are the thing to read. `base_body`
1434
+ authors the provider's own request — here Ollama's native `/api/chat` shape,
1435
+ with `think: false` and a `format` schema so the model returns
1436
+ schema-constrained JSON instead of prose with JSON in it. `response_extractor`
1437
+ maps the provider's reply back onto the canonical
1438
+ `{ output_json, input_tokens, … }` the platform expects, and its
1439
+ `output_schema` is required for an `llm_enrichment` tool.
1440
+
1441
+ Swapping providers is a matter of rewriting those two templates. Nothing
1442
+ downstream — not the agent, not the workflow — knows which model answered.
1443
+
1444
+ ```typescript
1445
+ const ENRICHMENT_OUTPUT_SCHEMA = {
1446
+ type: 'object',
1447
+ properties: {
1448
+ output_json: {},
1449
+ input_tokens: { type: ['integer', 'null'] },
1450
+ output_tokens: { type: ['integer', 'null'] },
1451
+ total_tokens: { type: ['integer', 'null'] },
1452
+ explanation: { type: ['string', 'null'] },
1453
+ },
1454
+ required: ['output_json'],
1455
+ }
1456
+
1457
+ const OLLAMA_BASE_BODY =
1458
+ '{"model": "{{ model }}", "messages": [{"role": "user", "content": "{{ rendered_prompt | json_escape }}", ' +
1459
+ '"images": [{% for img in images %}{% unless forloop.first %}, {% endunless %}"{{ img.data }}"{% endfor %}]}], ' +
1460
+ '"stream": false, "think": false, "options": {"temperature": {{ temperature }}, "num_predict": {{ max_tokens }}, "num_ctx": 40960}, "format": {{ schema | to_json }}}'
1461
+
1462
+ const OLLAMA_EXTRACTOR =
1463
+ '{"output_json": "{{ msg.message.content | json_escape }}", ' +
1464
+ '"explanation": "{{ msg.message.thinking | json_escape }}", ' +
1465
+ '"input_tokens": {{ msg.prompt_eval_count | default: 0 }}, ' +
1466
+ '"output_tokens": {{ msg.eval_count | default: 0 }}, ' +
1467
+ '"total_tokens": {{ msg.prompt_eval_count | default: 0 | plus: msg.eval_count }}}'
1468
+
1469
+ const llmTool = await ctx.ensure(
1470
+ 'LLM tool',
1471
+ async () => {
1472
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
1473
+ return ctx.firstNamed(data.data, 'Cookbook Billing LLM Tool')
1474
+ },
1475
+ async () => {
1476
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
1477
+ name: 'Cookbook Billing LLM Tool',
1478
+ description: 'Ollama-backed chat-completion adapter for account-priority triage.',
1479
+ intent: 'llm_enrichment',
1480
+ status: 'active',
1481
+ datalake_id: ctx.datalakeId,
1482
+ response_extractor: { type: 'custom', body: OLLAMA_EXTRACTOR, output_schema: ENRICHMENT_OUTPUT_SCHEMA },
1483
+ body: {
1484
+ tool_body_type: 'rest_api',
1485
+ base_url: 'http://localhost:11434',
1486
+ base_path: { type: 'custom', body: '/api/chat' },
1487
+ auth_method: 'api_key',
1488
+ api_key: 'stub-key',
1489
+ api_key_name: 'Authorization',
1490
+ api_key_location: 'header',
1491
+ request_type: 'json',
1492
+ response_type: 'json',
1493
+ timeout_ms: 60_000,
1494
+ base_body: { type: 'custom', body: OLLAMA_BASE_BODY },
1495
+ },
1496
+ })
1497
+ return data
1498
+ },
1499
+ )
1500
+ ctx.llmToolId = llmTool.id!
1501
+ ```
1502
+
1503
+ ## 025 — the Priority Triage agent
1504
+
1505
+ The agent binds three things the workflow cannot supply for itself: the
1506
+ model, an `input_schema` the workflow's context mapping must satisfy, and an
1507
+ `llm_response_schema` whose `enum` pins the vocabulary.
1508
+
1509
+ **The enum is the load-bearing part.** The workflow interpolates the band
1510
+ straight into its decision array, so a model that answers `"high priority"`
1511
+ instead of `priority_high` would produce a decision key no action is
1512
+ registered for, and the row would silently do nothing. Constraining the
1513
+ schema means the failure happens at the model boundary where it is legible,
1514
+ not three layers down as an empty result.
1515
+
1516
+ `temperature: 0.0` makes identical inputs classify identically, which is what
1517
+ lets §027 assert on the outcome at all.
1518
+
1519
+ `data_access: 'raw'` is the right call **on this surface and only on this
1520
+ surface**. There is no derived lake here to read from, and the agent's prompt
1521
+ takes the account number and the account type — no name, no email, no phone.
1522
+ On `payments-compliance` the equivalent agent runs `tokenized`, because there
1523
+ the model is reading a record that carries identity and the point is that it
1524
+ never sees it. The tier follows what the prompt actually contains.
1525
+
1526
+ ```typescript
1527
+ const TRIAGE_PROMPT_BODY = `You are an account-priority triage assistant for a subscription billing team. Classify the customer into EXACTLY ONE band by account type:
1528
+
1529
+ - "priority_high" — enterprise (large account, high balance, white-glove outreach)
1530
+ - "priority_medium" — individual (standard self-serve account)
1531
+ - "priority_low" — government (long payment cycles, low urgency)
1532
+
1533
+ Customer number: {{ customer_number }}
1534
+ Customer type: {{ customer_type }}
1535
+
1536
+ Respond with JSON: {"priority_band": "<one of priority_high|priority_medium|priority_low>"}`
1537
+
1538
+ const triageAgent = await ctx.ensure(
1539
+ 'Priority Triage agent',
1540
+ async () => {
1541
+ const { data } = await api.aiAgents.list(tenantSlug, datalakeSlug, { page_size: 100 })
1542
+ return ctx.firstNamed(data.data, 'Cookbook Priority Triage Agent')
1543
+ },
1544
+ async () => {
1545
+ const { data } = await api.aiAgents.create(tenantSlug, datalakeSlug, {
1546
+ name: 'Cookbook Priority Triage Agent',
1547
+ tool_id: ctx.llmToolId,
1548
+ model: 'qwen3-vl:8b-instruct',
1549
+ data_access: 'raw',
1550
+ temperature: 0.0,
1551
+ max_tokens: 1024,
1552
+ enabled: true,
1553
+ input_schema: {
1554
+ type: 'object',
1555
+ properties: {
1556
+ customer_number: { type: 'string' },
1557
+ customer_type: { type: 'string' },
1558
+ },
1559
+ required: ['customer_number', 'customer_type'],
1560
+ },
1561
+ llm_response_schema: {
1562
+ type: 'object',
1563
+ properties: {
1564
+ priority_band: {
1565
+ type: 'string',
1566
+ enum: ['priority_high', 'priority_medium', 'priority_low'],
1567
+ },
1568
+ },
1569
+ required: ['priority_band'],
1570
+ },
1571
+ prompt_config: { type: 'custom', body: TRIAGE_PROMPT_BODY },
1572
+ })
1573
+ return data
1574
+ },
1575
+ )
1576
+ aiAgentId = triageAgent.id!
1577
+ ctx.triageAgentSlug = triageAgent.slug!
1578
+ ```
1579
+
1580
+ ## 026 — the Priority Triage workflow
1581
+
1582
+ Two things make this workflow agent-driven rather than rule-driven.
1583
+
1584
+ The agent is **nested in the create body** — `workflow_ai_agents` is a
1585
+ `cast_assoc`, not a separate attach call — together with a Liquid
1586
+ `context_mapping_config` that projects each row into the agent's input
1587
+ schema. A workflow join omits `output_schema`; the server pins it from the
1588
+ agent's own `input_schema`, so the two cannot drift apart.
1589
+
1590
+ The `decision_config` then interpolates the agent's answer into a
1591
+ single-element decision array, and three actions are registered, one per
1592
+ band. Whichever band comes back picks one and leaves the other two
1593
+ `:skipped`.
1594
+
1595
+ **Bracket access is required, not stylistic.** The agent's slug contains
1596
+ hyphens, and Liquid's dot parser would read
1597
+ `additional_context.cookbook-priority-triage-agent` as a subtraction.
1598
+
1599
+ ```typescript
1600
+ const BANDS = ['priority_high', 'priority_medium', 'priority_low'] as const
1601
+
1602
+ const CONTEXT_MAPPING_BODY = JSON.stringify({
1603
+ customer_number: '{{ event_dataset.customer_number }}',
1604
+ customer_type: '{{ event_dataset.customer_type }}',
1605
+ })
1606
+
1607
+ const triageWorkflow = await ctx.ensure(
1608
+ 'Priority Triage workflow',
1609
+ async () => {
1610
+ const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
1611
+ return ctx.firstNamed(data.data, 'Cookbook Priority Triage Workflow')
1612
+ },
1613
+ async () => {
1614
+ const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
1615
+ name: 'Cookbook Priority Triage Workflow',
1616
+ description: 'Bands accounts high/medium/low via an LLM agent; one SMS action per band.',
1617
+ dataset_type: 'generic_table',
1618
+ generic_table_id: ctx.customersTableId,
1619
+ status: 'live',
1620
+ tags: ['billing', 'triage'],
1621
+ skip_mdm_resolution: false,
1622
+ filter_config: {
1623
+ type: 'custom',
1624
+ body:
1625
+ '{% if event_dataset.phone and event_dataset.phone != "" %}' +
1626
+ '{% if event_dataset.tax_id and event_dataset.tax_id != "" %}true{% endif %}' +
1627
+ '{% endif %}',
1628
+ },
1629
+ decision_config: {
1630
+ type: 'custom',
1631
+ body: `["{{ additional_context["${ctx.triageAgentSlug}"].priority_band }}"]`,
1632
+ output_schema: { type: 'array', items: { type: 'string' } },
1633
+ },
1634
+ actions: BANDS.map((band) => ({
1635
+ decision_key: band,
1636
+ action_type: 'sms' as const,
1637
+ tool_id: ctx.smsToolId,
1638
+ position: 0,
1639
+ trigger_template: 'now',
1640
+ idempotency_template: `{{ event_dataset.customer_number }}-${band}`,
1641
+ tool_call: {
1642
+ tool_call_type: 'sms_request' as const,
1643
+ to: { type: 'custom' as const, body: '{{ event_dataset.phone }}' },
1644
+ body: {
1645
+ type: 'custom' as const,
1646
+ body: `[${band}] {{ event_dataset.name }}, a note about account {{ event_dataset.customer_number }}.`,
1647
+ },
1648
+ sms_type: 'transactional' as const,
1649
+ },
1650
+ })),
1651
+ workflow_ai_agents: [
1652
+ {
1653
+ ai_agent_id: aiAgentId,
1654
+ position: 0,
1655
+ context_mapping_config: { type: 'custom', body: CONTEXT_MAPPING_BODY },
1656
+ },
1657
+ ],
1658
+ })
1659
+ return data
1660
+ },
1661
+ )
1662
+ ctx.triageWorkflowSlug = triageWorkflow.slug!
1663
+ ```
1664
+
1665
+ ## 027 — the agent's band steered the fan-out
1666
+
1667
+ The scope is the individual and the enterprise account — both reachable, both
1668
+ KYC-complete, so both clear the filter and reach the agent. They differ only
1669
+ in `customer_type`, which is the one thing the agent is asked about.
1670
+
1671
+ Each row produces an execution log carrying **three** action logs, one per
1672
+ band. On the first run for a given customer exactly **one** is matched
1673
+ (`:pending` or `:completed` — the band the agent chose) and **two** are
1674
+ `:skipped`. That one-of-three is the proof the model's output drove the
1675
+ decision rather than every action firing.
1676
+
1677
+ **On a later run all three are `:skipped`, and that is correct.** The
1678
+ idempotency key is `<customer number>-<band>`, the customer numbers are
1679
+ stable, so the key that already fired refuses to fire again — the platform
1680
+ will not send the same message about the same account twice. Both shapes are
1681
+ accepted here; anything else is a real failure. Two matched means the
1682
+ decision matched more than one band. Zero matched with zero skipped means the
1683
+ fan-out never happened at all.
1684
+
1685
+ ```typescript
1686
+ const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, {
1687
+ sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.enterpriseNumber}')`,
1688
+ mode: 'live',
1689
+ manual_override: false,
1690
+ })
1691
+ const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
1692
+ ctx.triageRunBatchId = fired.batchId!
1693
+
1694
+ const refreshDeadline = Date.now() + 180_000
1695
+ let status: string | null = null
1696
+ while (Date.now() < refreshDeadline) {
1697
+ const { data: log } = await api.workflows.batchLogs.refresh(
1698
+ tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, fired.workflowRunLogId,
1699
+ )
1700
+ status = log.status ?? null
1701
+ if (status && status !== 'pending') break
1702
+ await new Promise((r) => setTimeout(r, 2_000))
1703
+ }
1704
+ if (status === 'failed') throw new Error('triage run reached :failed')
1705
+ if (!status || status === 'pending') throw new Error('triage run did not leave :pending within 180s')
1706
+
1707
+ const welDeadline = Date.now() + 180_000
1708
+ let ourLogs: Array<Record<string, unknown>> = []
1709
+ while (Date.now() < welDeadline) {
1710
+ const { data: wfLogs } = await api.workflows.workflowLogs.list(
1711
+ tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, { page_size: 100 },
1712
+ )
1713
+ ourLogs = (wfLogs.data ?? [])
1714
+ .map((w) => w as Record<string, unknown>)
1715
+ .filter((w) => w.batch_id === ctx.triageRunBatchId)
1716
+ const settled = ourLogs.filter(
1717
+ (w) => ((w.action_execution_logs as unknown[] | undefined) ?? []).length === 3,
1718
+ )
1719
+ if (ourLogs.length === 2 && settled.length === 2) break
1720
+ await new Promise((r) => setTimeout(r, 2_000))
1721
+ }
1722
+ if (ourLogs.length !== 2) {
1723
+ throw new Error(`expected 2 execution logs for the triage run, got ${ourLogs.length}`)
1724
+ }
1725
+
1726
+ for (const wel of ourLogs) {
1727
+ const welId = (wel as { id?: string }).id
1728
+ const aels = (wel as { action_execution_logs?: Array<{ status?: string }> }).action_execution_logs ?? []
1729
+ if (aels.length !== 3) {
1730
+ throw new Error(`log ${welId}: expected 3 action logs (one per band), got ${aels.length}`)
1731
+ }
1732
+ const byStatus: Record<string, number> = {}
1733
+ for (const ael of aels) {
1734
+ const st = ael.status ?? 'unknown'
1735
+ byStatus[st] = (byStatus[st] ?? 0) + 1
1736
+ }
1737
+ const matched = (byStatus.pending ?? 0) + (byStatus.completed ?? 0)
1738
+ const skipped = byStatus.skipped ?? 0
1739
+ if (matched === 1 && skipped === 2) continue // first run for this account
1740
+ if (matched === 0 && skipped === 3) continue // already fired, correctly refusing
1741
+ throw new Error(
1742
+ `log ${welId}: expected 1 matched + 2 skipped, or 3 skipped on a repeat run — got ${JSON.stringify(byStatus)}`,
1743
+ )
1744
+ }
1745
+ ```
1746
+
1747
+ ## 028 — the enterprise account was banded high, and the SMS says so
1748
+
1749
+ §027 proved the fan-out chose one action per row. This proves *which* one,
1750
+ and it is the assertion that survives every rerun because it reads the
1751
+ durable message rather than the run that made it.
1752
+
1753
+ The enterprise account is the one asserted, because it is the only row whose
1754
+ band is genuinely unambiguous — the prompt says enterprise means
1755
+ `priority_high` in as many words, and a model that gets this wrong at
1756
+ temperature zero is a model worth knowing about.
1757
+
1758
+ The search is on the idempotency key, as always. Note the shape of it here:
1759
+ `<customer number>-priority_high` names the message the *high* action would
1760
+ have sent, so finding it is itself the band assertion. There is no need to
1761
+ parse the band back out of the body — though the body is checked too, since a
1762
+ message under the right key carrying the wrong copy would mean the actions
1763
+ are mis-wired to each other's templates.
1764
+
1765
+ ```typescript
1766
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
1767
+ search_query: `m.idempotency_key = '${ctx.enterpriseNumber}-priority_high'`,
1768
+ })
1769
+ if (search.status !== 'completed') {
1770
+ throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
1771
+ }
1772
+
1773
+ const msgDeadline = Date.now() + 90_000
1774
+ let bandBody: string | undefined
1775
+ while (Date.now() < msgDeadline && !bandBody) {
1776
+ const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
1777
+ userSearchId: search.id!,
1778
+ dataAccessMode: 'raw',
1779
+ })
1780
+ bandBody = ((data.data ?? []) as Array<Record<string, unknown>>)
1781
+ .map((m) => String(m.body ?? ''))
1782
+ .find((body) => body.includes('[priority_high]'))
1783
+ if (!bandBody) await new Promise((r) => setTimeout(r, 2_000))
1784
+ }
1785
+ if (!bandBody) {
1786
+ throw new Error(
1787
+ `no priority_high SMS for ${ctx.enterpriseNumber} within 90s — ` +
1788
+ 'the agent banded the enterprise account somewhere else, or the action never fired',
1789
+ )
1790
+ }
1791
+ if (!bandBody.includes(ctx.enterpriseNumber)) {
1792
+ throw new Error(`the banded SMS does not name the account: ${bandBody}`)
1793
+ }
1794
+ ```
1795
+
1796
+ ## 029 — a whole file at once: mint an upload link and PUT the CSV
1797
+
1798
+ Everything so far arrived one row at a time in an ingest call's body. That is
1799
+ the wrong shape for a backfill or a nightly export, where the file already
1800
+ exists and has four thousand rows in it.
1801
+
1802
+ The bulk path is three calls plus a discovery step:
1803
+
1804
+ 1. `datalakes.createUploadLink` returns a presigned PUT `url` and the storage
1805
+ `key` the platform will reference.
1806
+ 2. A **raw `fetch`** PUTs the bytes to that URL. This one goes straight to
1807
+ object storage and is deliberately not an SDK call — the SDK talks to the
1808
+ platform, and the platform is not in this hop.
1809
+ 3. `dataActivationClients.ingestFile` enqueues a job against the key.
1810
+ 4. The worker reads the file, allocates a **fresh `batch_id` it does not
1811
+ return**, and fans out per-row jobs. You discover that id by diffing the
1812
+ client's logs.
1813
+
1814
+ The file is the vendored four-row Stripe export. Its customer numbers are
1815
+ `CUS-10001`–`CUS-10004`, disjoint from the four inline rows, so this scenario
1816
+ adds rows rather than colliding with them.
1817
+
1818
+ ```typescript
1819
+ const { readFileSync } = await import('node:fs')
1820
+ const { join } = await import('node:path')
1821
+ const csvBody = readFileSync(
1822
+ join(process.env.COOKBOOK_FIXTURES_DIR!, 'subscription-saas', 'stripe_customers_batch1.csv'),
1823
+ 'utf8',
1824
+ )
1825
+
1826
+ const { data: link } = await api.datalakes.createUploadLink(tenantSlug, datalakeSlug, {
1827
+ content_type: 'text/csv',
1828
+ filename: 'stripe_customers_batch1.csv',
1829
+ })
1830
+ ctx.uploadKey = link.key!
1831
+
1832
+ const put = await fetch(link.url!, {
1833
+ method: 'PUT',
1834
+ headers: { 'Content-Type': 'text/csv' },
1835
+ body: csvBody,
1836
+ })
1837
+ if (put.status !== 200) {
1838
+ throw new Error(`presigned PUT failed: ${put.status}`)
1839
+ }
1840
+ ```
1841
+
1842
+ ## 030 — enqueue the same file twice, identity first
1843
+
1844
+ The split from §012 applies here too, and the file makes it cheaper rather
1845
+ than more expensive: **one upload, two enqueues**. The key names an object in
1846
+ storage, and nothing about it is bound to a client, so the identity client and
1847
+ the data client can each be pointed at the same bytes.
1848
+
1849
+ The order is the whole point. The identity enqueue runs and is waited on to
1850
+ completion; only then does the data enqueue go, and by that time every subject
1851
+ the second pass will resolve is already committed.
1852
+
1853
+ `ingestFile` returns a `job_id` and no `batch_id`, so each pass snapshots the
1854
+ client's existing batch ids first and then watches for the new one. The gate
1855
+ is `dataset_updated`, not `rows_ingested` — a file whose rows were all read
1856
+ and none written reports the former and not the latter.
1857
+
1858
+ ```typescript
1859
+ ctx.ingestFileAndWait = async (dacSlug: string, expectRows: number): Promise<string> => {
1860
+ const { data: pre } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
1861
+ const seen = new Set(
1862
+ (pre.data ?? [])
1863
+ .map((r) => (r as { batch_id?: string }).batch_id)
1864
+ .filter((b): b is string => typeof b === 'string'),
1865
+ )
1866
+
1867
+ const { data: job } = await api.dataActivationClients.ingestFile(tenantSlug, datalakeSlug, dacSlug, {
1868
+ key: ctx.uploadKey,
1869
+ })
1870
+ if (!job.job_id) throw new Error(`ingestFile on ${dacSlug} did not return a job_id`)
1871
+
1872
+ const deadline = Date.now() + 180_000
1873
+ while (Date.now() < deadline) {
1874
+ const { data } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
1875
+ for (const row of (data.data ?? []) as Array<Record<string, unknown>>) {
1876
+ const b = row.batch_id
1877
+ if (typeof b !== 'string' || seen.has(b)) continue
1878
+ if (row.status === 'partial' || row.status === 'failed') {
1879
+ throw new Error(
1880
+ `bulk batch ${b} on ${dacSlug} did not persist its rows (status: ${row.status}): ` +
1881
+ `${String(row.error ?? 'no reason given')}`,
1882
+ )
1883
+ }
1884
+ if (typeof row.dataset_updated !== 'number' || row.dataset_updated < expectRows) continue
1885
+ if (!Array.isArray(row.output_files) || row.output_files.length === 0) continue
1886
+ return b
1887
+ }
1888
+ await new Promise((r) => setTimeout(r, 1_000))
1889
+ }
1890
+ throw new Error(`no new fully-merged batch appeared on ${dacSlug} within 180s`)
1891
+ }
1892
+
1893
+ // Stage one — four legal entities.
1894
+ await ctx.ingestFileAndWait(ctx.identityDacSlug, 4)
1895
+ // Stage two — four table rows, stamped with the subjects stage one wrote.
1896
+ ctx.bulkBatchId = await ctx.ingestFileAndWait(ctx.dataDacSlug, 4)
1897
+ ```
1898
+
1899
+ ## 031 — every uploaded row landed, and every one is stamped
1900
+
1901
+ Both halves, scoped to the batch the worker allocated. The table side is read
1902
+ through read-only SQL; the identity side through the dataset search, where
1903
+ `le` is the legal-entity alias the query compiler expects.
1904
+
1905
+ The stamp check is the one that matters, and it is the same check §014 ran on
1906
+ the inline rows. Whether a row arrived in a JSON body or in a CSV, it is the
1907
+ same contract pair on the far side, and the proof that both contracts ran is a
1908
+ non-null `legal_entity_id` on a column no contract writes.
1909
+
1910
+ Note what is *not* asserted: an exact legal-entity count. This lake already
1911
+ holds subjects from the inline rows, and a fresh-tenant assumption is exactly
1912
+ the kind that goes red on the second run. The query is scoped to this batch,
1913
+ and four is the floor.
1914
+
1915
+ ```typescript
1916
+ const sqlDeadline = Date.now() + 120_000
1917
+ let tableRows: unknown[][] = []
1918
+ while (Date.now() < sqlDeadline && tableRows.length < 4) {
1919
+ const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
1920
+ sql: `SELECT customer_number, legal_entity_id
1921
+ FROM ${ctx.customersTableName}
1922
+ WHERE customer_number IN ('CUS-10001', 'CUS-10002', 'CUS-10003', 'CUS-10004')
1923
+ ORDER BY customer_number`,
1924
+ mode: 'raw',
1925
+ })
1926
+ if (typeof result !== 'string') tableRows = result.data
1927
+ if (tableRows.length < 4) await new Promise((r) => setTimeout(r, 1_000))
1928
+ }
1929
+ if (tableRows.length !== 4) {
1930
+ throw new Error(`expected the 4 uploaded rows in the table, found ${tableRows.length} within 120s`)
1931
+ }
1932
+ const unstamped = tableRows.filter((row) => !row[1]).map((row) => String(row[0]))
1933
+ if (unstamped.length > 0) {
1934
+ throw new Error(`uploaded rows landed with no legal_entity_id stamp: ${unstamped.join(', ')}`)
1935
+ }
1936
+
1937
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'legal_entity', {
1938
+ search_query: `le.batch_id = '${ctx.bulkBatchId}'`,
1939
+ })
1940
+ if (search.status !== 'completed') {
1941
+ throw new Error(`legal_entity user-search did not compile: ${search.status}`)
1942
+ }
1943
+ const { data: page } = await api.datasets.search(tenantSlug, datalakeSlug, 'legal_entity', {
1944
+ userSearchId: search.id!,
1945
+ dataAccessMode: 'raw',
1946
+ })
1947
+ if (!Array.isArray(page.data)) {
1948
+ throw new Error('legal_entity search returned no rows array')
1949
+ }
1950
+ ```
1951
+
1952
+ ## 032 — the REST tool: let the platform do the fetching
1953
+
1954
+ The third way in does not wait for a file or a call. A **pull-based** client
1955
+ goes and gets the rows itself, on a schedule or on demand, and the platform
1956
+ holds the credential rather than your code.
1957
+
1958
+ The tool carries the API's base URL and its auth. `auth_method: 'bearer'`
1959
+ with a static token is the simplest case; the OAuth2 shape is §036.
1960
+ `status: 'active'` is not decoration — a draft tool is skipped at fetch time,
1961
+ and the symptom is a run that reports success and ingests nothing.
1962
+
1963
+ Locally the target is the integration stack's WireMock on `:8080`, which
1964
+ answers `GET /stripe/v1/customers` with three customers: two individuals and
1965
+ one enterprise, numbered `CUS-20001`–`CUS-20003`.
1966
+
1967
+ ```typescript
1968
+ const restTool = await ctx.ensure(
1969
+ 'Stripe REST tool',
1970
+ async () => {
1971
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
1972
+ return ctx.firstNamed(data.data, 'Cookbook Stripe REST Tool')
1973
+ },
1974
+ async () => {
1975
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
1976
+ name: 'Cookbook Stripe REST Tool',
1977
+ description: 'Mocked Stripe REST API — the pull-based customer fetch.',
1978
+ intent: 'data_exchange',
1979
+ status: 'active',
1980
+ datalake_id: ctx.datalakeId,
1981
+ data_source_id: ctx.dataSourceId,
1982
+ body: {
1983
+ tool_body_type: 'rest_api',
1984
+ auth_method: 'bearer',
1985
+ base_url: 'http://localhost:8080/stripe',
1986
+ bearer_token: 'sk_test_vitest_stripe_token',
1987
+ request_type: 'json',
1988
+ response_type: 'json',
1989
+ timeout_ms: 30_000,
1990
+ },
1991
+ })
1992
+ return data
1993
+ },
1994
+ )
1995
+ ctx.restToolId = restTool.id!
1996
+ ```
1997
+
1998
+ ## 033 — two fetch clients, for the same reason as before
1999
+
2000
+ The client's `tool_call` is what turns a REST tool into a fetch: the HTTP
2001
+ `method` and the `path` to append to the tool's `base_url`, plus a
2002
+ `pagination_context_template` declaring whether there is more to get. This
2003
+ endpoint answers in one page, so it declares `has_next: false`.
2004
+
2005
+ The `response_extractor` unwraps the API's envelope. Stripe answers
2006
+ `{ object: "list", data: [...] }`, and each element of that array has to
2007
+ become one row — `{{ msg.data | to_json }}` is the whole of it.
2008
+
2009
+ And the pair is split across two clients again. Nothing about the fetch
2010
+ changes the reason: two contracts on one client resolve the same subject at
2011
+ the same instant, and one of them loses. Both clients hit the same endpoint
2012
+ and get the same three customers; they differ only in what they do with them.
2013
+
2014
+ ```typescript
2015
+ const restFetchCall = {
2016
+ tool_call_type: 'restapi_request' as const,
2017
+ method: 'get' as const,
2018
+ path: { type: 'custom' as const, body: '/v1/customers' },
2019
+ pagination_context_template: { type: 'custom' as const, body: '{"has_next": false}' },
2020
+ }
2021
+ const restExtractor = { type: 'custom' as const, body: '{{ msg.data | to_json }}' }
2022
+
2023
+ const restIdentityDac = await ctx.ensure(
2024
+ 'REST identity client',
2025
+ async () => {
2026
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
2027
+ return ctx.firstNamed(data.data, 'Cookbook Stripe REST Identity DAC')
2028
+ },
2029
+ async () => {
2030
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
2031
+ name: 'Cookbook Stripe REST Identity DAC',
2032
+ description: 'Fetches Stripe customers and writes their legal entities. Runs FIRST.',
2033
+ tool_id: ctx.restToolId,
2034
+ data_source_id: ctx.dataSourceId,
2035
+ tool_call: restFetchCall,
2036
+ response_extractor: restExtractor,
2037
+ interoperability_contract_ids: [ctx.leContractId],
2038
+ })
2039
+ return data
2040
+ },
2041
+ )
2042
+ ctx.restIdentityDacSlug = restIdentityDac.slug!
2043
+
2044
+ const restDataDac = await ctx.ensure(
2045
+ 'REST data client',
2046
+ async () => {
2047
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
2048
+ return ctx.firstNamed(data.data, 'Cookbook Stripe REST Data DAC')
2049
+ },
2050
+ async () => {
2051
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
2052
+ name: 'Cookbook Stripe REST Data DAC',
2053
+ description: 'Fetches the same customers and writes the table rows. Runs SECOND.',
2054
+ tool_id: ctx.restToolId,
2055
+ data_source_id: ctx.dataSourceId,
2056
+ tool_call: restFetchCall,
2057
+ response_extractor: restExtractor,
2058
+ interoperability_contract_ids: [ctx.gtContractId],
2059
+ })
2060
+ return data
2061
+ },
2062
+ )
2063
+ ctx.restDataDacSlug = restDataDac.slug!
2064
+ ```
2065
+
2066
+ ## 034 — trigger both fetches, in order
2067
+
2068
+ `runManually` differs from `ingest` and `ingestFile` in a useful way: it
2069
+ returns the `batch_id` immediately, because the platform allocates it when it
2070
+ schedules the fetch rather than when a worker opens a file. The HTTP call and
2071
+ the ingestion still happen in the background.
2072
+
2073
+ Same staging, same reason. The identity fetch is waited on to completion
2074
+ before the data fetch is triggered.
2075
+
2076
+ ```typescript
2077
+ const { data: identityRun } = await api.dataActivationClients.runManually(
2078
+ tenantSlug, datalakeSlug, ctx.restIdentityDacSlug,
2079
+ )
2080
+ if (!identityRun.batch_id) throw new Error('the identity fetch did not return a batch_id')
2081
+ await ctx.waitForBatches(ctx.restIdentityDacSlug, [identityRun.batch_id], 180_000)
2082
+
2083
+ const { data: dataRun } = await api.dataActivationClients.runManually(
2084
+ tenantSlug, datalakeSlug, ctx.restDataDacSlug,
2085
+ )
2086
+ if (!dataRun.batch_id) throw new Error('the data fetch did not return a batch_id')
2087
+ ctx.restBatchId = dataRun.batch_id
2088
+ await ctx.waitForBatches(ctx.restDataDacSlug, [ctx.restBatchId], 180_000)
2089
+ ```
2090
+
2091
+ ## 035 — three fetched customers, in the same table as everything else
2092
+
2093
+ This is the point of the whole scenario, and it is worth saying plainly:
2094
+ these three rows are indistinguishable from the four that arrived inline and
2095
+ the four that arrived in a CSV. Same table, same contracts, same stamp. The
2096
+ entry point is the only thing that differed, and it stopped mattering the
2097
+ moment the row was extracted.
2098
+
2099
+ The read filters on `customer_number`, a `none` column. The tokenized columns
2100
+ would answer here too — this lake has no derived copy, so raw is all there is
2101
+ — but filtering on a column declared `tokenize` is a habit that breaks the day
2102
+ someone provisions the pair.
2103
+
2104
+ ```typescript
2105
+ const sqlDeadline = Date.now() + 120_000
2106
+ let fetched: unknown[][] = []
2107
+ while (Date.now() < sqlDeadline && fetched.length < 3) {
2108
+ const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
2109
+ sql: `SELECT customer_number, customer_type, legal_entity_id
2110
+ FROM ${ctx.customersTableName}
2111
+ WHERE customer_number IN ('CUS-20001', 'CUS-20002', 'CUS-20003')
2112
+ ORDER BY customer_number`,
2113
+ mode: 'raw',
2114
+ })
2115
+ if (typeof result !== 'string') fetched = result.data
2116
+ if (fetched.length < 3) await new Promise((r) => setTimeout(r, 1_000))
2117
+ }
2118
+ if (fetched.length !== 3) {
2119
+ throw new Error(`expected the 3 fetched rows in the table, found ${fetched.length} within 120s`)
2120
+ }
2121
+ const unstamped = fetched.filter((row) => !row[2]).map((row) => String(row[0]))
2122
+ if (unstamped.length > 0) {
2123
+ throw new Error(`fetched rows landed with no legal_entity_id stamp: ${unstamped.join(', ')}`)
2124
+ }
2125
+ const enterprise = fetched.find((row) => String(row[1]) === 'enterprise')
2126
+ if (!enterprise) {
2127
+ throw new Error('the fetch did not produce the enterprise customer the endpoint returns')
2128
+ }
2129
+ ```
2130
+
2131
+ ## 036 — the OAuth2 variant, for an API that will not take a static token
2132
+
2133
+ Most real APIs will not. When the credential is an OAuth2 grant, the tool
2134
+ carries the whole grant config instead of a token, and **the platform
2135
+ exchanges the stored refresh token for an access token server-side before
2136
+ each fetch**. Your code never holds one, never refreshes one, never logs one.
2137
+
2138
+ `oauth2_grant_type` names the flow the credential came from —
2139
+ `authorization_code` here, the ordinary three-legged flow whose long-lived
2140
+ refresh token you keep. It is not `refresh_token`: the enum is
2141
+ `client_credentials | authorization_code`, and refreshing is what the
2142
+ platform does with the grant rather than a grant of its own.
2143
+
2144
+ Everything else is unchanged — the client, its `tool_call`, the contracts,
2145
+ the table. Swapping auth is a tool-level edit.
2146
+
2147
+ This step creates the tool to show the shape and stops there; the mock
2148
+ endpoint has no OAuth2 server behind it, and a fetch that cannot succeed
2149
+ proves less than a create that validates.
2150
+
2151
+ ```typescript
2152
+ const oauthTool = await ctx.ensure(
2153
+ 'OAuth2 REST tool',
2154
+ async () => {
2155
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
2156
+ return ctx.firstNamed(data.data, 'Cookbook Stripe OAuth2 Tool')
2157
+ },
2158
+ async () => {
2159
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
2160
+ name: 'Cookbook Stripe OAuth2 Tool',
2161
+ description: 'Shape reference — the same fetch behind an OAuth2 refresh-token grant.',
2162
+ intent: 'data_exchange',
2163
+ status: 'active',
2164
+ datalake_id: ctx.datalakeId,
2165
+ data_source_id: ctx.dataSourceId,
2166
+ body: {
2167
+ tool_body_type: 'rest_api',
2168
+ auth_method: 'oauth2',
2169
+ base_url: 'http://localhost:8080/stripe',
2170
+ oauth2_grant_type: 'authorization_code',
2171
+ oauth2_token_url: 'http://localhost:8080/oauth/token',
2172
+ oauth2_client_id: 'cookbook-client',
2173
+ oauth2_client_secret: 'cookbook-secret',
2174
+ oauth2_refresh_token: 'cookbook-refresh-token',
2175
+ oauth2_token_ttl: 3300,
2176
+ request_type: 'json',
2177
+ response_type: 'json',
2178
+ timeout_ms: 30_000,
2179
+ },
2180
+ })
2181
+ return data
2182
+ },
2183
+ )
2184
+ if (!oauthTool.id) {
2185
+ throw new Error('the OAuth2 tool was not persisted')
2186
+ }
2187
+ ```
2188
+
2189
+ ## 037 — a billing team is not one person
2190
+
2191
+ Everything so far ran as one admin. The last scenario adds a second person to
2192
+ the tenant, entirely through the SDK — no console, no out-of-band step.
2193
+
2194
+ The flow touches all three session scopes, and mixing them up is the single
2195
+ most common mistake here:
2196
+
2197
+ - **tenant-scoped** — the admin creates the invitation. Only an admin can.
2198
+ - **root** — signs the new user up and confirms them. A tenant admin can
2199
+ invite, but only root can create the underlying account.
2200
+ - **tenantless** — the invitee, who belongs to no tenant yet, lists the
2201
+ invitations waiting for them and accepts one.
2202
+
2203
+ Start by confirming the current session really is a tenant-scoped admin of
2204
+ the expected tenant. `sessions.verify` echoes both.
2205
+
2206
+ ```typescript
2207
+ const { data: who } = await api.sessions.verify()
2208
+ if (who.tenant?.slug !== tenantSlug) {
2209
+ throw new Error(`expected an admin session on ${tenantSlug}, got ${who.tenant?.slug}`)
2210
+ }
2211
+ if (!/admin/i.test(who.role?.name ?? '')) {
2212
+ throw new Error(`expected an admin role, got ${who.role?.name}`)
2213
+ }
2214
+ ```
2215
+
2216
+ ## 038 — the teammate's account, created by root
2217
+
2218
+ The invitee needs an account before they can accept anything. Signup and
2219
+ confirmation are root-scoped, so this runs on the root client from §001 —
2220
+ `confirmUser` is what activates the account so they can sign in at all.
2221
+
2222
+ The email is stable, like every other name in this walk, so the second run
2223
+ finds the account already there. Same catch-and-inspect as §001: a refusal
2224
+ that says the address is taken is the idempotent path, and anything else is
2225
+ re-raised.
2226
+
2227
+ ```typescript
2228
+ ctx.teammateEmail = 'cookbook-subscription-teammate@dev.local'
2229
+ ctx.teammatePassword = 'CookbookPass1!'
2230
+
2231
+ try {
2232
+ const { data: teammate } = await ctx.rootApi.admin.signUp({
2233
+ email: ctx.teammateEmail,
2234
+ password: ctx.teammatePassword,
2235
+ first_name: 'Emma',
2236
+ last_name: 'Wilson',
2237
+ })
2238
+ await ctx.rootApi.admin.confirmUser(teammate.id!)
2239
+ } catch (err) {
2240
+ const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
2241
+ if (!/taken|already|exist/i.test(detail)) throw err
2242
+ }
2243
+ ```
2244
+
2245
+ ## 039 — the invitation
2246
+
2247
+ The admin invites by email and role. The invite is keyed on
2248
+ `(tenant, email)`, so a re-invitation is refused — and it is refused with
2249
+ **two different messages depending on how far the first run got**. While the
2250
+ invitation is still pending, `/email` says *"already invited"*. Once it has
2251
+ been accepted and consumed, the same call says *"already in this
2252
+ organization"*, because at that point the address does not belong to a
2253
+ pending invite but to a member.
2254
+
2255
+ Both are the idempotent path and both are caught. What is **not** caught is
2256
+ anything else, and that distinction is the whole reason this is a `try`
2257
+ rather than a bare call: an invitation that fails because the tenant is
2258
+ misconfigured must not be read as "already there".
2259
+
2260
+ ```typescript
2261
+ try {
2262
+ const { data: invite } = await api.invitations.create(tenantSlug, {
2263
+ email: ctx.teammateEmail,
2264
+ role: 'member',
2265
+ })
2266
+ if (invite.email !== ctx.teammateEmail || invite.role !== 'member') {
2267
+ throw new Error(`unexpected invitation: ${JSON.stringify(invite)}`)
2268
+ }
2269
+ } catch (err) {
2270
+ const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
2271
+ // Both refusals, verbatim from the server. Matching loosely here is how a
2272
+ // real failure gets swallowed, so the two known messages are named.
2273
+ if (!/already invited|already in this organization/i.test(detail)) throw err
2274
+ }
2275
+ ```
2276
+
2277
+ ## 040 — the teammate signs in with no tenant, and accepts
2278
+
2279
+ The invitee has an account and belongs nowhere, so they sign in **tenantless**
2280
+ — `createBootstrapSession`, not `createSession`, which is tenant-login only.
2281
+ The resulting session carries `tenant: null`, and that is the assertion worth
2282
+ making: a client that quietly came back tenant-scoped would list the wrong
2283
+ invitations.
2284
+
2285
+ Then the branch that makes this step idempotent. On the first run there is a
2286
+ pending invitation for this tenant and it is accepted. On every run after, the
2287
+ invitation has already been consumed and the list is empty — which is not a
2288
+ failure but the same end state reached earlier. **The proof is deferred to
2289
+ §041**, which asserts the end state directly rather than trusting either
2290
+ branch.
2291
+
2292
+ Note that accepting does not upgrade the tenantless session in place. It
2293
+ creates the membership; the session that uses it is minted fresh in the next
2294
+ step.
2295
+
2296
+ ```typescript
2297
+ const teammateTenantless = await createBootstrapSession({
2298
+ baseUrl: process.env.ALVERA_BASE_URL!,
2299
+ email: ctx.teammateEmail,
2300
+ password: ctx.teammatePassword,
2301
+ })
2302
+ if (teammateTenantless.tenant !== null) {
2303
+ throw new Error('expected a tenantless session — tenant should be null')
2304
+ }
2305
+ const teammateTenantlessApi = createIsolatedPlatformApi({
2306
+ baseUrl: process.env.ALVERA_BASE_URL!,
2307
+ sessionToken: teammateTenantless.sessionToken,
2308
+ apiKey: '',
2309
+ })
2310
+
2311
+ const { data: invites } = await teammateTenantlessApi.invitations.list()
2312
+ const pending = (invites.data ?? []).find(
2313
+ (i: { tenant?: { slug?: string } }) => i.tenant?.slug === tenantSlug,
2314
+ )
2315
+
2316
+ if (pending?.id) {
2317
+ const { data: membership } = await teammateTenantlessApi.invitations.accept(pending.id)
2318
+ if (membership.tenant?.slug !== tenantSlug || membership.role !== 'member') {
2319
+ throw new Error(`unexpected membership: ${JSON.stringify(membership)}`)
2320
+ }
2321
+ } else {
2322
+ console.log(' ↻ no pending invitation — already accepted on an earlier run')
2323
+ }
2324
+ ```
2325
+
2326
+ ## 041 — the teammate signs in tenant-scoped, and is a member
2327
+
2328
+ This is the assertion the whole scenario is for, and it is written against the
2329
+ end state rather than against the path taken to reach it — which is what makes
2330
+ it identical on run one and run fifty.
2331
+
2332
+ The tenant-scoped sign-in needs the tenant's API key, the one minted in §004.
2333
+ Tenant login without a resolvable `X-API-Key` is a 401 regardless of how good
2334
+ the password is; the key identifies the tenant before the credentials
2335
+ identify the person.
2336
+
2337
+ The role is the other half. A session that came back with an admin role would
2338
+ mean the invitation's `role: 'member'` was ignored, and everything downstream
2339
+ that trusts the platform to enforce a role would be resting on nothing.
2340
+
2341
+ ```typescript
2342
+ const teammateScoped = await createSession({
2343
+ baseUrl: process.env.ALVERA_BASE_URL!,
2344
+ email: ctx.teammateEmail,
2345
+ password: ctx.teammatePassword,
2346
+ tenantSlug,
2347
+ apiKey: ctx.tenantApiKey,
2348
+ })
2349
+ if (teammateScoped.tenant?.slug !== tenantSlug) {
2350
+ throw new Error(`expected a tenant-scoped session on ${tenantSlug}, got ${teammateScoped.tenant?.slug}`)
2351
+ }
2352
+ if (!/member/i.test(teammateScoped.role?.name ?? '')) {
2353
+ throw new Error(`expected a member role, got ${teammateScoped.role?.name}`)
2354
+ }
2355
+ ```
2356
+
2357
+ # Outcome
2358
+
2359
+ After all forty-one steps run green, on one tenant and **one raw datalake**:
2360
+
2361
+ - **One customers table** holds eleven rows that arrived three different
2362
+ ways — four inline through an ingest call, four from a CSV through a
2363
+ presigned upload, three pulled from a REST API — and every one of them
2364
+ carries a `legal_entity_id` stamp written by a contract that never
2365
+ declares the column.
2366
+ - **Three workflows** run over that one table and disagree about which rows
2367
+ matter. The welcome workflow gates on reachability, the dunning workflow on
2368
+ reachability *and* KYC, and the triage workflow hands the question to a
2369
+ model and fans out on its answer.
2370
+ - **Three outbound messages** are readable from the `message` dataset, each
2371
+ under its own idempotency key, two carrying deep links that resolve against
2372
+ two separate connected apps on two separate routes.
2373
+ - **A second person** is on the tenant with a `member` role, invited and
2374
+ accepted entirely through the SDK.
2375
+
2376
+ Four things this walk teaches that are easy to learn the expensive way:
2377
+
2378
+ 1. **Never bind a contract pair to one activation client.** The fan-out
2379
+ staggers by row, so a pair on one client resolves the same subject twice
2380
+ at the same instant and the unique index refuses the loser. Split them and
2381
+ sequence them — §012, and again at §030 and §033.
2382
+ 2. **Gate on `dataset_updated`, not `rows_ingested`.** Received is not
2383
+ written. A refused row reports the first and not the second, and a walk
2384
+ gated on the wrong one goes green over an empty table.
2385
+ 3. **Search the idempotency key, never the workflow id.** The key is built
2386
+ from a stable business id and names the same message forever; a workflow
2387
+ id changes the moment the resource is recreated.
2388
+ 4. **Assert the end state, not the transition.** An idempotency key fires
2389
+ once ever, so every assertion here accepts both the first run's shape and
2390
+ the repeat run's — §027 spells that out, and §040 defers its proof to
2391
+ §041 for the same reason.
2392
+
2393
+ # See also
2394
+
2395
+ - `payments-compliance.md` — the same contract-pair shape on a tokenized
2396
+ surface, with the derived lakes provisioned
2397
+ - `organic-marketing.md` — three scenarios on one raw lake, the sibling of
2398
+ this walk
2399
+ - `../workflows.md` — the standard workflow primitive, `context_datasets`,
2400
+ and `workflows.run`
2401
+ - `../data_activation_clients.md` — data source → tool → contract → client
2402
+ - `../interoperability_contracts.md` — `template_config` vs `mdm_input_config`
2403
+ - `../connected_apps.md` — registration, `resolvePage`, message tracking