@alvera-ai/platform-sdk 0.17.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/.agent/AGENTS.md +82 -144
  2. package/.agent/account_management.md +2 -2
  3. package/.agent/action_logs.md +4 -4
  4. package/.agent/ai_agents.md +28 -21
  5. package/.agent/ai_sandbox.md +49 -39
  6. package/.agent/connected_apps.md +3 -3
  7. package/.agent/cookbook/_fixtures/README.md +1 -1
  8. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_generic_table.liquid +1 -1
  9. package/.agent/cookbook/_fixtures/organic-marketing/_lead_submissions_foundation_legal_entity.liquid +80 -0
  10. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_mdm.liquid +2 -1
  11. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_generic_table.liquid +57 -0
  12. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_legal_entity.liquid +30 -0
  13. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_mdm.liquid +44 -0
  14. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_generic_table.liquid +57 -0
  15. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_legal_entity.liquid +36 -0
  16. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_mdm.liquid +41 -0
  17. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_generic_table.liquid +70 -0
  18. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_legal_entity.liquid +52 -0
  19. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_mdm.liquid +42 -0
  20. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_generic_table.liquid +38 -0
  21. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_legal_entity.liquid +48 -0
  22. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_mdm.liquid +49 -0
  23. package/.agent/cookbook/organic-marketing.md +2801 -0
  24. package/.agent/cookbook/payments-compliance.md +2180 -0
  25. package/.agent/cookbook/primary-care.md +2175 -0
  26. package/.agent/cookbook/subscription-saas.md +2403 -0
  27. package/.agent/data_activation_clients.md +65 -52
  28. package/.agent/datalakes.md +338 -171
  29. package/.agent/errors.md +3 -3
  30. package/.agent/generic_tables.md +151 -62
  31. package/.agent/interoperability_contracts.md +57 -22
  32. package/.agent/mdm.md +136 -153
  33. package/.agent/messages.md +36 -34
  34. package/.agent/mock-services.md +1 -1
  35. package/.agent/mutations.md +2 -2
  36. package/.agent/templates.md +14 -13
  37. package/.agent/tools.md +63 -21
  38. package/.agent/type_naming.md +13 -13
  39. package/.agent/workflows.md +99 -53
  40. package/README.md +2 -2
  41. package/dist/bin/platform-sdk.mjs +33 -47
  42. package/dist/bin/platform-sdk.mjs.map +1 -1
  43. package/dist/index.d.mts +565 -379
  44. package/dist/index.d.mts.map +1 -1
  45. package/dist/index.mjs +494 -59
  46. package/dist/index.mjs.map +1 -1
  47. package/package.json +4 -3
  48. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_legal_entity.liquid +0 -88
  49. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_appointment.liquid +0 -47
  50. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_mdm.liquid +0 -24
  51. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_patient.liquid +0 -38
  52. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_compliance_screening.liquid +0 -59
  53. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_mdm.liquid +0 -36
  54. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_mdm.liquid +0 -30
  55. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_payment_account.liquid +0 -55
  56. package/.agent/cookbook/_fixtures/subscription/_customers_subscription_mdm.liquid +0 -20
  57. package/.agent/cookbook/_setup/foundation.md +0 -359
  58. package/.agent/cookbook/_setup/healthcare.md +0 -361
  59. package/.agent/cookbook/_setup/payments.md +0 -365
  60. package/.agent/cookbook/_setup/subscription.md +0 -364
  61. package/.agent/cookbook/action-status-updaters.md +0 -278
  62. package/.agent/cookbook/ai-agent-invoke.md +0 -279
  63. package/.agent/cookbook/appointment-review-sms-workflow.md +0 -801
  64. package/.agent/cookbook/birthday-greeting-sms-trigger.md +0 -696
  65. package/.agent/cookbook/bulk-ingest.md +0 -302
  66. package/.agent/cookbook/contact-us-triage-with-llm.md +0 -663
  67. package/.agent/cookbook/dunning-sms-for-delinquent.md +0 -659
  68. package/.agent/cookbook/generic-tables.md +0 -244
  69. package/.agent/cookbook/invite-team.md +0 -200
  70. package/.agent/cookbook/kyc-notification-on-account-activation.md +0 -661
  71. package/.agent/cookbook/marketing-campaign-send.md +0 -1044
  72. package/.agent/cookbook/paginated-restapi-poller.md +0 -383
  73. package/.agent/cookbook/rest-fetch.md +0 -273
  74. package/.agent/cookbook/sanctions-screening-with-agent-review.md +0 -773
  75. package/.agent/cookbook/score-leads-with-llm-categorization.md +0 -665
  76. package/.agent/cookbook/system-templates.md +0 -165
  77. package/.agent/cookbook/talk-to-data.md +0 -178
  78. package/.agent/cookbook/triage-prospects-by-priority.md +0 -571
  79. package/.agent/cookbook/welcome-sms-for-customers.md +0 -647
  80. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/memorandum-of-association-01.png +0 -0
  81. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/sample_two_page.pdf +0 -0
  82. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/_customers_subscription_customer.liquid +0 -0
  83. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/stripe_customers_batch1.csv +0 -0
@@ -0,0 +1,2175 @@
1
+ ---
2
+ title: "Primary care: review, triage, extract — on a tokenized lake"
3
+ summary: "The whole primary-care surface in one walk. Stand up a tenant, a raw datalake and the tokenized and redacted copies, then run three scenarios end to end — a post-visit review request that only reaches patients whose appointment was actually fulfilled, an LLM agent triaging inbound contact-us messages into a category-keyed fan-out, and a vision agent reading a scanned document. Ends by reading one appointment back through all three lakes to show what each reader sees, then inspecting how far the tokenized copy has got and repairing one of its tables."
4
+ use_case: primary-care-feedback
5
+ slug: primary-care
6
+ vitest_source:
7
+ - integration-tests/tests/primary-care-feedback/standard-workflow.test.ts
8
+ - integration-tests/tests/primary-care-feedback/agent-driven-workflow.test.ts
9
+ - integration-tests/tests/primary-care-feedback/ai-agent-invoke.test.ts
10
+ - integration-tests/tests/workspace/bootstrap.test.ts
11
+ status: draft
12
+ ---
13
+
14
+ # Problem
15
+
16
+ A primary-care practice generates three kinds of work that all touch the same
17
+ patient and none of which a person should have to remember.
18
+
19
+ A visit finishes and the practice wants to ask how it went — but only about
20
+ visits that actually happened. A patient who cancelled must not be asked to
21
+ review the appointment they did not attend, and the difference between those
22
+ two cases is one field on one row.
23
+
24
+ Messages arrive through the website and most of them are not clinical. Some
25
+ are. Something has to read each one and route it before a human decides
26
+ whether it was urgent.
27
+
28
+ And documents arrive as scans. A referral, a registration form, an insurance
29
+ letter — pixels, from which someone has to type the same six fields into the
30
+ same system every time.
31
+
32
+ **This surface is tokenized, and that is the whole reason it is separate from
33
+ the other three cookbooks.** Everything here is patient data. The triage agent
34
+ in §023 reads inbound messages and has no business knowing who sent them; the
35
+ extraction agent in §031 reads documents. Both are machines, and the tokenized
36
+ copy exists so a machine can be given the signal without the identity. The
37
+ redacted copy exists for the other reader — a person on a support screen who
38
+ needs to see that a message came in without seeing whose it was.
39
+
40
+ # Composition
41
+
42
+ | Resource | Why it exists here |
43
+ |-----------------------------------|--------------------------------------------|
44
+ | Tenant + raw, tokenized, redacted | Three lakes, because two readers are masked|
45
+ | Appointments table | The review workflow's event source |
46
+ | Contact Us table | The triage workflow's event source |
47
+ | SMS tool, LLM tool, vision tool | One dispatcher, two models |
48
+ | Two AI agents | Triage, and document extraction |
49
+ | Two workflows | Review request, and category fan-out |
50
+ | Connected app | The review form the SMS links to |
51
+
52
+ # Walkthrough
53
+
54
+ This cookbook is **self-reliant**: it stands up everything it uses and depends
55
+ on no other file. It is also **idempotent** — every creating step looks first
56
+ and creates only what is missing, so running it twice costs what running it
57
+ once cost. That matters more here than on the raw surfaces: this walk
58
+ provisions three datalakes rather than one, and nothing in the platform
59
+ reclaims an abandoned lake.
60
+
61
+ ## 001 — sign in as root, and make sure the admin user exists
62
+
63
+ Authenticate as the platform's root admin (`admin@dev.local` /
64
+ `devpassword` in local dev) via the tenantless bootstrap login — keyless by
65
+ structural necessity, since no tenant exists yet to scope a key to, and a
66
+ dev/test-only surface.
67
+
68
+ Then make sure the admin user this walk runs as exists. **The email is
69
+ stable, not per-run**, which is what makes the step idempotent — and it means
70
+ the second run finds the user already signed up. A duplicate signup is
71
+ refused, so the refusal is caught and inspected: if it says the email is
72
+ taken, that is the idempotent path and the walk continues. Any other failure
73
+ is re-raised, because swallowing it would turn a real auth problem into a
74
+ confusing failure three steps later.
75
+
76
+ ```typescript
77
+ ctx.rootSession = await createBootstrapSession({
78
+ baseUrl: process.env.ALVERA_BASE_URL!,
79
+ email: process.env.ALVERA_ROOT_EMAIL!,
80
+ password: process.env.ALVERA_ROOT_PASSWORD!,
81
+ })
82
+ ctx.rootApi = createIsolatedPlatformApi({
83
+ baseUrl: process.env.ALVERA_BASE_URL!,
84
+ sessionToken: ctx.rootSession.sessionToken,
85
+ apiKey: '',
86
+ })
87
+
88
+ ctx.practiceEmail = 'cookbook-primary-care@dev.local'
89
+ ctx.practicePassword = 'CookbookPass1!'
90
+
91
+ try {
92
+ const signUpResp = await ctx.rootApi.admin.signUp({
93
+ email: ctx.practiceEmail,
94
+ password: ctx.practicePassword,
95
+ first_name: 'Cookbook',
96
+ last_name: 'Practice',
97
+ })
98
+ await ctx.rootApi.admin.confirmUser(signUpResp.data.id!)
99
+ } catch (err) {
100
+ const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
101
+ if (!/taken|already|exist/i.test(detail)) throw err
102
+ }
103
+ ```
104
+
105
+ ## 002 — the find-or-create helper every later step uses
106
+
107
+ Idempotence is one question asked over and over: *is this already here?*
108
+ Rather than answer it a dozen different ways, the walk defines it once.
109
+
110
+ `ensure` takes a label, a lookup and a create. It runs the lookup, returns
111
+ what it finds, and only creates when the lookup comes back empty. The label
112
+ is not decoration — when a run reuses something you did not expect it to, the
113
+ log line naming it is how you find out.
114
+
115
+ Alongside it, `firstNamed` — the lookup half. **A list endpoint pages at
116
+ twenty and is not newest-first**, so a `.find()` over the default page
117
+ silently stops finding things as soon as a lake has a few of them, and the
118
+ failure looks like the resource was never created. `page_size: 100` is the
119
+ cheap fix.
120
+
121
+ Both live on `ctx` rather than as bare functions because each numbered step
122
+ compiles into its own `it()` block.
123
+
124
+ ```typescript
125
+ ctx.ensure = async <T>(
126
+ label: string,
127
+ find: () => Promise<T | undefined>,
128
+ create: () => Promise<T>,
129
+ ): Promise<T> => {
130
+ const existing = await find()
131
+ if (existing !== undefined) {
132
+ console.log(` ↻ reusing ${label}`)
133
+ return existing
134
+ }
135
+ console.log(` + creating ${label}`)
136
+ return await create()
137
+ }
138
+
139
+ ctx.firstNamed = <T extends { name?: string }>(
140
+ rows: readonly T[] | undefined,
141
+ name: string,
142
+ ): T | undefined => (rows ?? []).find((r) => r.name === name)
143
+ ```
144
+
145
+ ## 003 — the tenant
146
+
147
+ The practice admin signs in without a tenant scope — they may not belong to
148
+ one yet — and the walk finds or creates the tenant by its stable name. The
149
+ server derives the slug; capture it, because every later call is addressed by
150
+ it.
151
+
152
+ ```typescript
153
+ ctx.practiceTenantlessSession = await createBootstrapSession({
154
+ baseUrl: process.env.ALVERA_BASE_URL!,
155
+ email: ctx.practiceEmail,
156
+ password: ctx.practicePassword,
157
+ })
158
+ ctx.practiceTenantlessApi = createIsolatedPlatformApi({
159
+ baseUrl: process.env.ALVERA_BASE_URL!,
160
+ sessionToken: ctx.practiceTenantlessSession.sessionToken,
161
+ apiKey: '',
162
+ })
163
+
164
+ const TENANT_NAME = 'Cookbook Primary Care'
165
+
166
+ const tenant = await ctx.ensure(
167
+ `tenant ${TENANT_NAME}`,
168
+ async () => {
169
+ const { data } = await ctx.practiceTenantlessApi.tenants.list()
170
+ return (data.data ?? []).find((t: { name?: string }) => t.name === TENANT_NAME)
171
+ },
172
+ async () => {
173
+ const { data } = await ctx.practiceTenantlessApi.tenants.create({ name: TENANT_NAME })
174
+ return data
175
+ },
176
+ )
177
+ tenantSlug = tenant.slug!
178
+ ```
179
+
180
+ ## 004 — the tenant-scoped client, at the raw ceiling
181
+
182
+ A tenant-scoped login requires `X-API-Key`, so a key has to exist before the
183
+ admin can sign in against the tenant. Mint one through the platform-admin
184
+ side door with the root bearer; in the web console this is *Settings → API
185
+ Keys*.
186
+
187
+ `data_access_mode: 'raw'` is the **ceiling**, not the default view, and on a
188
+ tokenized surface the distinction finally matters. The key says how far this
189
+ session is *allowed* to see; each read then names the tier it wants. A raw
190
+ ceiling is what lets §033 ask the same row for all three answers and compare
191
+ them. A tokenized-ceiling key asking for `raw` is refused — which is the
192
+ right behaviour, and the reason a support console gets its own key.
193
+
194
+ **This is the one step in the walk that is not idempotent.** There is no
195
+ endpoint that lists a tenant's API keys, so `ensure` has no lookup to run and
196
+ each run mints another. A key row is cheap where a datalake is not.
197
+
198
+ ```typescript
199
+ const { data: mintedKey } = await ctx.rootApi.admin.createTenantApiKey(tenantSlug, {
200
+ name: 'Cookbook Primary Care Key',
201
+ data_access_mode: 'raw',
202
+ })
203
+ ctx.tenantApiKey = mintedKey.api_key
204
+
205
+ const practiceTenantSession = await createSession({
206
+ baseUrl: process.env.ALVERA_BASE_URL!,
207
+ email: ctx.practiceEmail,
208
+ password: ctx.practicePassword,
209
+ tenantSlug,
210
+ apiKey: ctx.tenantApiKey,
211
+ })
212
+ api = createIsolatedPlatformApi({
213
+ baseUrl: process.env.ALVERA_BASE_URL!,
214
+ sessionToken: practiceTenantSession.sessionToken,
215
+ apiKey: ctx.tenantApiKey,
216
+ })
217
+ ```
218
+
219
+ ## 005 — the raw datalake
220
+
221
+ A datalake is created `raw` and stays raw. The derived copies are a separate
222
+ call, in §007, and this one has to be `ready` before that call is made.
223
+
224
+ The lookup filters on `type === 'raw'`, and here that is load-bearing rather
225
+ than defensive: from §007 onward `datalakes.list` returns **three** rows, two
226
+ of which are copies. A lookup that matched on name alone would start
227
+ returning whichever came first and the walk would begin writing to a lake
228
+ that is meant to be read-only.
229
+
230
+ Local-dev defaults match the seeded `dev.exs` setup — `postgres` on
231
+ `localhost:5432`, database `alvera_dev_foundation`, LocalStack S3 on
232
+ `localhost:4566` — and the schema name is stable, because a fresh schema per
233
+ run is the same leak as a fresh tenant per run.
234
+
235
+ ```typescript
236
+ const DB = { host: 'localhost', port: 5432, user: 'postgres', pass: 'postgres', name: 'alvera_dev_foundation' }
237
+ const DB_SCHEMA = 'cookbook_primary_care'
238
+ const LAKE_NAME = 'Cookbook Primary Care Datalake'
239
+
240
+ const S3 = {
241
+ cloud_storage_type: 'aws' as const,
242
+ region: 'us-east-1',
243
+ access_key_id: 'test',
244
+ secret_access_key: 'test',
245
+ endpoint: 'http://localhost:4566',
246
+ }
247
+
248
+ const datalake = await ctx.ensure(
249
+ `datalake ${LAKE_NAME}`,
250
+ async () => {
251
+ const { data } = await api.datalakes.list(tenantSlug)
252
+ return (data.data ?? []).find(
253
+ (l: { name?: string; type?: string }) => l.name === LAKE_NAME && l.type === 'raw',
254
+ )
255
+ },
256
+ async () => {
257
+ const { data } = await api.datalakes.create(tenantSlug, {
258
+ name: LAKE_NAME,
259
+ description: 'Primary care datalake provisioned by the cookbook doctest.',
260
+ timezone: 'America/New_York',
261
+ pool_size: 3,
262
+ type: 'raw',
263
+
264
+ db_writer_host: DB.host,
265
+ db_writer_port: DB.port,
266
+ db_writer_name: DB.name,
267
+ db_writer_schema: DB_SCHEMA,
268
+ db_writer_auth_method: 'password',
269
+ db_writer_user: DB.user,
270
+ db_writer_pass: DB.pass,
271
+ db_writer_enable_ssl: false,
272
+ db_reader_host: DB.host,
273
+ db_reader_port: DB.port,
274
+ db_reader_name: DB.name,
275
+ db_reader_schema: DB_SCHEMA,
276
+ db_reader_auth_method: 'password',
277
+ db_reader_user: DB.user,
278
+ db_reader_pass: DB.pass,
279
+ db_reader_enable_ssl: false,
280
+
281
+ cloud_storage: { ...S3, bucket: 'alvera-platform-dev', base_path: 'cookbook/primary-care' },
282
+ })
283
+ return data
284
+ },
285
+ )
286
+ datalakeSlug = datalake.slug!
287
+ ctx.datalakeId = datalake.id!
288
+ ```
289
+
290
+ ## 006 — run the migrations, and wait for ready
291
+
292
+ `datalakes.create` persists the row at `status: 'new'`; it does not run the
293
+ schema DDL. Migration is triggered separately so the operator decides when
294
+ the potentially-slow part happens. `datalakes.migrate` enqueues the job and
295
+ returns immediately with `status: 'enqueued'`; the poll after it is what
296
+ waits for the worker.
297
+
298
+ **This wait is a precondition for §007, not politeness.** A derived lake is
299
+ provisioned empty and its own migration is what builds its tables and starts
300
+ the copy. Provision before the raw lake is ready and you get two lakes that
301
+ exist and stay empty forever.
302
+
303
+ ```typescript
304
+ const migrateResp = await api.datalakes.migrate(tenantSlug, datalakeSlug)
305
+ if (migrateResp.data.status !== 'enqueued') {
306
+ throw new Error(`datalake migration not enqueued (status: ${migrateResp.data.status})`)
307
+ }
308
+
309
+ const READY_TIMEOUT_MS = 5 * 60_000
310
+ const readyDeadline = Date.now() + READY_TIMEOUT_MS
311
+ let datalakeStatus: string | undefined
312
+ while (Date.now() < readyDeadline) {
313
+ const { data } = await api.datalakes.get(tenantSlug, ctx.datalakeId)
314
+ datalakeStatus = data.status
315
+ if (datalakeStatus === 'ready') break
316
+ await new Promise((r) => setTimeout(r, 5_000))
317
+ }
318
+ if (datalakeStatus !== 'ready') {
319
+ throw new Error(`datalake did not reach :ready within ${READY_TIMEOUT_MS}ms (last: ${datalakeStatus})`)
320
+ }
321
+ ```
322
+
323
+ ## 007 — turn tokenization on, because two machines are going to read this
324
+
325
+ `provisionTokenization` creates both derived lakes and enqueues a migration
326
+ for each. It is **both-or-neither**: a run that provisions the tokenized lake
327
+ and cannot provision the redacted one reports the failure rather than leaving
328
+ half a pair. It is **idempotent**, so this step needs no `ensure` around it —
329
+ already-provisioned lakes come back as they are.
330
+
331
+ Poll the **derived** ids. The raw lake's own status says nothing about
332
+ whether its copies are ready, and a walk that waits on the wrong row gets a
333
+ `ready` it did not earn.
334
+
335
+ The statuses a derived lake passes through on the way are worth watching
336
+ rather than merely waiting out, because they are the difference between *this
337
+ run built these lakes* and *this run found them* — and on a cookbook whose
338
+ whole point is masking, that is the difference between having tested the copy
339
+ and having tested a copy someone else made.
340
+
341
+ The enum is `provisioning | new | processing | ready`, and **a derived lake
342
+ does not use the first of those.** `:provisioning` is set by exactly one
343
+ function, `Platform.Datalakes.provision_datalake/2`, which is deliberately
344
+ unreachable from `attrs` — and the tokenization path does not call it.
345
+ `provision_derived_datalakes/2` reuses the ordinary `create_datalake/2`
346
+ rather than opening a second create path, so a derived lake starts at the
347
+ default `:new` and climbs from there. A client watching for `provisioning`
348
+ would wait forever and conclude nothing was being built.
349
+
350
+ **That is the settled design, not an oversight to route around.** The stage
351
+ is reserved for a topology where each copy is a machine of its own and
352
+ standing one up takes minutes — there, a row exists while its database is
353
+ still being built, and `provisioning` is the honest thing to show. Today
354
+ every derived database is created *ahead* of its row, so no such moment
355
+ exists: the stage would be unreachable by any operator, and wiring it would
356
+ cost the connection probe on the one path that deliberately keeps it. A
357
+ missing database would become a row that quietly sits there instead of a
358
+ refusal you can read.
359
+
360
+ So the observation is keyed on `new` or `processing`, and it is a **report,
361
+ not an assertion**. A fast enough provision could have its first poll already
362
+ read `ready`, and failing the walk on that would be failing it for finishing
363
+ quickly.
364
+
365
+ ```typescript
366
+ const { data: provisioned } = await api.datalakes.provisionTokenization(tenantSlug, datalakeSlug)
367
+ const derivedLakes = provisioned.derived_datalakes ?? []
368
+ if (derivedLakes.length !== 2) {
369
+ throw new Error(`expected a tokenized + redacted pair, got ${derivedLakes.length} derived lake(s)`)
370
+ }
371
+
372
+ const DERIVED_TIMEOUT_MS = 5 * 60_000
373
+ const derivedDeadline = Date.now() + DERIVED_TIMEOUT_MS
374
+ const seenStatuses = new Set<string>()
375
+ let pendingLakes = derivedLakes.map((d) => d.id!)
376
+ while (Date.now() < derivedDeadline && pendingLakes.length > 0) {
377
+ const stillPending: string[] = []
378
+ for (const id of pendingLakes) {
379
+ const { data } = await api.datalakes.get(tenantSlug, id)
380
+ if (data.status) seenStatuses.add(data.status)
381
+ if (data.status !== 'ready') stillPending.push(id)
382
+ }
383
+ pendingLakes = stillPending
384
+ if (pendingLakes.length > 0) await new Promise((r) => setTimeout(r, 5_000))
385
+ }
386
+ if (pendingLakes.length > 0) {
387
+ throw new Error(`derived datalakes did not reach :ready: ${pendingLakes.join(', ')}`)
388
+ }
389
+
390
+ ctx.provisionedFresh = seenStatuses.has('new') || seenStatuses.has('processing')
391
+ console.log(
392
+ ` derived lakes ${ctx.provisionedFresh ? 'BUILT by this run' : 'already ready when first polled'} ` +
393
+ `— statuses observed: ${[...seenStatuses].join(', ')}`,
394
+ )
395
+
396
+ // Capture the tokenized lake's slug — §012 reads its physical column types,
397
+ // and §033 reads a row through it.
398
+ const tokenizedLake = derivedLakes.find((d) => d.type === 'tokenized')
399
+ if (!tokenizedLake?.slug) throw new Error('no tokenized lake in the provisioned pair')
400
+ ctx.tokenizedLakeSlug = tokenizedLake.slug
401
+ ```
402
+
403
+ ## 008 — three helpers the scenarios below share
404
+
405
+ **`waitForFiredRun`.** `workflows.run` only *schedules* a run. It returns
406
+ immediately with a `workflow_run_id`, and the `workflow_run_log_id` and
407
+ `batch_id` a scenario needs are written later, when the run actually fires.
408
+ Poll until `workflow_run_log_id` is a **string** — not until `status` leaves
409
+ `'scheduled'`. Those are different moments, and because `typeof null ===
410
+ 'object'` a predicate on status alone produces a baffling *"expected string,
411
+ got object"* rather than an obvious nil. Raise on `failed` carrying
412
+ `failure_reason` rather than polling to the deadline.
413
+
414
+ **`waitForBatches`.** Ingestion is async. **Gate on `dataset_updated`, never
415
+ on `rows_ingested`** — received is not written, and a batch whose row is
416
+ refused reports `rows_ingested: 1, dataset_updated: 0, status: 'partial'`
417
+ with the reason in `error`.
418
+
419
+ **`deployGenericTable`.** Creating a generic table does not deploy it. On
420
+ this surface the migration has more to do than on a raw one: it builds the
421
+ table in the raw lake **and** in both copies. The wait that means anything
422
+ asks for the table the same way the next step will, and retries until it
423
+ stops erroring.
424
+
425
+ ```typescript
426
+ ctx.waitForFiredRun = async (
427
+ runDatalakeSlug: string,
428
+ runId: string,
429
+ timeoutMs = 120_000,
430
+ ): Promise<{ workflowRunLogId: string; batchId: string | null }> => {
431
+ const deadline = Date.now() + timeoutMs
432
+ let lastStatus: string | undefined
433
+ while (Date.now() < deadline) {
434
+ const { data } = await api.workflowRuns.get(tenantSlug, runDatalakeSlug, runId)
435
+ lastStatus = data.status
436
+ if (data.status === 'failed') {
437
+ throw new Error(`workflow run ${runId} failed: ${data.failure_reason ?? 'no failure_reason given'}`)
438
+ }
439
+ if (typeof data.workflow_run_log_id === 'string') {
440
+ return { workflowRunLogId: data.workflow_run_log_id, batchId: data.batch_id ?? null }
441
+ }
442
+ await new Promise((r) => setTimeout(r, 1_000))
443
+ }
444
+ throw new Error(`workflow run ${runId} never fired within ${timeoutMs}ms (last status: ${lastStatus})`)
445
+ }
446
+
447
+ ctx.waitForBatches = async (
448
+ dacSlug: string,
449
+ batchIds: readonly string[],
450
+ timeoutMs = 120_000,
451
+ ): Promise<void> => {
452
+ const targets = new Set(batchIds)
453
+ const deadline = Date.now() + timeoutMs
454
+ let greenCount = 0
455
+ while (Date.now() < deadline) {
456
+ const { data } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
457
+ const green = new Set<string>()
458
+ for (const row of (data.data ?? []) as Array<Record<string, unknown>>) {
459
+ const b = row.batch_id
460
+ if (typeof b !== 'string' || !targets.has(b)) continue
461
+ if (row.status === 'partial' || row.status === 'failed') {
462
+ throw new Error(
463
+ `batch ${b} on ${dacSlug} did not persist its rows (status: ${row.status}): ` +
464
+ `${String(row.error ?? 'no reason given')}`,
465
+ )
466
+ }
467
+ if (typeof row.dataset_updated !== 'number' || row.dataset_updated < 1) continue
468
+ const files = row.output_files
469
+ if (!Array.isArray(files) || files.length === 0) continue
470
+ green.add(b)
471
+ }
472
+ greenCount = green.size
473
+ if (greenCount === targets.size) return
474
+ await new Promise((r) => setTimeout(r, 1_000))
475
+ }
476
+ throw new Error(`only ${greenCount}/${targets.size} batches on ${dacSlug} persisted within ${timeoutMs}ms`)
477
+ }
478
+
479
+ ctx.deployGenericTable = async (tableName: string, timeoutMs = 180_000): Promise<void> => {
480
+ await api.datalakes.migrate(tenantSlug, datalakeSlug)
481
+ const deadline = Date.now() + timeoutMs
482
+ let lastError: unknown
483
+ while (Date.now() < deadline) {
484
+ try {
485
+ await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
486
+ sql: `SELECT 1 FROM ${tableName} LIMIT 1`,
487
+ mode: 'raw',
488
+ })
489
+ return
490
+ } catch (err) {
491
+ lastError = err
492
+ await new Promise((r) => setTimeout(r, 1_000))
493
+ }
494
+ }
495
+ throw new Error(
496
+ `generic table ${tableName} was not readable within ${timeoutMs}ms ` +
497
+ `(last error: ${lastError instanceof Error ? lastError.message : String(lastError)})`,
498
+ )
499
+ }
500
+ ```
501
+
502
+ ## 009 — the SMS tool and the review connected app
503
+
504
+ One SMS tool dispatches every outbound message in this walk. **The
505
+ per-message body lives on the workflow action, never on the tool**, which is
506
+ what lets one tool serve both the review request and the triage fan-out.
507
+ Local dev points at LocalStack's SNS on `:4566`, so nothing leaves the
508
+ machine.
509
+
510
+ The connected app is the review form the SMS links to — a thin registration
511
+ of the page's URL and hosting mode, with the page itself living outside the
512
+ platform. Capture the server-derived slug; resolving the link in §019 is
513
+ addressed by it.
514
+
515
+ ```typescript
516
+ const smsTool = await ctx.ensure(
517
+ 'SMS tool',
518
+ async () => {
519
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
520
+ return ctx.firstNamed(data.data, 'Cookbook Practice SMS Tool')
521
+ },
522
+ async () => {
523
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
524
+ name: 'Cookbook Practice SMS Tool',
525
+ description: 'SNS-backed SMS dispatcher for review requests and triage, wired to LocalStack.',
526
+ intent: 'sms',
527
+ status: 'active',
528
+ datalake_id: ctx.datalakeId,
529
+ body: {
530
+ tool_body_type: 'sns',
531
+ auth_method: 'access_key',
532
+ region: 'us-east-1',
533
+ phone_number: '+15551234567',
534
+ endpoint_url: 'http://localhost:4566',
535
+ access_key_id: 'test',
536
+ secret_access_key: 'test',
537
+ },
538
+ })
539
+ return data
540
+ },
541
+ )
542
+ toolId = smsTool.id!
543
+ ctx.smsToolId = smsTool.id!
544
+
545
+ const reviewApp = await ctx.ensure(
546
+ 'review connected app',
547
+ async () => {
548
+ const { data } = await api.connectedApps.list(tenantSlug, datalakeSlug, { page_size: 100 })
549
+ return ctx.firstNamed(data.data, 'Cookbook Visit Review Form')
550
+ },
551
+ async () => {
552
+ const { data } = await api.connectedApps.create(tenantSlug, datalakeSlug, {
553
+ name: 'Cookbook Visit Review Form',
554
+ description: 'Post-visit review form linked from the outbound review SMS.',
555
+ mode: 'self_hosted',
556
+ urls: [{ url: 'https://review.example.local', is_primary: true, label: 'production' }],
557
+ })
558
+ return data
559
+ },
560
+ )
561
+ connectedAppId = reviewApp.id!
562
+ ctx.reviewAppId = reviewApp.id!
563
+ ctx.reviewAppSlug = reviewApp.slug!
564
+ ```
565
+
566
+ ## 010 — the appointments table, and the one column that exercises everything
567
+
568
+ GH-859 removed the `appointment` dataset along with every other per-domain
569
+ resource, so an appointment is a table you declare and the columns are yours
570
+ rather than FHIR's.
571
+
572
+ `appointment_id` is unique, so re-ingesting the same export updates those
573
+ rows instead of doubling them. `status` is the canonical value the filter
574
+ reads — the vendor's `appt_slot_status` string is mapped onto it by the
575
+ contract in §012, so the filter never has to know about `"3 - checked out"`.
576
+ `patient_name` and `patient_mobile` are `tokenize`; they ride on the row
577
+ because that is what lets the action templates read `event_dataset.*` rather
578
+ than reaching back through the resolved subject.
579
+
580
+ **`care_gaps` is the column worth reading closely.** It is `jsonb`, **and
581
+ `is_array: true`, and `redact`** — all three at once, which no other column
582
+ here is. Each half matters:
583
+
584
+ - `is_array: true` is not optional and its absence does not look like a type
585
+ error. Without it the column is a single JSON document, and a template that
586
+ renders a list into it is refused per row with
587
+ `upsert_row_failed %{care_gaps: ["is invalid"]}` while the batch reports
588
+ itself green.
589
+ - `redact` is what puts it in front of the masking machinery. A column that
590
+ is only ever empty proves nothing about masking, which is why §014 ingests
591
+ gaps that actually have entries in them.
592
+
593
+ Do **not** declare a `legal_entity_id` column. The platform stamps it from
594
+ Contract B's resolve, and the name is reserved — declaring it does not shadow
595
+ the stamp, it 422s the create, so you lose the table until you drop it.
596
+
597
+ ```typescript
598
+ const appointments = await ctx.ensure(
599
+ 'appointments table',
600
+ async () => {
601
+ const { data } = await api.genericTables.list(tenantSlug, datalakeSlug, { page_size: 100 })
602
+ return (data.data ?? []).find((t: { title?: string }) => t.title === 'Cookbook CAHPS Appointments')
603
+ },
604
+ async () => {
605
+ const { data } = await api.genericTables.create(tenantSlug, datalakeSlug, {
606
+ title: 'Cookbook CAHPS Appointments',
607
+ description: 'Appointments loaded from the practice management CAHPS export.',
608
+ columns: [
609
+ { name: 'appointment_id', title: 'Appointment ID', type: 'string', description: 'Vendor appointment id — unique per appointment', is_unique: true, privacy_requirement: 'none' },
610
+ { name: 'status', title: 'Status', type: 'string', description: 'Canonical status: fulfilled / cancelled / arrived / booked / pending', is_unique: false, privacy_requirement: 'none' },
611
+ { name: 'appt_start', title: 'Appointment Start', type: 'string', description: 'Local start timestamp, ISO 8601', is_unique: false, privacy_requirement: 'none' },
612
+ { name: 'minutes_duration', title: 'Duration (minutes)', type: 'integer', description: 'Slot length in minutes', is_unique: false, privacy_requirement: 'none' },
613
+ { name: 'description', title: 'Description', type: 'string', description: 'Appointment type as the practice records it', is_unique: false, privacy_requirement: 'none' },
614
+ { name: 'department', title: 'Department', type: 'string', description: 'Servicing department', is_unique: false, privacy_requirement: 'none' },
615
+ { name: 'provider_name', title: 'Provider Name', type: 'string', description: 'Rendering provider display name', is_unique: false, privacy_requirement: 'none' },
616
+ { name: 'patient_name', title: 'Patient Name', type: 'string', description: 'Patient display name — rendered into the SMS greeting', is_unique: false, privacy_requirement: 'tokenize' },
617
+ { name: 'patient_mobile', title: 'Patient Mobile', type: 'string', description: 'Patient mobile — the SMS destination', is_unique: false, privacy_requirement: 'tokenize' },
618
+ { name: 'patient_ref', title: 'Patient Ref', type: 'string', description: "The vendor's own patient id, for tracing a row back to the export", is_unique: false, privacy_requirement: 'tokenize' },
619
+ { name: 'care_gaps', title: 'Care Gaps', type: 'jsonb', description: 'Open care gaps, as a JSON array', is_unique: false, is_array: true, privacy_requirement: 'redact' },
620
+ ],
621
+ })
622
+ return data
623
+ },
624
+ )
625
+ genericTableId = appointments.id!
626
+ ctx.appointmentsTableId = appointments.id!
627
+ ctx.appointmentsTableName = appointments.name!
628
+
629
+ await ctx.deployGenericTable(ctx.appointmentsTableName)
630
+ ```
631
+
632
+ ## 011 — the data source and the manual-upload tool
633
+
634
+ The chain's first link is a **data source** — a registration of where the
635
+ rows originate. Its `uri` is the system-of-record address, and both contracts
636
+ render it onto the identifier, so it is half of the key that collapses them
637
+ onto one patient.
638
+
639
+ The **manual-upload tool** needs no endpoint and no credentials, because the
640
+ rows arrive in the ingest call's own body rather than being fetched.
641
+
642
+ ```typescript
643
+ const dataSource = await ctx.ensure(
644
+ 'athenahealth data source',
645
+ async () => {
646
+ const { data } = await api.dataSources.list(tenantSlug, datalakeSlug, { page_size: 100 })
647
+ return ctx.firstNamed(data.data, 'Cookbook Athenahealth Source')
648
+ },
649
+ async () => {
650
+ const { data } = await api.dataSources.create(tenantSlug, datalakeSlug, {
651
+ name: 'Cookbook Athenahealth Source',
652
+ uri: 'athenahealth.example.com/cookbook-primary-care',
653
+ description: 'Practice management system — origin of the CAHPS appointment export.',
654
+ status: 'active',
655
+ is_default: false,
656
+ })
657
+ return data
658
+ },
659
+ )
660
+ dataSourceId = dataSource.id!
661
+
662
+ const manualUploadTool = await ctx.ensure(
663
+ 'manual-upload tool',
664
+ async () => {
665
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
666
+ return ctx.firstNamed(data.data, 'Cookbook Practice Manual Upload Tool')
667
+ },
668
+ async () => {
669
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
670
+ name: 'Cookbook Practice Manual Upload Tool',
671
+ description: 'Manual-upload data-exchange tool — backs the appointment ingest.',
672
+ intent: 'data_exchange',
673
+ status: 'active',
674
+ datalake_id: ctx.datalakeId,
675
+ data_source_id: dataSourceId,
676
+ body: { tool_body_type: 'manual_upload' },
677
+ })
678
+ return data
679
+ },
680
+ )
681
+ ctx.manualUploadToolId = manualUploadTool.id!
682
+ ```
683
+
684
+ ## 012 — the CAHPS contract pair, on two clients
685
+
686
+ One inbound CAHPS row becomes two things. **Contract A**
687
+ (`resource_type: 'legal_entity'`) writes the patient — the subject — and its
688
+ `mdm_input_config` is `{ type: 'null' }`, because a template that already
689
+ *is* the subject has nothing to resolve. **Contract B**
690
+ (`resource_type: 'generic_table'`) writes the appointment row, and its
691
+ `mdm_input_config` emits the same `(uri, patient_id)` pair so both converge
692
+ on one patient and the row earns its `legal_entity_id` stamp.
693
+
694
+ **The pair is split across two clients, and the ingest in §014 is staged.** A
695
+ client fans out per `(row, contract)` and the job key staggers by *row*, so a
696
+ pair bound to one client runs both contracts on the same row at the same
697
+ instant. Both find-or-create the same patient, both miss the other's
698
+ uncommitted write, and the unique index on `(uri, id_type, id_number)`
699
+ refuses the loser. The legal entity is deliberately not in that key, which is
700
+ what makes the refusal correct rather than a bug.
701
+
702
+ **Watch the identifier shapes.** `Platform.MDMInput` takes `uri` / `type` /
703
+ `value`; a legal-entity identification takes `id_type` / `id_number` / `uri`.
704
+ `system` is a genuine alias on the MDM side — cast and normalised onto `uri`
705
+ — but `id_type` and `id_number` are not cast at all, and unknown keys are
706
+ dropped rather than refused. And the reason the two look interchangeable is
707
+ that one becomes the other: `MDMInput.to_identification/1` takes
708
+ `(uri, type, value)` and returns `(uri, id_type, id_number)`, so the
709
+ identifier reads back under names it will not accept on the way in.
710
+
711
+ ```typescript
712
+ const { readFileSync } = await import('node:fs')
713
+ const { join } = await import('node:path')
714
+ const read = (f: string) =>
715
+ readFileSync(join(process.env.COOKBOOK_FIXTURES_DIR!, 'primary-care-feedback', f), 'utf8')
716
+
717
+ const leContract = await ctx.ensure(
718
+ 'patient contract',
719
+ async () => {
720
+ const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
721
+ return ctx.firstNamed(data.data, 'Cookbook CAHPS Patient Contract')
722
+ },
723
+ async () => {
724
+ const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
725
+ name: 'Cookbook CAHPS Patient Contract',
726
+ description: 'CAHPS appointment row → the patient LegalEntity.',
727
+ resource_type: 'legal_entity',
728
+ type: 'identity',
729
+ generic_table_id: null,
730
+ template_config: { type: 'custom', body: read('_cahps_appointments_legal_entity.liquid') },
731
+ mdm_input_config: { type: 'null' },
732
+ })
733
+ return data
734
+ },
735
+ )
736
+ ctx.leContractId = leContract.id!
737
+
738
+ const gtContract = await ctx.ensure(
739
+ 'appointment contract',
740
+ async () => {
741
+ const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
742
+ return ctx.firstNamed(data.data, 'Cookbook CAHPS Appointment Contract')
743
+ },
744
+ async () => {
745
+ const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
746
+ name: 'Cookbook CAHPS Appointment Contract',
747
+ description: 'CAHPS appointment row → appointments table row, stamped with its patient.',
748
+ resource_type: 'generic_table',
749
+ type: 'identity',
750
+ generic_table_id: ctx.appointmentsTableId,
751
+ template_config: { type: 'custom', body: read('_cahps_appointments_generic_table.liquid') },
752
+ mdm_input_config: { type: 'custom', body: read('_cahps_appointments_mdm.liquid') },
753
+ })
754
+ return data
755
+ },
756
+ )
757
+ interopContractId = gtContract.id!
758
+ ctx.gtContractId = gtContract.id!
759
+ ```
760
+
761
+ ## 013 — the two activation clients
762
+
763
+ Both share the tool and the data source. Only the contract list differs, and
764
+ the order they run in is the whole point: identity first, then the row that
765
+ resolves against it.
766
+
767
+ ```typescript
768
+ const identityDac = await ctx.ensure(
769
+ 'patient identity client',
770
+ async () => {
771
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
772
+ return ctx.firstNamed(data.data, 'Cookbook CAHPS Patient Identity DAC')
773
+ },
774
+ async () => {
775
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
776
+ name: 'Cookbook CAHPS Patient Identity DAC',
777
+ description: 'Writes the patient LegalEntity for each appointment row. Runs FIRST.',
778
+ tool_id: ctx.manualUploadToolId,
779
+ data_source_id: dataSourceId,
780
+ tool_call: { tool_call_type: 'manual_upload' },
781
+ interoperability_contract_ids: [ctx.leContractId],
782
+ })
783
+ return data
784
+ },
785
+ )
786
+ ctx.identityDacSlug = identityDac.slug!
787
+
788
+ const dataDac = await ctx.ensure(
789
+ 'appointment data client',
790
+ async () => {
791
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
792
+ return ctx.firstNamed(data.data, 'Cookbook CAHPS Appointment Data DAC')
793
+ },
794
+ async () => {
795
+ const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
796
+ name: 'Cookbook CAHPS Appointment Data DAC',
797
+ description: 'Writes the appointment row and stamps it with the patient. Runs SECOND.',
798
+ tool_id: ctx.manualUploadToolId,
799
+ data_source_id: dataSourceId,
800
+ tool_call: { tool_call_type: 'manual_upload' },
801
+ interoperability_contract_ids: [ctx.gtContractId],
802
+ })
803
+ return data
804
+ },
805
+ )
806
+ dacId = dataDac.id!
807
+ ctx.dataDacSlug = dataDac.slug!
808
+ ```
809
+
810
+ ## 014 — two appointments, one attended and one not
811
+
812
+ Two rows in **the vendor's own shape**, which is why every key below looks
813
+ slightly wrong: `appt_slot_status` rather than `status`,
814
+ `patient_mobile_no` rather than `patient_mobile`, `rndrng_provider` rather
815
+ than `provider_name`. Renaming those onto the table's columns is Contract B's
816
+ entire job, and ingesting the canonical names instead would bypass the
817
+ contract and prove nothing about it.
818
+
819
+ The status mapping is the visible half — the export says `"3 - checked out"`
820
+ and `"x - cancelled"`, and the contract maps them onto `fulfilled` and
821
+ `cancelled` so the filter never has to know about vendor strings. The field
822
+ renames are the quiet half, and they fail differently: a key the template does
823
+ not read is not an error. It is simply absent, the column lands `null`, the
824
+ ingest reports itself green, and the failure surfaces much later as an SMS
825
+ action rejecting an empty recipient. §015 exists to catch that on the row
826
+ rather than three steps downstream.
827
+
828
+ **Both rows carry `care_gaps` with real entries.** That is deliberate. A
829
+ jsonb array column that is only ever `[]` still exists and still replicates,
830
+ so a walk that never populates it will go green while proving nothing about
831
+ whether a *populated* array survives the copy — which is the only version of
832
+ the question anyone cares about. §016 reads them back.
833
+
834
+ ```typescript
835
+ ctx.fulfilledId = 'APPT-COOKBOOK-0001'
836
+ ctx.cancelledId = 'APPT-COOKBOOK-0002'
837
+
838
+ const appointmentRows = [
839
+ {
840
+ appointment_id: ctx.fulfilledId,
841
+ appt_slot_status: '3 - checked out',
842
+ appt_date: '3/12/2026',
843
+ appt_start_time: '2:30 PM',
844
+ appt_slot_duration: 30,
845
+ appt_type: 'Annual wellness visit',
846
+ svc_department: 'Family Medicine',
847
+ rndrng_provider_id: 'PROV-114',
848
+ rndrng_provider: 'Dr. Amara Osei',
849
+ patient_name: 'Rosalind Achterberg',
850
+ patient_mobile_no: '+12025554101',
851
+ patient_id: 'PAT-COOKBOOK-0001',
852
+ care_gaps: [
853
+ { code: 'A1C', description: 'HbA1c not recorded in 12 months', due_date: '2026-04-30' },
854
+ { code: 'BP', description: 'Blood pressure follow-up overdue', due_date: '2026-03-15' },
855
+ ],
856
+ },
857
+ {
858
+ appointment_id: ctx.cancelledId,
859
+ appt_slot_status: 'x - cancelled',
860
+ appt_date: '3/13/2026',
861
+ appt_start_time: '9:00 AM',
862
+ appt_slot_duration: 20,
863
+ appt_type: 'Follow-up',
864
+ svc_department: 'Family Medicine',
865
+ rndrng_provider_id: 'PROV-114',
866
+ rndrng_provider: 'Dr. Amara Osei',
867
+ patient_name: 'Teodoro Vasquez-Lindqvist',
868
+ patient_mobile_no: '+12025554102',
869
+ patient_id: 'PAT-COOKBOOK-0002',
870
+ care_gaps: [{ code: 'FLU', description: 'Influenza vaccination not on file', due_date: '2026-10-01' }],
871
+ },
872
+ ]
873
+
874
+ // Stage one — the patients.
875
+ const identityIngests = await Promise.all(
876
+ appointmentRows.map((row) =>
877
+ api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.identityDacSlug, { data: row }),
878
+ ),
879
+ )
880
+ await ctx.waitForBatches(
881
+ ctx.identityDacSlug,
882
+ identityIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!),
883
+ )
884
+
885
+ // Stage two — the appointment rows, now that every patient exists.
886
+ const dataIngests = await Promise.all(
887
+ appointmentRows.map((row) =>
888
+ api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.dataDacSlug, { data: row }),
889
+ ),
890
+ )
891
+ await ctx.waitForBatches(
892
+ ctx.dataDacSlug,
893
+ dataIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!),
894
+ )
895
+ ```
896
+
897
+ ## 015 — the rows landed, the mapping ran, and the stamp is on
898
+
899
+ Four assertions in one read, and each catches a different failure.
900
+
901
+ The **row count** catches an ingest that reported green over nothing. The
902
+ **status values** catch a contract whose mapping silently fell through to the
903
+ `pending` default — which would leave the filter in §018 with nothing to
904
+ route and make the whole scenario a false green. The **stamp** catches a
905
+ resolve that never ran, which would leave every link minted against nobody.
906
+
907
+ And **`patient_mobile` must not be null**, which is the one that catches a
908
+ misspelled vendor key. A template reading a field the payload does not carry
909
+ is not an error — Liquid renders nothing, the column lands `null`, and the
910
+ batch reports itself green. The consequence appears two steps later as an SMS
911
+ action failing `output_schema_violation` on an empty recipient, which reads
912
+ like a broken tool rather than a missing letter in an ingest key. Assert the
913
+ column on the row, where the answer is unambiguous.
914
+
915
+ ```typescript
916
+ const stampDeadline = Date.now() + 120_000
917
+ let rows: unknown[][] = []
918
+ while (Date.now() < stampDeadline && rows.length < 2) {
919
+ const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
920
+ sql: `SELECT appointment_id, status, legal_entity_id, patient_mobile
921
+ FROM ${ctx.appointmentsTableName}
922
+ WHERE appointment_id IN ('${ctx.fulfilledId}', '${ctx.cancelledId}')
923
+ ORDER BY appointment_id`,
924
+ mode: 'raw',
925
+ })
926
+ if (typeof result !== 'string') rows = result.data
927
+ if (rows.length < 2) await new Promise((r) => setTimeout(r, 1_000))
928
+ }
929
+ if (rows.length !== 2) {
930
+ throw new Error(`expected 2 appointment rows, found ${rows.length} within 120s`)
931
+ }
932
+
933
+ const byId = new Map(
934
+ rows.map((r) => [String(r[0]), { status: String(r[1]), stamp: r[2], mobile: r[3] }]),
935
+ )
936
+ if (byId.get(ctx.fulfilledId)?.status !== 'fulfilled') {
937
+ throw new Error(
938
+ `the contract did not map "3 - checked out" onto fulfilled — got ` +
939
+ `${byId.get(ctx.fulfilledId)?.status}`,
940
+ )
941
+ }
942
+ if (byId.get(ctx.cancelledId)?.status !== 'cancelled') {
943
+ throw new Error(
944
+ `the contract did not map "x - cancelled" onto cancelled — got ` +
945
+ `${byId.get(ctx.cancelledId)?.status}`,
946
+ )
947
+ }
948
+ for (const [id, row] of byId) {
949
+ if (!row.stamp) throw new Error(`appointment ${id} landed with no legal_entity_id stamp`)
950
+ if (!row.mobile) {
951
+ throw new Error(
952
+ `appointment ${id} landed with a null patient_mobile — the contract template reads ` +
953
+ `the vendor key 'patient_mobile_no', so a payload using the column name writes nothing`,
954
+ )
955
+ }
956
+ }
957
+ ```
958
+
959
+ ## 016 — a populated masked array survives the copy
960
+
961
+ This is the step that proves the derived lakes are real, and it is the
962
+ cheapest one in the file to get wrong by writing something weaker.
963
+
964
+ `care_gaps` is `jsonb` + `is_array` + `redact`, and what happens to that
965
+ column across the three lakes is worth stating exactly, because the obvious
966
+ guess is wrong.
967
+
968
+ **The column is `jsonb` in the raw lake and `text` in both copies. The types
969
+ are deliberately different, and that is what makes replication work.** Logical
970
+ replication ships the publisher's text representation and hands it to the
971
+ subscriber column's input function. Any jsonb value renders to a string and
972
+ any string is valid `text`, so `jsonb → text` is always accepted. What is
973
+ refused is `jsonb → text[]`, because `[{"code":"A1C",…}]` is JSON and not a
974
+ Postgres array literal — and a masking policy that widened the column to an
975
+ array type in the copies would fail the apply worker on the first row
976
+ carrying one.
977
+
978
+ So the invariant is **scalar-ness, not type parity**: a jsonb column is never
979
+ a Postgres array type, in any lake. `is_array: true` describes the JSON
980
+ document the column holds. It is not an instruction to make the Postgres
981
+ column an array, and reading it that way is the mistake that produces a
982
+ derived lake which exists, reports `ready`, and silently never fills — a
983
+ failure with no error surface a client can see.
984
+
985
+ None of that is checkable from here. `executeSql` refuses any query touching
986
+ `information_schema` or `pg_catalog` — they are blocked tables — so a client
987
+ cannot read a column's physical type at all, in any mode. The mechanism is
988
+ therefore something to know rather than something to assert, and **the walk
989
+ asserts its consequence instead**: a row carrying a populated array must
990
+ arrive in both copies. That is directly observable, and it is the claim
991
+ anyone actually depends on.
992
+
993
+ It also explains the shape of the masked value below, which is otherwise
994
+ baffling. The copies hand back a masked *string*, not an array, because on
995
+ that side the column really is text and carries no declared semantic kind —
996
+ so the anonymizer masks its string rendering the way it masks any free text.
997
+
998
+ Read the same row through all three lakes. `mode` is a parameter of the read,
999
+ not a different slug: one lake is addressed and `mode` selects which copy
1000
+ answers. This works only because §004 built its client at the `raw` ceiling.
1001
+
1002
+ ```typescript
1003
+ const gapsSql =
1004
+ `SELECT appointment_id, care_gaps, minutes_duration ` +
1005
+ `FROM ${ctx.appointmentsTableName} WHERE appointment_id = '${ctx.fulfilledId}'`
1006
+
1007
+ const readGapsAs = async (mode: 'raw' | 'tokenized' | 'redacted') => {
1008
+ const deadline = Date.now() + 120_000
1009
+ while (Date.now() < deadline) {
1010
+ const { data } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, { sql: gapsSql, mode })
1011
+ const row = (data.data ?? [])[0]
1012
+ if (row) {
1013
+ const cell = (column: string): unknown => row[data.meta.columns.indexOf(column)]
1014
+ return { gaps: cell('care_gaps'), minutes: cell('minutes_duration') }
1015
+ }
1016
+ await new Promise((r) => setTimeout(r, 2_000))
1017
+ }
1018
+ throw new Error(`the appointment never became readable through the ${mode} lake within 120s`)
1019
+ }
1020
+
1021
+ const rawGaps = await readGapsAs('raw')
1022
+ const tokenizedGaps = await readGapsAs('tokenized')
1023
+ const redactedGaps = await readGapsAs('redacted')
1024
+
1025
+ console.log(' raw →', JSON.stringify(rawGaps))
1026
+ console.log(' tokenized →', JSON.stringify(tokenizedGaps))
1027
+ console.log(' redacted →', JSON.stringify(redactedGaps))
1028
+
1029
+ // The array is populated in the primary. A walk whose array is always empty
1030
+ // proves nothing about arrays.
1031
+ const rawText = JSON.stringify(rawGaps.gaps)
1032
+ if (!rawText.includes('A1C') || !rawText.includes('BP')) {
1033
+ throw new Error(`the raw lake does not hold the two care gaps that were ingested: ${rawText}`)
1034
+ }
1035
+
1036
+ // The row REACHED both copies. Had the masked column been widened to a
1037
+ // Postgres array type, the apply worker would have stopped on this row and
1038
+ // these reads would time out against a lake that reports itself ready.
1039
+ for (const [mode, read] of [['tokenized', tokenizedGaps], ['redacted', redactedGaps]] as const) {
1040
+ if (read.gaps === undefined || read.gaps === null) {
1041
+ throw new Error(`the ${mode} lake has no care_gaps value — replication did not apply the row`)
1042
+ }
1043
+ if (Number(read.minutes) !== Number(rawGaps.minutes)) {
1044
+ throw new Error(
1045
+ `the ${mode} lake changed a 'none' column: minutes_duration ` +
1046
+ `${read.minutes} vs ${rawGaps.minutes}`,
1047
+ )
1048
+ }
1049
+ }
1050
+
1051
+ // And the clinical detail does NOT survive it. `redact` is the strictest
1052
+ // tier, so the copies must not hand back the gap codes.
1053
+ for (const [mode, read] of [['tokenized', tokenizedGaps], ['redacted', redactedGaps]] as const) {
1054
+ const masked = JSON.stringify(read.gaps)
1055
+ if (masked.includes('A1C') || masked.includes('HbA1c')) {
1056
+ throw new Error(`the ${mode} lake handed back the real care gaps — redaction is not in force: ${masked}`)
1057
+ }
1058
+ }
1059
+ ```
1060
+
1061
+ ## 017 — the Review SMS workflow
1062
+
1063
+ Standard shape — filter, decision, one action — targeting §010's table.
1064
+
1065
+ `skip_mdm_resolution: false` is set explicitly, and must be. A generic-table
1066
+ row is never its own subject; the subject is the patient §012's pair stamped
1067
+ onto it. With resolution on, the platform loads that subject and mints a page
1068
+ token against it, which is what makes `{{ connected_app_form_url }}` *this
1069
+ patient's* review link. Skip it and the token is never minted and the link
1070
+ renders empty.
1071
+
1072
+ The filter is one comparison — `status == "fulfilled"` — and it carries the
1073
+ whole clinical rule this scenario exists for: never ask someone to review a
1074
+ visit they did not attend. That the rule is one line of Liquid on the
1075
+ workflow, rather than a branch in a service, is the point of the primitive.
1076
+ A reviewer can read it.
1077
+
1078
+ ```typescript
1079
+ const REVIEW_KEY = 'send_review_request'
1080
+
1081
+ const reviewWorkflow = await ctx.ensure(
1082
+ 'Review SMS workflow',
1083
+ async () => {
1084
+ const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
1085
+ return ctx.firstNamed(data.data, 'Cookbook Visit Review Workflow')
1086
+ },
1087
+ async () => {
1088
+ const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
1089
+ name: 'Cookbook Visit Review Workflow',
1090
+ description: 'Asks for a review after a visit that actually happened.',
1091
+ dataset_type: 'generic_table',
1092
+ generic_table_id: ctx.appointmentsTableId,
1093
+ status: 'live',
1094
+ tags: ['cahps', 'review'],
1095
+ skip_mdm_resolution: false,
1096
+ filter_config: {
1097
+ type: 'custom',
1098
+ body: '{% if event_dataset.status == "fulfilled" %}true{% endif %}',
1099
+ },
1100
+ decision_config: {
1101
+ type: 'custom',
1102
+ body: `["${REVIEW_KEY}"]`,
1103
+ output_schema: { type: 'array', items: { type: 'string' } },
1104
+ },
1105
+ context_datasets: [
1106
+ {
1107
+ dataset_type: 'message',
1108
+ where_clause:
1109
+ `m.legal_entity_id = '{{ legal_entity_id }}' AND m.decision_key = '${REVIEW_KEY}' ` +
1110
+ "AND m.sent_at > NOW() - INTERVAL '90 days'",
1111
+ limit: 1,
1112
+ position: 0,
1113
+ },
1114
+ ],
1115
+ actions: [
1116
+ {
1117
+ action_type: 'sms',
1118
+ tool_id: ctx.smsToolId,
1119
+ decision_key: REVIEW_KEY,
1120
+ position: 0,
1121
+ trigger_template: 'now',
1122
+ idempotency_template: `{{ event_dataset.appointment_id }}-${REVIEW_KEY}`,
1123
+ connected_app_id: ctx.reviewAppId,
1124
+ connected_app_route: '/review/visit',
1125
+ connected_app_metadata_template: '{"appointment_id":"{{ event_dataset.appointment_id }}"}',
1126
+ tool_call: {
1127
+ tool_call_type: 'sms_request',
1128
+ to: { type: 'custom', body: '{{ event_dataset.patient_mobile }}' },
1129
+ body: {
1130
+ type: 'custom',
1131
+ body:
1132
+ 'Hi {{ event_dataset.patient_name }}, thanks for visiting ' +
1133
+ '{{ event_dataset.department }}. How did we do? ' +
1134
+ '{{ connected_app_form_url }}',
1135
+ },
1136
+ sms_type: 'transactional',
1137
+ },
1138
+ },
1139
+ ],
1140
+ })
1141
+ return data
1142
+ },
1143
+ )
1144
+ workflowId = reviewWorkflow.id!
1145
+ ctx.reviewWorkflowSlug = reviewWorkflow.slug!
1146
+
1147
+ if (reviewWorkflow.skip_mdm_resolution !== false) {
1148
+ throw new Error('a workflow that mints a per-patient link must keep MDM resolution ON')
1149
+ }
1150
+ ```
1151
+
1152
+ ## 018 — run it, and the cancelled visit is filtered
1153
+
1154
+ `manual_override: false` is what makes this a real test — the filter is
1155
+ evaluated, so `status` genuinely routes each row. With `true` the filter is
1156
+ bypassed and the cancelled appointment would get a review request, which is
1157
+ the exact outcome the scenario exists to prevent.
1158
+
1159
+ The assertion is at the execution-log level, which is what lets it survive a
1160
+ rerun: the action inside the passing row's log refuses to fire a second time
1161
+ because its idempotency key already fired, but the row still passes the
1162
+ filter, so the one-passed-one-filtered split is identical on every run.
1163
+
1164
+ ```typescript
1165
+ const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.reviewWorkflowSlug, {
1166
+ sql_where_clause: `appointment_id IN ('${ctx.fulfilledId}', '${ctx.cancelledId}')`,
1167
+ mode: 'live',
1168
+ manual_override: false,
1169
+ })
1170
+ const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
1171
+ ctx.reviewRunBatchId = fired.batchId!
1172
+
1173
+ const refreshDeadline = Date.now() + 120_000
1174
+ let status: string | null = null
1175
+ while (Date.now() < refreshDeadline) {
1176
+ const { data: log } = await api.workflows.batchLogs.refresh(
1177
+ tenantSlug, datalakeSlug, ctx.reviewWorkflowSlug, fired.workflowRunLogId,
1178
+ )
1179
+ status = log.status ?? null
1180
+ if (status && status !== 'pending') break
1181
+ await new Promise((r) => setTimeout(r, 2_000))
1182
+ }
1183
+ if (status === 'failed') throw new Error('the review run reached :failed')
1184
+ if (!status || status === 'pending') throw new Error('the review run did not leave :pending within 120s')
1185
+
1186
+ const welDeadline = Date.now() + 90_000
1187
+ let ourLogs: Array<Record<string, unknown>> = []
1188
+ let byStatus: Record<string, number> = {}
1189
+ while (Date.now() < welDeadline) {
1190
+ const { data: wfLogs } = await api.workflows.workflowLogs.list(
1191
+ tenantSlug, datalakeSlug, ctx.reviewWorkflowSlug, { page_size: 100 },
1192
+ )
1193
+ ourLogs = (wfLogs.data ?? [])
1194
+ .map((w) => w as Record<string, unknown>)
1195
+ .filter((w) => w.batch_id === ctx.reviewRunBatchId)
1196
+ byStatus = {}
1197
+ for (const w of ourLogs) {
1198
+ const st = (w.status as string | undefined) ?? 'unknown'
1199
+ byStatus[st] = (byStatus[st] ?? 0) + 1
1200
+ }
1201
+ const settled = (byStatus.completed ?? 0) + (byStatus.executing ?? 0)
1202
+ if (ourLogs.length === 2 && settled === 1 && (byStatus.filtered ?? 0) === 1) break
1203
+ await new Promise((r) => setTimeout(r, 2_000))
1204
+ }
1205
+ if (ourLogs.length !== 2) {
1206
+ throw new Error(`expected 2 execution logs for the review run, got ${ourLogs.length}`)
1207
+ }
1208
+ if ((byStatus.executing ?? 0) + (byStatus.completed ?? 0) !== 1) {
1209
+ throw new Error(`expected 1 fulfilled visit past the filter — ${JSON.stringify(byStatus)}`)
1210
+ }
1211
+ if ((byStatus.filtered ?? 0) !== 1) {
1212
+ throw new Error(`expected the cancelled visit to be :filtered — ${JSON.stringify(byStatus)}`)
1213
+ }
1214
+ ```
1215
+
1216
+ ## 019 — the review request went out, and its link resolves
1217
+
1218
+ Search on the **idempotency key**, never on `workflow_id`. The key is
1219
+ `<appointment id>-send_review_request` and the appointment id is stable, so it
1220
+ names the same message forever. A workflow id is stable only as long as the
1221
+ workflow resource is.
1222
+
1223
+ Then resolve the `/t/<token>` the platform baked into the body when it minted
1224
+ the page token, and post the tracking a connected-app frontend would post
1225
+ when the patient opens the form.
1226
+
1227
+ ```typescript
1228
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
1229
+ search_query: `m.idempotency_key = '${ctx.fulfilledId}-send_review_request'`,
1230
+ })
1231
+ if (search.status !== 'completed') {
1232
+ throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
1233
+ }
1234
+
1235
+ const msgDeadline = Date.now() + 90_000
1236
+ let reviewBody: string | undefined
1237
+ while (Date.now() < msgDeadline && !reviewBody) {
1238
+ const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
1239
+ userSearchId: search.id!,
1240
+ dataAccessMode: 'raw',
1241
+ })
1242
+ reviewBody = ((data.data ?? []) as Array<Record<string, unknown>>)
1243
+ .map((m) => String(m.body ?? ''))
1244
+ .find((b) => b.includes('/t/') && b.includes('How did we do?'))
1245
+ if (!reviewBody) await new Promise((r) => setTimeout(r, 2_000))
1246
+ }
1247
+ if (!reviewBody) {
1248
+ throw new Error(`no rendered review SMS for ${ctx.fulfilledId} within 90s`)
1249
+ }
1250
+
1251
+ const token = reviewBody.match(/\/t\/([A-Za-z0-9_-]+)/)
1252
+ if (!token) throw new Error(`no /t/<token> in the review body: ${reviewBody}`)
1253
+
1254
+ const { data: resolved } = await api.connectedApps.resolvePage(
1255
+ tenantSlug, datalakeSlug, ctx.reviewAppSlug,
1256
+ { short_path: token[1]!, user_agent: 'cookbook-doctest/primary-care' },
1257
+ )
1258
+ if (resolved.route_path !== '/review/visit') {
1259
+ throw new Error(`the review link resolved to the wrong route: ${resolved.route_path}`)
1260
+ }
1261
+
1262
+ const now = new Date().toISOString()
1263
+ const { data: tracked } = await api.connectedApps.updateMessageTracking(
1264
+ tenantSlug, datalakeSlug, ctx.reviewAppSlug,
1265
+ { short_path: token[1]!, opened_at: now, form_submitted_at: now },
1266
+ )
1267
+ if (!tracked.message?.opened_at || !tracked.message?.form_submitted_at) {
1268
+ throw new Error('message tracking did not persist opened_at + form_submitted_at')
1269
+ }
1270
+ ```
1271
+
1272
+ ## 020 — the Contact Us table, and its client you did not create
1273
+
1274
+ Inbound website messages do not fit any canonical entity, so they are a table
1275
+ you declare. `message` is what the agent reads; `category` is what a
1276
+ production system writes back.
1277
+
1278
+ **Creating a generic table provisions a default Data Activation Client for
1279
+ you**, named after the table, with `tool_call: manual_upload`. For plain row
1280
+ ingestion into a table you declared there is nothing to build — no data
1281
+ source, no tool, no contract, no client. That is the whole reason this
1282
+ scenario has none of the machinery §011–§013 needed: those existed because
1283
+ the appointment rows also had to become *patients*, and this table's rows are
1284
+ their own subject.
1285
+
1286
+ **Scope the lookup server-side.** The listing pages and is not newest-first,
1287
+ so a `.find()` over page one starts missing the client as soon as the lake
1288
+ has a few tables — and the failure looks like it was never provisioned.
1289
+
1290
+ ```typescript
1291
+ const contactUs = await ctx.ensure(
1292
+ 'contact-us table',
1293
+ async () => {
1294
+ const { data } = await api.genericTables.list(tenantSlug, datalakeSlug, { page_size: 100 })
1295
+ return (data.data ?? []).find((t: { title?: string }) => t.title === 'Cookbook Contact Us')
1296
+ },
1297
+ async () => {
1298
+ const { data } = await api.genericTables.create(tenantSlug, datalakeSlug, {
1299
+ title: 'Cookbook Contact Us',
1300
+ description: 'Inbound contact-us form submissions the triage agent classifies.',
1301
+ columns: [
1302
+ { name: 'submission_id', title: 'Submission ID', type: 'string', description: 'Vendor-supplied unique submission id', is_unique: true, privacy_requirement: 'none' },
1303
+ { name: 'name', title: 'Name', type: 'string', description: 'Submitter full name', is_unique: false, privacy_requirement: 'tokenize' },
1304
+ { name: 'email', title: 'Email', type: 'string', description: 'Submitter email', is_unique: false, privacy_requirement: 'tokenize' },
1305
+ { name: 'message', title: 'Message', type: 'string', description: 'Free-text message — the agent classifies this', is_unique: false, privacy_requirement: 'none' },
1306
+ { name: 'category', title: 'Category', type: 'string', description: 'Triage bucket assigned by the agent', is_unique: false, privacy_requirement: 'none' },
1307
+ ],
1308
+ })
1309
+ return data
1310
+ },
1311
+ )
1312
+ ctx.contactTableId = contactUs.id!
1313
+ ctx.contactTableName = contactUs.name!
1314
+
1315
+ await ctx.deployGenericTable(ctx.contactTableName)
1316
+
1317
+ const dacDeadline = Date.now() + 120_000
1318
+ let defaultDac: { slug?: string | null } | undefined
1319
+ while (Date.now() < dacDeadline && !defaultDac) {
1320
+ const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, {
1321
+ filters: [{ field: 'name', op: 'ilike', value: ctx.contactTableName }],
1322
+ })
1323
+ defaultDac = (data.data ?? [])[0]
1324
+ if (!defaultDac) await new Promise((r) => setTimeout(r, 1_000))
1325
+ }
1326
+ if (!defaultDac?.slug) {
1327
+ throw new Error(`no default client found for table ${ctx.contactTableName} within 120s`)
1328
+ }
1329
+ ctx.contactDacSlug = defaultDac.slug
1330
+ ```
1331
+
1332
+ ## 021 — the LLM tool and the triage agent
1333
+
1334
+ The tool is a **provider adapter**: `base_body` authors Ollama's native
1335
+ `/api/chat` request with `think: false` and a `format` schema so the model
1336
+ returns schema-constrained JSON, and `response_extractor` maps the reply back
1337
+ onto the canonical `{ output_json, … }`.
1338
+
1339
+ The agent's `llm_response_schema` pins the vocabulary with an `enum`, and
1340
+ that is load-bearing rather than decorative: the workflow interpolates the
1341
+ category straight into its decision array, so a model answering
1342
+ `"a clinical question"` instead of `clinical` would produce a decision key no
1343
+ action is registered for and the row would silently do nothing.
1344
+
1345
+ **`data_access: 'tokenized'` is the decision this whole surface exists for.**
1346
+ The agent reads inbound messages to decide how urgent they are. It does not
1347
+ need to know who sent them, and on this lake it cannot: the tokenized
1348
+ projection is what it sees. Compare `subscription-saas`, where the equivalent
1349
+ agent runs `raw` — there the prompt carries an account number and an account
1350
+ type and no identity at all, so there is nothing to hold back. The tier
1351
+ follows what the prompt can reach, not a house style.
1352
+
1353
+ ```typescript
1354
+ const ENRICHMENT_OUTPUT_SCHEMA = {
1355
+ type: 'object',
1356
+ properties: {
1357
+ output_json: {},
1358
+ input_tokens: { type: ['integer', 'null'] },
1359
+ output_tokens: { type: ['integer', 'null'] },
1360
+ total_tokens: { type: ['integer', 'null'] },
1361
+ explanation: { type: ['string', 'null'] },
1362
+ },
1363
+ required: ['output_json'],
1364
+ }
1365
+
1366
+ const OLLAMA_BASE_BODY =
1367
+ '{"model": "{{ model }}", "messages": [{"role": "user", "content": "{{ rendered_prompt | json_escape }}", ' +
1368
+ '"images": [{% for img in images %}{% unless forloop.first %}, {% endunless %}"{{ img.data }}"{% endfor %}]}], ' +
1369
+ '"stream": false, "think": false, "options": {"temperature": {{ temperature }}, "num_predict": {{ max_tokens }}, "num_ctx": 40960}, "format": {{ schema | to_json }}}'
1370
+
1371
+ const OLLAMA_EXTRACTOR =
1372
+ '{"output_json": "{{ msg.message.content | json_escape }}", ' +
1373
+ '"explanation": "{{ msg.message.thinking | json_escape }}", ' +
1374
+ '"input_tokens": {{ msg.prompt_eval_count | default: 0 }}, ' +
1375
+ '"output_tokens": {{ msg.eval_count | default: 0 }}, ' +
1376
+ '"total_tokens": {{ msg.prompt_eval_count | default: 0 | plus: msg.eval_count }}}'
1377
+
1378
+ const llmTool = await ctx.ensure(
1379
+ 'LLM tool',
1380
+ async () => {
1381
+ const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
1382
+ return ctx.firstNamed(data.data, 'Cookbook Practice LLM Tool')
1383
+ },
1384
+ async () => {
1385
+ const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
1386
+ name: 'Cookbook Practice LLM Tool',
1387
+ description: 'Ollama-backed chat-completion adapter for contact-us triage.',
1388
+ intent: 'llm_enrichment',
1389
+ status: 'active',
1390
+ datalake_id: ctx.datalakeId,
1391
+ response_extractor: { type: 'custom', body: OLLAMA_EXTRACTOR, output_schema: ENRICHMENT_OUTPUT_SCHEMA },
1392
+ body: {
1393
+ tool_body_type: 'rest_api',
1394
+ base_url: 'http://localhost:11434',
1395
+ base_path: { type: 'custom', body: '/api/chat' },
1396
+ auth_method: 'api_key',
1397
+ api_key: 'stub-key',
1398
+ api_key_name: 'Authorization',
1399
+ api_key_location: 'header',
1400
+ request_type: 'json',
1401
+ response_type: 'json',
1402
+ timeout_ms: 60_000,
1403
+ base_body: { type: 'custom', body: OLLAMA_BASE_BODY },
1404
+ },
1405
+ })
1406
+ return data
1407
+ },
1408
+ )
1409
+ ctx.llmToolId = llmTool.id!
1410
+
1411
+ const TRIAGE_PROMPT = `You are a triage assistant for a primary-care practice website. Read the inbound message and classify it into EXACTLY ONE category:
1412
+
1413
+ - "clinical" — a symptom, a medication question, anything a clinician must see
1414
+ - "billing" — an invoice, a statement, insurance or payment
1415
+ - "appointment" — booking, rescheduling or cancelling a visit
1416
+ - "spam" — marketing, solicitation, or nothing to do with the practice
1417
+
1418
+ Message: {{ message }}
1419
+
1420
+ Respond with JSON: {"category": "<one of clinical|billing|appointment|spam>"}`
1421
+
1422
+ const triageAgent = await ctx.ensure(
1423
+ 'Contact Us triage agent',
1424
+ async () => {
1425
+ const { data } = await api.aiAgents.list(tenantSlug, datalakeSlug, { page_size: 100 })
1426
+ return ctx.firstNamed(data.data, 'Cookbook Contact Us Triage Agent')
1427
+ },
1428
+ async () => {
1429
+ const { data } = await api.aiAgents.create(tenantSlug, datalakeSlug, {
1430
+ name: 'Cookbook Contact Us Triage Agent',
1431
+ tool_id: ctx.llmToolId,
1432
+ model: 'qwen3-vl:8b-instruct',
1433
+ data_access: 'tokenized',
1434
+ temperature: 0.0,
1435
+ max_tokens: 1024,
1436
+ enabled: true,
1437
+ input_schema: {
1438
+ type: 'object',
1439
+ properties: { message: { type: 'string' } },
1440
+ required: ['message'],
1441
+ },
1442
+ llm_response_schema: {
1443
+ type: 'object',
1444
+ properties: {
1445
+ category: { type: 'string', enum: ['clinical', 'billing', 'appointment', 'spam'] },
1446
+ },
1447
+ required: ['category'],
1448
+ },
1449
+ prompt_config: { type: 'custom', body: TRIAGE_PROMPT },
1450
+ })
1451
+ return data
1452
+ },
1453
+ )
1454
+ aiAgentId = triageAgent.id!
1455
+ ctx.triageAgentSlug = triageAgent.slug!
1456
+ ```
1457
+
1458
+ ## 022 — the triage workflow, and three submissions to run it on
1459
+
1460
+ The agent is **nested in the create body** — `workflow_ai_agents` is a
1461
+ `cast_assoc`, not a separate attach call — with a `context_mapping_config`
1462
+ that projects each row into the agent's input schema. **Bracket access is
1463
+ required, not stylistic**: the agent's slug contains hyphens and Liquid's dot
1464
+ parser would read `additional_context.cookbook-contact-us-triage-agent` as a
1465
+ subtraction.
1466
+
1467
+ The three submissions are deliberately unambiguous — chest pain, an invoice
1468
+ query, and obvious solicitation — so the classification is not a coin toss
1469
+ and the assertion in §023 means something. `category` is left unset on
1470
+ ingest; assigning it is the agent's job.
1471
+
1472
+ ```typescript
1473
+ const CATEGORIES = ['clinical', 'billing', 'appointment', 'spam'] as const
1474
+
1475
+ const triageWorkflow = await ctx.ensure(
1476
+ 'Contact Us triage workflow',
1477
+ async () => {
1478
+ const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
1479
+ return ctx.firstNamed(data.data, 'Cookbook Contact Us Triage Workflow')
1480
+ },
1481
+ async () => {
1482
+ const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
1483
+ name: 'Cookbook Contact Us Triage Workflow',
1484
+ description: 'Classifies inbound messages into four buckets; one SMS action per bucket.',
1485
+ dataset_type: 'generic_table',
1486
+ generic_table_id: ctx.contactTableId,
1487
+ status: 'live',
1488
+ tags: ['triage', 'llm'],
1489
+ skip_mdm_resolution: true,
1490
+ filter_config: { type: 'custom', body: 'true' },
1491
+ decision_config: {
1492
+ type: 'custom',
1493
+ body: `["{{ additional_context["${ctx.triageAgentSlug}"].category }}"]`,
1494
+ output_schema: { type: 'array', items: { type: 'string' } },
1495
+ },
1496
+ actions: CATEGORIES.map((category) => ({
1497
+ decision_key: category,
1498
+ action_type: 'sms' as const,
1499
+ tool_id: ctx.smsToolId,
1500
+ position: 0,
1501
+ trigger_template: 'now',
1502
+ idempotency_template: `{{ event_dataset.submission_id }}-${category}`,
1503
+ tool_call: {
1504
+ tool_call_type: 'sms_request' as const,
1505
+ to: { type: 'custom' as const, body: '+15555550100' },
1506
+ body: {
1507
+ type: 'custom' as const,
1508
+ body: `[${category}] contact-us submission {{ event_dataset.submission_id }} needs routing.`,
1509
+ },
1510
+ sms_type: 'transactional' as const,
1511
+ },
1512
+ })),
1513
+ workflow_ai_agents: [
1514
+ {
1515
+ ai_agent_id: aiAgentId,
1516
+ position: 0,
1517
+ context_mapping_config: {
1518
+ type: 'custom',
1519
+ body: JSON.stringify({ message: '{{ event_dataset.message }}' }),
1520
+ },
1521
+ },
1522
+ ],
1523
+ })
1524
+ return data
1525
+ },
1526
+ )
1527
+ ctx.triageWorkflowSlug = triageWorkflow.slug!
1528
+
1529
+ ctx.clinicalId = 'SUB-COOKBOOK-0001'
1530
+ ctx.billingId = 'SUB-COOKBOOK-0002'
1531
+ ctx.spamId = 'SUB-COOKBOOK-0003'
1532
+
1533
+ const submissions = [
1534
+ {
1535
+ submission_id: ctx.clinicalId,
1536
+ name: 'Rosalind Achterberg',
1537
+ email: 'rosalind.achterberg@example.com',
1538
+ message:
1539
+ 'I have had chest tightness and shortness of breath since yesterday evening, ' +
1540
+ 'and it gets worse when I walk up stairs. Should I come in?',
1541
+ },
1542
+ {
1543
+ submission_id: ctx.billingId,
1544
+ name: 'Teodoro Vasquez-Lindqvist',
1545
+ email: 'teodoro.vl@example.com',
1546
+ message:
1547
+ 'I received an invoice for my visit last month but my insurance should have ' +
1548
+ 'covered it. Can someone check the claim and re-send the statement?',
1549
+ },
1550
+ {
1551
+ submission_id: ctx.spamId,
1552
+ name: 'Growth Partners',
1553
+ email: 'outreach@growth-partners.example',
1554
+ message:
1555
+ 'Boost your clinic revenue 300% with our SEO package! Limited slots. ' +
1556
+ 'Reply STOP to opt out. Click here for a free audit.',
1557
+ },
1558
+ ]
1559
+
1560
+ const ingests = await Promise.all(
1561
+ submissions.map((row) =>
1562
+ api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.contactDacSlug, { data: row }),
1563
+ ),
1564
+ )
1565
+ await ctx.waitForBatches(
1566
+ ctx.contactDacSlug,
1567
+ ingests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!),
1568
+ )
1569
+ ```
1570
+
1571
+ ## 023 — the agent's category steered the fan-out
1572
+
1573
+ Each submission produces an execution log carrying **four** action logs, one
1574
+ per category. On the first run for a given submission exactly **one** is
1575
+ matched (`:pending` or `:completed` — the category the agent chose) and
1576
+ **three** are `:skipped`.
1577
+
1578
+ **On a later run all four are `:skipped`, and that is correct.** The
1579
+ idempotency key is `<submission id>-<category>` and the submission ids are
1580
+ stable, so a key that already fired refuses to fire again. Both shapes are
1581
+ accepted; anything else is a real failure — two matched means the decision
1582
+ matched more than one category, zero matched with zero skipped means the
1583
+ fan-out never happened.
1584
+
1585
+ ```typescript
1586
+ const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, {
1587
+ sql_where_clause: `submission_id IN ('${ctx.clinicalId}', '${ctx.billingId}', '${ctx.spamId}')`,
1588
+ mode: 'live',
1589
+ manual_override: false,
1590
+ })
1591
+ const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
1592
+ ctx.triageRunBatchId = fired.batchId!
1593
+
1594
+ const refreshDeadline = Date.now() + 180_000
1595
+ let status: string | null = null
1596
+ while (Date.now() < refreshDeadline) {
1597
+ const { data: log } = await api.workflows.batchLogs.refresh(
1598
+ tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, fired.workflowRunLogId,
1599
+ )
1600
+ status = log.status ?? null
1601
+ if (status && status !== 'pending') break
1602
+ await new Promise((r) => setTimeout(r, 2_000))
1603
+ }
1604
+ if (status === 'failed') throw new Error('the triage run reached :failed')
1605
+ if (!status || status === 'pending') throw new Error('the triage run did not leave :pending within 180s')
1606
+
1607
+ const welDeadline = Date.now() + 180_000
1608
+ let ourLogs: Array<Record<string, unknown>> = []
1609
+ while (Date.now() < welDeadline) {
1610
+ const { data: wfLogs } = await api.workflows.workflowLogs.list(
1611
+ tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, { page_size: 100 },
1612
+ )
1613
+ ourLogs = (wfLogs.data ?? [])
1614
+ .map((w) => w as Record<string, unknown>)
1615
+ .filter((w) => w.batch_id === ctx.triageRunBatchId)
1616
+ const settled = ourLogs.filter(
1617
+ (w) => ((w.action_execution_logs as unknown[] | undefined) ?? []).length === 4,
1618
+ )
1619
+ if (ourLogs.length === 3 && settled.length === 3) break
1620
+ await new Promise((r) => setTimeout(r, 2_000))
1621
+ }
1622
+ if (ourLogs.length !== 3) {
1623
+ throw new Error(`expected 3 execution logs for the triage run, got ${ourLogs.length}`)
1624
+ }
1625
+
1626
+ for (const wel of ourLogs) {
1627
+ const welId = (wel as { id?: string }).id
1628
+ const aels = (wel as { action_execution_logs?: Array<{ status?: string }> }).action_execution_logs ?? []
1629
+ if (aels.length !== 4) {
1630
+ throw new Error(`log ${welId}: expected 4 action logs (one per category), got ${aels.length}`)
1631
+ }
1632
+ const byStatus: Record<string, number> = {}
1633
+ for (const ael of aels) {
1634
+ const st = ael.status ?? 'unknown'
1635
+ byStatus[st] = (byStatus[st] ?? 0) + 1
1636
+ }
1637
+ const matched = (byStatus.pending ?? 0) + (byStatus.completed ?? 0)
1638
+ const skipped = byStatus.skipped ?? 0
1639
+ if (matched === 1 && skipped === 3) continue // first run for this submission
1640
+ if (matched === 0 && skipped === 4) continue // already fired, correctly refusing
1641
+ throw new Error(
1642
+ `log ${welId}: expected 1 matched + 3 skipped, or 4 skipped on a repeat run — ${JSON.stringify(byStatus)}`,
1643
+ )
1644
+ }
1645
+ ```
1646
+
1647
+ ## 024 — the chest-pain message was routed clinical
1648
+
1649
+ §023 proved the fan-out chose one action per row. This proves *which* one,
1650
+ and it is the assertion that survives every rerun because it reads the
1651
+ durable message rather than the run that made it.
1652
+
1653
+ The chest-pain submission is the one asserted, because it is the only one
1654
+ whose category is genuinely beyond argument — tightness, breathlessness and
1655
+ exertional worsening is clinical to any classifier worth running, and a model
1656
+ that files it as billing is a model to find out about immediately.
1657
+
1658
+ The key is `<submission id>-clinical`, so **finding the message at all is the
1659
+ category assertion**. There is no need to parse the bucket back out of the
1660
+ body.
1661
+
1662
+ ```typescript
1663
+ const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
1664
+ search_query: `m.idempotency_key = '${ctx.clinicalId}-clinical'`,
1665
+ })
1666
+ if (search.status !== 'completed') {
1667
+ throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
1668
+ }
1669
+
1670
+ const msgDeadline = Date.now() + 90_000
1671
+ let clinicalBody: string | undefined
1672
+ while (Date.now() < msgDeadline && !clinicalBody) {
1673
+ const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
1674
+ userSearchId: search.id!,
1675
+ dataAccessMode: 'raw',
1676
+ })
1677
+ clinicalBody = ((data.data ?? []) as Array<Record<string, unknown>>)
1678
+ .map((m) => String(m.body ?? ''))
1679
+ .find((b) => b.includes('[clinical]'))
1680
+ if (!clinicalBody) await new Promise((r) => setTimeout(r, 2_000))
1681
+ }
1682
+ if (!clinicalBody) {
1683
+ throw new Error(
1684
+ `no clinical routing message for ${ctx.clinicalId} within 90s — ` +
1685
+ 'the agent filed the chest-pain message somewhere else, or the action never fired',
1686
+ )
1687
+ }
1688
+ if (!clinicalBody.includes(ctx.clinicalId)) {
1689
+ throw new Error(`the routing message does not name the submission: ${clinicalBody}`)
1690
+ }
1691
+ ```
1692
+
1693
+ ## 025 — upload a scanned document
1694
+
1695
+ The practice's back office receives documents as pixels — an employer group's
1696
+ registration papers, an insurance letter, a referral — and someone keys the
1697
+ same handful of fields out of each one. This scenario hands that to a model.
1698
+
1699
+ Mint a presigned URL per file and PUT the bytes with a **raw `fetch`**. That
1700
+ upload goes straight to object storage and is deliberately not an SDK call:
1701
+ the SDK talks to the platform, and the platform is not in this hop. Keep each
1702
+ returned `key` — that is how a file is handed to an agent.
1703
+
1704
+ ```typescript
1705
+ const { readFileSync } = await import('node:fs')
1706
+ const { join } = await import('node:path')
1707
+
1708
+ const uploadFixture = async (filename: string, contentType: string): Promise<string> => {
1709
+ const bytes = readFileSync(
1710
+ join(process.env.COOKBOOK_FIXTURES_DIR!, 'primary-care-feedback', filename),
1711
+ )
1712
+ const { data: link } = await api.datalakes.createUploadLink(tenantSlug, datalakeSlug, {
1713
+ filename,
1714
+ content_type: contentType,
1715
+ })
1716
+ const put = await fetch(link.url!, {
1717
+ method: 'PUT',
1718
+ headers: { 'Content-Type': contentType },
1719
+ body: bytes,
1720
+ })
1721
+ if (put.status !== 200) throw new Error(`upload of ${filename} failed: ${put.status}`)
1722
+ return link.key!
1723
+ }
1724
+
1725
+ ctx.scanKey = await uploadFixture('memorandum-of-association-01.png', 'image/png')
1726
+ ```
1727
+
1728
+ ## 026 — the vision agent
1729
+
1730
+ Same tool as §021 — one provider adapter serves both agents, because the
1731
+ adapter is about the provider's wire shape and not about the task. Its
1732
+ `base_body` already appends `images` to the message, so a multimodal call
1733
+ needs no second tool.
1734
+
1735
+ The agent differs in three ways that matter. Its `input_schema` takes
1736
+ `instructions` rather than a message, because the caller steers this one per
1737
+ invocation. Its `llm_response_schema` describes the fields to pull out of the
1738
+ document, and it is load-bearing: the platform constrains the model to it, so
1739
+ a document with no company name comes back with `company_name: null` rather
1740
+ than with the model's apology. And `additionalProperties: false` means a
1741
+ model that volunteers a field nobody asked for is refused rather than
1742
+ silently widening the record.
1743
+
1744
+ `data_access: 'tokenized'` again. It changes nothing about *this* call — the
1745
+ input is an uploaded file, not a lake row — and it is set anyway, because the
1746
+ tier is a property of the agent rather than of the invocation, and an agent
1747
+ that is later pointed at a workflow inherits whatever it was created with. A
1748
+ tokenized agent cannot be made raw by the caller; that is the point of
1749
+ putting the tier on the resource.
1750
+
1751
+ ```typescript
1752
+ const extractionAgent = await ctx.ensure(
1753
+ 'document extraction agent',
1754
+ async () => {
1755
+ const { data } = await api.aiAgents.list(tenantSlug, datalakeSlug, { page_size: 100 })
1756
+ return ctx.firstNamed(data.data, 'Cookbook Document Extraction Agent')
1757
+ },
1758
+ async () => {
1759
+ const { data } = await api.aiAgents.create(tenantSlug, datalakeSlug, {
1760
+ name: 'Cookbook Document Extraction Agent',
1761
+ description: 'Reads scanned back-office documents and returns structured fields.',
1762
+ tool_id: ctx.llmToolId,
1763
+ model: 'qwen3-vl:8b-instruct',
1764
+ data_access: 'tokenized',
1765
+ temperature: 0.0,
1766
+ max_tokens: 2048,
1767
+ enabled: true,
1768
+ input_schema: {
1769
+ type: 'object',
1770
+ properties: { instructions: { type: 'string' } },
1771
+ required: ['instructions'],
1772
+ },
1773
+ llm_response_schema: {
1774
+ type: 'object',
1775
+ properties: {
1776
+ company_name: { type: ['string', 'null'] },
1777
+ registered_address: { type: ['string', 'null'] },
1778
+ date_of_formation: { type: ['string', 'null'] },
1779
+ },
1780
+ required: ['company_name'],
1781
+ additionalProperties: false,
1782
+ },
1783
+ prompt_config: {
1784
+ type: 'custom',
1785
+ body:
1786
+ 'You are a document data-extraction specialist. Read the attached ' +
1787
+ 'document image and return the requested fields as JSON. Use null for ' +
1788
+ 'anything the document does not state.\n\n{{ instructions }}',
1789
+ },
1790
+ })
1791
+ return data
1792
+ },
1793
+ )
1794
+ ctx.extractionAgentId = extractionAgent.id!
1795
+ ```
1796
+
1797
+ ## 027 — invoke it against the scan
1798
+
1799
+ `invoke` runs the agent against the uploaded files: `input` satisfies the
1800
+ `input_schema`, and `files` lists the storage keys plus their content types.
1801
+ What comes back is the parsed JSON matching `llm_response_schema`, an
1802
+ `explanation`, and token `usage`.
1803
+
1804
+ The assertion is deliberately weak about *content* and strict about *shape*.
1805
+ A live model reading a scan will not return the same string twice, so
1806
+ asserting the company's exact name would make this step a coin toss that
1807
+ fails on a model upgrade rather than on a regression. What must hold is that
1808
+ the call round-trips: structured output, conforming to the schema the agent
1809
+ pins, with usage accounted for. That is the contract the platform owns; the
1810
+ transcription accuracy is the model's.
1811
+
1812
+ ```typescript
1813
+ const { data: result } = await api.aiAgents.invoke(tenantSlug, datalakeSlug, ctx.extractionAgentId, {
1814
+ input: {
1815
+ instructions:
1816
+ 'Extract the company name, its registered address, and the date of formation ' +
1817
+ 'from the attached document.',
1818
+ },
1819
+ files: [{ key: ctx.scanKey, content_type: 'image/png' }],
1820
+ })
1821
+
1822
+ if (!result.output || typeof result.output !== 'object') {
1823
+ throw new Error('invoke did not return a structured output object')
1824
+ }
1825
+ const output = result.output as Record<string, unknown>
1826
+ if (!('company_name' in output)) {
1827
+ throw new Error(
1828
+ `the response does not carry the required company_name key: ${JSON.stringify(output)}`,
1829
+ )
1830
+ }
1831
+ // additionalProperties: false — nothing outside the pinned schema may appear.
1832
+ const allowed = new Set(['company_name', 'registered_address', 'date_of_formation'])
1833
+ const extra = Object.keys(output).filter((k) => !allowed.has(k))
1834
+ if (extra.length > 0) {
1835
+ throw new Error(`the model returned fields outside llm_response_schema: ${extra.join(', ')}`)
1836
+ }
1837
+ if (!result.usage || (result.usage as { input_tokens?: number }).input_tokens === undefined) {
1838
+ throw new Error('expected token usage on the invoke result')
1839
+ }
1840
+ console.log(' extracted →', JSON.stringify(output))
1841
+ ```
1842
+
1843
+ ## 028 — a malformed input is refused before the model is called
1844
+
1845
+ `invoke` validates `input` against the agent's `input_schema` **before**
1846
+ calling the provider. Omitting the required `instructions` returns a 422, and
1847
+ no tokens are spent.
1848
+
1849
+ That ordering is the point. The schema is not documentation about what the
1850
+ caller ought to send — it is a gate, and it fails closed at the cheapest
1851
+ possible moment. An agent whose input gate ran after the model call would
1852
+ turn every caller's typo into a billed request.
1853
+
1854
+ The refusal is caught so the walk proves the gate without aborting, and the
1855
+ status is checked rather than the message: a 500 here would also throw, and
1856
+ treating it as a pass would be exactly the kind of test that proves nothing.
1857
+
1858
+ ```typescript
1859
+ let rejected = false
1860
+ try {
1861
+ await api.aiAgents.invoke(tenantSlug, datalakeSlug, ctx.extractionAgentId, {
1862
+ input: { wrong_field: 'no instructions here' },
1863
+ files: [],
1864
+ })
1865
+ } catch (err) {
1866
+ const status = (err as { _httpStatus?: number })._httpStatus
1867
+ if (status !== 422) throw err
1868
+ rejected = true
1869
+ }
1870
+ if (!rejected) {
1871
+ throw new Error('expected a 422 for input that violates the agent input_schema')
1872
+ }
1873
+ ```
1874
+
1875
+ ## 029 — one patient, three lakes, three answers
1876
+
1877
+ The last step is the one this whole surface exists for. §016 proved a masked
1878
+ *array* survives the copy; this proves what masking does to the thing anyone
1879
+ would actually recognise — a person's name — and it shows all three answers
1880
+ side by side.
1881
+
1882
+ `mode` is a parameter of the read, not a different slug. One lake is
1883
+ addressed and `mode` selects which copy answers, which works only because
1884
+ §004 built its client at the `raw` ceiling. A tokenized-ceiling session
1885
+ asking for `raw` is refused rather than quietly downgraded, and that is how a
1886
+ support console is kept honest: you give it a key it cannot exceed instead of
1887
+ trusting every screen to ask nicely.
1888
+
1889
+ Read the three and compare. The **name must differ** in both copies — if it
1890
+ does not, masking is not in force and the agents above have been reading
1891
+ identities all along. The **appointment id must be identical** in all three —
1892
+ it is a `none` column, and a masking policy that changed it would have broken
1893
+ every join in the system while looking like it was working.
1894
+
1895
+ The two copies are also different from *each other*, and that is the design
1896
+ rather than an accident: the tokenized copy carries a stable
1897
+ `tok_<hash>` — a machine can join on it across tables without ever resolving
1898
+ it — while the redacted copy carries `xxxx-<tail>`, which is unjoinable and
1899
+ meant for a human to glance at. One is for a model, the other is for a
1900
+ person.
1901
+
1902
+ ```typescript
1903
+ const identitySql =
1904
+ `SELECT appointment_id, patient_name, department ` +
1905
+ `FROM ${ctx.appointmentsTableName} WHERE appointment_id = '${ctx.fulfilledId}'`
1906
+
1907
+ const readIdentityAs = async (mode: 'raw' | 'tokenized' | 'redacted') => {
1908
+ const { data } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
1909
+ sql: identitySql,
1910
+ mode,
1911
+ })
1912
+ const row = (data.data ?? [])[0]
1913
+ if (!row) throw new Error(`the appointment was not readable through the ${mode} lake`)
1914
+ const cell = (column: string): unknown => row[data.meta.columns.indexOf(column)]
1915
+ return {
1916
+ id: String(cell('appointment_id')),
1917
+ name: String(cell('patient_name')),
1918
+ department: String(cell('department')),
1919
+ }
1920
+ }
1921
+
1922
+ const rawRead = await readIdentityAs('raw')
1923
+ const tokenizedRead = await readIdentityAs('tokenized')
1924
+ const redactedRead = await readIdentityAs('redacted')
1925
+
1926
+ console.log(' raw →', JSON.stringify(rawRead))
1927
+ console.log(' tokenized →', JSON.stringify(tokenizedRead))
1928
+ console.log(' redacted →', JSON.stringify(redactedRead))
1929
+
1930
+ if (rawRead.name !== 'Rosalind Achterberg') {
1931
+ throw new Error(`the raw lake should hold the real name, got ${JSON.stringify(rawRead.name)}`)
1932
+ }
1933
+ for (const [mode, read] of [['tokenized', tokenizedRead], ['redacted', redactedRead]] as const) {
1934
+ if (read.name === rawRead.name) {
1935
+ throw new Error(`the ${mode} lake handed back the real patient name — masking is not in force`)
1936
+ }
1937
+ if (read.id !== rawRead.id || read.department !== rawRead.department) {
1938
+ throw new Error(
1939
+ `the ${mode} lake changed a 'none' column: id ${read.id} vs ${rawRead.id}, ` +
1940
+ `department ${read.department} vs ${rawRead.department}`,
1941
+ )
1942
+ }
1943
+ }
1944
+ // The two copies serve different readers, so they must not be the same value.
1945
+ if (tokenizedRead.name === redactedRead.name) {
1946
+ throw new Error(
1947
+ 'the tokenized and redacted lakes returned the same masked value — ' +
1948
+ 'one is meant to be a stable join key and the other a human-readable placeholder',
1949
+ )
1950
+ }
1951
+ ```
1952
+
1953
+ ## 030 — how far the copy actually got, table by table
1954
+
1955
+ Every read above trusted the copy. This step asks it directly.
1956
+
1957
+ `derivedLakeTables` answers per table, and the addressing is the first thing
1958
+ to get right: you name the **primary** lake plus a mode. A derived lake does
1959
+ have a slug of its own, but that slug is not an address — the lookup resolves
1960
+ primaries only, and the walk proves it by asking with the tokenized lake's own
1961
+ slug and catching the refusal.
1962
+
1963
+ The state to read is `caught_up`, never the state's *name*. One of the five
1964
+ states reads as finished without being finished: `finished_copy` means the
1965
+ first pass is done but the changes made *during* that pass are not applied
1966
+ yet. A surface that treats it as ready shows a lake as complete while rows
1967
+ are still missing. The boolean exists so nobody has to pattern-match the word.
1968
+
1969
+ `resync_empties` carries a rule worth asserting rather than assuming: it is
1970
+ `null` for a table that is already caught up — the platform does not spend a
1971
+ query per table answering a question about tables nobody is going to repair —
1972
+ and an array otherwise. `null` and `[]` mean different things, and `[]` never
1973
+ occurs, because a table is always in its own set.
1974
+
1975
+ ```typescript
1976
+ const { data: pair } = await api.datalakes.provisionTokenization(tenantSlug, datalakeSlug)
1977
+ const tokenizedLake = (pair.derived_datalakes ?? []).find((d) => d.type === 'tokenized')
1978
+ if (!tokenizedLake) throw new Error('no tokenized copy on the lake — §007 should have provisioned it')
1979
+
1980
+ const { data: tables } = await api.datalakes.derivedLakeTables(
1981
+ tenantSlug, datalakeSlug, 'tokenized',
1982
+ )
1983
+ if (tables.data.length === 0) {
1984
+ throw new Error('the tokenized copy reported no tables at all')
1985
+ }
1986
+ console.log(` ${tables.data.length} tables in the tokenized copy`)
1987
+
1988
+ for (const t of tables.data) {
1989
+ if (typeof t.caught_up !== 'boolean') {
1990
+ throw new Error(`${t.table}: caught_up is not a boolean, got ${JSON.stringify(t.caught_up)}`)
1991
+ }
1992
+ if (!t.label || t.label.length === 0) {
1993
+ throw new Error(`${t.table}: label is empty — it is what a screen renders`)
1994
+ }
1995
+ // The agreement that makes the boolean trustworthy: `caught_up` is true for
1996
+ // exactly one of the five states. `finished_copy` reading true here would be
1997
+ // the defect the boolean exists to prevent. This cannot be asserted against a
1998
+ // mock — the schema cannot express a cross-field constraint, so Prism draws
1999
+ // the two independently — which is exactly why it is asserted here.
2000
+ if (t.caught_up !== (t.state === 'caught_up')) {
2001
+ throw new Error(
2002
+ `${t.table}: state '${t.state}' disagrees with caught_up=${t.caught_up} — ` +
2003
+ `only 'caught_up' means arrived`,
2004
+ )
2005
+ }
2006
+ // The null-vs-array rule, asserted in both directions.
2007
+ if (t.caught_up && t.resync_empties !== null && t.resync_empties !== undefined) {
2008
+ throw new Error(
2009
+ `${t.table} is caught up, so resync_empties should not be computed — got ` +
2010
+ JSON.stringify(t.resync_empties),
2011
+ )
2012
+ }
2013
+ if (!t.caught_up) {
2014
+ if (!Array.isArray(t.resync_empties) || t.resync_empties.length === 0) {
2015
+ throw new Error(`${t.table} is behind, so resync_empties should be a non-empty array`)
2016
+ }
2017
+ if (t.resync_empties[0] !== t.table) {
2018
+ throw new Error(`${t.table}: resync_empties should name the table itself first`)
2019
+ }
2020
+ }
2021
+ }
2022
+ ctx.copyTables = tables.data.map((t) => t.table)
2023
+
2024
+ // A derived lake's own slug is not an address. Prove the refusal rather than
2025
+ // assuming it — a 200 here would mean two ways to name the same thing.
2026
+ let refusedDerivedSlug = false
2027
+ try {
2028
+ await api.datalakes.derivedLakeTables(tenantSlug, tokenizedLake.slug!, 'tokenized')
2029
+ } catch (err) {
2030
+ const status = (err as { _httpStatus?: number })._httpStatus
2031
+ if (status !== 404) throw err
2032
+ refusedDerivedSlug = true
2033
+ }
2034
+ if (!refusedDerivedSlug) {
2035
+ throw new Error("the derived lake's own slug was accepted as an address — it should 404")
2036
+ }
2037
+ ```
2038
+
2039
+ ## 031 — empty one table and watch it come back
2040
+
2041
+ `resyncTable` is the repair: it empties a table in the copy and copies it
2042
+ again from the start. Two things about it are easy to get wrong, and both are
2043
+ asserted here rather than described.
2044
+
2045
+ **It empties more than the table you name.** Postgres refuses to truncate any
2046
+ table with an inbound foreign key — whether or not the referencing tables
2047
+ hold rows — so repairing a parent takes its whole dependent closure with it,
2048
+ transitively, in one statement. All of them are re-copied by the same run.
2049
+ `legal_entities` is chosen deliberately: patient identities were resolved
2050
+ through it in §008, and `legal_entity_identifications` points at it, so the
2051
+ `emptied` list here is genuinely longer than one. A surface offering this
2052
+ action should read `resync_empties` and say the number before the button is
2053
+ pressed.
2054
+
2055
+ **`202` means started, not done.** There is no completion callback. The walk
2056
+ polls `derivedLakeTables` until the table reads `caught_up` again, which is
2057
+ also the honest way to prove the repair worked rather than merely dispatched.
2058
+
2059
+ The table name is gated against what the copy carries, so a typo is a `422`
2060
+ and **nothing is emptied**. That gate is the only thing standing between a
2061
+ mistyped name and an emptied table, so it is proved here too — and it is
2062
+ proved *first*, because a walk that emptied the wrong table and then checked
2063
+ the gate would have the order backwards.
2064
+
2065
+ ```typescript
2066
+ const TARGET = 'legal_entities'
2067
+ if (!(ctx.copyTables as string[]).includes(TARGET)) {
2068
+ throw new Error(`${TARGET} is not in the copy; it carries: ${(ctx.copyTables as string[]).join(', ')}`)
2069
+ }
2070
+
2071
+ // The gate first — nothing should be emptied by a name the copy does not carry.
2072
+ let refusedBadTable = false
2073
+ try {
2074
+ await api.datalakes.resyncTable(tenantSlug, datalakeSlug, 'tokenized', {
2075
+ table: 'no_such_table_here',
2076
+ })
2077
+ } catch (err) {
2078
+ const status = (err as { _httpStatus?: number })._httpStatus
2079
+ if (status !== 422) throw err
2080
+ refusedBadTable = true
2081
+ }
2082
+ if (!refusedBadTable) throw new Error('an unknown table name was accepted — the gate is not closed')
2083
+
2084
+ const { data: repair } = await api.datalakes.resyncTable(
2085
+ tenantSlug, datalakeSlug, 'tokenized', { table: TARGET },
2086
+ )
2087
+ if (repair.table !== TARGET) throw new Error(`repaired ${repair.table}, asked for ${TARGET}`)
2088
+ if (repair.mode !== 'tokenized') throw new Error(`repaired the ${repair.mode} copy, asked for tokenized`)
2089
+ if (repair.status !== 'copying') throw new Error(`expected status 'copying', got ${repair.status}`)
2090
+ if (!repair.emptied.includes(TARGET)) {
2091
+ throw new Error(`emptied ${JSON.stringify(repair.emptied)} — the named table is not among them`)
2092
+ }
2093
+ console.log(` emptied ${repair.emptied.length} table(s): ${repair.emptied.join(', ')}`)
2094
+
2095
+ // The dependent closure is the point: identifications point at legal entities,
2096
+ // so a repair of the parent cannot leave the child behind.
2097
+ if (repair.emptied.length < 2) {
2098
+ throw new Error(
2099
+ `${TARGET} has dependents, so a repair should empty more than itself — got ` +
2100
+ JSON.stringify(repair.emptied),
2101
+ )
2102
+ }
2103
+
2104
+ // 202 = started. Wait for the copy to arrive again.
2105
+ const RESYNC_TIMEOUT_MS = 5 * 60_000
2106
+ const resyncDeadline = Date.now() + RESYNC_TIMEOUT_MS
2107
+ let backOnline = false
2108
+ while (Date.now() < resyncDeadline && !backOnline) {
2109
+ const { data: after } = await api.datalakes.derivedLakeTables(
2110
+ tenantSlug, datalakeSlug, 'tokenized',
2111
+ )
2112
+ backOnline = (repair.emptied as string[]).every(
2113
+ (name) => after.data.find((t) => t.table === name)?.caught_up === true,
2114
+ )
2115
+ if (!backOnline) await new Promise((r) => setTimeout(r, 5_000))
2116
+ }
2117
+ if (!backOnline) {
2118
+ throw new Error(`the repaired tables did not come back within ${RESYNC_TIMEOUT_MS / 1000}s`)
2119
+ }
2120
+ console.log(' every emptied table is caught up again')
2121
+ ```
2122
+
2123
+ # Outcome
2124
+
2125
+ After all thirty-one steps run green, on one tenant and **three datalakes**:
2126
+
2127
+ - **A raw lake and its tokenized and redacted copies**, provisioned as a pair
2128
+ and each carrying every table the primary has.
2129
+ - **Two generic tables** — appointments and contact-us submissions — one
2130
+ ingested through a contract pair on two sequenced clients, the other
2131
+ through the default client the platform provisioned when the table was
2132
+ declared.
2133
+ - **Two workflows**: a review request that reaches only the patient whose
2134
+ visit was fulfilled, and a four-way category fan-out driven by an LLM
2135
+ agent reading the tokenized projection.
2136
+ - **Two agents**, one classifying text and one reading a scanned document,
2137
+ both pinned to `data_access: 'tokenized'` and both constrained by an
2138
+ `llm_response_schema` the platform enforces.
2139
+ - **The same row, read three ways**, showing what each reader sees.
2140
+ - **The tokenized copy inspected and repaired** — every table's arrival state
2141
+ read back, one table emptied and copied again, and every table the repair
2142
+ took with it waited back to `caught_up`.
2143
+
2144
+ What this surface teaches that the raw cookbooks cannot:
2145
+
2146
+ 1. **A masked jsonb column is `jsonb` in raw and `text` in both copies, and
2147
+ the types are supposed to differ.** Logical replication ships a text
2148
+ rendering and feeds it to the subscriber's input function, so
2149
+ `jsonb → text` always works and `jsonb → text[]` never does. The
2150
+ invariant is scalar-ness, not type parity — and `is_array: true`
2151
+ describes the JSON document, not the Postgres column.
2152
+ 2. **The tier follows what the prompt can reach.** The triage agent here runs
2153
+ `tokenized` because it reads patient messages; the equivalent agent in
2154
+ `subscription-saas` runs `raw` because its prompt carries an account
2155
+ number and nothing else. Neither is a house style.
2156
+ 3. **The two copies are different from each other on purpose.** `tok_<hash>`
2157
+ is a stable join key for a machine; `xxxx-<tail>` is an unjoinable
2158
+ placeholder for a person.
2159
+ 4. **A key the contract does not read is not an error.** It renders nothing,
2160
+ the column lands `null`, the batch reports green, and the failure appears
2161
+ two steps later as an action rejecting an empty recipient. Assert columns
2162
+ on the row.
2163
+ 5. **"The copy exists" and "the copy is complete" are different questions.**
2164
+ `finished_copy` has not finished — read `caught_up`, not the state's name.
2165
+ And a repair is not scoped to the table you name: Postgres will not
2166
+ truncate a table anything references, so the dependent closure goes with
2167
+ it. Both are why `emptied` is returned rather than assumed.
2168
+
2169
+ # See also
2170
+
2171
+ - `payments-compliance.md` — the other tokenized surface, and the one that
2172
+ shows a masked read against an agent's verdict
2173
+ - `subscription-saas.md` — the same contract-pair discipline on a raw lake
2174
+ - `../datalakes.md` — the tier split, and what each copy is for
2175
+ - `../ai_agents.md` — `data_access`, `input_schema`, `llm_response_schema`