@alvera-ai/platform-sdk 0.17.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (83) hide show
  1. package/.agent/AGENTS.md +82 -144
  2. package/.agent/account_management.md +2 -2
  3. package/.agent/action_logs.md +4 -4
  4. package/.agent/ai_agents.md +28 -21
  5. package/.agent/ai_sandbox.md +49 -39
  6. package/.agent/connected_apps.md +3 -3
  7. package/.agent/cookbook/_fixtures/README.md +1 -1
  8. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_generic_table.liquid +1 -1
  9. package/.agent/cookbook/_fixtures/organic-marketing/_lead_submissions_foundation_legal_entity.liquid +80 -0
  10. package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_mdm.liquid +2 -1
  11. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_generic_table.liquid +57 -0
  12. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_legal_entity.liquid +30 -0
  13. package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_mdm.liquid +44 -0
  14. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_generic_table.liquid +57 -0
  15. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_legal_entity.liquid +36 -0
  16. package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_mdm.liquid +41 -0
  17. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_generic_table.liquid +70 -0
  18. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_legal_entity.liquid +52 -0
  19. package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_mdm.liquid +42 -0
  20. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_generic_table.liquid +38 -0
  21. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_legal_entity.liquid +48 -0
  22. package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_mdm.liquid +49 -0
  23. package/.agent/cookbook/organic-marketing.md +2801 -0
  24. package/.agent/cookbook/payments-compliance.md +2180 -0
  25. package/.agent/cookbook/primary-care.md +2175 -0
  26. package/.agent/cookbook/subscription-saas.md +2403 -0
  27. package/.agent/data_activation_clients.md +65 -52
  28. package/.agent/datalakes.md +338 -171
  29. package/.agent/errors.md +3 -3
  30. package/.agent/generic_tables.md +151 -62
  31. package/.agent/interoperability_contracts.md +57 -22
  32. package/.agent/mdm.md +136 -153
  33. package/.agent/messages.md +36 -34
  34. package/.agent/mock-services.md +1 -1
  35. package/.agent/mutations.md +2 -2
  36. package/.agent/templates.md +14 -13
  37. package/.agent/tools.md +63 -21
  38. package/.agent/type_naming.md +13 -13
  39. package/.agent/workflows.md +99 -53
  40. package/README.md +2 -2
  41. package/dist/bin/platform-sdk.mjs +33 -47
  42. package/dist/bin/platform-sdk.mjs.map +1 -1
  43. package/dist/index.d.mts +565 -379
  44. package/dist/index.d.mts.map +1 -1
  45. package/dist/index.mjs +494 -59
  46. package/dist/index.mjs.map +1 -1
  47. package/package.json +4 -3
  48. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_legal_entity.liquid +0 -88
  49. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_appointment.liquid +0 -47
  50. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_mdm.liquid +0 -24
  51. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_patient.liquid +0 -38
  52. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_compliance_screening.liquid +0 -59
  53. package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_mdm.liquid +0 -36
  54. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_mdm.liquid +0 -30
  55. package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_payment_account.liquid +0 -55
  56. package/.agent/cookbook/_fixtures/subscription/_customers_subscription_mdm.liquid +0 -20
  57. package/.agent/cookbook/_setup/foundation.md +0 -359
  58. package/.agent/cookbook/_setup/healthcare.md +0 -361
  59. package/.agent/cookbook/_setup/payments.md +0 -365
  60. package/.agent/cookbook/_setup/subscription.md +0 -364
  61. package/.agent/cookbook/action-status-updaters.md +0 -278
  62. package/.agent/cookbook/ai-agent-invoke.md +0 -279
  63. package/.agent/cookbook/appointment-review-sms-workflow.md +0 -801
  64. package/.agent/cookbook/birthday-greeting-sms-trigger.md +0 -696
  65. package/.agent/cookbook/bulk-ingest.md +0 -302
  66. package/.agent/cookbook/contact-us-triage-with-llm.md +0 -663
  67. package/.agent/cookbook/dunning-sms-for-delinquent.md +0 -659
  68. package/.agent/cookbook/generic-tables.md +0 -244
  69. package/.agent/cookbook/invite-team.md +0 -200
  70. package/.agent/cookbook/kyc-notification-on-account-activation.md +0 -661
  71. package/.agent/cookbook/marketing-campaign-send.md +0 -1044
  72. package/.agent/cookbook/paginated-restapi-poller.md +0 -383
  73. package/.agent/cookbook/rest-fetch.md +0 -273
  74. package/.agent/cookbook/sanctions-screening-with-agent-review.md +0 -773
  75. package/.agent/cookbook/score-leads-with-llm-categorization.md +0 -665
  76. package/.agent/cookbook/system-templates.md +0 -165
  77. package/.agent/cookbook/talk-to-data.md +0 -178
  78. package/.agent/cookbook/triage-prospects-by-priority.md +0 -571
  79. package/.agent/cookbook/welcome-sms-for-customers.md +0 -647
  80. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/memorandum-of-association-01.png +0 -0
  81. /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/sample_two_page.pdf +0 -0
  82. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/_customers_subscription_customer.liquid +0 -0
  83. /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/stripe_customers_batch1.csv +0 -0
@@ -1,17 +1,20 @@
1
1
  # Datalakes
2
2
 
3
- A **datalake** is the storage layer for one tenant slice. It holds
4
- two databases under one roof:
3
+ A **datalake** is the storage layer for one tenant slice: one
4
+ Postgres database — a writer profile and a reader profile onto it
5
+ — plus one cloud-storage bucket.
5
6
 
6
- - an **unregulated** database for tokenized data, exposed to AI
7
- workflows and downstream agents
8
- - a **regulated** database for raw values that never leave the
9
- trust boundary
7
+ Masking is **not** a second database under the same roof. A
8
+ datalake carries a `type` (`raw`, `tokenized`, or `redacted`), and
9
+ the derived tiers are *separate datalake rows* pointing back at the
10
+ raw one through `primary_datalake_id`, filled by replication rather
11
+ than written directly. §2b covers which tier to read, and why
12
+ reading the wrong one is the subtlest bug in this guide.
10
13
 
11
14
  Every other resource in the platform (data sources, tools, AI
12
15
  agents, workflows, etc.) scopes under a datalake. A tenant can own
13
- many datalakes; the canonical pattern is one per data domain (e.g.
14
- `healthcare`, `foundation`, `payments`).
16
+ many datalakes; group them however the tenant's data actually
17
+ divides.
15
18
 
16
19
  ```typescript
17
20
  import type { PlatformApi } from '@alvera-ai/platform-sdk'
@@ -26,7 +29,7 @@ SDK namespace: `api.datalakes`.
26
29
  ## 1. Wire shape
27
30
 
28
31
  The request type splits across two layers: top-level database +
29
- identity fields, plus two polymorphic cloud-storage embeds.
32
+ identity fields, plus one polymorphic cloud-storage embed.
30
33
 
31
34
  ```typescript
32
35
  import type {
@@ -41,8 +44,8 @@ there is no runtime validator; validation is server-authoritative.)
41
44
 
42
45
  ### Polymorphic cloud-storage embed
43
46
 
44
- `unregulated_cloud_storage` and `regulated_cloud_storage` are
45
- polymorphic. The discriminator field is **`cloud_storage_type`**
47
+ `cloud_storage` is polymorphic. The discriminator field is
48
+ **`cloud_storage_type`**
46
49
  (see `mutations.md` "Polymorphic discriminator naming") with three
47
50
  branches:
48
51
 
@@ -63,16 +66,17 @@ credentials the same way the DB-profile `auth_method` does. See
63
66
  §2 "AWS cloud-storage `auth_method = iam_role` forbids
64
67
  credentials" for the conditional the type cannot encode.
65
68
 
66
- The two embeds are independent — a datalake can pair an `aws`
67
- unregulated bucket with a `custom` regulated bucket. Both embeds
68
- ship in the same POST/PUT body.
69
+ There is one embed, not a pair. The old
70
+ `unregulated_cloud_storage` / `regulated_cloud_storage` pairing
71
+ died with the two-tier lake — a derived lake is its own datalake
72
+ row and carries its own `cloud_storage`.
69
73
 
70
74
  ### Cloud storage is also the cold-archive destination
71
75
 
72
76
  Beyond being the platform's primary blob store for inline payloads
73
77
  (image uploads, large context payloads, presigned-URL ingest), the
74
- configured `regulated_cloud_storage` and `unregulated_cloud_storage`
75
- buckets are the **cold-archive destination** for the datalake. As
78
+ configured `cloud_storage` bucket is the **cold-archive
79
+ destination** for the datalake. As
76
80
  data flows through the pipeline, two NDJSON streams land under
77
81
  well-defined key prefixes inside the bucket the credentials point at:
78
82
 
@@ -111,20 +115,13 @@ data lives in the cloud-storage buckets configured here.
111
115
 
112
116
  ### Canonical create body
113
117
 
114
- Paste-ready TypeScript body. The four database-role groups
115
- (unregulated writer + reader, regulated writer + reader) share
116
- an identical eight-field shape; the type lists all 32 explicitly
117
- with no role-level union:
118
+ Paste-ready TypeScript body. There are two database-role groups —
119
+ writer and reader — sharing an identical eight-field shape; the
120
+ type lists all 16 explicitly with no role-level union.
118
121
 
119
122
  ```typescript
120
123
  import type { DatalakeRequestWritable } from '@alvera-ai/platform-sdk'
121
124
 
122
- // Pick one of four supported industries. The body shape below is
123
- // identical across them; only `data_domain` and the names it
124
- // interpolates differ.
125
- const dataDomain = 'healthcare'
126
- // ^ swap to 'foundation' | 'subscription' | 'payments'
127
-
128
125
  // Secrets resolved upstream by your code (secret store, config loader, etc.)
129
126
  const pgHost = '...'
130
127
  const pgUser = '...'
@@ -137,56 +134,40 @@ const awsSecretAccessKey = '...'
137
134
 
138
135
  const body: DatalakeRequestWritable = {
139
136
  name: 'Production Lake',
140
- description: `Primary ${dataDomain} datalake`,
141
- data_domain: dataDomain,
137
+ description: 'Primary datalake',
142
138
  timezone: 'America/New_York',
143
139
  pool_size: 5,
144
140
 
145
- unregulated_db_writer_host: pgHost,
146
- unregulated_db_writer_port: 5432,
147
- unregulated_db_writer_name: `alvera_${dataDomain}`,
148
- unregulated_db_writer_schema: 'tenant_acme_unreg',
149
- unregulated_db_writer_auth_method: 'password',
150
- unregulated_db_writer_user: pgUser,
151
- unregulated_db_writer_pass: pgPass,
152
- unregulated_db_writer_enable_ssl: true,
153
- unregulated_db_reader_host: pgReaderHost,
154
- unregulated_db_reader_port: 5432,
155
- unregulated_db_reader_name: `alvera_${dataDomain}`,
156
- unregulated_db_reader_schema: 'tenant_acme_unreg',
157
- unregulated_db_reader_auth_method: 'password',
158
- unregulated_db_reader_user: pgReaderUser,
159
- unregulated_db_reader_pass: pgReaderPass,
160
- unregulated_db_reader_enable_ssl: true,
161
-
162
- regulated_data_db_writer_host: pgHost,
163
- regulated_data_db_writer_port: 5432,
164
- regulated_data_db_writer_name: `alvera_${dataDomain}`,
165
- regulated_data_db_writer_schema: 'tenant_acme_reg',
166
- regulated_data_db_writer_auth_method: 'password',
167
- regulated_data_db_writer_user: pgUser,
168
- regulated_data_db_writer_pass: pgPass,
169
- regulated_data_db_writer_enable_ssl: true,
170
- regulated_data_db_reader_host: pgReaderHost,
171
- regulated_data_db_reader_port: 5432,
172
- regulated_data_db_reader_name: `alvera_${dataDomain}`,
173
- regulated_data_db_reader_schema: 'tenant_acme_reg',
174
- regulated_data_db_reader_auth_method: 'password',
175
- regulated_data_db_reader_user: pgReaderUser,
176
- regulated_data_db_reader_pass: pgReaderPass,
177
- regulated_data_db_reader_enable_ssl: true,
178
-
179
- unregulated_cloud_storage: {
141
+ // `type` defaults to 'raw' — the lake the platform writes.
142
+ // A derived lake sets type: 'tokenized' | 'redacted' AND
143
+ // primary_datalake_id; see §2b before creating one.
144
+ type: 'raw',
145
+
146
+ db_writer_host: pgHost,
147
+ db_writer_port: 5432,
148
+ db_writer_name: 'alvera_production',
149
+ db_writer_schema: 'tenant_acme',
150
+ db_writer_auth_method: 'password',
151
+ db_writer_user: pgUser,
152
+ db_writer_pass: pgPass,
153
+ db_writer_enable_ssl: true,
154
+
155
+ db_reader_host: pgReaderHost,
156
+ db_reader_port: 5432,
157
+ db_reader_name: 'alvera_production',
158
+ db_reader_schema: 'tenant_acme',
159
+ db_reader_auth_method: 'password',
160
+ db_reader_user: pgReaderUser,
161
+ db_reader_pass: pgReaderPass,
162
+ db_reader_enable_ssl: true,
163
+
164
+ // One bucket. Scope multiple lakes onto a shared bucket with
165
+ // `base_path` rather than provisioning a bucket per lake.
166
+ cloud_storage: {
180
167
  cloud_storage_type: 'aws',
181
168
  region: 'us-east-1',
182
- bucket: `acme-${dataDomain}-unregulated`,
183
- access_key_id: awsAccessKeyId,
184
- secret_access_key: awsSecretAccessKey,
185
- },
186
- regulated_cloud_storage: {
187
- cloud_storage_type: 'aws',
188
- region: 'us-east-1',
189
- bucket: `acme-${dataDomain}-regulated`,
169
+ bucket: 'acme-platform',
170
+ base_path: 'production',
190
171
  access_key_id: awsAccessKeyId,
191
172
  secret_access_key: awsSecretAccessKey,
192
173
  },
@@ -226,16 +207,15 @@ embed:
226
207
  are NOT silently server-cleared — the platform treats their
227
208
  presence as a configuration conflict.
228
209
 
229
- The two embeds (`unregulated_cloud_storage` and
230
- `regulated_cloud_storage`) are validated independently, so a body
231
- can pair `iam_role` on one side with `access_key` on the other.
210
+ There is a single embed, so this conditional is evaluated once
211
+ per body.
232
212
 
233
213
  ### Database identifiers follow strict format rules
234
214
 
235
- The four `*_db_*_schema` and four `*_db_*_name` fields must be
215
+ The two `db_*_schema` and two `db_*_name` fields must be
236
216
  lowercase identifiers (letters, digits, underscore) — uppercase
237
- letters and dashes are rejected with a `/<role>_schema` or
238
- `/<role>_name` pointer. The eight `*_db_*_host` fields must match
217
+ letters and dashes are rejected with a `/db_<role>_schema` or
218
+ `/db_<role>_name` pointer. Both `db_*_host` fields must match
239
219
  a hostname pattern; values like `"invalid_host!"` (with
240
220
  non-hostname punctuation) are rejected.
241
221
 
@@ -245,14 +225,15 @@ malformed identifier never gets as far as a connection attempt.
245
225
 
246
226
  ### Schema names must be unique within a database
247
227
 
248
- `*_db_*_schema` fields name the PostgreSQL schema the platform
228
+ `db_*_schema` fields name the PostgreSQL schema the platform
249
229
  creates inside the target database. Sibling datalakes sharing a
250
- `*_db_*_name` (same physical database) **must** use distinct
230
+ `db_*_name` (same physical database) **must** use distinct
251
231
  schema values. Collision produces a 422 at create time.
252
232
 
253
- The reader and writer for the same role typically share a schema
254
- (they're two connection profiles into the same logical schema);
255
- unregulated and regulated must not (different trust boundaries).
233
+ The reader and writer normally share a schema — they are two
234
+ connection profiles into the same logical schema. A derived
235
+ (tokenized or redacted) lake is a separate datalake row and so
236
+ brings its own schema; it must not collide with its raw parent's.
256
237
 
257
238
  ### The discriminator can't change on `update()`
258
239
 
@@ -261,13 +242,180 @@ existing row's discriminator. To change the embed type (e.g. `aws`
261
242
  → `r2`), `delete()` and `create()` fresh — you can't change body
262
243
  type on `update()`.
263
244
 
264
- ### `data_domain` is a closed enum
245
+ ## 2b. The three lakes, and which one to read
246
+
247
+ A datalake is not one database. It is three, and the other two are
248
+ **derived**: fully managed, read-only, and written only by
249
+ replication from the raw lake.
250
+
251
+ ```
252
+ raw every column, every real value. the superset.
253
+ │
254
+ ├── tokenized columns declared privacy_requirement: "tokenize"
255
+ │ come back as a token; everything else passes through.
256
+ │
257
+ └── redacted tokenize AND redact columns are scrubbed.
258
+ ```
259
+
260
+ The tier is chosen per call, not per datalake — `mode` on
261
+ `executeSql`, `dataAccessMode` on a dataset search. Nothing about
262
+ the datalake decides it.
263
+
264
+ ### What decides what a reader sees
265
+
266
+ **`privacy_requirement` on the column.** That single declaration is
267
+ the whole masking policy — `tokenize` for PII, `redact` for free
268
+ text, `none` to pass through. There is no second schema to write
269
+ sensitive values into, and no changeset quietly masking on the way
270
+ past: the raw value goes in, and the tier the reader asked for
271
+ decides what comes out. Dropping or changing that declaration
272
+ silently changes who can see what.
265
273
 
266
- Industry values are enumerated in the OpenAPI spec
267
- (`DatalakeDataDomain`); the SDK exports them as a const-style
268
- value (see `type_naming.md` "Enum types"). The enum drifts as
269
- new industries land — consult the SDK's `DatalakeDataDomain`
270
- const for the live set rather than hardcoding values.
274
+ ### Which tier to read
275
+
276
+ **Read `raw` unless you have a reason not to.** It is the superset:
277
+ every column is present with its real value, so any assertion about
278
+ what landed — a row exists, a foreign key was stamped, a count is
279
+ right — is answerable there. It also does not depend on replication
280
+ having caught up, which matters more than it sounds like: a derived
281
+ lake is populated asynchronously, so a read against one is a question
282
+ about *replication* as much as about your data. Every read-back in
283
+ the cookbook corpus reads `raw` for exactly this reason.
284
+
285
+ **Read `tokenized` when the point IS the masking** — proving a
286
+ `tokenize` column really is masked for a consumer that should not
287
+ see it, or reproducing what an AI agent sees. An agent's
288
+ `data_access` is a separate, stricter thing: it takes
289
+ `'raw' | 'tokenized'` only, because an agent is never pointed at the
290
+ redacted lake.
291
+
292
+ **Read `redacted` when you are checking the strictest projection.**
293
+
294
+ ### Two traps
295
+
296
+ **A derived lake is not there until it is provisioned.** A datalake
297
+ is created as the raw tier alone; the tokenized and redacted copies
298
+ are provisioned separately and migrate on their own jobs. A read
299
+ against a tier that does not exist yet fails with a missing-relation
300
+ error, not an empty result.
301
+
302
+ **Filter and assert on a `none` column.** In `tokenized` a
303
+ `tokenize` column returns its token, so a `WHERE` on one matches
304
+ nothing and an equality assertion on one compares a token to a
305
+ plaintext value. Pick a `none` column — a business key, a status —
306
+ for anything you filter or assert on.
307
+
308
+ ### Has the copy actually arrived?
309
+
310
+ Because a derived lake is populated asynchronously, "the copy
311
+ exists" and "the copy is complete" are different questions.
312
+ `derivedLakeTables` answers the second, per table.
313
+
314
+ ```typescript
315
+ const { data: tables } = await api.datalakes.derivedLakeTables(
316
+ tenantSlug,
317
+ datalakeSlug, // the PRIMARY lake — see the addressing note below
318
+ 'tokenized', // which copy: 'tokenized' | 'redacted'
319
+ )
320
+
321
+ for (const t of tables.data) {
322
+ // `caught_up` is the answer. `label` is the same state in words.
323
+ console.log(t.table, t.label, t.caught_up, t.resync_empties)
324
+ }
325
+
326
+ const behind = tables.data.filter((t) => !t.caught_up)
327
+ ```
328
+
329
+ **You address a derived lake by naming its PRIMARY lake plus a
330
+ mode.** A derived lake does have a slug of its own, but that slug is
331
+ not an address: the lookup resolves primaries only and answers `404`
332
+ for a derived slug deliberately. Passing the tokenized lake's own
333
+ slug is not a working alternative spelling.
334
+
335
+ **Read `caught_up`, never the state's name.** There are five states,
336
+ and one of them reads as finished without being finished:
337
+
338
+ | `state` | `caught_up` | what it means |
339
+ |---|---|---|
340
+ | `waiting_to_start` | `false` | queued |
341
+ | `copying` | `false` | first pass running |
342
+ | `finished_copy` | **`false`** | first pass done — changes made *during* it are NOT applied yet |
343
+ | `catching_up` | `false` | applying those changes |
344
+ | `caught_up` | **`true`** | arrived |
345
+
346
+ `finished_copy` is how a lake gets called ready while rows are still
347
+ missing. Only `caught_up` means arrived, and the boolean is there so
348
+ you never have to pattern-match the word.
349
+
350
+ `label` is that same state in words, identical to what the
351
+ platform's own Tokenization screen shows. Use it verbatim rather
352
+ than inventing a second vocabulary — a table reading "Copying" in
353
+ one surface must not read something else in another.
354
+
355
+ `resync_empties` says what repairing *that* table would empty, so a
356
+ surface can state the cost before offering the action. It is `null`
357
+ for tables already caught up — the platform does not spend a query
358
+ per table answering a question about tables nobody will repair — and
359
+ an array otherwise, with the named table first. **`null` and `[]`
360
+ are different answers**, and `[]` never occurs, since a table is
361
+ always in its own set. Type it as optional-nullable, not as an empty
362
+ array default.
363
+
364
+ ### Repairing one table
365
+
366
+ When a table is stuck or wrong, `resyncTable` empties it in the copy
367
+ and copies it again from the start.
368
+
369
+ ```typescript
370
+ const { data: repair } = await api.datalakes.resyncTable(
371
+ tenantSlug, datalakeSlug, 'tokenized',
372
+ { table: 'legal_entities' },
373
+ )
374
+
375
+ repair.emptied // EVERY table this emptied — often more than the one asked for
376
+ repair.status // always 'copying'
377
+
378
+ // 202 = started. There is no completion callback, so poll.
379
+ let done = false
380
+ while (!done) {
381
+ const { data: after } = await api.datalakes.derivedLakeTables(
382
+ tenantSlug, datalakeSlug, 'tokenized',
383
+ )
384
+ done = after.data.some((t) => t.table === repair.table && t.caught_up)
385
+ }
386
+ ```
387
+
388
+ **It can empty more than the table you named.** Postgres refuses to
389
+ truncate any table with an inbound foreign key — whether or not the
390
+ referencing tables hold rows — so repairing a parent empties its
391
+ whole dependent closure, transitively, in one statement. All of them
392
+ are re-copied by the same run. On real data one request has reached
393
+ nine tables. `emptied` is returned rather than assumed for exactly
394
+ this reason: a caller that emptied nine tables should be told nine.
395
+ Read `resync_empties` first if you are going to ask someone to
396
+ confirm.
397
+
398
+ **`202` means started, not done.** The repair empties the table and
399
+ a fresh first pass begins; it finishes later. Poll
400
+ `derivedLakeTables` until that table reads `caught_up`.
401
+
402
+ **It does not roll back.** A failure part-way leaves the table empty
403
+ and says so. That is deliberate — the rows to put back are the ones
404
+ that broke the copy in the first place. The recovery is to call it
405
+ again once the cause is addressed; the endpoint handles the case
406
+ where a previous attempt left the table off the publication.
407
+
408
+ **The table name is gated.** `table` must be one the copy carries or
409
+ one the original can send. Anything else is a `422` and **nothing is
410
+ emptied** — that gate is what stands between a typo and an emptied
411
+ table. Read `derivedLakeTables` for the list rather than
412
+ constructing names.
413
+
414
+ The `422` bodies here are sentences written for a person, not codes
415
+ — a lake with no such copy, a name that is neither carried nor
416
+ sendable, a dependent that could be emptied but not copied back
417
+ (ending `Nothing has been emptied.`). Surface them; do not parse
418
+ them.
271
419
 
272
420
  ## 3. Field ownership
273
421
 
@@ -285,14 +433,16 @@ resources (see §5).
285
433
  `DatalakeRequestWritable` and `DatalakeResponse`:
286
434
 
287
435
  ```
288
- identification name, description, data_domain
436
+ identification name, description
437
+ tiering type, primary_datalake_id
289
438
  global settings timezone, pool_size
290
- db profiles (×4) <role>_{host, port, name, schema, user,
291
- auth_method, enable_ssl}
292
- cloud storage (×2) <role>_cloud_storage.{cloud_storage_type,
293
- bucket, region,
294
- + branch-specific
295
- fields per §1}
439
+ db profiles (×2) db_{writer,reader}_{host, port, name,
440
+ schema, user, auth_method,
441
+ enable_ssl}
442
+ cloud storage cloud_storage.{cloud_storage_type, bucket,
443
+ region, base_path,
444
+ + branch-specific fields
445
+ per §1}
296
446
  ```
297
447
 
298
448
  **Write-only (Request-only).** Present on
@@ -301,15 +451,11 @@ consumer must re-supply these on every PUT from their secret
301
451
  store (see §5 Update for the canonical pattern):
302
452
 
303
453
  ```
304
- db credentials (×4) unregulated_db_writer_pass
305
- unregulated_db_reader_pass
306
- regulated_data_db_writer_pass
307
- regulated_data_db_reader_pass
308
-
309
- cloud storage unregulated_cloud_storage.access_key_id
310
- credentials (×4) unregulated_cloud_storage.secret_access_key
311
- regulated_cloud_storage.access_key_id
312
- regulated_cloud_storage.secret_access_key
454
+ db credentials (×2) db_writer_pass
455
+ db_reader_pass
456
+
457
+ cloud storage cloud_storage.access_key_id
458
+ credentials (×2) cloud_storage.secret_access_key
313
459
  ```
314
460
 
315
461
  The `slug` returned in the Response is the canonical reference
@@ -332,13 +478,14 @@ Datalake-specific common rejections:
332
478
  | `source.pointer` | Typical cause |
333
479
  |-------------------------------------------------|------------------------------------------------------|
334
480
  | `/name` | name collision within tenant (uniqueness) |
335
- | `/data_domain` | value not in the live enum |
336
- | `/unregulated_cloud_storage/cloud_storage_type` | discriminator value doesn't match any oneOf branch |
337
- | `/unregulated_db_writer_schema` | schema-name collision within the target DB; or identifier format violation (uppercase / dash) |
338
- | `/unregulated_db_writer_name` | identifier format violation |
339
- | `/unregulated_db_writer_host` | hostname format violation, OR synchronous probe could not reach the host |
340
- | `/unregulated_cloud_storage` | synchronous probe could not reach the bucket with the supplied credentials |
341
- | `/unregulated_cloud_storage/access_key_id` | supplied under `auth_method: "iam_role"` (must be omitted) |
481
+ | `/type` | not one of `raw` / `tokenized` / `redacted` |
482
+ | `/primary_datalake_id` | missing on a derived lake, or supplied on a raw one |
483
+ | `/cloud_storage/cloud_storage_type` | discriminator value doesn't match any oneOf branch |
484
+ | `/db_writer_schema` | schema-name collision within the target DB; or identifier format violation (uppercase / dash) |
485
+ | `/db_writer_name` | identifier format violation |
486
+ | `/db_writer_host` | hostname format violation, OR synchronous probe could not reach the host |
487
+ | `/cloud_storage` | synchronous probe could not reach the bucket with the supplied credentials |
488
+ | `/cloud_storage/access_key_id` | supplied under `auth_method: "iam_role"` (must be omitted) |
342
489
  | `/pool_size` | out of the inclusive range 1–50 |
343
490
  | `/timezone` | not in the closed 8-zone US whitelist (valid IANA values like `UTC` are also rejected — see gotcha 11) |
344
491
 
@@ -470,17 +617,10 @@ const next = {
470
617
  ...current,
471
618
 
472
619
  // Re-supply every write-only field from your secret store:
473
- unregulated_db_writer_pass: pgPass,
474
- unregulated_db_reader_pass: pgReaderPass,
475
- regulated_data_db_writer_pass: pgPass,
476
- regulated_data_db_reader_pass: pgReaderPass,
477
- unregulated_cloud_storage: {
478
- ...current.unregulated_cloud_storage,
479
- access_key_id: awsAccessKeyId,
480
- secret_access_key: awsSecretAccessKey,
481
- },
482
- regulated_cloud_storage: {
483
- ...current.regulated_cloud_storage,
620
+ db_writer_pass: pgPass,
621
+ db_reader_pass: pgReaderPass,
622
+ cloud_storage: {
623
+ ...current.cloud_storage,
484
624
  access_key_id: awsAccessKeyId,
485
625
  secret_access_key: awsSecretAccessKey,
486
626
  },
@@ -492,7 +632,7 @@ const next = {
492
632
  const { data: updated } = await api.datalakes.update(
493
633
  tenantSlug,
494
634
  datalakeId,
495
- next,
635
+ next as unknown as DatalakeRequestWritable, // see the note below
496
636
  )
497
637
  // updated.id === datalakeId — path-key stable
498
638
  // updated.description === 'Now with revised description'
@@ -500,12 +640,12 @@ const { data: updated } = await api.datalakes.update(
500
640
 
501
641
  The canonical pattern: caller re-supplies the write-only DB +
502
642
  cloud-storage fields verbatim, mutates only what changes (e.g.
503
- `description`). The eight write-only fields are enumerated in
643
+ `description`). The four write-only fields are enumerated in
504
644
  §3 Field ownership.
505
645
 
506
646
  **`update()` fires server-side connectivity probes before persisting.**
507
647
  Submitting a body produces a successful 200 only after two
508
- probes pass for each role profile:
648
+ probes pass:
509
649
 
510
650
  - **Cloud-storage probe** — verifies the bucket exists + the
511
651
  supplied `access_key_id` / `secret_access_key` (or IAM role)
@@ -517,8 +657,8 @@ probes pass for each role profile:
517
657
  A body that's structurally valid but carries stale or wrong
518
658
  credentials therefore fails at `update()` time, not at later
519
659
  write time. The 422 envelope (see `errors.md`) names the failing
520
- probe via `source.pointer` (e.g. `/unregulated_cloud_storage` or
521
- `/regulated_data_db_writer_host`). This makes drift-then-update
660
+ probe via `source.pointer` (e.g. `/cloud_storage` or
661
+ `/db_writer_host`). This makes drift-then-update
522
662
  workflows fail fast and locally rather than mid-pipeline.
523
663
 
524
664
  Most field changes do NOT re-trigger migration. Schema-name or
@@ -548,7 +688,7 @@ returning a different shape:
548
688
  | `.get(tenantSlug, id)` | One row: `DatalakeResponse` — UUID only (slug 422s) |
549
689
  | `.metadata(tenantSlug, slug)` | **Markdown string** describing the lake |
550
690
  | `.indexMetadata(tenantSlug)` | **Markdown string** listing every lake |
551
- | `.systemDatasets(tenantSlug, slug)` | `{ datasets: string[] }` — industry-built dataset names |
691
+ | `.systemDatasets(tenantSlug, slug)` | `{ datasets: string[] }` — platform dataset names |
552
692
 
553
693
  The two `metadata` methods return `string`, not a structured
554
694
  object — they're agent-facing summaries (`data: string`, not `data:
@@ -559,14 +699,14 @@ tenant` and lists every datalake by slug — useful for agent-side
559
699
  discovery without paging through `.list`.
560
700
 
561
701
  `.systemDatasets(...)` enumerates the dataset names the platform
562
- auto-creates for the datalake's `data_domain`. Each industry's
563
- catalog mixes **domain-anchored names** (the primary subject of
564
- the industry plus related entities) with **cross-cutting concepts**
565
- (audit logs,
566
- messages, documents, etc.) that the platform composes per
567
- industry — agents iterating the catalog should not hardcode a
568
- closed set of expected names; discover them at runtime via
569
- this call. The array is sorted, unique, and **excludes
702
+ auto-creates for the datalake. GH-859 removed the per-industry
703
+ catalogs: every datalake now gets the same platform-defined set
704
+ (`legal_entity`, `message`, `action_log`, `document`,
705
+ `beneficial_owner`), and the subject-specific tables an industry
706
+ used to bring are modelled as generic tables instead. Still
707
+ discover the catalog at runtime rather than hardcoding it —
708
+ the set is server-owned and this call is the only authority on
709
+ it. The array is sorted, unique, and **excludes
570
710
  operator-defined generic tables** — those live under
571
711
  `api.genericTables.list(...)`. The
572
712
  literal placeholder `"generic_table"` is also absent. Each name
@@ -575,7 +715,7 @@ can be drilled into via
575
715
  agent-facing markdown, or the whole-domain catalog can be fetched
576
716
  in one shot via `api.datasets.metadata(tenantSlug, datalakeSlug)`. Pair `.systemDatasets()` with a
577
717
  `page_size: 1` probe per name as a post-migration smoke (proves
578
- each industry-built table is queryable; the post-ready probe
718
+ each platform-built table is queryable; the post-ready probe
579
719
  code is in §8 Gotcha 3).
580
720
 
581
721
  ```typescript
@@ -583,7 +723,7 @@ const { data: catalog } = await api.datalakes.systemDatasets(
583
723
  tenantSlug,
584
724
  datalakeSlug,
585
725
  )
586
- // catalog.datasets: sorted, unique array of industry-built dataset names
726
+ // catalog.datasets: sorted, unique array of platform dataset names
587
727
 
588
728
  for (const name of catalog.datasets) {
589
729
  const { data: md } = await api.datasets.metadataDetails(tenantSlug, datalakeSlug, name)
@@ -602,8 +742,8 @@ the SQL between the two calls.
602
742
  | `.textToSql(tenantSlug, datalakeSlug, body)` | `{ sql, model, provider, explanation }` |
603
743
  | `.executeSql(tenantSlug, datalakeSlug, body, options?)` | `{ data, meta }` — or a CSV **string** with `{ format: 'csv' }` |
604
744
 
605
- - **`textToSql`** body is `{ prompt, mode }` (`mode` ∈ `'regulated' | 'unregulated'`,
606
- selecting which schema to target). It runs ordered multi-provider LLM failover and
745
+ - **`textToSql`** body is `{ prompt, mode }` (`mode` ∈ `'raw' | 'tokenized' |
746
+ 'redacted'`, selecting which lake tier to target — see §2b). It runs ordered multi-provider LLM failover and
607
747
  returns the generated `sql`, the winning `model` + `provider`, and a best-effort
608
748
  `explanation` (`null` when the explainer is unavailable). **Only the prompt + the
609
749
  schema cross the LLM boundary — never datalake rows**, so it is safe in both modes.
@@ -612,22 +752,40 @@ the SQL between the two calls.
612
752
  live — e.g. `WHERE some_boolean = 1` 422s on Postgres (booleans compare to
613
753
  `true`/`false`, never `1`). Surface the SQL to a human (or a checking step)
614
754
  and expect to edit before `executeSql`.
755
+ - **`executeSql` returns rows as POSITIONAL ARRAYS, not objects keyed by
756
+ column.** `data` is `unknown[][]`; the column names are in `meta.columns`, in
757
+ the same order. So `row.legal_entity_id` is silently `undefined` — not an
758
+ error, just absent — and a probe written against it fails with a message
759
+ built from its own bug while the data is perfectly correct. Zip them:
760
+
761
+ ```ts
762
+ const { data: res } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
763
+ sql, mode: 'raw',
764
+ })
765
+ const cols = res.meta.columns
766
+ const rows = res.data.map((r) => Object.fromEntries(cols.map((c, i) => [c, r[i]])))
767
+ // now rows[0].legal_entity_id is real
768
+ ```
769
+
770
+ Worth knowing before you write a richer check than the scaffolds ship: their
771
+ seeded example only asserts `data.length > 0`, which behaves identically for
772
+ both shapes, so nothing in the example reveals this.
615
773
  - **`executeSql`** body is `{ sql, mode, page?, page_size? }`. It runs the SQL
616
- **read-only** (INSERT/UPDATE/DELETE/DDL are rejected) on the mode-appropriate
617
- schema, with Flop-inspired pagination (`page_size` capped server-side). The JSON
774
+ **read-only** (INSERT/UPDATE/DELETE/DDL are rejected) against the
775
+ mode-selected lake tier, with Flop-inspired pagination (`page_size` capped server-side). The JSON
618
776
  shape is `{ data, meta }`; pass `options = { format: 'csv' }` to get the page as a
619
777
  CSV string attachment instead.
620
778
 
621
779
  ```typescript
622
780
  const { data: gen } = await api.datalakes.textToSql(tenantSlug, datalakeSlug, {
623
- prompt: 'count contacts created this month',
624
- mode: 'unregulated',
781
+ prompt: 'count legal entities created this month',
782
+ mode: 'raw',
625
783
  })
626
784
  // gen.sql — review / edit before running
627
785
 
628
786
  const { data: page } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
629
787
  sql: gen.sql,
630
- mode: 'unregulated',
788
+ mode: 'raw',
631
789
  page: 1,
632
790
  page_size: 100,
633
791
  })
@@ -636,7 +794,7 @@ const { data: page } = await api.datalakes.executeSql(tenantSlug, datalakeSlug,
636
794
 
637
795
  const { data: csv } = await api.datalakes.executeSql(
638
796
  tenantSlug, datalakeSlug,
639
- { sql: gen.sql, mode: 'unregulated' },
797
+ { sql: gen.sql, mode: 'raw' },
640
798
  { format: 'csv' },
641
799
  ) // csv is a string, not the JSON envelope
642
800
  ```
@@ -646,15 +804,26 @@ const { data: csv } = await api.datalakes.executeSql(
646
804
  The name you author is not the name you query. Datasets are addressed
647
805
  **logically** everywhere in a manifest or SDK call, and **physically** in SQL:
648
806
 
649
- | You author | You query in SQL | Schema |
650
- |---|---|---|
651
- | `message` | `regulated_messages` | `app_regulated` |
652
- | `action_log` | `action_logs` | `app_unregulated` |
653
- | a generic table `clients` | `regulated_alvera_custom_clients` | `app_regulated` |
807
+ For a **generic table**, the physical name is the `name` the server
808
+ returns from `genericTables.create` / `.get` — query it directly, with
809
+ no prefix and no schema qualifier:
810
+
811
+ ```typescript
812
+ const { data: table } = await api.genericTables.create(tenantSlug, datalakeSlug, body)
813
+ await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
814
+ sql: `SELECT * FROM ${table.name} LIMIT 1`,
815
+ mode: 'raw',
816
+ })
817
+ ```
818
+
819
+ Do not reconstruct that name from the slug you authored — the server
820
+ derives it, and the derivation is not part of the contract.
654
821
 
655
- Custom (generic) tables carry an `alvera_custom_` prefix, and the regulated side
656
- adds a further `regulated_` prefix on top. The schema follows `execute-sql`'s
657
- `--mode` (`regulated` → `app_regulated`, `unregulated` → `app_unregulated`).
822
+ For the **platform datasets** (`message`, `action_log`, and the rest of
823
+ §5's `systemDatasets` catalog) the physical name and schema are likewise
824
+ server-owned. Read them, don't assume them: the physical mapping moved
825
+ with the GH-859 tier rename and the old `app_regulated` /
826
+ `app_unregulated` schema pair no longer follows from `mode`.
658
827
 
659
828
  **Don't derive these by hand — read them.** Guessing column and table names is
660
829
  how six turns get spent on `column "error_message" does not exist` (the real
@@ -695,10 +864,10 @@ rows downstream reveal the mismatch. Check the derivation, not the metadata.
695
864
  const { data: created } = await api.datalakes.create(tenantSlug, body)
696
865
  // created.status === 'new'
697
866
 
698
- await api.datalakes.migrate(tenantSlug, created.slug) // migrate is slug-keyed
867
+ await api.datalakes.migrate(tenantSlug, created.slug!) // migrate is slug-keyed
699
868
 
700
869
  const ready = await waitUntilReady(
701
- () => api.datalakes.get(tenantSlug, created.id), // get is UUID-id-keyed
870
+ () => api.datalakes.get(tenantSlug, created.id!), // get is UUID-id-keyed
702
871
  (s) => s === 'ready' ? 'ready' : 'pending',
703
872
  { intervalMs: 15_000, timeoutMs: 5 * 60_000 },
704
873
  )
@@ -722,7 +891,7 @@ rows downstream reveal the mismatch. Check the derivation, not the metadata.
722
891
  const { data: catalog } = await api.datalakes.systemDatasets(
723
892
  tenantSlug, datalakeSlug,
724
893
  )
725
- // catalog.datasets is the industry-built name list —
894
+ // catalog.datasets is the platform dataset name list —
726
895
  // discovered at runtime; don't hardcode names.
727
896
 
728
897
  for (const dataset of catalog.datasets) {
@@ -755,11 +924,10 @@ rows downstream reveal the mismatch. Check the derivation, not the metadata.
755
924
  a required string. Supply a sentinel literal (e.g.
756
925
  `"unused-iam-role"`); the platform clears it server-side.
757
926
 
758
- 7. **`cloud_storage_type` lives inside each embed, not at the
759
- top level.** A single body has two discriminators
760
- (`unregulated_cloud_storage.cloud_storage_type` and
761
- `regulated_cloud_storage.cloud_storage_type`); they're
762
- independently validated and can carry different values.
927
+ 7. **`cloud_storage_type` lives inside the embed, not at the
928
+ top level.** The discriminator is
929
+ `cloud_storage.cloud_storage_type` — one per body, since a
930
+ datalake carries one bucket.
763
931
 
764
932
  8. **`delete()` is rare in practice.** Datalakes accrete dependents
765
933
  quickly; routine cleanup is via operator-driven console
@@ -768,15 +936,14 @@ rows downstream reveal the mismatch. Check the derivation, not the metadata.
768
936
 
769
937
  9. **Create and update run synchronous reachability probes.** Before
770
938
  the row is written, the platform opens connections to every
771
- declared endpoint — all four DB profiles (unregulated reader +
772
- writer, regulated reader + writer) AND both cloud-storage
773
- buckets — and only proceeds if every probe succeeds. A
774
- structurally valid body can still fail with field-level
775
- rejections on `*_db_*_host` (unreachable, wrong credentials, or
776
- SSL-mismatch) or `*_cloud_storage` (bucket inaccessible). The
777
- same probe runs on update; an update that only changes the
778
- datalake's `name` still re-probes if the body re-supplies any
779
- `*_host` or `*_cloud_storage` field.
939
+ declared endpoint — both DB profiles (reader and writer) AND the
940
+ cloud-storage bucket — and only proceeds if every probe
941
+ succeeds. A structurally valid body can still fail with
942
+ field-level rejections on `db_*_host` (unreachable, wrong
943
+ credentials, or SSL-mismatch) or `cloud_storage` (bucket
944
+ inaccessible). The same probe runs on update; an update that
945
+ only changes the datalake's `name` still re-probes if the body
946
+ re-supplies any `db_*_host` or `cloud_storage` field.
780
947
 
781
948
  Two consumer-relevant consequences:
782
949