@alvera-ai/platform-sdk 0.17.0 → 0.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agent/AGENTS.md +82 -144
- package/.agent/account_management.md +2 -2
- package/.agent/action_logs.md +4 -4
- package/.agent/ai_agents.md +28 -21
- package/.agent/ai_sandbox.md +49 -39
- package/.agent/connected_apps.md +3 -3
- package/.agent/cookbook/_fixtures/README.md +1 -1
- package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_generic_table.liquid +1 -1
- package/.agent/cookbook/_fixtures/organic-marketing/_lead_submissions_foundation_legal_entity.liquid +80 -0
- package/.agent/cookbook/_fixtures/{foundation → organic-marketing}/_lead_submissions_foundation_mdm.liquid +2 -1
- package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_generic_table.liquid +57 -0
- package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_legal_entity.liquid +30 -0
- package/.agent/cookbook/_fixtures/payments-compliance/_compliance_screenings_mdm.liquid +44 -0
- package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_generic_table.liquid +57 -0
- package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_legal_entity.liquid +36 -0
- package/.agent/cookbook/_fixtures/payments-compliance/_payment_accounts_mdm.liquid +41 -0
- package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_generic_table.liquid +70 -0
- package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_legal_entity.liquid +52 -0
- package/.agent/cookbook/_fixtures/primary-care-feedback/_cahps_appointments_mdm.liquid +42 -0
- package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_generic_table.liquid +38 -0
- package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_legal_entity.liquid +48 -0
- package/.agent/cookbook/_fixtures/subscription-saas/_customers_subscription_mdm.liquid +49 -0
- package/.agent/cookbook/organic-marketing.md +2801 -0
- package/.agent/cookbook/payments-compliance.md +2180 -0
- package/.agent/cookbook/primary-care.md +2175 -0
- package/.agent/cookbook/subscription-saas.md +2403 -0
- package/.agent/data_activation_clients.md +65 -52
- package/.agent/datalakes.md +338 -171
- package/.agent/errors.md +3 -3
- package/.agent/generic_tables.md +151 -62
- package/.agent/interoperability_contracts.md +57 -22
- package/.agent/mdm.md +136 -153
- package/.agent/messages.md +36 -34
- package/.agent/mock-services.md +1 -1
- package/.agent/mutations.md +2 -2
- package/.agent/templates.md +14 -13
- package/.agent/tools.md +63 -21
- package/.agent/type_naming.md +13 -13
- package/.agent/workflows.md +99 -53
- package/README.md +2 -2
- package/dist/bin/platform-sdk.mjs +33 -47
- package/dist/bin/platform-sdk.mjs.map +1 -1
- package/dist/index.d.mts +565 -379
- package/dist/index.d.mts.map +1 -1
- package/dist/index.mjs +494 -59
- package/dist/index.mjs.map +1 -1
- package/package.json +4 -3
- package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_legal_entity.liquid +0 -88
- package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_appointment.liquid +0 -47
- package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_mdm.liquid +0 -24
- package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_patient.liquid +0 -38
- package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_compliance_screening.liquid +0 -59
- package/.agent/cookbook/_fixtures/payments/_compliance_screenings_payments_mdm.liquid +0 -36
- package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_mdm.liquid +0 -30
- package/.agent/cookbook/_fixtures/payments/_payment_accounts_payments_payment_account.liquid +0 -55
- package/.agent/cookbook/_fixtures/subscription/_customers_subscription_mdm.liquid +0 -20
- package/.agent/cookbook/_setup/foundation.md +0 -359
- package/.agent/cookbook/_setup/healthcare.md +0 -361
- package/.agent/cookbook/_setup/payments.md +0 -365
- package/.agent/cookbook/_setup/subscription.md +0 -364
- package/.agent/cookbook/action-status-updaters.md +0 -278
- package/.agent/cookbook/ai-agent-invoke.md +0 -279
- package/.agent/cookbook/appointment-review-sms-workflow.md +0 -801
- package/.agent/cookbook/birthday-greeting-sms-trigger.md +0 -696
- package/.agent/cookbook/bulk-ingest.md +0 -302
- package/.agent/cookbook/contact-us-triage-with-llm.md +0 -663
- package/.agent/cookbook/dunning-sms-for-delinquent.md +0 -659
- package/.agent/cookbook/generic-tables.md +0 -244
- package/.agent/cookbook/invite-team.md +0 -200
- package/.agent/cookbook/kyc-notification-on-account-activation.md +0 -661
- package/.agent/cookbook/marketing-campaign-send.md +0 -1044
- package/.agent/cookbook/paginated-restapi-poller.md +0 -383
- package/.agent/cookbook/rest-fetch.md +0 -273
- package/.agent/cookbook/sanctions-screening-with-agent-review.md +0 -773
- package/.agent/cookbook/score-leads-with-llm-categorization.md +0 -665
- package/.agent/cookbook/system-templates.md +0 -165
- package/.agent/cookbook/talk-to-data.md +0 -178
- package/.agent/cookbook/triage-prospects-by-priority.md +0 -571
- package/.agent/cookbook/welcome-sms-for-customers.md +0 -647
- /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/memorandum-of-association-01.png +0 -0
- /package/.agent/cookbook/_fixtures/{healthcare → primary-care-feedback}/sample_two_page.pdf +0 -0
- /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/_customers_subscription_customer.liquid +0 -0
- /package/.agent/cookbook/_fixtures/{subscription → subscription-saas}/stripe_customers_batch1.csv +0 -0
|
@@ -0,0 +1,2403 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: "Subscription SaaS: one customer spine, three workflows, three ways in"
|
|
3
|
+
summary: "The whole subscription-billing surface in one walk. Stand up a tenant and a raw datalake, declare one customers table, and run three workflows over it — a welcome SMS gated on reachability, a dunning reminder gated on reachability and KYC, and an LLM agent that bands accounts by priority and fans out one SMS per band. Then fill the same table three more ways — a bulk CSV upload, a pull-based REST fetch, and finally a teammate invited onto the tenant."
|
|
4
|
+
use_case: subscription-saas
|
|
5
|
+
slug: subscription-saas
|
|
6
|
+
vitest_source:
|
|
7
|
+
- integration-tests/tests/subscription-saas/standard-workflow.test.ts
|
|
8
|
+
- integration-tests/tests/subscription-saas/dunning-sms-workflow.test.ts
|
|
9
|
+
- integration-tests/tests/subscription-saas/agent-lead-triage.test.ts
|
|
10
|
+
- integration-tests/tests/subscription-saas/run-dac-bulk.test.ts
|
|
11
|
+
- integration-tests/tests/subscription-saas/run-dac-fetch.test.ts
|
|
12
|
+
- integration-tests/tests/subscription-saas/invite-team.test.ts
|
|
13
|
+
- integration-tests/tests/workspace/bootstrap.test.ts
|
|
14
|
+
status: draft
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
# Problem
|
|
18
|
+
|
|
19
|
+
A subscription business talks to its customers on a schedule set by the
|
|
20
|
+
billing system. A contract is signed and a welcome message goes out with a
|
|
21
|
+
link to the self-serve portal. An invoice goes unpaid and a reminder goes
|
|
22
|
+
out with a link to pay. An account needs chasing and *how* it is chased
|
|
23
|
+
depends on what kind of account it is — a white-glove note to an enterprise,
|
|
24
|
+
a standard nudge to a self-serve individual.
|
|
25
|
+
|
|
26
|
+
Three jobs, one audience. The classic implementation is three services that
|
|
27
|
+
each query a customers table and each carry their own copy of the
|
|
28
|
+
reachability rule, the KYC rule and the account-tier rule. This walk builds
|
|
29
|
+
all three on **one customers table**, because the customer is the same
|
|
30
|
+
customer in all three and the rules belong to the platform rather than to
|
|
31
|
+
three codebases.
|
|
32
|
+
|
|
33
|
+
The last three scenarios are about how rows get *in*. A row can arrive
|
|
34
|
+
one at a time through an inline ingest, four thousand at a time through a
|
|
35
|
+
file upload, or be pulled from someone else's API on demand — and all three
|
|
36
|
+
land in the same table through the same contract pair. The walk ends by
|
|
37
|
+
inviting a second person onto the tenant, because a billing team is rarely
|
|
38
|
+
one person.
|
|
39
|
+
|
|
40
|
+
**This surface stays raw.** No tokenized or redacted copy is provisioned.
|
|
41
|
+
Every read-back here is a question about this walk's own rows, and the
|
|
42
|
+
priority agent classifies on account type rather than on anyone's name. The
|
|
43
|
+
PII columns are still declared `tokenize` — that declaration is the policy,
|
|
44
|
+
and it goes live the moment someone provisions the pair. If you want a model
|
|
45
|
+
held away from identities today, that is what `payments-compliance` §007
|
|
46
|
+
shows.
|
|
47
|
+
|
|
48
|
+
# Composition
|
|
49
|
+
|
|
50
|
+
| Resource | Why it exists here |
|
|
51
|
+
|-----------------------------------|-------------------------------------------|
|
|
52
|
+
| Tenant + raw datalake | The billing team's own world |
|
|
53
|
+
| One customers generic table | The spine all six scenarios write to |
|
|
54
|
+
| SMS tool, LLM tool | Shared by every workflow below |
|
|
55
|
+
| Contract pair, on two clients | Identity and table row, deliberately split|
|
|
56
|
+
| Three workflows | Welcome, dunning, priority triage |
|
|
57
|
+
| Two connected apps | The portal link, and the pay link |
|
|
58
|
+
| Manual-upload, bulk and REST paths| Three ways into the same table |
|
|
59
|
+
| An invitation | A second person on the tenant |
|
|
60
|
+
|
|
61
|
+
# Walkthrough
|
|
62
|
+
|
|
63
|
+
This cookbook is **self-reliant**: it stands up everything it uses and
|
|
64
|
+
depends on no other file. It is also **idempotent** — every creating step
|
|
65
|
+
looks first and creates only what is missing, so running it twice costs what
|
|
66
|
+
running it once cost. That is not a nicety. Nothing in the platform reclaims
|
|
67
|
+
an abandoned datalake, and a walk that mints a fresh tenant per run leaves
|
|
68
|
+
every previous run's lakes behind forever.
|
|
69
|
+
|
|
70
|
+
## 001 — sign in as root, and make sure the admin user exists
|
|
71
|
+
|
|
72
|
+
Authenticate as the platform's root admin (`admin@dev.local` /
|
|
73
|
+
`devpassword` in local dev) via the tenantless bootstrap login — keyless by
|
|
74
|
+
structural necessity, since no tenant exists yet to scope a key to, and a
|
|
75
|
+
dev/test-only surface.
|
|
76
|
+
|
|
77
|
+
Then make sure the admin user this walk runs as exists. **The email is
|
|
78
|
+
stable, not per-run**, which is what makes the step idempotent — and it means
|
|
79
|
+
the second run finds the user already signed up. A duplicate signup is
|
|
80
|
+
refused, so the refusal is caught and inspected: if it says the email is
|
|
81
|
+
taken, that is the idempotent path and the walk continues. Any other failure
|
|
82
|
+
is re-raised, because swallowing it would turn a real auth problem into a
|
|
83
|
+
confusing failure three steps later.
|
|
84
|
+
|
|
85
|
+
```typescript
|
|
86
|
+
ctx.rootSession = await createBootstrapSession({
|
|
87
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
88
|
+
email: process.env.ALVERA_ROOT_EMAIL!,
|
|
89
|
+
password: process.env.ALVERA_ROOT_PASSWORD!,
|
|
90
|
+
})
|
|
91
|
+
ctx.rootApi = createIsolatedPlatformApi({
|
|
92
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
93
|
+
sessionToken: ctx.rootSession.sessionToken,
|
|
94
|
+
apiKey: '',
|
|
95
|
+
})
|
|
96
|
+
|
|
97
|
+
ctx.billingEmail = 'cookbook-subscription-saas@dev.local'
|
|
98
|
+
ctx.billingPassword = 'CookbookPass1!'
|
|
99
|
+
|
|
100
|
+
try {
|
|
101
|
+
const signUpResp = await ctx.rootApi.admin.signUp({
|
|
102
|
+
email: ctx.billingEmail,
|
|
103
|
+
password: ctx.billingPassword,
|
|
104
|
+
first_name: 'Cookbook',
|
|
105
|
+
last_name: 'Billing',
|
|
106
|
+
})
|
|
107
|
+
await ctx.rootApi.admin.confirmUser(signUpResp.data.id!)
|
|
108
|
+
} catch (err) {
|
|
109
|
+
// Already provisioned by a previous run. Confirm that is what happened
|
|
110
|
+
// rather than assuming it — a genuine signup failure must not be read as
|
|
111
|
+
// "already there".
|
|
112
|
+
const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
|
|
113
|
+
if (!/taken|already|exist/i.test(detail)) throw err
|
|
114
|
+
}
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
## 002 — the find-or-create helper every later step uses
|
|
118
|
+
|
|
119
|
+
Idempotence is one question asked over and over: *is this already here?*
|
|
120
|
+
Rather than answer it a dozen different ways, the walk defines it once.
|
|
121
|
+
|
|
122
|
+
`ensure` takes a label, a lookup and a create. It runs the lookup, returns
|
|
123
|
+
what it finds, and only creates when the lookup comes back empty. The label
|
|
124
|
+
is not decoration — when a run reuses something you did not expect it to, the
|
|
125
|
+
log line naming it is how you find out.
|
|
126
|
+
|
|
127
|
+
Alongside it, `firstNamed` — the lookup half, written once. **A list endpoint
|
|
128
|
+
pages at twenty and is not newest-first**, so a `.find()` over the default
|
|
129
|
+
page silently stops finding things as soon as a lake has a few of them, and
|
|
130
|
+
the failure looks like the resource was never created. Passing
|
|
131
|
+
`page_size: 100` is the cheap fix, and this walk creates enough tools,
|
|
132
|
+
contracts and clients on one lake to need it.
|
|
133
|
+
|
|
134
|
+
Both live on `ctx` rather than as bare functions because each numbered step
|
|
135
|
+
compiles into its own `it()` block, so a plain `function` here would not be
|
|
136
|
+
in scope for the steps that call it.
|
|
137
|
+
|
|
138
|
+
```typescript
|
|
139
|
+
ctx.ensure = async <T>(
|
|
140
|
+
label: string,
|
|
141
|
+
find: () => Promise<T | undefined>,
|
|
142
|
+
create: () => Promise<T>,
|
|
143
|
+
): Promise<T> => {
|
|
144
|
+
const existing = await find()
|
|
145
|
+
if (existing !== undefined) {
|
|
146
|
+
console.log(` ↻ reusing ${label}`)
|
|
147
|
+
return existing
|
|
148
|
+
}
|
|
149
|
+
console.log(` + creating ${label}`)
|
|
150
|
+
return await create()
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
// Look one resource up by name across the WHOLE listing, not page one.
|
|
154
|
+
ctx.firstNamed = <T extends { name?: string }>(
|
|
155
|
+
rows: readonly T[] | undefined,
|
|
156
|
+
name: string,
|
|
157
|
+
): T | undefined => (rows ?? []).find((r) => r.name === name)
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
## 003 — the tenant
|
|
161
|
+
|
|
162
|
+
The billing admin signs in without a tenant scope — they may not belong to
|
|
163
|
+
one yet — and the walk finds or creates the tenant by its stable name. The
|
|
164
|
+
server derives the slug; capture it, because every later call is addressed by
|
|
165
|
+
it.
|
|
166
|
+
|
|
167
|
+
```typescript
|
|
168
|
+
ctx.billingTenantlessSession = await createBootstrapSession({
|
|
169
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
170
|
+
email: ctx.billingEmail,
|
|
171
|
+
password: ctx.billingPassword,
|
|
172
|
+
})
|
|
173
|
+
ctx.billingTenantlessApi = createIsolatedPlatformApi({
|
|
174
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
175
|
+
sessionToken: ctx.billingTenantlessSession.sessionToken,
|
|
176
|
+
apiKey: '',
|
|
177
|
+
})
|
|
178
|
+
|
|
179
|
+
const TENANT_NAME = 'Cookbook Subscription SaaS'
|
|
180
|
+
|
|
181
|
+
const tenant = await ctx.ensure(
|
|
182
|
+
`tenant ${TENANT_NAME}`,
|
|
183
|
+
async () => {
|
|
184
|
+
const { data } = await ctx.billingTenantlessApi.tenants.list()
|
|
185
|
+
return (data.data ?? []).find((t: { name?: string }) => t.name === TENANT_NAME)
|
|
186
|
+
},
|
|
187
|
+
async () => {
|
|
188
|
+
const { data } = await ctx.billingTenantlessApi.tenants.create({ name: TENANT_NAME })
|
|
189
|
+
return data
|
|
190
|
+
},
|
|
191
|
+
)
|
|
192
|
+
tenantSlug = tenant.slug!
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
## 004 — the tenant-scoped client
|
|
196
|
+
|
|
197
|
+
A tenant-scoped login requires `X-API-Key`, so a key has to exist before the
|
|
198
|
+
admin can sign in against the tenant. Mint one through the platform-admin
|
|
199
|
+
side door with the root bearer; in the web console this is *Settings → API
|
|
200
|
+
Keys*.
|
|
201
|
+
|
|
202
|
+
`data_access_mode: 'raw'` is the ceiling, and on this surface it is the only
|
|
203
|
+
tier there is — no derived lake is provisioned, so raw is what every read
|
|
204
|
+
answers from.
|
|
205
|
+
|
|
206
|
+
**This is the one step in the walk that is not idempotent**, and it is worth
|
|
207
|
+
knowing why rather than discovering it. There is no endpoint that lists a
|
|
208
|
+
tenant's API keys, so there is nothing to look the existing one up with —
|
|
209
|
+
`ensure` has no lookup to run. Each run therefore mints another key. A key
|
|
210
|
+
row is cheap where a datalake is not, so the walk accepts it; if you are
|
|
211
|
+
counting rows in a shared environment, this is the one to count.
|
|
212
|
+
|
|
213
|
+
```typescript
|
|
214
|
+
const { data: mintedKey } = await ctx.rootApi.admin.createTenantApiKey(tenantSlug, {
|
|
215
|
+
name: 'Cookbook Subscription SaaS Key',
|
|
216
|
+
data_access_mode: 'raw',
|
|
217
|
+
})
|
|
218
|
+
ctx.tenantApiKey = mintedKey.api_key
|
|
219
|
+
|
|
220
|
+
const billingTenantSession = await createSession({
|
|
221
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
222
|
+
email: ctx.billingEmail,
|
|
223
|
+
password: ctx.billingPassword,
|
|
224
|
+
tenantSlug,
|
|
225
|
+
apiKey: ctx.tenantApiKey,
|
|
226
|
+
})
|
|
227
|
+
api = createIsolatedPlatformApi({
|
|
228
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
229
|
+
sessionToken: billingTenantSession.sessionToken,
|
|
230
|
+
apiKey: ctx.tenantApiKey,
|
|
231
|
+
})
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
## 005 — the datalake
|
|
235
|
+
|
|
236
|
+
One datalake, created `raw`, and on this surface that is the whole story —
|
|
237
|
+
there is no derived pair to provision and nothing to wait for beyond the
|
|
238
|
+
migration in §006.
|
|
239
|
+
|
|
240
|
+
The lookup still filters on `type === 'raw'`. That costs nothing here and
|
|
241
|
+
keeps the step correct if anyone ever turns tokenization on: from that moment
|
|
242
|
+
`datalakes.list` returns three rows, two of which are copies that must never
|
|
243
|
+
be mistaken for the primary.
|
|
244
|
+
|
|
245
|
+
Local-dev defaults match the seeded `dev.exs` setup — `postgres` on
|
|
246
|
+
`localhost:5432`, database `alvera_dev_foundation`, LocalStack S3 on
|
|
247
|
+
`localhost:4566` — and the schema name is stable, because a fresh schema per
|
|
248
|
+
run is the same leak as a fresh tenant per run.
|
|
249
|
+
|
|
250
|
+
```typescript
|
|
251
|
+
const DB = { host: 'localhost', port: 5432, user: 'postgres', pass: 'postgres', name: 'alvera_dev_foundation' }
|
|
252
|
+
const DB_SCHEMA = 'cookbook_subscription_saas'
|
|
253
|
+
const LAKE_NAME = 'Cookbook Subscription SaaS Datalake'
|
|
254
|
+
|
|
255
|
+
const S3 = {
|
|
256
|
+
cloud_storage_type: 'aws' as const,
|
|
257
|
+
region: 'us-east-1',
|
|
258
|
+
access_key_id: 'test',
|
|
259
|
+
secret_access_key: 'test',
|
|
260
|
+
endpoint: 'http://localhost:4566',
|
|
261
|
+
}
|
|
262
|
+
|
|
263
|
+
const datalake = await ctx.ensure(
|
|
264
|
+
`datalake ${LAKE_NAME}`,
|
|
265
|
+
async () => {
|
|
266
|
+
const { data } = await api.datalakes.list(tenantSlug)
|
|
267
|
+
return (data.data ?? []).find(
|
|
268
|
+
(l: { name?: string; type?: string }) => l.name === LAKE_NAME && l.type === 'raw',
|
|
269
|
+
)
|
|
270
|
+
},
|
|
271
|
+
async () => {
|
|
272
|
+
const { data } = await api.datalakes.create(tenantSlug, {
|
|
273
|
+
name: LAKE_NAME,
|
|
274
|
+
description: 'Subscription SaaS datalake provisioned by the cookbook doctest.',
|
|
275
|
+
timezone: 'America/New_York',
|
|
276
|
+
pool_size: 3,
|
|
277
|
+
type: 'raw',
|
|
278
|
+
|
|
279
|
+
db_writer_host: DB.host,
|
|
280
|
+
db_writer_port: DB.port,
|
|
281
|
+
db_writer_name: DB.name,
|
|
282
|
+
db_writer_schema: DB_SCHEMA,
|
|
283
|
+
db_writer_auth_method: 'password',
|
|
284
|
+
db_writer_user: DB.user,
|
|
285
|
+
db_writer_pass: DB.pass,
|
|
286
|
+
db_writer_enable_ssl: false,
|
|
287
|
+
db_reader_host: DB.host,
|
|
288
|
+
db_reader_port: DB.port,
|
|
289
|
+
db_reader_name: DB.name,
|
|
290
|
+
db_reader_schema: DB_SCHEMA,
|
|
291
|
+
db_reader_auth_method: 'password',
|
|
292
|
+
db_reader_user: DB.user,
|
|
293
|
+
db_reader_pass: DB.pass,
|
|
294
|
+
db_reader_enable_ssl: false,
|
|
295
|
+
|
|
296
|
+
cloud_storage: { ...S3, bucket: 'alvera-platform-dev', base_path: 'cookbook/subscription-saas' },
|
|
297
|
+
})
|
|
298
|
+
return data
|
|
299
|
+
},
|
|
300
|
+
)
|
|
301
|
+
datalakeSlug = datalake.slug!
|
|
302
|
+
ctx.datalakeId = datalake.id!
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
## 006 — run the migrations, and wait for ready
|
|
306
|
+
|
|
307
|
+
`datalakes.create` persists the row at `status: 'new'`; it does not run the
|
|
308
|
+
schema DDL. Migration is triggered separately so the operator decides when
|
|
309
|
+
the potentially-slow part happens. `datalakes.migrate` enqueues the job and
|
|
310
|
+
returns immediately with `status: 'enqueued'`; the poll after it is what
|
|
311
|
+
waits for the worker.
|
|
312
|
+
|
|
313
|
+
Migrating is safe to repeat, which is what lets this step stay unguarded on a
|
|
314
|
+
second run. The wait exits the moment the lake reports `ready`, so the
|
|
315
|
+
five-minute ceiling is only ever paid in failure.
|
|
316
|
+
|
|
317
|
+
```typescript
|
|
318
|
+
const migrateResp = await api.datalakes.migrate(tenantSlug, datalakeSlug)
|
|
319
|
+
if (migrateResp.data.status !== 'enqueued') {
|
|
320
|
+
throw new Error(`datalake migration not enqueued (status: ${migrateResp.data.status})`)
|
|
321
|
+
}
|
|
322
|
+
|
|
323
|
+
const READY_TIMEOUT_MS = 5 * 60_000
|
|
324
|
+
const readyDeadline = Date.now() + READY_TIMEOUT_MS
|
|
325
|
+
let datalakeStatus: string | undefined
|
|
326
|
+
while (Date.now() < readyDeadline) {
|
|
327
|
+
const { data } = await api.datalakes.get(tenantSlug, ctx.datalakeId)
|
|
328
|
+
datalakeStatus = data.status
|
|
329
|
+
if (datalakeStatus === 'ready') break
|
|
330
|
+
await new Promise((r) => setTimeout(r, 5_000))
|
|
331
|
+
}
|
|
332
|
+
if (datalakeStatus !== 'ready') {
|
|
333
|
+
throw new Error(`datalake did not reach :ready within ${READY_TIMEOUT_MS}ms (last: ${datalakeStatus})`)
|
|
334
|
+
}
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
## 007 — three helpers the scenarios below share
|
|
338
|
+
|
|
339
|
+
**`waitForFiredRun`.** `workflows.run` only *schedules* a run. It returns
|
|
340
|
+
immediately with a `workflow_run_id`, and the `workflow_run_log_id` and
|
|
341
|
+
`batch_id` a scenario needs are written later, when the run actually fires.
|
|
342
|
+
Two traps live in that gap:
|
|
343
|
+
|
|
344
|
+
- **Poll until `workflow_run_log_id` is a string — not until `status` leaves
|
|
345
|
+
`'scheduled'`.** Those are different moments; the run reaches `processing`
|
|
346
|
+
first and writes the log id a beat later. A predicate on status alone
|
|
347
|
+
releases you to read a `null`, and because `typeof null === 'object'` the
|
|
348
|
+
symptom is a baffling *"expected string, got object"* rather than an
|
|
349
|
+
obvious nil.
|
|
350
|
+
- **Raise on `failed` carrying `failure_reason`** rather than polling to the
|
|
351
|
+
deadline. A scenario blocked on a run that will never fire should say why
|
|
352
|
+
on the first read, not thirty seconds later behind a generic timeout.
|
|
353
|
+
|
|
354
|
+
**`waitForBatches`.** Ingestion is async: `ingest` returns a `batch_id` the
|
|
355
|
+
instant the rows are accepted, and the per-row jobs drain afterwards. This
|
|
356
|
+
polls a client's activation logs until every named batch has actually
|
|
357
|
+
written.
|
|
358
|
+
|
|
359
|
+
**Gate on `dataset_updated`, never on `rows_ingested`.** They answer
|
|
360
|
+
different questions — received versus written — and a batch whose row is
|
|
361
|
+
refused reports `rows_ingested: 1, dataset_updated: 0, status: 'partial'`
|
|
362
|
+
with the reason in `error`. A gate on `rows_ingested` calls that green, the
|
|
363
|
+
next step then runs against a table with nothing in it, and the failure
|
|
364
|
+
surfaces three steps later as *"expected 2 execution logs, got 0"* — which
|
|
365
|
+
reads like a broken workflow rather than a row that never landed.
|
|
366
|
+
|
|
367
|
+
**`deployGenericTable`.** Creating a generic table does not deploy it; it
|
|
368
|
+
rests at `status: 'new'` until a migration runs. The wait that means anything
|
|
369
|
+
asks for the table the same way the next step is about to, and retries until
|
|
370
|
+
it stops erroring — polling the table's own `status: 'deployed'` goes true
|
|
371
|
+
earlier and proves less.
|
|
372
|
+
|
|
373
|
+
```typescript
|
|
374
|
+
ctx.waitForFiredRun = async (
|
|
375
|
+
runDatalakeSlug: string,
|
|
376
|
+
runId: string,
|
|
377
|
+
timeoutMs = 120_000,
|
|
378
|
+
): Promise<{ workflowRunLogId: string; batchId: string | null }> => {
|
|
379
|
+
const deadline = Date.now() + timeoutMs
|
|
380
|
+
let lastStatus: string | undefined
|
|
381
|
+
while (Date.now() < deadline) {
|
|
382
|
+
const { data } = await api.workflowRuns.get(tenantSlug, runDatalakeSlug, runId)
|
|
383
|
+
lastStatus = data.status
|
|
384
|
+
if (data.status === 'failed') {
|
|
385
|
+
throw new Error(`workflow run ${runId} failed: ${data.failure_reason ?? 'no failure_reason given'}`)
|
|
386
|
+
}
|
|
387
|
+
if (typeof data.workflow_run_log_id === 'string') {
|
|
388
|
+
return { workflowRunLogId: data.workflow_run_log_id, batchId: data.batch_id ?? null }
|
|
389
|
+
}
|
|
390
|
+
await new Promise((r) => setTimeout(r, 1_000))
|
|
391
|
+
}
|
|
392
|
+
throw new Error(`workflow run ${runId} never fired within ${timeoutMs}ms (last status: ${lastStatus})`)
|
|
393
|
+
}
|
|
394
|
+
|
|
395
|
+
ctx.waitForBatches = async (
|
|
396
|
+
dacSlug: string,
|
|
397
|
+
batchIds: readonly string[],
|
|
398
|
+
timeoutMs = 90_000,
|
|
399
|
+
): Promise<void> => {
|
|
400
|
+
const targets = new Set(batchIds)
|
|
401
|
+
const deadline = Date.now() + timeoutMs
|
|
402
|
+
let greenCount = 0
|
|
403
|
+
while (Date.now() < deadline) {
|
|
404
|
+
const { data } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
|
|
405
|
+
const green = new Set<string>()
|
|
406
|
+
for (const row of (data.data ?? []) as Array<Record<string, unknown>>) {
|
|
407
|
+
const b = row.batch_id
|
|
408
|
+
if (typeof b !== 'string' || !targets.has(b)) continue
|
|
409
|
+
if (row.status === 'partial' || row.status === 'failed') {
|
|
410
|
+
throw new Error(
|
|
411
|
+
`batch ${b} on ${dacSlug} did not persist its rows (status: ${row.status}): ` +
|
|
412
|
+
`${String(row.error ?? 'no reason given')}`,
|
|
413
|
+
)
|
|
414
|
+
}
|
|
415
|
+
if (typeof row.dataset_updated !== 'number' || row.dataset_updated < 1) continue
|
|
416
|
+
const files = row.output_files
|
|
417
|
+
if (!Array.isArray(files) || files.length === 0) continue
|
|
418
|
+
green.add(b)
|
|
419
|
+
}
|
|
420
|
+
greenCount = green.size
|
|
421
|
+
if (greenCount === targets.size) return
|
|
422
|
+
await new Promise((r) => setTimeout(r, 1_000))
|
|
423
|
+
}
|
|
424
|
+
throw new Error(`only ${greenCount}/${targets.size} batches on ${dacSlug} persisted within ${timeoutMs}ms`)
|
|
425
|
+
}
|
|
426
|
+
|
|
427
|
+
ctx.deployGenericTable = async (tableName: string, timeoutMs = 120_000): Promise<void> => {
|
|
428
|
+
await api.datalakes.migrate(tenantSlug, datalakeSlug)
|
|
429
|
+
const deadline = Date.now() + timeoutMs
|
|
430
|
+
let lastError: unknown
|
|
431
|
+
while (Date.now() < deadline) {
|
|
432
|
+
try {
|
|
433
|
+
await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
|
|
434
|
+
sql: `SELECT 1 FROM ${tableName} LIMIT 1`,
|
|
435
|
+
mode: 'raw',
|
|
436
|
+
})
|
|
437
|
+
return
|
|
438
|
+
} catch (err) {
|
|
439
|
+
lastError = err
|
|
440
|
+
await new Promise((r) => setTimeout(r, 1_000))
|
|
441
|
+
}
|
|
442
|
+
}
|
|
443
|
+
throw new Error(
|
|
444
|
+
`generic table ${tableName} was not readable within ${timeoutMs}ms ` +
|
|
445
|
+
`(last error: ${lastError instanceof Error ? lastError.message : String(lastError)})`,
|
|
446
|
+
)
|
|
447
|
+
}
|
|
448
|
+
```
|
|
449
|
+
|
|
450
|
+
## 008 — the SMS tool
|
|
451
|
+
|
|
452
|
+
One SMS tool dispatches every outbound message in this walk — a welcome, a
|
|
453
|
+
dunning reminder and three priority bands. **The per-message body lives on
|
|
454
|
+
the workflow action, never on the tool**, which is what lets one tool serve
|
|
455
|
+
all of them. `intent: 'sms'` tags it for workflow-action use.
|
|
456
|
+
|
|
457
|
+
Local dev points at LocalStack's SNS on `:4566`, so nothing leaves the
|
|
458
|
+
machine.
|
|
459
|
+
|
|
460
|
+
```typescript
|
|
461
|
+
const smsTool = await ctx.ensure(
|
|
462
|
+
'SMS tool',
|
|
463
|
+
async () => {
|
|
464
|
+
const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
465
|
+
return ctx.firstNamed(data.data, 'Cookbook Billing SMS Tool')
|
|
466
|
+
},
|
|
467
|
+
async () => {
|
|
468
|
+
const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
|
|
469
|
+
name: 'Cookbook Billing SMS Tool',
|
|
470
|
+
description: 'SNS-backed SMS dispatcher for every billing scenario, wired to LocalStack.',
|
|
471
|
+
intent: 'sms',
|
|
472
|
+
status: 'active',
|
|
473
|
+
datalake_id: ctx.datalakeId,
|
|
474
|
+
body: {
|
|
475
|
+
tool_body_type: 'sns',
|
|
476
|
+
auth_method: 'access_key',
|
|
477
|
+
region: 'us-east-1',
|
|
478
|
+
phone_number: '+15551234567',
|
|
479
|
+
endpoint_url: 'http://localhost:4566',
|
|
480
|
+
access_key_id: 'test',
|
|
481
|
+
secret_access_key: 'test',
|
|
482
|
+
},
|
|
483
|
+
})
|
|
484
|
+
return data
|
|
485
|
+
},
|
|
486
|
+
)
|
|
487
|
+
toolId = smsTool.id!
|
|
488
|
+
ctx.smsToolId = smsTool.id!
|
|
489
|
+
```
|
|
490
|
+
|
|
491
|
+
## 009 — the customers table, which all three workflows share
|
|
492
|
+
|
|
493
|
+
A customer is not one of the datasets the platform ships. The platform models
|
|
494
|
+
exactly one identity — the legal entity — and everything else a build needs
|
|
495
|
+
is a generic table you declare, plus the legal entity the rows resolve to.
|
|
496
|
+
There is no `customer` dataset to reach for.
|
|
497
|
+
|
|
498
|
+
**One table serves all three workflows**, and the column list is the union of
|
|
499
|
+
what they gate on rather than three near-identical tables. `phone` is the
|
|
500
|
+
welcome workflow's reachability gate. `tax_id` is the dunning workflow's KYC
|
|
501
|
+
gate. `customer_type` is what the priority agent bands on. `status` is what
|
|
502
|
+
makes a customer newly contracted.
|
|
503
|
+
|
|
504
|
+
The PII columns are declared `tokenize` and the rest `none`. **That
|
|
505
|
+
declaration is the masking policy, and it is correct whether or not a
|
|
506
|
+
tokenized lake exists today** — this walk provisions none, so nothing masks
|
|
507
|
+
here, and the declaration goes live the moment someone calls
|
|
508
|
+
`provisionTokenization`. Declaring it now is how the policy survives that
|
|
509
|
+
call rather than being remembered afterwards.
|
|
510
|
+
|
|
511
|
+
Do **not** declare a `legal_entity_id` column. The platform stamps it from
|
|
512
|
+
the MDM resolve in §012, and the name is reserved — declaring it does not
|
|
513
|
+
shadow the stamp, it 422s the create, so you lose the table until you drop it.
|
|
514
|
+
|
|
515
|
+
```typescript
|
|
516
|
+
const customers = await ctx.ensure(
|
|
517
|
+
'customers table',
|
|
518
|
+
async () => {
|
|
519
|
+
const { data } = await api.genericTables.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
520
|
+
return (data.data ?? []).find((t: { title?: string }) => t.title === 'Cookbook Customers')
|
|
521
|
+
},
|
|
522
|
+
async () => {
|
|
523
|
+
const { data } = await api.genericTables.create(tenantSlug, datalakeSlug, {
|
|
524
|
+
title: 'Cookbook Customers',
|
|
525
|
+
description: 'Stripe customers the welcome, dunning and triage workflows all run on.',
|
|
526
|
+
columns: [
|
|
527
|
+
{ name: 'customer_number', title: 'Customer Number', type: 'string', description: 'Stripe customer id — unique per customer', is_unique: true, privacy_requirement: 'none' },
|
|
528
|
+
{ name: 'customer_type', title: 'Customer Type', type: 'string', description: 'individual / enterprise — the legal-entity branch, and what the priority agent bands on', is_unique: false, privacy_requirement: 'none' },
|
|
529
|
+
{ name: 'status', title: 'Status', type: 'string', description: 'Commercial state — contracted is what earns a welcome', is_unique: false, privacy_requirement: 'none' },
|
|
530
|
+
{ name: 'name', title: 'Name', type: 'string', description: 'Customer display name', is_unique: false, privacy_requirement: 'tokenize' },
|
|
531
|
+
{ name: 'email', title: 'Email', type: 'string', description: 'Billing email', is_unique: false, privacy_requirement: 'tokenize' },
|
|
532
|
+
{ name: 'phone', title: 'Phone', type: 'string', description: 'Billing phone — the SMS destination and the reachability gate', is_unique: false, privacy_requirement: 'tokenize' },
|
|
533
|
+
{ name: 'tax_id', title: 'Tax ID', type: 'string', description: 'Tax identifier — its presence is the KYC gate', is_unique: false, privacy_requirement: 'tokenize' },
|
|
534
|
+
{ name: 'currency', title: 'Currency', type: 'string', description: 'Settlement currency', is_unique: false, privacy_requirement: 'none' },
|
|
535
|
+
{ name: 'delinquent', title: 'Delinquent', type: 'boolean', description: 'Whether the latest invoice is past due', is_unique: false, privacy_requirement: 'none' },
|
|
536
|
+
{ name: 'address_city', title: 'Address City', type: 'string', description: 'Billing city', is_unique: false, privacy_requirement: 'none' },
|
|
537
|
+
{ name: 'address_country', title: 'Address Country', type: 'string', description: 'Billing country', is_unique: false, privacy_requirement: 'none' },
|
|
538
|
+
],
|
|
539
|
+
})
|
|
540
|
+
return data
|
|
541
|
+
},
|
|
542
|
+
)
|
|
543
|
+
genericTableId = customers.id!
|
|
544
|
+
ctx.customersTableId = customers.id!
|
|
545
|
+
ctx.customersTableName = customers.name!
|
|
546
|
+
|
|
547
|
+
// Creating the table records it; it is not deployed until a migration runs.
|
|
548
|
+
await ctx.deployGenericTable(ctx.customersTableName)
|
|
549
|
+
```
|
|
550
|
+
|
|
551
|
+
## 010 — the Stripe data source and the manual-upload tool
|
|
552
|
+
|
|
553
|
+
A workflow runs on rows, and rows arrive through the data-activation chain.
|
|
554
|
+
The chain's first link is a **data source** — a registration of where the
|
|
555
|
+
rows originate. Its `uri` is the system-of-record address, and it matters
|
|
556
|
+
more than it looks: the client injects it into every ingested row as
|
|
557
|
+
`source_uri`, and both contracts below render it onto the identifier. It is
|
|
558
|
+
half of the key that collapses two contracts onto one legal entity.
|
|
559
|
+
|
|
560
|
+
The **manual-upload tool** is the minimal data-exchange tool. It needs no
|
|
561
|
+
endpoint and no credentials, because the rows arrive in the ingest call's own
|
|
562
|
+
body rather than being fetched. `intent: 'data_exchange'` is what
|
|
563
|
+
distinguishes it from the SMS tool above.
|
|
564
|
+
|
|
565
|
+
```typescript
|
|
566
|
+
const dataSource = await ctx.ensure(
|
|
567
|
+
'Stripe data source',
|
|
568
|
+
async () => {
|
|
569
|
+
const { data } = await api.dataSources.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
570
|
+
return ctx.firstNamed(data.data, 'Cookbook Stripe Source')
|
|
571
|
+
},
|
|
572
|
+
async () => {
|
|
573
|
+
const { data } = await api.dataSources.create(tenantSlug, datalakeSlug, {
|
|
574
|
+
name: 'Cookbook Stripe Source',
|
|
575
|
+
uri: 'stripe.example.com/cookbook-subscription-saas',
|
|
576
|
+
description: 'Stripe billing system — origin of every customer row in this walk.',
|
|
577
|
+
status: 'active',
|
|
578
|
+
is_default: false,
|
|
579
|
+
})
|
|
580
|
+
return data
|
|
581
|
+
},
|
|
582
|
+
)
|
|
583
|
+
dataSourceId = dataSource.id!
|
|
584
|
+
ctx.dataSourceId = dataSource.id!
|
|
585
|
+
|
|
586
|
+
const manualUploadTool = await ctx.ensure(
|
|
587
|
+
'manual-upload tool',
|
|
588
|
+
async () => {
|
|
589
|
+
const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
590
|
+
return ctx.firstNamed(data.data, 'Cookbook Manual Upload Tool')
|
|
591
|
+
},
|
|
592
|
+
async () => {
|
|
593
|
+
const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
|
|
594
|
+
name: 'Cookbook Manual Upload Tool',
|
|
595
|
+
description: 'Manual-upload data-exchange tool — backs the inline and bulk ingest paths.',
|
|
596
|
+
intent: 'data_exchange',
|
|
597
|
+
status: 'active',
|
|
598
|
+
datalake_id: ctx.datalakeId,
|
|
599
|
+
data_source_id: dataSourceId,
|
|
600
|
+
body: { tool_body_type: 'manual_upload' },
|
|
601
|
+
})
|
|
602
|
+
return data
|
|
603
|
+
},
|
|
604
|
+
)
|
|
605
|
+
ctx.manualUploadToolId = manualUploadTool.id!
|
|
606
|
+
```
|
|
607
|
+
|
|
608
|
+
## 011 — the contract pair
|
|
609
|
+
|
|
610
|
+
One inbound Stripe row becomes two things, so it takes two contracts.
|
|
611
|
+
|
|
612
|
+
**Contract A** (`resource_type: 'legal_entity'`) writes the identity — the
|
|
613
|
+
name, the type, the identification triple. Its `mdm_input_config` is
|
|
614
|
+
`{ type: 'null' }`, because a template that already *is* the subject has
|
|
615
|
+
nothing to resolve.
|
|
616
|
+
|
|
617
|
+
**Contract B** (`resource_type: 'generic_table'`, pinned to §009's table)
|
|
618
|
+
writes the table row, and its `mdm_input_config` emits the same
|
|
619
|
+
`(uri, customer_number)` pair. That is what collapses both onto one legal
|
|
620
|
+
entity per customer, and what earns the row the `legal_entity_id` stamp every
|
|
621
|
+
connected-app link is minted against.
|
|
622
|
+
|
|
623
|
+
**The identifier shape is where this goes wrong, and there are three of them
|
|
624
|
+
on this platform.** `Platform.MDMInput` — what a `mdm_input_config` renders —
|
|
625
|
+
takes `uri` / `type` / `value`. A legal-entity identification, which is
|
|
626
|
+
Contract A's own body, takes `id_type` / `id_number` / `uri`. And
|
|
627
|
+
`POST /mdm/verify` takes `system` / `id_type` / `value`. The overlap is
|
|
628
|
+
uneven: `system` is a genuine alias on the MDM side — it is cast and
|
|
629
|
+
normalised onto `uri` — but `id_type` and `id_number` are not cast at all.
|
|
630
|
+
Unknown keys are dropped rather than refused, so borrowing the legal-entity
|
|
631
|
+
names for the MDM template resolves nothing and fails as a bare
|
|
632
|
+
`mdm_dispatcher_error` naming no field.
|
|
633
|
+
|
|
634
|
+
And the reason the two look interchangeable is that **one becomes the
|
|
635
|
+
other**. `MDMInput.to_identification/1` takes `(uri, type, value)` and
|
|
636
|
+
returns `(uri, id_type, id_number)` — the rename happens inside the platform,
|
|
637
|
+
on the way through. So the identifier reads back under names it will not
|
|
638
|
+
accept on the way in, and that asymmetry is invisible from the stored shape
|
|
639
|
+
alone.
|
|
640
|
+
|
|
641
|
+
The templates are loaded from the vendored fixtures rather than inlined; they
|
|
642
|
+
are long, and a reader who wants them can open the file.
|
|
643
|
+
|
|
644
|
+
```typescript
|
|
645
|
+
const { readFileSync } = await import('node:fs')
|
|
646
|
+
const { join } = await import('node:path')
|
|
647
|
+
const read = (f: string) =>
|
|
648
|
+
readFileSync(join(process.env.COOKBOOK_FIXTURES_DIR!, 'subscription-saas', f), 'utf8')
|
|
649
|
+
|
|
650
|
+
const leContract = await ctx.ensure(
|
|
651
|
+
'legal-entity contract',
|
|
652
|
+
async () => {
|
|
653
|
+
const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
654
|
+
return ctx.firstNamed(data.data, 'Cookbook Customer LE Contract')
|
|
655
|
+
},
|
|
656
|
+
async () => {
|
|
657
|
+
const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
|
|
658
|
+
name: 'Cookbook Customer LE Contract',
|
|
659
|
+
description: 'Stripe customer → LegalEntity (individual or business).',
|
|
660
|
+
resource_type: 'legal_entity',
|
|
661
|
+
type: 'identity',
|
|
662
|
+
generic_table_id: null,
|
|
663
|
+
template_config: { type: 'custom', body: read('_customers_subscription_legal_entity.liquid') },
|
|
664
|
+
mdm_input_config: { type: 'null' },
|
|
665
|
+
})
|
|
666
|
+
return data
|
|
667
|
+
},
|
|
668
|
+
)
|
|
669
|
+
ctx.leContractId = leContract.id!
|
|
670
|
+
|
|
671
|
+
const gtContract = await ctx.ensure(
|
|
672
|
+
'generic-table contract',
|
|
673
|
+
async () => {
|
|
674
|
+
const { data } = await api.interoperabilityContracts.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
675
|
+
return ctx.firstNamed(data.data, 'Cookbook Customer GT Contract')
|
|
676
|
+
},
|
|
677
|
+
async () => {
|
|
678
|
+
const { data } = await api.interoperabilityContracts.create(tenantSlug, datalakeSlug, {
|
|
679
|
+
name: 'Cookbook Customer GT Contract',
|
|
680
|
+
description: 'Stripe customer → customers table row, stamped with its subject.',
|
|
681
|
+
resource_type: 'generic_table',
|
|
682
|
+
type: 'identity',
|
|
683
|
+
generic_table_id: ctx.customersTableId,
|
|
684
|
+
template_config: { type: 'custom', body: read('_customers_subscription_generic_table.liquid') },
|
|
685
|
+
mdm_input_config: { type: 'custom', body: read('_customers_subscription_mdm.liquid') },
|
|
686
|
+
})
|
|
687
|
+
return data
|
|
688
|
+
},
|
|
689
|
+
)
|
|
690
|
+
interopContractId = gtContract.id!
|
|
691
|
+
ctx.gtContractId = gtContract.id!
|
|
692
|
+
```
|
|
693
|
+
|
|
694
|
+
## 012 — two clients, because the pair must not share one
|
|
695
|
+
|
|
696
|
+
Here is the trap, and it is the most expensive one in this file.
|
|
697
|
+
|
|
698
|
+
The obvious move is to bind both contracts to a single client and ingest
|
|
699
|
+
once. **Do not.** A client fans out per `(row, contract)` at enqueue time,
|
|
700
|
+
and the job key is staggered by *row* — so a pair bound to one client runs
|
|
701
|
+
both contracts on the same row **at the same instant**. Both resolve the same
|
|
702
|
+
subject, both find-or-create it, both miss the other's uncommitted write, and
|
|
703
|
+
the unique index on `(uri, id_type, id_number)` refuses the loser. The legal
|
|
704
|
+
entity is deliberately not part of that key, which is what makes the refusal
|
|
705
|
+
correct rather than a bug: two writers, one subject, one winner.
|
|
706
|
+
|
|
707
|
+
The stagger is the platform's designed guard, and it only helps when the two
|
|
708
|
+
contracts are on different *rows*. So split them across two clients and
|
|
709
|
+
sequence the ingest: identity first, wait, then the table row. Contract B's
|
|
710
|
+
resolve then finds the subject Contract A already committed, and stamps it.
|
|
711
|
+
|
|
712
|
+
Both clients share the tool and the data source. Only the contract list
|
|
713
|
+
differs.
|
|
714
|
+
|
|
715
|
+
```typescript
|
|
716
|
+
const identityDac = await ctx.ensure(
|
|
717
|
+
'identity activation client',
|
|
718
|
+
async () => {
|
|
719
|
+
const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
720
|
+
return ctx.firstNamed(data.data, 'Cookbook Customer Identity DAC')
|
|
721
|
+
},
|
|
722
|
+
async () => {
|
|
723
|
+
const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
|
|
724
|
+
name: 'Cookbook Customer Identity DAC',
|
|
725
|
+
description: 'Writes the legal entity for each Stripe customer. Runs FIRST.',
|
|
726
|
+
tool_id: ctx.manualUploadToolId,
|
|
727
|
+
data_source_id: dataSourceId,
|
|
728
|
+
tool_call: { tool_call_type: 'manual_upload' },
|
|
729
|
+
interoperability_contract_ids: [ctx.leContractId],
|
|
730
|
+
})
|
|
731
|
+
return data
|
|
732
|
+
},
|
|
733
|
+
)
|
|
734
|
+
ctx.identityDacSlug = identityDac.slug!
|
|
735
|
+
|
|
736
|
+
const dataDac = await ctx.ensure(
|
|
737
|
+
'data activation client',
|
|
738
|
+
async () => {
|
|
739
|
+
const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
740
|
+
return ctx.firstNamed(data.data, 'Cookbook Customer Data DAC')
|
|
741
|
+
},
|
|
742
|
+
async () => {
|
|
743
|
+
const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
|
|
744
|
+
name: 'Cookbook Customer Data DAC',
|
|
745
|
+
description: 'Writes the customers table row and stamps it with the subject. Runs SECOND.',
|
|
746
|
+
tool_id: ctx.manualUploadToolId,
|
|
747
|
+
data_source_id: dataSourceId,
|
|
748
|
+
tool_call: { tool_call_type: 'manual_upload' },
|
|
749
|
+
interoperability_contract_ids: [ctx.gtContractId],
|
|
750
|
+
})
|
|
751
|
+
return data
|
|
752
|
+
},
|
|
753
|
+
)
|
|
754
|
+
dacId = dataDac.id!
|
|
755
|
+
ctx.dataDacSlug = dataDac.slug!
|
|
756
|
+
```
|
|
757
|
+
|
|
758
|
+
## 013 — four customers, identities first
|
|
759
|
+
|
|
760
|
+
Four rows, chosen so that each of the three workflows below can be scoped to
|
|
761
|
+
a pair that splits cleanly:
|
|
762
|
+
|
|
763
|
+
| customer | phone | tax_id | type | proves |
|
|
764
|
+
|-------------------|-------|--------|------------|------------------------------|
|
|
765
|
+
| `CUS-SS-REACH` | yes | yes | individual | passes all three filters |
|
|
766
|
+
| `CUS-SS-NOPHONE` | — | yes | individual | welcome filters it out |
|
|
767
|
+
| `CUS-SS-NOKYC` | yes | — | individual | dunning filters it out |
|
|
768
|
+
| `CUS-SS-ENTERPRISE`| yes | yes | enterprise | the agent bands it high |
|
|
769
|
+
|
|
770
|
+
The customer numbers are **stable, not per-run**. That is what makes the walk
|
|
771
|
+
idempotent — a second run upserts the same four rows on `customer_number`
|
|
772
|
+
rather than adding four more — and it is also what makes the outbound
|
|
773
|
+
messages idempotent, since the action keys are built from these ids.
|
|
774
|
+
|
|
775
|
+
The ingest is **staged**, per §012. Every identity row goes through the
|
|
776
|
+
identity client and is waited on to completion; only then do the table rows
|
|
777
|
+
go through the data client.
|
|
778
|
+
|
|
779
|
+
```typescript
|
|
780
|
+
ctx.reachNumber = 'CUS-SS-REACH'
|
|
781
|
+
ctx.noPhoneNumber = 'CUS-SS-NOPHONE'
|
|
782
|
+
ctx.noKycNumber = 'CUS-SS-NOKYC'
|
|
783
|
+
ctx.enterpriseNumber = 'CUS-SS-ENTERPRISE'
|
|
784
|
+
|
|
785
|
+
ctx.customerRows = [
|
|
786
|
+
{
|
|
787
|
+
customer_number: ctx.reachNumber,
|
|
788
|
+
customer_type: 'individual',
|
|
789
|
+
status: 'contracted',
|
|
790
|
+
name: 'Olivia Hartmann',
|
|
791
|
+
email: 'olivia.hartmann@example.com',
|
|
792
|
+
phone: '+12025553101',
|
|
793
|
+
tax_id: '311-22-7890',
|
|
794
|
+
currency: 'USD',
|
|
795
|
+
delinquent: 'true',
|
|
796
|
+
address_city: 'Boston',
|
|
797
|
+
address_country: 'US',
|
|
798
|
+
},
|
|
799
|
+
{
|
|
800
|
+
customer_number: ctx.noPhoneNumber,
|
|
801
|
+
customer_type: 'individual',
|
|
802
|
+
status: 'contracted',
|
|
803
|
+
name: 'Noah Pemberton',
|
|
804
|
+
email: 'noah.pemberton@example.com',
|
|
805
|
+
tax_id: '412-88-1200',
|
|
806
|
+
currency: 'USD',
|
|
807
|
+
delinquent: 'false',
|
|
808
|
+
address_city: 'Denver',
|
|
809
|
+
address_country: 'US',
|
|
810
|
+
// phone deliberately absent — the welcome filter must reject this row
|
|
811
|
+
},
|
|
812
|
+
{
|
|
813
|
+
customer_number: ctx.noKycNumber,
|
|
814
|
+
customer_type: 'individual',
|
|
815
|
+
status: 'contracted',
|
|
816
|
+
name: 'Priya Raghunathan',
|
|
817
|
+
email: 'priya.raghunathan@example.com',
|
|
818
|
+
phone: '+12025553103',
|
|
819
|
+
currency: 'USD',
|
|
820
|
+
delinquent: 'true',
|
|
821
|
+
address_city: 'Austin',
|
|
822
|
+
address_country: 'US',
|
|
823
|
+
// tax_id deliberately absent — the dunning filter must reject this row
|
|
824
|
+
},
|
|
825
|
+
{
|
|
826
|
+
customer_number: ctx.enterpriseNumber,
|
|
827
|
+
customer_type: 'enterprise',
|
|
828
|
+
status: 'contracted',
|
|
829
|
+
name: 'Pinnacle Financial Group',
|
|
830
|
+
email: 'ap@pinnaclefg.example.com',
|
|
831
|
+
phone: '+12125559100',
|
|
832
|
+
tax_id: '84-2156789',
|
|
833
|
+
currency: 'USD',
|
|
834
|
+
delinquent: 'true',
|
|
835
|
+
address_city: 'New York',
|
|
836
|
+
address_country: 'US',
|
|
837
|
+
},
|
|
838
|
+
]
|
|
839
|
+
|
|
840
|
+
// Stage one — the identities.
|
|
841
|
+
const identityIngests = await Promise.all(
|
|
842
|
+
ctx.customerRows.map((row: Record<string, unknown>) =>
|
|
843
|
+
api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.identityDacSlug, { data: row }),
|
|
844
|
+
),
|
|
845
|
+
)
|
|
846
|
+
await ctx.waitForBatches(
|
|
847
|
+
ctx.identityDacSlug,
|
|
848
|
+
identityIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!),
|
|
849
|
+
)
|
|
850
|
+
|
|
851
|
+
// Stage two — the table rows, now that every subject exists.
|
|
852
|
+
const dataIngests = await Promise.all(
|
|
853
|
+
ctx.customerRows.map((row: Record<string, unknown>) =>
|
|
854
|
+
api.dataActivationClients.ingest(tenantSlug, datalakeSlug, ctx.dataDacSlug, { data: row }),
|
|
855
|
+
),
|
|
856
|
+
)
|
|
857
|
+
ctx.customerBatchIds = dataIngests.map((r: { data: { batch_id?: string } }) => r.data.batch_id!)
|
|
858
|
+
await ctx.waitForBatches(ctx.dataDacSlug, ctx.customerBatchIds)
|
|
859
|
+
```
|
|
860
|
+
|
|
861
|
+
## 014 — the stamp is the proof that resolution ran
|
|
862
|
+
|
|
863
|
+
The activation log said the rows were written. This asks the table itself,
|
|
864
|
+
which is the independent answer, and it asks for the one column no contract
|
|
865
|
+
writes: `legal_entity_id`.
|
|
866
|
+
|
|
867
|
+
A row present with a null stamp is the failure this step exists to catch. It
|
|
868
|
+
means Contract B's template landed but its `mdm_input_config` resolved
|
|
869
|
+
nothing — the wrong identifier shape, or a subject that was not there yet —
|
|
870
|
+
and every link minted from that row afterwards would point at nobody.
|
|
871
|
+
|
|
872
|
+
```typescript
|
|
873
|
+
const stampDeadline = Date.now() + 90_000
|
|
874
|
+
let stamped: unknown[][] = []
|
|
875
|
+
while (Date.now() < stampDeadline && stamped.length < 4) {
|
|
876
|
+
const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
|
|
877
|
+
sql: `SELECT customer_number, legal_entity_id
|
|
878
|
+
FROM ${ctx.customersTableName}
|
|
879
|
+
WHERE customer_number IN (
|
|
880
|
+
'${ctx.reachNumber}', '${ctx.noPhoneNumber}',
|
|
881
|
+
'${ctx.noKycNumber}', '${ctx.enterpriseNumber}'
|
|
882
|
+
)
|
|
883
|
+
ORDER BY customer_number`,
|
|
884
|
+
mode: 'raw',
|
|
885
|
+
})
|
|
886
|
+
if (typeof result !== 'string') stamped = result.data
|
|
887
|
+
if (stamped.length < 4) await new Promise((r) => setTimeout(r, 1_000))
|
|
888
|
+
}
|
|
889
|
+
if (stamped.length !== 4) {
|
|
890
|
+
throw new Error(`expected 4 customer rows, found ${stamped.length} within 90s`)
|
|
891
|
+
}
|
|
892
|
+
const unstamped = stamped.filter((row) => !row[1]).map((row) => String(row[0]))
|
|
893
|
+
if (unstamped.length > 0) {
|
|
894
|
+
throw new Error(
|
|
895
|
+
`these rows landed with no legal_entity_id stamp — the MDM resolve did not run: ${unstamped.join(', ')}`,
|
|
896
|
+
)
|
|
897
|
+
}
|
|
898
|
+
```
|
|
899
|
+
|
|
900
|
+
## 015 — the billing-portal connected app
|
|
901
|
+
|
|
902
|
+
The welcome SMS carries a deep link to the self-serve billing portal. A
|
|
903
|
+
connected app is a thin registration of that page's URL and hosting mode; the
|
|
904
|
+
page itself lives outside the platform (`mode: 'self_hosted'`). Capture the
|
|
905
|
+
server-derived slug — resolving the link in §019 is addressed by it.
|
|
906
|
+
|
|
907
|
+
```typescript
|
|
908
|
+
const portalApp = await ctx.ensure(
|
|
909
|
+
'billing-portal connected app',
|
|
910
|
+
async () => {
|
|
911
|
+
const { data } = await api.connectedApps.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
912
|
+
return ctx.firstNamed(data.data, 'Cookbook Billing Portal')
|
|
913
|
+
},
|
|
914
|
+
async () => {
|
|
915
|
+
const { data } = await api.connectedApps.create(tenantSlug, datalakeSlug, {
|
|
916
|
+
name: 'Cookbook Billing Portal',
|
|
917
|
+
description: 'Self-serve billing portal linked from the outbound welcome SMS.',
|
|
918
|
+
mode: 'self_hosted',
|
|
919
|
+
urls: [{ url: 'https://billing.example.local', is_primary: true, label: 'production' }],
|
|
920
|
+
})
|
|
921
|
+
return data
|
|
922
|
+
},
|
|
923
|
+
)
|
|
924
|
+
connectedAppId = portalApp.id!
|
|
925
|
+
ctx.portalAppId = portalApp.id!
|
|
926
|
+
ctx.portalAppSlug = portalApp.slug!
|
|
927
|
+
```
|
|
928
|
+
|
|
929
|
+
## 016 — the Welcome SMS workflow
|
|
930
|
+
|
|
931
|
+
Standard shape — a filter, a decision, one action — with three things worth
|
|
932
|
+
reading closely.
|
|
933
|
+
|
|
934
|
+
`dataset_type: 'generic_table'` with `generic_table_id` makes §009's table
|
|
935
|
+
the event source. The row is injected under the `event_dataset` key, so the
|
|
936
|
+
filter reads `event_dataset.phone` directly.
|
|
937
|
+
|
|
938
|
+
`skip_mdm_resolution: false` is set explicitly, and must be. A generic-table
|
|
939
|
+
row is never its own subject — the subject is the legal entity §011's pair
|
|
940
|
+
stamped onto it. With resolution on, the platform loads that subject and
|
|
941
|
+
mints a page token against it, which is what makes
|
|
942
|
+
`{{ connected_app_form_url }}` a *per-customer* link. Skip it and the token
|
|
943
|
+
is never minted and the link renders empty.
|
|
944
|
+
|
|
945
|
+
The `context_datasets` entry queries the `message` dataset for a welcome
|
|
946
|
+
already sent to this legal entity in the last six months — the platform's
|
|
947
|
+
built-in guard against welcoming someone twice, expressed as a context query
|
|
948
|
+
rather than as application code. On a fresh lake it returns empty, which is
|
|
949
|
+
harmless; on the second run of this walk it is one of two guards that stop a
|
|
950
|
+
duplicate message, the other being the idempotency key on the action.
|
|
951
|
+
|
|
952
|
+
```typescript
|
|
953
|
+
const WELCOME_KEY = 'send_welcome_sms'
|
|
954
|
+
|
|
955
|
+
const welcomeWorkflow = await ctx.ensure(
|
|
956
|
+
'Welcome SMS workflow',
|
|
957
|
+
async () => {
|
|
958
|
+
const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
959
|
+
return ctx.firstNamed(data.data, 'Cookbook Welcome SMS Workflow')
|
|
960
|
+
},
|
|
961
|
+
async () => {
|
|
962
|
+
const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
|
|
963
|
+
name: 'Cookbook Welcome SMS Workflow',
|
|
964
|
+
description: 'Sends a welcome SMS to newly contracted customers with a self-serve billing link.',
|
|
965
|
+
dataset_type: 'generic_table',
|
|
966
|
+
generic_table_id: ctx.customersTableId,
|
|
967
|
+
status: 'live',
|
|
968
|
+
tags: ['lifecycle', 'welcome'],
|
|
969
|
+
skip_mdm_resolution: false,
|
|
970
|
+
filter_config: {
|
|
971
|
+
type: 'custom',
|
|
972
|
+
body:
|
|
973
|
+
'{% if event_dataset.status == "contracted" %}' +
|
|
974
|
+
'{% if event_dataset.phone and event_dataset.phone != "" %}true{% endif %}' +
|
|
975
|
+
'{% endif %}',
|
|
976
|
+
},
|
|
977
|
+
decision_config: {
|
|
978
|
+
type: 'custom',
|
|
979
|
+
body: `["${WELCOME_KEY}"]`,
|
|
980
|
+
output_schema: { type: 'array', items: { type: 'string' } },
|
|
981
|
+
},
|
|
982
|
+
context_datasets: [
|
|
983
|
+
{
|
|
984
|
+
dataset_type: 'message',
|
|
985
|
+
where_clause:
|
|
986
|
+
`m.legal_entity_id = '{{ legal_entity_id }}' AND m.decision_key = '${WELCOME_KEY}' ` +
|
|
987
|
+
"AND m.sent_at > NOW() - INTERVAL '6 months'",
|
|
988
|
+
limit: 1,
|
|
989
|
+
position: 0,
|
|
990
|
+
},
|
|
991
|
+
],
|
|
992
|
+
actions: [
|
|
993
|
+
{
|
|
994
|
+
action_type: 'sms',
|
|
995
|
+
tool_id: ctx.smsToolId,
|
|
996
|
+
decision_key: WELCOME_KEY,
|
|
997
|
+
position: 0,
|
|
998
|
+
trigger_template: 'now',
|
|
999
|
+
idempotency_template: `{{ event_dataset.customer_number }}-${WELCOME_KEY}`,
|
|
1000
|
+
connected_app_id: ctx.portalAppId,
|
|
1001
|
+
connected_app_route: '/portal/welcome',
|
|
1002
|
+
connected_app_metadata_template: '{"customer_number":"{{ event_dataset.customer_number }}"}',
|
|
1003
|
+
tool_call: {
|
|
1004
|
+
tool_call_type: 'sms_request',
|
|
1005
|
+
to: { type: 'custom', body: '{{ event_dataset.phone }}' },
|
|
1006
|
+
body: {
|
|
1007
|
+
type: 'custom',
|
|
1008
|
+
body:
|
|
1009
|
+
'Hi {{ event_dataset.name }}, welcome to Alvera Billing. ' +
|
|
1010
|
+
'Self-serve billing portal: {{ connected_app_form_url }}',
|
|
1011
|
+
},
|
|
1012
|
+
sms_type: 'transactional',
|
|
1013
|
+
},
|
|
1014
|
+
},
|
|
1015
|
+
],
|
|
1016
|
+
})
|
|
1017
|
+
return data
|
|
1018
|
+
},
|
|
1019
|
+
)
|
|
1020
|
+
workflowId = welcomeWorkflow.id!
|
|
1021
|
+
ctx.welcomeWorkflowId = welcomeWorkflow.id!
|
|
1022
|
+
ctx.welcomeWorkflowSlug = welcomeWorkflow.slug!
|
|
1023
|
+
|
|
1024
|
+
if (welcomeWorkflow.skip_mdm_resolution !== false) {
|
|
1025
|
+
throw new Error('a generic-table workflow that mints a per-customer link must keep MDM resolution ON')
|
|
1026
|
+
}
|
|
1027
|
+
```
|
|
1028
|
+
|
|
1029
|
+
## 017 — run it across the reachable and unreachable pair
|
|
1030
|
+
|
|
1031
|
+
`manual_override: false` is what makes this a real test — the filter is
|
|
1032
|
+
evaluated, so the `event_dataset.phone` gate genuinely routes each row. With
|
|
1033
|
+
`true` the filter is bypassed and both rows would pass, which proves nothing.
|
|
1034
|
+
|
|
1035
|
+
The where-clause scopes the run to exactly two of the four rows. A
|
|
1036
|
+
generic-table run selects on the table's own columns, so it reads
|
|
1037
|
+
`customer_number` with no alias to prefix.
|
|
1038
|
+
|
|
1039
|
+
A `:partial` batch status is the expected outcome here, not a warning: one
|
|
1040
|
+
row passed and one was filtered, which is precisely the split this scenario
|
|
1041
|
+
is about.
|
|
1042
|
+
|
|
1043
|
+
```typescript
|
|
1044
|
+
const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, {
|
|
1045
|
+
sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.noPhoneNumber}')`,
|
|
1046
|
+
mode: 'live',
|
|
1047
|
+
manual_override: false,
|
|
1048
|
+
})
|
|
1049
|
+
const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
|
|
1050
|
+
ctx.welcomeRunLogId = fired.workflowRunLogId
|
|
1051
|
+
ctx.welcomeRunBatchId = fired.batchId!
|
|
1052
|
+
|
|
1053
|
+
const deadline = Date.now() + 120_000
|
|
1054
|
+
let status: string | null = null
|
|
1055
|
+
while (Date.now() < deadline) {
|
|
1056
|
+
const { data: log } = await api.workflows.batchLogs.refresh(
|
|
1057
|
+
tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, ctx.welcomeRunLogId,
|
|
1058
|
+
)
|
|
1059
|
+
status = log.status ?? null
|
|
1060
|
+
if (status && status !== 'pending') break
|
|
1061
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1062
|
+
}
|
|
1063
|
+
if (status === 'failed') throw new Error('welcome run reached :failed')
|
|
1064
|
+
if (!status || status === 'pending') {
|
|
1065
|
+
throw new Error('welcome run did not leave :pending within 120s')
|
|
1066
|
+
}
|
|
1067
|
+
```
|
|
1068
|
+
|
|
1069
|
+
## 018 — the filter routed, and that is asserted at the log level
|
|
1070
|
+
|
|
1071
|
+
Each row produced a Workflow Execution Log. The reachable customer passed, so
|
|
1072
|
+
its log is `:executing` or `:completed`. The customer with no phone failed the
|
|
1073
|
+
filter, so its log is `:filtered` — a distinct terminal status, and a
|
|
1074
|
+
*correct* outcome rather than an error. The platform looked at the row, the
|
|
1075
|
+
gate rendered empty, and it recorded that the row was intentionally skipped.
|
|
1076
|
+
|
|
1077
|
+
A run where both passed, or both were filtered, would mean the filter is not
|
|
1078
|
+
actually reading `event_dataset.phone`.
|
|
1079
|
+
|
|
1080
|
+
**This assertion is at the execution-log level on purpose, and that is what
|
|
1081
|
+
makes it survive a rerun.** The action inside the passing row's log will
|
|
1082
|
+
refuse to fire a second time — its idempotency key already fired — but the
|
|
1083
|
+
row still passes the filter, so the one-passed-one-filtered split is the same
|
|
1084
|
+
on every run. An assertion written against the action instead would go red on
|
|
1085
|
+
run two for a reason that is not a defect.
|
|
1086
|
+
|
|
1087
|
+
```typescript
|
|
1088
|
+
const deadline = Date.now() + 90_000
|
|
1089
|
+
let ourLogs: Array<Record<string, unknown>> = []
|
|
1090
|
+
let byStatus: Record<string, number> = {}
|
|
1091
|
+
while (Date.now() < deadline) {
|
|
1092
|
+
const { data: wfLogs } = await api.workflows.workflowLogs.list(
|
|
1093
|
+
tenantSlug, datalakeSlug, ctx.welcomeWorkflowSlug, { page_size: 100 },
|
|
1094
|
+
)
|
|
1095
|
+
ourLogs = (wfLogs.data ?? [])
|
|
1096
|
+
.map((w) => w as Record<string, unknown>)
|
|
1097
|
+
.filter((w) => w.batch_id === ctx.welcomeRunBatchId)
|
|
1098
|
+
byStatus = {}
|
|
1099
|
+
for (const w of ourLogs) {
|
|
1100
|
+
const st = (w.status as string | undefined) ?? 'unknown'
|
|
1101
|
+
byStatus[st] = (byStatus[st] ?? 0) + 1
|
|
1102
|
+
}
|
|
1103
|
+
const settled = (byStatus.completed ?? 0) + (byStatus.executing ?? 0)
|
|
1104
|
+
if (ourLogs.length === 2 && settled === 1 && (byStatus.filtered ?? 0) === 1) break
|
|
1105
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1106
|
+
}
|
|
1107
|
+
if (ourLogs.length !== 2) {
|
|
1108
|
+
throw new Error(`expected 2 execution logs for the run, got ${ourLogs.length}`)
|
|
1109
|
+
}
|
|
1110
|
+
const passed = (byStatus.executing ?? 0) + (byStatus.completed ?? 0)
|
|
1111
|
+
if (passed !== 1) {
|
|
1112
|
+
throw new Error(`expected 1 row past the filter, got ${passed} — distribution ${JSON.stringify(byStatus)}`)
|
|
1113
|
+
}
|
|
1114
|
+
if ((byStatus.filtered ?? 0) !== 1) {
|
|
1115
|
+
throw new Error(
|
|
1116
|
+
`expected the no-phone customer to be :filtered — distribution ${JSON.stringify(byStatus)}`,
|
|
1117
|
+
)
|
|
1118
|
+
}
|
|
1119
|
+
```
|
|
1120
|
+
|
|
1121
|
+
## 019 — the welcome SMS exists, and it carries a working link
|
|
1122
|
+
|
|
1123
|
+
§018 proved the routing. This proves something was actually *sent*, and it is
|
|
1124
|
+
the assertion that survives every rerun.
|
|
1125
|
+
|
|
1126
|
+
**Search on the idempotency key, never on `workflow_id`.** The key is
|
|
1127
|
+
`<customer number>-send_welcome_sms` and the customer number is stable, so it
|
|
1128
|
+
names the same message forever. A workflow id is stable only as long as the
|
|
1129
|
+
workflow resource is: recreate it against a lake that still holds its rows
|
|
1130
|
+
and the id changes while the key does not — the action then correctly refuses
|
|
1131
|
+
to fire, and a search for the new id finds nothing at all.
|
|
1132
|
+
|
|
1133
|
+
Then resolve the `/t/<token>` the platform baked into the body when it minted
|
|
1134
|
+
the page token, and post the tracking a connected-app frontend would post
|
|
1135
|
+
when the customer opens the portal. That closes the loop: the message went
|
|
1136
|
+
out, the link in it resolves to this customer's page, and the open was
|
|
1137
|
+
recorded.
|
|
1138
|
+
|
|
1139
|
+
```typescript
|
|
1140
|
+
const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
|
|
1141
|
+
search_query: `m.idempotency_key = '${ctx.reachNumber}-send_welcome_sms'`,
|
|
1142
|
+
})
|
|
1143
|
+
if (search.status !== 'completed') {
|
|
1144
|
+
throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
|
|
1145
|
+
}
|
|
1146
|
+
|
|
1147
|
+
const msgDeadline = Date.now() + 90_000
|
|
1148
|
+
let welcomeBody: string | undefined
|
|
1149
|
+
while (Date.now() < msgDeadline && !welcomeBody) {
|
|
1150
|
+
const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
|
|
1151
|
+
userSearchId: search.id!,
|
|
1152
|
+
dataAccessMode: 'raw',
|
|
1153
|
+
})
|
|
1154
|
+
welcomeBody = ((data.data ?? []) as Array<Record<string, unknown>>)
|
|
1155
|
+
.map((m) => String(m.body ?? ''))
|
|
1156
|
+
.find((body) => body.includes('/t/') && body.includes('welcome to Alvera Billing'))
|
|
1157
|
+
if (!welcomeBody) await new Promise((r) => setTimeout(r, 2_000))
|
|
1158
|
+
}
|
|
1159
|
+
if (!welcomeBody) {
|
|
1160
|
+
throw new Error(
|
|
1161
|
+
`no rendered welcome SMS for ${ctx.reachNumber} within 90s — ` +
|
|
1162
|
+
'the action never fired, or it fired under a different idempotency key',
|
|
1163
|
+
)
|
|
1164
|
+
}
|
|
1165
|
+
const tokenMatch = welcomeBody.match(/\/t\/([A-Za-z0-9_-]+)/)
|
|
1166
|
+
if (!tokenMatch) throw new Error(`no /t/<token> in the rendered body: ${welcomeBody}`)
|
|
1167
|
+
|
|
1168
|
+
const { data: resolved } = await api.connectedApps.resolvePage(
|
|
1169
|
+
tenantSlug, datalakeSlug, ctx.portalAppSlug,
|
|
1170
|
+
{ short_path: tokenMatch[1]!, user_agent: 'cookbook-doctest/subscription-saas' },
|
|
1171
|
+
)
|
|
1172
|
+
if (resolved.route_path !== '/portal/welcome') {
|
|
1173
|
+
throw new Error(`resolvePage route_path mismatch: ${resolved.route_path}`)
|
|
1174
|
+
}
|
|
1175
|
+
if (!(resolved.message?.body ?? '').includes('welcome to Alvera Billing')) {
|
|
1176
|
+
throw new Error('the resolved page carries the wrong message')
|
|
1177
|
+
}
|
|
1178
|
+
|
|
1179
|
+
const now = new Date().toISOString()
|
|
1180
|
+
const { data: tracked } = await api.connectedApps.updateMessageTracking(
|
|
1181
|
+
tenantSlug, datalakeSlug, ctx.portalAppSlug,
|
|
1182
|
+
{ short_path: tokenMatch[1]!, opened_at: now, form_submitted_at: now },
|
|
1183
|
+
)
|
|
1184
|
+
if (!tracked.message?.opened_at || !tracked.message?.form_submitted_at) {
|
|
1185
|
+
throw new Error('message tracking did not persist opened_at + form_submitted_at')
|
|
1186
|
+
}
|
|
1187
|
+
```
|
|
1188
|
+
|
|
1189
|
+
## 020 — the pay-portal connected app
|
|
1190
|
+
|
|
1191
|
+
The dunning SMS carries a different link to a different page: not the
|
|
1192
|
+
welcome portal but the one where an overdue invoice gets paid. A second
|
|
1193
|
+
connected app, registered the same way.
|
|
1194
|
+
|
|
1195
|
+
Two apps rather than two routes on one app is a deliberate choice worth
|
|
1196
|
+
naming. A connected app is the unit a page token is minted against, so
|
|
1197
|
+
separating them means a token issued for a payment page cannot be replayed
|
|
1198
|
+
against the welcome page, and revoking one does not take the other down.
|
|
1199
|
+
|
|
1200
|
+
```typescript
|
|
1201
|
+
const payApp = await ctx.ensure(
|
|
1202
|
+
'pay-portal connected app',
|
|
1203
|
+
async () => {
|
|
1204
|
+
const { data } = await api.connectedApps.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1205
|
+
return ctx.firstNamed(data.data, 'Cookbook Payment Portal')
|
|
1206
|
+
},
|
|
1207
|
+
async () => {
|
|
1208
|
+
const { data } = await api.connectedApps.create(tenantSlug, datalakeSlug, {
|
|
1209
|
+
name: 'Cookbook Payment Portal',
|
|
1210
|
+
description: 'Self-serve payment page linked from the outbound dunning SMS.',
|
|
1211
|
+
mode: 'self_hosted',
|
|
1212
|
+
urls: [{ url: 'https://pay.example.local', is_primary: true, label: 'production' }],
|
|
1213
|
+
})
|
|
1214
|
+
return data
|
|
1215
|
+
},
|
|
1216
|
+
)
|
|
1217
|
+
ctx.payAppId = payApp.id!
|
|
1218
|
+
ctx.payAppSlug = payApp.slug!
|
|
1219
|
+
```
|
|
1220
|
+
|
|
1221
|
+
## 021 — the Dunning SMS workflow: two gates, not one
|
|
1222
|
+
|
|
1223
|
+
Same table, same shape, one more gate. The welcome workflow asks whether the
|
|
1224
|
+
platform *can* reach this customer. The dunning workflow asks that and
|
|
1225
|
+
whether it is *allowed* to: a customer whose KYC is incomplete — no tax
|
|
1226
|
+
identifier on file — is not chased, however overdue the invoice.
|
|
1227
|
+
|
|
1228
|
+
Both gates are one Liquid filter. That is the whole point of the primitive:
|
|
1229
|
+
the reachability rule and the KYC rule live next to each other on the
|
|
1230
|
+
workflow, where a compliance reviewer can read them, rather than in two
|
|
1231
|
+
branches of a service nobody reviews.
|
|
1232
|
+
|
|
1233
|
+
The recency guard is thirty days rather than the welcome's six months,
|
|
1234
|
+
because that is the shape of the business rule — one reminder per billing
|
|
1235
|
+
cycle, not one reminder per lifetime.
|
|
1236
|
+
|
|
1237
|
+
```typescript
|
|
1238
|
+
const DUNNING_KEY = 'send_dunning_sms'
|
|
1239
|
+
|
|
1240
|
+
const dunningWorkflow = await ctx.ensure(
|
|
1241
|
+
'Dunning SMS workflow',
|
|
1242
|
+
async () => {
|
|
1243
|
+
const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1244
|
+
return ctx.firstNamed(data.data, 'Cookbook Dunning SMS Workflow')
|
|
1245
|
+
},
|
|
1246
|
+
async () => {
|
|
1247
|
+
const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
|
|
1248
|
+
name: 'Cookbook Dunning SMS Workflow',
|
|
1249
|
+
description: 'Sends a payment reminder to reachable, KYC-complete customers with a self-serve pay link.',
|
|
1250
|
+
dataset_type: 'generic_table',
|
|
1251
|
+
generic_table_id: ctx.customersTableId,
|
|
1252
|
+
status: 'live',
|
|
1253
|
+
tags: ['billing', 'dunning'],
|
|
1254
|
+
skip_mdm_resolution: false,
|
|
1255
|
+
filter_config: {
|
|
1256
|
+
type: 'custom',
|
|
1257
|
+
body:
|
|
1258
|
+
'{% if event_dataset.phone and event_dataset.phone != "" %}' +
|
|
1259
|
+
'{% if event_dataset.tax_id and event_dataset.tax_id != "" %}true{% endif %}' +
|
|
1260
|
+
'{% endif %}',
|
|
1261
|
+
},
|
|
1262
|
+
decision_config: {
|
|
1263
|
+
type: 'custom',
|
|
1264
|
+
body: `["${DUNNING_KEY}"]`,
|
|
1265
|
+
output_schema: { type: 'array', items: { type: 'string' } },
|
|
1266
|
+
},
|
|
1267
|
+
context_datasets: [
|
|
1268
|
+
{
|
|
1269
|
+
dataset_type: 'message',
|
|
1270
|
+
where_clause:
|
|
1271
|
+
`m.legal_entity_id = '{{ legal_entity_id }}' AND m.decision_key = '${DUNNING_KEY}' ` +
|
|
1272
|
+
"AND m.sent_at > NOW() - INTERVAL '30 days'",
|
|
1273
|
+
limit: 1,
|
|
1274
|
+
position: 0,
|
|
1275
|
+
},
|
|
1276
|
+
],
|
|
1277
|
+
actions: [
|
|
1278
|
+
{
|
|
1279
|
+
action_type: 'sms',
|
|
1280
|
+
tool_id: ctx.smsToolId,
|
|
1281
|
+
decision_key: DUNNING_KEY,
|
|
1282
|
+
position: 0,
|
|
1283
|
+
trigger_template: 'now',
|
|
1284
|
+
idempotency_template: `{{ event_dataset.customer_number }}-${DUNNING_KEY}`,
|
|
1285
|
+
connected_app_id: ctx.payAppId,
|
|
1286
|
+
connected_app_route: '/portal/pay',
|
|
1287
|
+
connected_app_metadata_template: '{"customer_number":"{{ event_dataset.customer_number }}"}',
|
|
1288
|
+
tool_call: {
|
|
1289
|
+
tool_call_type: 'sms_request',
|
|
1290
|
+
to: { type: 'custom', body: '{{ event_dataset.phone }}' },
|
|
1291
|
+
body: {
|
|
1292
|
+
type: 'custom',
|
|
1293
|
+
body:
|
|
1294
|
+
'Hi {{ event_dataset.name }}, an invoice on account ' +
|
|
1295
|
+
'{{ event_dataset.customer_number }} needs your attention. ' +
|
|
1296
|
+
'Pay now: {{ connected_app_form_url }}',
|
|
1297
|
+
},
|
|
1298
|
+
sms_type: 'transactional',
|
|
1299
|
+
},
|
|
1300
|
+
},
|
|
1301
|
+
],
|
|
1302
|
+
})
|
|
1303
|
+
return data
|
|
1304
|
+
},
|
|
1305
|
+
)
|
|
1306
|
+
ctx.dunningWorkflowSlug = dunningWorkflow.slug!
|
|
1307
|
+
```
|
|
1308
|
+
|
|
1309
|
+
## 022 — run it, and prove the KYC gate is the one that bit
|
|
1310
|
+
|
|
1311
|
+
The scope is the reachable customer and the one missing a tax identifier.
|
|
1312
|
+
**Both have a phone**, which is what makes this run prove something the
|
|
1313
|
+
welcome run could not: the row that gets filtered here is filtered by the
|
|
1314
|
+
second gate, not the first. Had the scope reused the no-phone customer, a
|
|
1315
|
+
filter that ignored `tax_id` entirely would still have produced the same
|
|
1316
|
+
one-passed-one-filtered split and looked green.
|
|
1317
|
+
|
|
1318
|
+
That is the general lesson about filter tests. A pass/fail split only proves
|
|
1319
|
+
the gate you meant to test if the two rows differ in exactly that field.
|
|
1320
|
+
|
|
1321
|
+
```typescript
|
|
1322
|
+
const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, {
|
|
1323
|
+
sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.noKycNumber}')`,
|
|
1324
|
+
mode: 'live',
|
|
1325
|
+
manual_override: false,
|
|
1326
|
+
})
|
|
1327
|
+
const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
|
|
1328
|
+
ctx.dunningRunBatchId = fired.batchId!
|
|
1329
|
+
|
|
1330
|
+
const refreshDeadline = Date.now() + 120_000
|
|
1331
|
+
let status: string | null = null
|
|
1332
|
+
while (Date.now() < refreshDeadline) {
|
|
1333
|
+
const { data: log } = await api.workflows.batchLogs.refresh(
|
|
1334
|
+
tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, fired.workflowRunLogId,
|
|
1335
|
+
)
|
|
1336
|
+
status = log.status ?? null
|
|
1337
|
+
if (status && status !== 'pending') break
|
|
1338
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1339
|
+
}
|
|
1340
|
+
if (status === 'failed') throw new Error('dunning run reached :failed')
|
|
1341
|
+
if (!status || status === 'pending') throw new Error('dunning run did not leave :pending within 120s')
|
|
1342
|
+
|
|
1343
|
+
const welDeadline = Date.now() + 90_000
|
|
1344
|
+
let ourLogs: Array<Record<string, unknown>> = []
|
|
1345
|
+
let byStatus: Record<string, number> = {}
|
|
1346
|
+
while (Date.now() < welDeadline) {
|
|
1347
|
+
const { data: wfLogs } = await api.workflows.workflowLogs.list(
|
|
1348
|
+
tenantSlug, datalakeSlug, ctx.dunningWorkflowSlug, { page_size: 100 },
|
|
1349
|
+
)
|
|
1350
|
+
ourLogs = (wfLogs.data ?? [])
|
|
1351
|
+
.map((w) => w as Record<string, unknown>)
|
|
1352
|
+
.filter((w) => w.batch_id === ctx.dunningRunBatchId)
|
|
1353
|
+
byStatus = {}
|
|
1354
|
+
for (const w of ourLogs) {
|
|
1355
|
+
const st = (w.status as string | undefined) ?? 'unknown'
|
|
1356
|
+
byStatus[st] = (byStatus[st] ?? 0) + 1
|
|
1357
|
+
}
|
|
1358
|
+
const settled = (byStatus.completed ?? 0) + (byStatus.executing ?? 0)
|
|
1359
|
+
if (ourLogs.length === 2 && settled === 1 && (byStatus.filtered ?? 0) === 1) break
|
|
1360
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1361
|
+
}
|
|
1362
|
+
if (ourLogs.length !== 2) {
|
|
1363
|
+
throw new Error(`expected 2 execution logs for the dunning run, got ${ourLogs.length}`)
|
|
1364
|
+
}
|
|
1365
|
+
if ((byStatus.executing ?? 0) + (byStatus.completed ?? 0) !== 1) {
|
|
1366
|
+
throw new Error(`expected 1 row past the KYC gate — distribution ${JSON.stringify(byStatus)}`)
|
|
1367
|
+
}
|
|
1368
|
+
if ((byStatus.filtered ?? 0) !== 1) {
|
|
1369
|
+
throw new Error(
|
|
1370
|
+
`expected the customer with no tax_id to be :filtered — distribution ${JSON.stringify(byStatus)}`,
|
|
1371
|
+
)
|
|
1372
|
+
}
|
|
1373
|
+
```
|
|
1374
|
+
|
|
1375
|
+
## 023 — the reminder went out, and it is a different message on a different app
|
|
1376
|
+
|
|
1377
|
+
Same read-back idiom as §019, and worth doing twice for one reason: it proves
|
|
1378
|
+
the two workflows are genuinely independent. The same customer received two
|
|
1379
|
+
messages, under two idempotency keys, carrying two links that resolve against
|
|
1380
|
+
two connected apps and two routes.
|
|
1381
|
+
|
|
1382
|
+
The key here is `<customer number>-send_dunning_sms`, and it names a row that
|
|
1383
|
+
persists in the lake long after the run that made it. That is what a
|
|
1384
|
+
read-back should ask for — the durable record, not the run.
|
|
1385
|
+
|
|
1386
|
+
```typescript
|
|
1387
|
+
const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
|
|
1388
|
+
search_query: `m.idempotency_key = '${ctx.reachNumber}-send_dunning_sms'`,
|
|
1389
|
+
})
|
|
1390
|
+
if (search.status !== 'completed') {
|
|
1391
|
+
throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
|
|
1392
|
+
}
|
|
1393
|
+
|
|
1394
|
+
const msgDeadline = Date.now() + 90_000
|
|
1395
|
+
let dunningBody: string | undefined
|
|
1396
|
+
while (Date.now() < msgDeadline && !dunningBody) {
|
|
1397
|
+
const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
|
|
1398
|
+
userSearchId: search.id!,
|
|
1399
|
+
dataAccessMode: 'raw',
|
|
1400
|
+
})
|
|
1401
|
+
dunningBody = ((data.data ?? []) as Array<Record<string, unknown>>)
|
|
1402
|
+
.map((m) => String(m.body ?? ''))
|
|
1403
|
+
.find((body) => body.includes('/t/') && body.includes('needs your attention'))
|
|
1404
|
+
if (!dunningBody) await new Promise((r) => setTimeout(r, 2_000))
|
|
1405
|
+
}
|
|
1406
|
+
if (!dunningBody) {
|
|
1407
|
+
throw new Error(`no rendered dunning SMS for ${ctx.reachNumber} within 90s`)
|
|
1408
|
+
}
|
|
1409
|
+
if (!dunningBody.includes(ctx.reachNumber)) {
|
|
1410
|
+
throw new Error(`the dunning body does not name the overdue account: ${dunningBody}`)
|
|
1411
|
+
}
|
|
1412
|
+
|
|
1413
|
+
const token = dunningBody.match(/\/t\/([A-Za-z0-9_-]+)/)
|
|
1414
|
+
if (!token) throw new Error(`no /t/<token> in the dunning body: ${dunningBody}`)
|
|
1415
|
+
|
|
1416
|
+
const { data: resolved } = await api.connectedApps.resolvePage(
|
|
1417
|
+
tenantSlug, datalakeSlug, ctx.payAppSlug,
|
|
1418
|
+
{ short_path: token[1]!, user_agent: 'cookbook-doctest/subscription-saas' },
|
|
1419
|
+
)
|
|
1420
|
+
if (resolved.route_path !== '/portal/pay') {
|
|
1421
|
+
throw new Error(`the dunning link resolved to the wrong route: ${resolved.route_path}`)
|
|
1422
|
+
}
|
|
1423
|
+
```
|
|
1424
|
+
|
|
1425
|
+
## 024 — the LLM tool
|
|
1426
|
+
|
|
1427
|
+
The welcome and dunning workflows fire the *same* action for everything that
|
|
1428
|
+
passes their filter. The third scenario is the opposite shape: which message
|
|
1429
|
+
goes out depends on what kind of account it is, and that is a judgement over
|
|
1430
|
+
the row rather than a rule about it.
|
|
1431
|
+
|
|
1432
|
+
That judgement needs a model, and a model needs a tool. This one is a
|
|
1433
|
+
**provider adapter**, and the two halves are the thing to read. `base_body`
|
|
1434
|
+
authors the provider's own request — here Ollama's native `/api/chat` shape,
|
|
1435
|
+
with `think: false` and a `format` schema so the model returns
|
|
1436
|
+
schema-constrained JSON instead of prose with JSON in it. `response_extractor`
|
|
1437
|
+
maps the provider's reply back onto the canonical
|
|
1438
|
+
`{ output_json, input_tokens, … }` the platform expects, and its
|
|
1439
|
+
`output_schema` is required for an `llm_enrichment` tool.
|
|
1440
|
+
|
|
1441
|
+
Swapping providers is a matter of rewriting those two templates. Nothing
|
|
1442
|
+
downstream — not the agent, not the workflow — knows which model answered.
|
|
1443
|
+
|
|
1444
|
+
```typescript
|
|
1445
|
+
const ENRICHMENT_OUTPUT_SCHEMA = {
|
|
1446
|
+
type: 'object',
|
|
1447
|
+
properties: {
|
|
1448
|
+
output_json: {},
|
|
1449
|
+
input_tokens: { type: ['integer', 'null'] },
|
|
1450
|
+
output_tokens: { type: ['integer', 'null'] },
|
|
1451
|
+
total_tokens: { type: ['integer', 'null'] },
|
|
1452
|
+
explanation: { type: ['string', 'null'] },
|
|
1453
|
+
},
|
|
1454
|
+
required: ['output_json'],
|
|
1455
|
+
}
|
|
1456
|
+
|
|
1457
|
+
const OLLAMA_BASE_BODY =
|
|
1458
|
+
'{"model": "{{ model }}", "messages": [{"role": "user", "content": "{{ rendered_prompt | json_escape }}", ' +
|
|
1459
|
+
'"images": [{% for img in images %}{% unless forloop.first %}, {% endunless %}"{{ img.data }}"{% endfor %}]}], ' +
|
|
1460
|
+
'"stream": false, "think": false, "options": {"temperature": {{ temperature }}, "num_predict": {{ max_tokens }}, "num_ctx": 40960}, "format": {{ schema | to_json }}}'
|
|
1461
|
+
|
|
1462
|
+
const OLLAMA_EXTRACTOR =
|
|
1463
|
+
'{"output_json": "{{ msg.message.content | json_escape }}", ' +
|
|
1464
|
+
'"explanation": "{{ msg.message.thinking | json_escape }}", ' +
|
|
1465
|
+
'"input_tokens": {{ msg.prompt_eval_count | default: 0 }}, ' +
|
|
1466
|
+
'"output_tokens": {{ msg.eval_count | default: 0 }}, ' +
|
|
1467
|
+
'"total_tokens": {{ msg.prompt_eval_count | default: 0 | plus: msg.eval_count }}}'
|
|
1468
|
+
|
|
1469
|
+
const llmTool = await ctx.ensure(
|
|
1470
|
+
'LLM tool',
|
|
1471
|
+
async () => {
|
|
1472
|
+
const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1473
|
+
return ctx.firstNamed(data.data, 'Cookbook Billing LLM Tool')
|
|
1474
|
+
},
|
|
1475
|
+
async () => {
|
|
1476
|
+
const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
|
|
1477
|
+
name: 'Cookbook Billing LLM Tool',
|
|
1478
|
+
description: 'Ollama-backed chat-completion adapter for account-priority triage.',
|
|
1479
|
+
intent: 'llm_enrichment',
|
|
1480
|
+
status: 'active',
|
|
1481
|
+
datalake_id: ctx.datalakeId,
|
|
1482
|
+
response_extractor: { type: 'custom', body: OLLAMA_EXTRACTOR, output_schema: ENRICHMENT_OUTPUT_SCHEMA },
|
|
1483
|
+
body: {
|
|
1484
|
+
tool_body_type: 'rest_api',
|
|
1485
|
+
base_url: 'http://localhost:11434',
|
|
1486
|
+
base_path: { type: 'custom', body: '/api/chat' },
|
|
1487
|
+
auth_method: 'api_key',
|
|
1488
|
+
api_key: 'stub-key',
|
|
1489
|
+
api_key_name: 'Authorization',
|
|
1490
|
+
api_key_location: 'header',
|
|
1491
|
+
request_type: 'json',
|
|
1492
|
+
response_type: 'json',
|
|
1493
|
+
timeout_ms: 60_000,
|
|
1494
|
+
base_body: { type: 'custom', body: OLLAMA_BASE_BODY },
|
|
1495
|
+
},
|
|
1496
|
+
})
|
|
1497
|
+
return data
|
|
1498
|
+
},
|
|
1499
|
+
)
|
|
1500
|
+
ctx.llmToolId = llmTool.id!
|
|
1501
|
+
```
|
|
1502
|
+
|
|
1503
|
+
## 025 — the Priority Triage agent
|
|
1504
|
+
|
|
1505
|
+
The agent binds three things the workflow cannot supply for itself: the
|
|
1506
|
+
model, an `input_schema` the workflow's context mapping must satisfy, and an
|
|
1507
|
+
`llm_response_schema` whose `enum` pins the vocabulary.
|
|
1508
|
+
|
|
1509
|
+
**The enum is the load-bearing part.** The workflow interpolates the band
|
|
1510
|
+
straight into its decision array, so a model that answers `"high priority"`
|
|
1511
|
+
instead of `priority_high` would produce a decision key no action is
|
|
1512
|
+
registered for, and the row would silently do nothing. Constraining the
|
|
1513
|
+
schema means the failure happens at the model boundary where it is legible,
|
|
1514
|
+
not three layers down as an empty result.
|
|
1515
|
+
|
|
1516
|
+
`temperature: 0.0` makes identical inputs classify identically, which is what
|
|
1517
|
+
lets §027 assert on the outcome at all.
|
|
1518
|
+
|
|
1519
|
+
`data_access: 'raw'` is the right call **on this surface and only on this
|
|
1520
|
+
surface**. There is no derived lake here to read from, and the agent's prompt
|
|
1521
|
+
takes the account number and the account type — no name, no email, no phone.
|
|
1522
|
+
On `payments-compliance` the equivalent agent runs `tokenized`, because there
|
|
1523
|
+
the model is reading a record that carries identity and the point is that it
|
|
1524
|
+
never sees it. The tier follows what the prompt actually contains.
|
|
1525
|
+
|
|
1526
|
+
```typescript
|
|
1527
|
+
const TRIAGE_PROMPT_BODY = `You are an account-priority triage assistant for a subscription billing team. Classify the customer into EXACTLY ONE band by account type:
|
|
1528
|
+
|
|
1529
|
+
- "priority_high" — enterprise (large account, high balance, white-glove outreach)
|
|
1530
|
+
- "priority_medium" — individual (standard self-serve account)
|
|
1531
|
+
- "priority_low" — government (long payment cycles, low urgency)
|
|
1532
|
+
|
|
1533
|
+
Customer number: {{ customer_number }}
|
|
1534
|
+
Customer type: {{ customer_type }}
|
|
1535
|
+
|
|
1536
|
+
Respond with JSON: {"priority_band": "<one of priority_high|priority_medium|priority_low>"}`
|
|
1537
|
+
|
|
1538
|
+
const triageAgent = await ctx.ensure(
|
|
1539
|
+
'Priority Triage agent',
|
|
1540
|
+
async () => {
|
|
1541
|
+
const { data } = await api.aiAgents.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1542
|
+
return ctx.firstNamed(data.data, 'Cookbook Priority Triage Agent')
|
|
1543
|
+
},
|
|
1544
|
+
async () => {
|
|
1545
|
+
const { data } = await api.aiAgents.create(tenantSlug, datalakeSlug, {
|
|
1546
|
+
name: 'Cookbook Priority Triage Agent',
|
|
1547
|
+
tool_id: ctx.llmToolId,
|
|
1548
|
+
model: 'qwen3-vl:8b-instruct',
|
|
1549
|
+
data_access: 'raw',
|
|
1550
|
+
temperature: 0.0,
|
|
1551
|
+
max_tokens: 1024,
|
|
1552
|
+
enabled: true,
|
|
1553
|
+
input_schema: {
|
|
1554
|
+
type: 'object',
|
|
1555
|
+
properties: {
|
|
1556
|
+
customer_number: { type: 'string' },
|
|
1557
|
+
customer_type: { type: 'string' },
|
|
1558
|
+
},
|
|
1559
|
+
required: ['customer_number', 'customer_type'],
|
|
1560
|
+
},
|
|
1561
|
+
llm_response_schema: {
|
|
1562
|
+
type: 'object',
|
|
1563
|
+
properties: {
|
|
1564
|
+
priority_band: {
|
|
1565
|
+
type: 'string',
|
|
1566
|
+
enum: ['priority_high', 'priority_medium', 'priority_low'],
|
|
1567
|
+
},
|
|
1568
|
+
},
|
|
1569
|
+
required: ['priority_band'],
|
|
1570
|
+
},
|
|
1571
|
+
prompt_config: { type: 'custom', body: TRIAGE_PROMPT_BODY },
|
|
1572
|
+
})
|
|
1573
|
+
return data
|
|
1574
|
+
},
|
|
1575
|
+
)
|
|
1576
|
+
aiAgentId = triageAgent.id!
|
|
1577
|
+
ctx.triageAgentSlug = triageAgent.slug!
|
|
1578
|
+
```
|
|
1579
|
+
|
|
1580
|
+
## 026 — the Priority Triage workflow
|
|
1581
|
+
|
|
1582
|
+
Two things make this workflow agent-driven rather than rule-driven.
|
|
1583
|
+
|
|
1584
|
+
The agent is **nested in the create body** — `workflow_ai_agents` is a
|
|
1585
|
+
`cast_assoc`, not a separate attach call — together with a Liquid
|
|
1586
|
+
`context_mapping_config` that projects each row into the agent's input
|
|
1587
|
+
schema. A workflow join omits `output_schema`; the server pins it from the
|
|
1588
|
+
agent's own `input_schema`, so the two cannot drift apart.
|
|
1589
|
+
|
|
1590
|
+
The `decision_config` then interpolates the agent's answer into a
|
|
1591
|
+
single-element decision array, and three actions are registered, one per
|
|
1592
|
+
band. Whichever band comes back picks one and leaves the other two
|
|
1593
|
+
`:skipped`.
|
|
1594
|
+
|
|
1595
|
+
**Bracket access is required, not stylistic.** The agent's slug contains
|
|
1596
|
+
hyphens, and Liquid's dot parser would read
|
|
1597
|
+
`additional_context.cookbook-priority-triage-agent` as a subtraction.
|
|
1598
|
+
|
|
1599
|
+
```typescript
|
|
1600
|
+
const BANDS = ['priority_high', 'priority_medium', 'priority_low'] as const
|
|
1601
|
+
|
|
1602
|
+
const CONTEXT_MAPPING_BODY = JSON.stringify({
|
|
1603
|
+
customer_number: '{{ event_dataset.customer_number }}',
|
|
1604
|
+
customer_type: '{{ event_dataset.customer_type }}',
|
|
1605
|
+
})
|
|
1606
|
+
|
|
1607
|
+
const triageWorkflow = await ctx.ensure(
|
|
1608
|
+
'Priority Triage workflow',
|
|
1609
|
+
async () => {
|
|
1610
|
+
const { data } = await api.workflows.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1611
|
+
return ctx.firstNamed(data.data, 'Cookbook Priority Triage Workflow')
|
|
1612
|
+
},
|
|
1613
|
+
async () => {
|
|
1614
|
+
const { data } = await api.workflows.create(tenantSlug, datalakeSlug, {
|
|
1615
|
+
name: 'Cookbook Priority Triage Workflow',
|
|
1616
|
+
description: 'Bands accounts high/medium/low via an LLM agent; one SMS action per band.',
|
|
1617
|
+
dataset_type: 'generic_table',
|
|
1618
|
+
generic_table_id: ctx.customersTableId,
|
|
1619
|
+
status: 'live',
|
|
1620
|
+
tags: ['billing', 'triage'],
|
|
1621
|
+
skip_mdm_resolution: false,
|
|
1622
|
+
filter_config: {
|
|
1623
|
+
type: 'custom',
|
|
1624
|
+
body:
|
|
1625
|
+
'{% if event_dataset.phone and event_dataset.phone != "" %}' +
|
|
1626
|
+
'{% if event_dataset.tax_id and event_dataset.tax_id != "" %}true{% endif %}' +
|
|
1627
|
+
'{% endif %}',
|
|
1628
|
+
},
|
|
1629
|
+
decision_config: {
|
|
1630
|
+
type: 'custom',
|
|
1631
|
+
body: `["{{ additional_context["${ctx.triageAgentSlug}"].priority_band }}"]`,
|
|
1632
|
+
output_schema: { type: 'array', items: { type: 'string' } },
|
|
1633
|
+
},
|
|
1634
|
+
actions: BANDS.map((band) => ({
|
|
1635
|
+
decision_key: band,
|
|
1636
|
+
action_type: 'sms' as const,
|
|
1637
|
+
tool_id: ctx.smsToolId,
|
|
1638
|
+
position: 0,
|
|
1639
|
+
trigger_template: 'now',
|
|
1640
|
+
idempotency_template: `{{ event_dataset.customer_number }}-${band}`,
|
|
1641
|
+
tool_call: {
|
|
1642
|
+
tool_call_type: 'sms_request' as const,
|
|
1643
|
+
to: { type: 'custom' as const, body: '{{ event_dataset.phone }}' },
|
|
1644
|
+
body: {
|
|
1645
|
+
type: 'custom' as const,
|
|
1646
|
+
body: `[${band}] {{ event_dataset.name }}, a note about account {{ event_dataset.customer_number }}.`,
|
|
1647
|
+
},
|
|
1648
|
+
sms_type: 'transactional' as const,
|
|
1649
|
+
},
|
|
1650
|
+
})),
|
|
1651
|
+
workflow_ai_agents: [
|
|
1652
|
+
{
|
|
1653
|
+
ai_agent_id: aiAgentId,
|
|
1654
|
+
position: 0,
|
|
1655
|
+
context_mapping_config: { type: 'custom', body: CONTEXT_MAPPING_BODY },
|
|
1656
|
+
},
|
|
1657
|
+
],
|
|
1658
|
+
})
|
|
1659
|
+
return data
|
|
1660
|
+
},
|
|
1661
|
+
)
|
|
1662
|
+
ctx.triageWorkflowSlug = triageWorkflow.slug!
|
|
1663
|
+
```
|
|
1664
|
+
|
|
1665
|
+
## 027 — the agent's band steered the fan-out
|
|
1666
|
+
|
|
1667
|
+
The scope is the individual and the enterprise account — both reachable, both
|
|
1668
|
+
KYC-complete, so both clear the filter and reach the agent. They differ only
|
|
1669
|
+
in `customer_type`, which is the one thing the agent is asked about.
|
|
1670
|
+
|
|
1671
|
+
Each row produces an execution log carrying **three** action logs, one per
|
|
1672
|
+
band. On the first run for a given customer exactly **one** is matched
|
|
1673
|
+
(`:pending` or `:completed` — the band the agent chose) and **two** are
|
|
1674
|
+
`:skipped`. That one-of-three is the proof the model's output drove the
|
|
1675
|
+
decision rather than every action firing.
|
|
1676
|
+
|
|
1677
|
+
**On a later run all three are `:skipped`, and that is correct.** The
|
|
1678
|
+
idempotency key is `<customer number>-<band>`, the customer numbers are
|
|
1679
|
+
stable, so the key that already fired refuses to fire again — the platform
|
|
1680
|
+
will not send the same message about the same account twice. Both shapes are
|
|
1681
|
+
accepted here; anything else is a real failure. Two matched means the
|
|
1682
|
+
decision matched more than one band. Zero matched with zero skipped means the
|
|
1683
|
+
fan-out never happened at all.
|
|
1684
|
+
|
|
1685
|
+
```typescript
|
|
1686
|
+
const runResp = await api.workflows.run(tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, {
|
|
1687
|
+
sql_where_clause: `customer_number IN ('${ctx.reachNumber}', '${ctx.enterpriseNumber}')`,
|
|
1688
|
+
mode: 'live',
|
|
1689
|
+
manual_override: false,
|
|
1690
|
+
})
|
|
1691
|
+
const fired = await ctx.waitForFiredRun(datalakeSlug, runResp.data.workflow_run_id)
|
|
1692
|
+
ctx.triageRunBatchId = fired.batchId!
|
|
1693
|
+
|
|
1694
|
+
const refreshDeadline = Date.now() + 180_000
|
|
1695
|
+
let status: string | null = null
|
|
1696
|
+
while (Date.now() < refreshDeadline) {
|
|
1697
|
+
const { data: log } = await api.workflows.batchLogs.refresh(
|
|
1698
|
+
tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, fired.workflowRunLogId,
|
|
1699
|
+
)
|
|
1700
|
+
status = log.status ?? null
|
|
1701
|
+
if (status && status !== 'pending') break
|
|
1702
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1703
|
+
}
|
|
1704
|
+
if (status === 'failed') throw new Error('triage run reached :failed')
|
|
1705
|
+
if (!status || status === 'pending') throw new Error('triage run did not leave :pending within 180s')
|
|
1706
|
+
|
|
1707
|
+
const welDeadline = Date.now() + 180_000
|
|
1708
|
+
let ourLogs: Array<Record<string, unknown>> = []
|
|
1709
|
+
while (Date.now() < welDeadline) {
|
|
1710
|
+
const { data: wfLogs } = await api.workflows.workflowLogs.list(
|
|
1711
|
+
tenantSlug, datalakeSlug, ctx.triageWorkflowSlug, { page_size: 100 },
|
|
1712
|
+
)
|
|
1713
|
+
ourLogs = (wfLogs.data ?? [])
|
|
1714
|
+
.map((w) => w as Record<string, unknown>)
|
|
1715
|
+
.filter((w) => w.batch_id === ctx.triageRunBatchId)
|
|
1716
|
+
const settled = ourLogs.filter(
|
|
1717
|
+
(w) => ((w.action_execution_logs as unknown[] | undefined) ?? []).length === 3,
|
|
1718
|
+
)
|
|
1719
|
+
if (ourLogs.length === 2 && settled.length === 2) break
|
|
1720
|
+
await new Promise((r) => setTimeout(r, 2_000))
|
|
1721
|
+
}
|
|
1722
|
+
if (ourLogs.length !== 2) {
|
|
1723
|
+
throw new Error(`expected 2 execution logs for the triage run, got ${ourLogs.length}`)
|
|
1724
|
+
}
|
|
1725
|
+
|
|
1726
|
+
for (const wel of ourLogs) {
|
|
1727
|
+
const welId = (wel as { id?: string }).id
|
|
1728
|
+
const aels = (wel as { action_execution_logs?: Array<{ status?: string }> }).action_execution_logs ?? []
|
|
1729
|
+
if (aels.length !== 3) {
|
|
1730
|
+
throw new Error(`log ${welId}: expected 3 action logs (one per band), got ${aels.length}`)
|
|
1731
|
+
}
|
|
1732
|
+
const byStatus: Record<string, number> = {}
|
|
1733
|
+
for (const ael of aels) {
|
|
1734
|
+
const st = ael.status ?? 'unknown'
|
|
1735
|
+
byStatus[st] = (byStatus[st] ?? 0) + 1
|
|
1736
|
+
}
|
|
1737
|
+
const matched = (byStatus.pending ?? 0) + (byStatus.completed ?? 0)
|
|
1738
|
+
const skipped = byStatus.skipped ?? 0
|
|
1739
|
+
if (matched === 1 && skipped === 2) continue // first run for this account
|
|
1740
|
+
if (matched === 0 && skipped === 3) continue // already fired, correctly refusing
|
|
1741
|
+
throw new Error(
|
|
1742
|
+
`log ${welId}: expected 1 matched + 2 skipped, or 3 skipped on a repeat run — got ${JSON.stringify(byStatus)}`,
|
|
1743
|
+
)
|
|
1744
|
+
}
|
|
1745
|
+
```
|
|
1746
|
+
|
|
1747
|
+
## 028 — the enterprise account was banded high, and the SMS says so
|
|
1748
|
+
|
|
1749
|
+
§027 proved the fan-out chose one action per row. This proves *which* one,
|
|
1750
|
+
and it is the assertion that survives every rerun because it reads the
|
|
1751
|
+
durable message rather than the run that made it.
|
|
1752
|
+
|
|
1753
|
+
The enterprise account is the one asserted, because it is the only row whose
|
|
1754
|
+
band is genuinely unambiguous — the prompt says enterprise means
|
|
1755
|
+
`priority_high` in as many words, and a model that gets this wrong at
|
|
1756
|
+
temperature zero is a model worth knowing about.
|
|
1757
|
+
|
|
1758
|
+
The search is on the idempotency key, as always. Note the shape of it here:
|
|
1759
|
+
`<customer number>-priority_high` names the message the *high* action would
|
|
1760
|
+
have sent, so finding it is itself the band assertion. There is no need to
|
|
1761
|
+
parse the band back out of the body — though the body is checked too, since a
|
|
1762
|
+
message under the right key carrying the wrong copy would mean the actions
|
|
1763
|
+
are mis-wired to each other's templates.
|
|
1764
|
+
|
|
1765
|
+
```typescript
|
|
1766
|
+
const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'message', {
|
|
1767
|
+
search_query: `m.idempotency_key = '${ctx.enterpriseNumber}-priority_high'`,
|
|
1768
|
+
})
|
|
1769
|
+
if (search.status !== 'completed') {
|
|
1770
|
+
throw new Error(`message user-search status=${search.status} error=${search.error_message ?? '(none)'}`)
|
|
1771
|
+
}
|
|
1772
|
+
|
|
1773
|
+
const msgDeadline = Date.now() + 90_000
|
|
1774
|
+
let bandBody: string | undefined
|
|
1775
|
+
while (Date.now() < msgDeadline && !bandBody) {
|
|
1776
|
+
const { data } = await api.datasets.search(tenantSlug, datalakeSlug, 'message', {
|
|
1777
|
+
userSearchId: search.id!,
|
|
1778
|
+
dataAccessMode: 'raw',
|
|
1779
|
+
})
|
|
1780
|
+
bandBody = ((data.data ?? []) as Array<Record<string, unknown>>)
|
|
1781
|
+
.map((m) => String(m.body ?? ''))
|
|
1782
|
+
.find((body) => body.includes('[priority_high]'))
|
|
1783
|
+
if (!bandBody) await new Promise((r) => setTimeout(r, 2_000))
|
|
1784
|
+
}
|
|
1785
|
+
if (!bandBody) {
|
|
1786
|
+
throw new Error(
|
|
1787
|
+
`no priority_high SMS for ${ctx.enterpriseNumber} within 90s — ` +
|
|
1788
|
+
'the agent banded the enterprise account somewhere else, or the action never fired',
|
|
1789
|
+
)
|
|
1790
|
+
}
|
|
1791
|
+
if (!bandBody.includes(ctx.enterpriseNumber)) {
|
|
1792
|
+
throw new Error(`the banded SMS does not name the account: ${bandBody}`)
|
|
1793
|
+
}
|
|
1794
|
+
```
|
|
1795
|
+
|
|
1796
|
+
## 029 — a whole file at once: mint an upload link and PUT the CSV
|
|
1797
|
+
|
|
1798
|
+
Everything so far arrived one row at a time in an ingest call's body. That is
|
|
1799
|
+
the wrong shape for a backfill or a nightly export, where the file already
|
|
1800
|
+
exists and has four thousand rows in it.
|
|
1801
|
+
|
|
1802
|
+
The bulk path is three calls plus a discovery step:
|
|
1803
|
+
|
|
1804
|
+
1. `datalakes.createUploadLink` returns a presigned PUT `url` and the storage
|
|
1805
|
+
`key` the platform will reference.
|
|
1806
|
+
2. A **raw `fetch`** PUTs the bytes to that URL. This one goes straight to
|
|
1807
|
+
object storage and is deliberately not an SDK call — the SDK talks to the
|
|
1808
|
+
platform, and the platform is not in this hop.
|
|
1809
|
+
3. `dataActivationClients.ingestFile` enqueues a job against the key.
|
|
1810
|
+
4. The worker reads the file, allocates a **fresh `batch_id` it does not
|
|
1811
|
+
return**, and fans out per-row jobs. You discover that id by diffing the
|
|
1812
|
+
client's logs.
|
|
1813
|
+
|
|
1814
|
+
The file is the vendored four-row Stripe export. Its customer numbers are
|
|
1815
|
+
`CUS-10001`–`CUS-10004`, disjoint from the four inline rows, so this scenario
|
|
1816
|
+
adds rows rather than colliding with them.
|
|
1817
|
+
|
|
1818
|
+
```typescript
|
|
1819
|
+
const { readFileSync } = await import('node:fs')
|
|
1820
|
+
const { join } = await import('node:path')
|
|
1821
|
+
const csvBody = readFileSync(
|
|
1822
|
+
join(process.env.COOKBOOK_FIXTURES_DIR!, 'subscription-saas', 'stripe_customers_batch1.csv'),
|
|
1823
|
+
'utf8',
|
|
1824
|
+
)
|
|
1825
|
+
|
|
1826
|
+
const { data: link } = await api.datalakes.createUploadLink(tenantSlug, datalakeSlug, {
|
|
1827
|
+
content_type: 'text/csv',
|
|
1828
|
+
filename: 'stripe_customers_batch1.csv',
|
|
1829
|
+
})
|
|
1830
|
+
ctx.uploadKey = link.key!
|
|
1831
|
+
|
|
1832
|
+
const put = await fetch(link.url!, {
|
|
1833
|
+
method: 'PUT',
|
|
1834
|
+
headers: { 'Content-Type': 'text/csv' },
|
|
1835
|
+
body: csvBody,
|
|
1836
|
+
})
|
|
1837
|
+
if (put.status !== 200) {
|
|
1838
|
+
throw new Error(`presigned PUT failed: ${put.status}`)
|
|
1839
|
+
}
|
|
1840
|
+
```
|
|
1841
|
+
|
|
1842
|
+
## 030 — enqueue the same file twice, identity first
|
|
1843
|
+
|
|
1844
|
+
The split from §012 applies here too, and the file makes it cheaper rather
|
|
1845
|
+
than more expensive: **one upload, two enqueues**. The key names an object in
|
|
1846
|
+
storage, and nothing about it is bound to a client, so the identity client and
|
|
1847
|
+
the data client can each be pointed at the same bytes.
|
|
1848
|
+
|
|
1849
|
+
The order is the whole point. The identity enqueue runs and is waited on to
|
|
1850
|
+
completion; only then does the data enqueue go, and by that time every subject
|
|
1851
|
+
the second pass will resolve is already committed.
|
|
1852
|
+
|
|
1853
|
+
`ingestFile` returns a `job_id` and no `batch_id`, so each pass snapshots the
|
|
1854
|
+
client's existing batch ids first and then watches for the new one. The gate
|
|
1855
|
+
is `dataset_updated`, not `rows_ingested` — a file whose rows were all read
|
|
1856
|
+
and none written reports the former and not the latter.
|
|
1857
|
+
|
|
1858
|
+
```typescript
|
|
1859
|
+
ctx.ingestFileAndWait = async (dacSlug: string, expectRows: number): Promise<string> => {
|
|
1860
|
+
const { data: pre } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
|
|
1861
|
+
const seen = new Set(
|
|
1862
|
+
(pre.data ?? [])
|
|
1863
|
+
.map((r) => (r as { batch_id?: string }).batch_id)
|
|
1864
|
+
.filter((b): b is string => typeof b === 'string'),
|
|
1865
|
+
)
|
|
1866
|
+
|
|
1867
|
+
const { data: job } = await api.dataActivationClients.ingestFile(tenantSlug, datalakeSlug, dacSlug, {
|
|
1868
|
+
key: ctx.uploadKey,
|
|
1869
|
+
})
|
|
1870
|
+
if (!job.job_id) throw new Error(`ingestFile on ${dacSlug} did not return a job_id`)
|
|
1871
|
+
|
|
1872
|
+
const deadline = Date.now() + 180_000
|
|
1873
|
+
while (Date.now() < deadline) {
|
|
1874
|
+
const { data } = await api.dataActivationClients.logs.list(tenantSlug, datalakeSlug, dacSlug)
|
|
1875
|
+
for (const row of (data.data ?? []) as Array<Record<string, unknown>>) {
|
|
1876
|
+
const b = row.batch_id
|
|
1877
|
+
if (typeof b !== 'string' || seen.has(b)) continue
|
|
1878
|
+
if (row.status === 'partial' || row.status === 'failed') {
|
|
1879
|
+
throw new Error(
|
|
1880
|
+
`bulk batch ${b} on ${dacSlug} did not persist its rows (status: ${row.status}): ` +
|
|
1881
|
+
`${String(row.error ?? 'no reason given')}`,
|
|
1882
|
+
)
|
|
1883
|
+
}
|
|
1884
|
+
if (typeof row.dataset_updated !== 'number' || row.dataset_updated < expectRows) continue
|
|
1885
|
+
if (!Array.isArray(row.output_files) || row.output_files.length === 0) continue
|
|
1886
|
+
return b
|
|
1887
|
+
}
|
|
1888
|
+
await new Promise((r) => setTimeout(r, 1_000))
|
|
1889
|
+
}
|
|
1890
|
+
throw new Error(`no new fully-merged batch appeared on ${dacSlug} within 180s`)
|
|
1891
|
+
}
|
|
1892
|
+
|
|
1893
|
+
// Stage one — four legal entities.
|
|
1894
|
+
await ctx.ingestFileAndWait(ctx.identityDacSlug, 4)
|
|
1895
|
+
// Stage two — four table rows, stamped with the subjects stage one wrote.
|
|
1896
|
+
ctx.bulkBatchId = await ctx.ingestFileAndWait(ctx.dataDacSlug, 4)
|
|
1897
|
+
```
|
|
1898
|
+
|
|
1899
|
+
## 031 — every uploaded row landed, and every one is stamped
|
|
1900
|
+
|
|
1901
|
+
Both halves, scoped to the batch the worker allocated. The table side is read
|
|
1902
|
+
through read-only SQL; the identity side through the dataset search, where
|
|
1903
|
+
`le` is the legal-entity alias the query compiler expects.
|
|
1904
|
+
|
|
1905
|
+
The stamp check is the one that matters, and it is the same check §014 ran on
|
|
1906
|
+
the inline rows. Whether a row arrived in a JSON body or in a CSV, it is the
|
|
1907
|
+
same contract pair on the far side, and the proof that both contracts ran is a
|
|
1908
|
+
non-null `legal_entity_id` on a column no contract writes.
|
|
1909
|
+
|
|
1910
|
+
Note what is *not* asserted: an exact legal-entity count. This lake already
|
|
1911
|
+
holds subjects from the inline rows, and a fresh-tenant assumption is exactly
|
|
1912
|
+
the kind that goes red on the second run. The query is scoped to this batch,
|
|
1913
|
+
and four is the floor.
|
|
1914
|
+
|
|
1915
|
+
```typescript
|
|
1916
|
+
const sqlDeadline = Date.now() + 120_000
|
|
1917
|
+
let tableRows: unknown[][] = []
|
|
1918
|
+
while (Date.now() < sqlDeadline && tableRows.length < 4) {
|
|
1919
|
+
const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
|
|
1920
|
+
sql: `SELECT customer_number, legal_entity_id
|
|
1921
|
+
FROM ${ctx.customersTableName}
|
|
1922
|
+
WHERE customer_number IN ('CUS-10001', 'CUS-10002', 'CUS-10003', 'CUS-10004')
|
|
1923
|
+
ORDER BY customer_number`,
|
|
1924
|
+
mode: 'raw',
|
|
1925
|
+
})
|
|
1926
|
+
if (typeof result !== 'string') tableRows = result.data
|
|
1927
|
+
if (tableRows.length < 4) await new Promise((r) => setTimeout(r, 1_000))
|
|
1928
|
+
}
|
|
1929
|
+
if (tableRows.length !== 4) {
|
|
1930
|
+
throw new Error(`expected the 4 uploaded rows in the table, found ${tableRows.length} within 120s`)
|
|
1931
|
+
}
|
|
1932
|
+
const unstamped = tableRows.filter((row) => !row[1]).map((row) => String(row[0]))
|
|
1933
|
+
if (unstamped.length > 0) {
|
|
1934
|
+
throw new Error(`uploaded rows landed with no legal_entity_id stamp: ${unstamped.join(', ')}`)
|
|
1935
|
+
}
|
|
1936
|
+
|
|
1937
|
+
const { data: search } = await api.datasets.createUserSearch(tenantSlug, datalakeSlug, 'legal_entity', {
|
|
1938
|
+
search_query: `le.batch_id = '${ctx.bulkBatchId}'`,
|
|
1939
|
+
})
|
|
1940
|
+
if (search.status !== 'completed') {
|
|
1941
|
+
throw new Error(`legal_entity user-search did not compile: ${search.status}`)
|
|
1942
|
+
}
|
|
1943
|
+
const { data: page } = await api.datasets.search(tenantSlug, datalakeSlug, 'legal_entity', {
|
|
1944
|
+
userSearchId: search.id!,
|
|
1945
|
+
dataAccessMode: 'raw',
|
|
1946
|
+
})
|
|
1947
|
+
if (!Array.isArray(page.data)) {
|
|
1948
|
+
throw new Error('legal_entity search returned no rows array')
|
|
1949
|
+
}
|
|
1950
|
+
```
|
|
1951
|
+
|
|
1952
|
+
## 032 — the REST tool: let the platform do the fetching
|
|
1953
|
+
|
|
1954
|
+
The third way in does not wait for a file or a call. A **pull-based** client
|
|
1955
|
+
goes and gets the rows itself, on a schedule or on demand, and the platform
|
|
1956
|
+
holds the credential rather than your code.
|
|
1957
|
+
|
|
1958
|
+
The tool carries the API's base URL and its auth. `auth_method: 'bearer'`
|
|
1959
|
+
with a static token is the simplest case; the OAuth2 shape is §036.
|
|
1960
|
+
`status: 'active'` is not decoration — a draft tool is skipped at fetch time,
|
|
1961
|
+
and the symptom is a run that reports success and ingests nothing.
|
|
1962
|
+
|
|
1963
|
+
Locally the target is the integration stack's WireMock on `:8080`, which
|
|
1964
|
+
answers `GET /stripe/v1/customers` with three customers: two individuals and
|
|
1965
|
+
one enterprise, numbered `CUS-20001`–`CUS-20003`.
|
|
1966
|
+
|
|
1967
|
+
```typescript
|
|
1968
|
+
const restTool = await ctx.ensure(
|
|
1969
|
+
'Stripe REST tool',
|
|
1970
|
+
async () => {
|
|
1971
|
+
const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
1972
|
+
return ctx.firstNamed(data.data, 'Cookbook Stripe REST Tool')
|
|
1973
|
+
},
|
|
1974
|
+
async () => {
|
|
1975
|
+
const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
|
|
1976
|
+
name: 'Cookbook Stripe REST Tool',
|
|
1977
|
+
description: 'Mocked Stripe REST API — the pull-based customer fetch.',
|
|
1978
|
+
intent: 'data_exchange',
|
|
1979
|
+
status: 'active',
|
|
1980
|
+
datalake_id: ctx.datalakeId,
|
|
1981
|
+
data_source_id: ctx.dataSourceId,
|
|
1982
|
+
body: {
|
|
1983
|
+
tool_body_type: 'rest_api',
|
|
1984
|
+
auth_method: 'bearer',
|
|
1985
|
+
base_url: 'http://localhost:8080/stripe',
|
|
1986
|
+
bearer_token: 'sk_test_vitest_stripe_token',
|
|
1987
|
+
request_type: 'json',
|
|
1988
|
+
response_type: 'json',
|
|
1989
|
+
timeout_ms: 30_000,
|
|
1990
|
+
},
|
|
1991
|
+
})
|
|
1992
|
+
return data
|
|
1993
|
+
},
|
|
1994
|
+
)
|
|
1995
|
+
ctx.restToolId = restTool.id!
|
|
1996
|
+
```
|
|
1997
|
+
|
|
1998
|
+
## 033 — two fetch clients, for the same reason as before
|
|
1999
|
+
|
|
2000
|
+
The client's `tool_call` is what turns a REST tool into a fetch: the HTTP
|
|
2001
|
+
`method` and the `path` to append to the tool's `base_url`, plus a
|
|
2002
|
+
`pagination_context_template` declaring whether there is more to get. This
|
|
2003
|
+
endpoint answers in one page, so it declares `has_next: false`.
|
|
2004
|
+
|
|
2005
|
+
The `response_extractor` unwraps the API's envelope. Stripe answers
|
|
2006
|
+
`{ object: "list", data: [...] }`, and each element of that array has to
|
|
2007
|
+
become one row — `{{ msg.data | to_json }}` is the whole of it.
|
|
2008
|
+
|
|
2009
|
+
And the pair is split across two clients again. Nothing about the fetch
|
|
2010
|
+
changes the reason: two contracts on one client resolve the same subject at
|
|
2011
|
+
the same instant, and one of them loses. Both clients hit the same endpoint
|
|
2012
|
+
and get the same three customers; they differ only in what they do with them.
|
|
2013
|
+
|
|
2014
|
+
```typescript
|
|
2015
|
+
const restFetchCall = {
|
|
2016
|
+
tool_call_type: 'restapi_request' as const,
|
|
2017
|
+
method: 'get' as const,
|
|
2018
|
+
path: { type: 'custom' as const, body: '/v1/customers' },
|
|
2019
|
+
pagination_context_template: { type: 'custom' as const, body: '{"has_next": false}' },
|
|
2020
|
+
}
|
|
2021
|
+
const restExtractor = { type: 'custom' as const, body: '{{ msg.data | to_json }}' }
|
|
2022
|
+
|
|
2023
|
+
const restIdentityDac = await ctx.ensure(
|
|
2024
|
+
'REST identity client',
|
|
2025
|
+
async () => {
|
|
2026
|
+
const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
2027
|
+
return ctx.firstNamed(data.data, 'Cookbook Stripe REST Identity DAC')
|
|
2028
|
+
},
|
|
2029
|
+
async () => {
|
|
2030
|
+
const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
|
|
2031
|
+
name: 'Cookbook Stripe REST Identity DAC',
|
|
2032
|
+
description: 'Fetches Stripe customers and writes their legal entities. Runs FIRST.',
|
|
2033
|
+
tool_id: ctx.restToolId,
|
|
2034
|
+
data_source_id: ctx.dataSourceId,
|
|
2035
|
+
tool_call: restFetchCall,
|
|
2036
|
+
response_extractor: restExtractor,
|
|
2037
|
+
interoperability_contract_ids: [ctx.leContractId],
|
|
2038
|
+
})
|
|
2039
|
+
return data
|
|
2040
|
+
},
|
|
2041
|
+
)
|
|
2042
|
+
ctx.restIdentityDacSlug = restIdentityDac.slug!
|
|
2043
|
+
|
|
2044
|
+
const restDataDac = await ctx.ensure(
|
|
2045
|
+
'REST data client',
|
|
2046
|
+
async () => {
|
|
2047
|
+
const { data } = await api.dataActivationClients.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
2048
|
+
return ctx.firstNamed(data.data, 'Cookbook Stripe REST Data DAC')
|
|
2049
|
+
},
|
|
2050
|
+
async () => {
|
|
2051
|
+
const { data } = await api.dataActivationClients.create(tenantSlug, datalakeSlug, {
|
|
2052
|
+
name: 'Cookbook Stripe REST Data DAC',
|
|
2053
|
+
description: 'Fetches the same customers and writes the table rows. Runs SECOND.',
|
|
2054
|
+
tool_id: ctx.restToolId,
|
|
2055
|
+
data_source_id: ctx.dataSourceId,
|
|
2056
|
+
tool_call: restFetchCall,
|
|
2057
|
+
response_extractor: restExtractor,
|
|
2058
|
+
interoperability_contract_ids: [ctx.gtContractId],
|
|
2059
|
+
})
|
|
2060
|
+
return data
|
|
2061
|
+
},
|
|
2062
|
+
)
|
|
2063
|
+
ctx.restDataDacSlug = restDataDac.slug!
|
|
2064
|
+
```
|
|
2065
|
+
|
|
2066
|
+
## 034 — trigger both fetches, in order
|
|
2067
|
+
|
|
2068
|
+
`runManually` differs from `ingest` and `ingestFile` in a useful way: it
|
|
2069
|
+
returns the `batch_id` immediately, because the platform allocates it when it
|
|
2070
|
+
schedules the fetch rather than when a worker opens a file. The HTTP call and
|
|
2071
|
+
the ingestion still happen in the background.
|
|
2072
|
+
|
|
2073
|
+
Same staging, same reason. The identity fetch is waited on to completion
|
|
2074
|
+
before the data fetch is triggered.
|
|
2075
|
+
|
|
2076
|
+
```typescript
|
|
2077
|
+
const { data: identityRun } = await api.dataActivationClients.runManually(
|
|
2078
|
+
tenantSlug, datalakeSlug, ctx.restIdentityDacSlug,
|
|
2079
|
+
)
|
|
2080
|
+
if (!identityRun.batch_id) throw new Error('the identity fetch did not return a batch_id')
|
|
2081
|
+
await ctx.waitForBatches(ctx.restIdentityDacSlug, [identityRun.batch_id], 180_000)
|
|
2082
|
+
|
|
2083
|
+
const { data: dataRun } = await api.dataActivationClients.runManually(
|
|
2084
|
+
tenantSlug, datalakeSlug, ctx.restDataDacSlug,
|
|
2085
|
+
)
|
|
2086
|
+
if (!dataRun.batch_id) throw new Error('the data fetch did not return a batch_id')
|
|
2087
|
+
ctx.restBatchId = dataRun.batch_id
|
|
2088
|
+
await ctx.waitForBatches(ctx.restDataDacSlug, [ctx.restBatchId], 180_000)
|
|
2089
|
+
```
|
|
2090
|
+
|
|
2091
|
+
## 035 — three fetched customers, in the same table as everything else
|
|
2092
|
+
|
|
2093
|
+
This is the point of the whole scenario, and it is worth saying plainly:
|
|
2094
|
+
these three rows are indistinguishable from the four that arrived inline and
|
|
2095
|
+
the four that arrived in a CSV. Same table, same contracts, same stamp. The
|
|
2096
|
+
entry point is the only thing that differed, and it stopped mattering the
|
|
2097
|
+
moment the row was extracted.
|
|
2098
|
+
|
|
2099
|
+
The read filters on `customer_number`, a `none` column. The tokenized columns
|
|
2100
|
+
would answer here too — this lake has no derived copy, so raw is all there is
|
|
2101
|
+
— but filtering on a column declared `tokenize` is a habit that breaks the day
|
|
2102
|
+
someone provisions the pair.
|
|
2103
|
+
|
|
2104
|
+
```typescript
|
|
2105
|
+
const sqlDeadline = Date.now() + 120_000
|
|
2106
|
+
let fetched: unknown[][] = []
|
|
2107
|
+
while (Date.now() < sqlDeadline && fetched.length < 3) {
|
|
2108
|
+
const { data: result } = await api.datalakes.executeSql(tenantSlug, datalakeSlug, {
|
|
2109
|
+
sql: `SELECT customer_number, customer_type, legal_entity_id
|
|
2110
|
+
FROM ${ctx.customersTableName}
|
|
2111
|
+
WHERE customer_number IN ('CUS-20001', 'CUS-20002', 'CUS-20003')
|
|
2112
|
+
ORDER BY customer_number`,
|
|
2113
|
+
mode: 'raw',
|
|
2114
|
+
})
|
|
2115
|
+
if (typeof result !== 'string') fetched = result.data
|
|
2116
|
+
if (fetched.length < 3) await new Promise((r) => setTimeout(r, 1_000))
|
|
2117
|
+
}
|
|
2118
|
+
if (fetched.length !== 3) {
|
|
2119
|
+
throw new Error(`expected the 3 fetched rows in the table, found ${fetched.length} within 120s`)
|
|
2120
|
+
}
|
|
2121
|
+
const unstamped = fetched.filter((row) => !row[2]).map((row) => String(row[0]))
|
|
2122
|
+
if (unstamped.length > 0) {
|
|
2123
|
+
throw new Error(`fetched rows landed with no legal_entity_id stamp: ${unstamped.join(', ')}`)
|
|
2124
|
+
}
|
|
2125
|
+
const enterprise = fetched.find((row) => String(row[1]) === 'enterprise')
|
|
2126
|
+
if (!enterprise) {
|
|
2127
|
+
throw new Error('the fetch did not produce the enterprise customer the endpoint returns')
|
|
2128
|
+
}
|
|
2129
|
+
```
|
|
2130
|
+
|
|
2131
|
+
## 036 — the OAuth2 variant, for an API that will not take a static token
|
|
2132
|
+
|
|
2133
|
+
Most real APIs will not. When the credential is an OAuth2 grant, the tool
|
|
2134
|
+
carries the whole grant config instead of a token, and **the platform
|
|
2135
|
+
exchanges the stored refresh token for an access token server-side before
|
|
2136
|
+
each fetch**. Your code never holds one, never refreshes one, never logs one.
|
|
2137
|
+
|
|
2138
|
+
`oauth2_grant_type` names the flow the credential came from —
|
|
2139
|
+
`authorization_code` here, the ordinary three-legged flow whose long-lived
|
|
2140
|
+
refresh token you keep. It is not `refresh_token`: the enum is
|
|
2141
|
+
`client_credentials | authorization_code`, and refreshing is what the
|
|
2142
|
+
platform does with the grant rather than a grant of its own.
|
|
2143
|
+
|
|
2144
|
+
Everything else is unchanged — the client, its `tool_call`, the contracts,
|
|
2145
|
+
the table. Swapping auth is a tool-level edit.
|
|
2146
|
+
|
|
2147
|
+
This step creates the tool to show the shape and stops there; the mock
|
|
2148
|
+
endpoint has no OAuth2 server behind it, and a fetch that cannot succeed
|
|
2149
|
+
proves less than a create that validates.
|
|
2150
|
+
|
|
2151
|
+
```typescript
|
|
2152
|
+
const oauthTool = await ctx.ensure(
|
|
2153
|
+
'OAuth2 REST tool',
|
|
2154
|
+
async () => {
|
|
2155
|
+
const { data } = await api.tools.list(tenantSlug, datalakeSlug, { page_size: 100 })
|
|
2156
|
+
return ctx.firstNamed(data.data, 'Cookbook Stripe OAuth2 Tool')
|
|
2157
|
+
},
|
|
2158
|
+
async () => {
|
|
2159
|
+
const { data } = await api.tools.create(tenantSlug, datalakeSlug, {
|
|
2160
|
+
name: 'Cookbook Stripe OAuth2 Tool',
|
|
2161
|
+
description: 'Shape reference — the same fetch behind an OAuth2 refresh-token grant.',
|
|
2162
|
+
intent: 'data_exchange',
|
|
2163
|
+
status: 'active',
|
|
2164
|
+
datalake_id: ctx.datalakeId,
|
|
2165
|
+
data_source_id: ctx.dataSourceId,
|
|
2166
|
+
body: {
|
|
2167
|
+
tool_body_type: 'rest_api',
|
|
2168
|
+
auth_method: 'oauth2',
|
|
2169
|
+
base_url: 'http://localhost:8080/stripe',
|
|
2170
|
+
oauth2_grant_type: 'authorization_code',
|
|
2171
|
+
oauth2_token_url: 'http://localhost:8080/oauth/token',
|
|
2172
|
+
oauth2_client_id: 'cookbook-client',
|
|
2173
|
+
oauth2_client_secret: 'cookbook-secret',
|
|
2174
|
+
oauth2_refresh_token: 'cookbook-refresh-token',
|
|
2175
|
+
oauth2_token_ttl: 3300,
|
|
2176
|
+
request_type: 'json',
|
|
2177
|
+
response_type: 'json',
|
|
2178
|
+
timeout_ms: 30_000,
|
|
2179
|
+
},
|
|
2180
|
+
})
|
|
2181
|
+
return data
|
|
2182
|
+
},
|
|
2183
|
+
)
|
|
2184
|
+
if (!oauthTool.id) {
|
|
2185
|
+
throw new Error('the OAuth2 tool was not persisted')
|
|
2186
|
+
}
|
|
2187
|
+
```
|
|
2188
|
+
|
|
2189
|
+
## 037 — a billing team is not one person
|
|
2190
|
+
|
|
2191
|
+
Everything so far ran as one admin. The last scenario adds a second person to
|
|
2192
|
+
the tenant, entirely through the SDK — no console, no out-of-band step.
|
|
2193
|
+
|
|
2194
|
+
The flow touches all three session scopes, and mixing them up is the single
|
|
2195
|
+
most common mistake here:
|
|
2196
|
+
|
|
2197
|
+
- **tenant-scoped** — the admin creates the invitation. Only an admin can.
|
|
2198
|
+
- **root** — signs the new user up and confirms them. A tenant admin can
|
|
2199
|
+
invite, but only root can create the underlying account.
|
|
2200
|
+
- **tenantless** — the invitee, who belongs to no tenant yet, lists the
|
|
2201
|
+
invitations waiting for them and accepts one.
|
|
2202
|
+
|
|
2203
|
+
Start by confirming the current session really is a tenant-scoped admin of
|
|
2204
|
+
the expected tenant. `sessions.verify` echoes both.
|
|
2205
|
+
|
|
2206
|
+
```typescript
|
|
2207
|
+
const { data: who } = await api.sessions.verify()
|
|
2208
|
+
if (who.tenant?.slug !== tenantSlug) {
|
|
2209
|
+
throw new Error(`expected an admin session on ${tenantSlug}, got ${who.tenant?.slug}`)
|
|
2210
|
+
}
|
|
2211
|
+
if (!/admin/i.test(who.role?.name ?? '')) {
|
|
2212
|
+
throw new Error(`expected an admin role, got ${who.role?.name}`)
|
|
2213
|
+
}
|
|
2214
|
+
```
|
|
2215
|
+
|
|
2216
|
+
## 038 — the teammate's account, created by root
|
|
2217
|
+
|
|
2218
|
+
The invitee needs an account before they can accept anything. Signup and
|
|
2219
|
+
confirmation are root-scoped, so this runs on the root client from §001 —
|
|
2220
|
+
`confirmUser` is what activates the account so they can sign in at all.
|
|
2221
|
+
|
|
2222
|
+
The email is stable, like every other name in this walk, so the second run
|
|
2223
|
+
finds the account already there. Same catch-and-inspect as §001: a refusal
|
|
2224
|
+
that says the address is taken is the idempotent path, and anything else is
|
|
2225
|
+
re-raised.
|
|
2226
|
+
|
|
2227
|
+
```typescript
|
|
2228
|
+
ctx.teammateEmail = 'cookbook-subscription-teammate@dev.local'
|
|
2229
|
+
ctx.teammatePassword = 'CookbookPass1!'
|
|
2230
|
+
|
|
2231
|
+
try {
|
|
2232
|
+
const { data: teammate } = await ctx.rootApi.admin.signUp({
|
|
2233
|
+
email: ctx.teammateEmail,
|
|
2234
|
+
password: ctx.teammatePassword,
|
|
2235
|
+
first_name: 'Emma',
|
|
2236
|
+
last_name: 'Wilson',
|
|
2237
|
+
})
|
|
2238
|
+
await ctx.rootApi.admin.confirmUser(teammate.id!)
|
|
2239
|
+
} catch (err) {
|
|
2240
|
+
const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
|
|
2241
|
+
if (!/taken|already|exist/i.test(detail)) throw err
|
|
2242
|
+
}
|
|
2243
|
+
```
|
|
2244
|
+
|
|
2245
|
+
## 039 — the invitation
|
|
2246
|
+
|
|
2247
|
+
The admin invites by email and role. The invite is keyed on
|
|
2248
|
+
`(tenant, email)`, so a re-invitation is refused — and it is refused with
|
|
2249
|
+
**two different messages depending on how far the first run got**. While the
|
|
2250
|
+
invitation is still pending, `/email` says *"already invited"*. Once it has
|
|
2251
|
+
been accepted and consumed, the same call says *"already in this
|
|
2252
|
+
organization"*, because at that point the address does not belong to a
|
|
2253
|
+
pending invite but to a member.
|
|
2254
|
+
|
|
2255
|
+
Both are the idempotent path and both are caught. What is **not** caught is
|
|
2256
|
+
anything else, and that distinction is the whole reason this is a `try`
|
|
2257
|
+
rather than a bare call: an invitation that fails because the tenant is
|
|
2258
|
+
misconfigured must not be read as "already there".
|
|
2259
|
+
|
|
2260
|
+
```typescript
|
|
2261
|
+
try {
|
|
2262
|
+
const { data: invite } = await api.invitations.create(tenantSlug, {
|
|
2263
|
+
email: ctx.teammateEmail,
|
|
2264
|
+
role: 'member',
|
|
2265
|
+
})
|
|
2266
|
+
if (invite.email !== ctx.teammateEmail || invite.role !== 'member') {
|
|
2267
|
+
throw new Error(`unexpected invitation: ${JSON.stringify(invite)}`)
|
|
2268
|
+
}
|
|
2269
|
+
} catch (err) {
|
|
2270
|
+
const detail = JSON.stringify((err as { errors?: unknown }).errors ?? err)
|
|
2271
|
+
// Both refusals, verbatim from the server. Matching loosely here is how a
|
|
2272
|
+
// real failure gets swallowed, so the two known messages are named.
|
|
2273
|
+
if (!/already invited|already in this organization/i.test(detail)) throw err
|
|
2274
|
+
}
|
|
2275
|
+
```
|
|
2276
|
+
|
|
2277
|
+
## 040 — the teammate signs in with no tenant, and accepts
|
|
2278
|
+
|
|
2279
|
+
The invitee has an account and belongs nowhere, so they sign in **tenantless**
|
|
2280
|
+
— `createBootstrapSession`, not `createSession`, which is tenant-login only.
|
|
2281
|
+
The resulting session carries `tenant: null`, and that is the assertion worth
|
|
2282
|
+
making: a client that quietly came back tenant-scoped would list the wrong
|
|
2283
|
+
invitations.
|
|
2284
|
+
|
|
2285
|
+
Then the branch that makes this step idempotent. On the first run there is a
|
|
2286
|
+
pending invitation for this tenant and it is accepted. On every run after, the
|
|
2287
|
+
invitation has already been consumed and the list is empty — which is not a
|
|
2288
|
+
failure but the same end state reached earlier. **The proof is deferred to
|
|
2289
|
+
§041**, which asserts the end state directly rather than trusting either
|
|
2290
|
+
branch.
|
|
2291
|
+
|
|
2292
|
+
Note that accepting does not upgrade the tenantless session in place. It
|
|
2293
|
+
creates the membership; the session that uses it is minted fresh in the next
|
|
2294
|
+
step.
|
|
2295
|
+
|
|
2296
|
+
```typescript
|
|
2297
|
+
const teammateTenantless = await createBootstrapSession({
|
|
2298
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
2299
|
+
email: ctx.teammateEmail,
|
|
2300
|
+
password: ctx.teammatePassword,
|
|
2301
|
+
})
|
|
2302
|
+
if (teammateTenantless.tenant !== null) {
|
|
2303
|
+
throw new Error('expected a tenantless session — tenant should be null')
|
|
2304
|
+
}
|
|
2305
|
+
const teammateTenantlessApi = createIsolatedPlatformApi({
|
|
2306
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
2307
|
+
sessionToken: teammateTenantless.sessionToken,
|
|
2308
|
+
apiKey: '',
|
|
2309
|
+
})
|
|
2310
|
+
|
|
2311
|
+
const { data: invites } = await teammateTenantlessApi.invitations.list()
|
|
2312
|
+
const pending = (invites.data ?? []).find(
|
|
2313
|
+
(i: { tenant?: { slug?: string } }) => i.tenant?.slug === tenantSlug,
|
|
2314
|
+
)
|
|
2315
|
+
|
|
2316
|
+
if (pending?.id) {
|
|
2317
|
+
const { data: membership } = await teammateTenantlessApi.invitations.accept(pending.id)
|
|
2318
|
+
if (membership.tenant?.slug !== tenantSlug || membership.role !== 'member') {
|
|
2319
|
+
throw new Error(`unexpected membership: ${JSON.stringify(membership)}`)
|
|
2320
|
+
}
|
|
2321
|
+
} else {
|
|
2322
|
+
console.log(' ↻ no pending invitation — already accepted on an earlier run')
|
|
2323
|
+
}
|
|
2324
|
+
```
|
|
2325
|
+
|
|
2326
|
+
## 041 — the teammate signs in tenant-scoped, and is a member
|
|
2327
|
+
|
|
2328
|
+
This is the assertion the whole scenario is for, and it is written against the
|
|
2329
|
+
end state rather than against the path taken to reach it — which is what makes
|
|
2330
|
+
it identical on run one and run fifty.
|
|
2331
|
+
|
|
2332
|
+
The tenant-scoped sign-in needs the tenant's API key, the one minted in §004.
|
|
2333
|
+
Tenant login without a resolvable `X-API-Key` is a 401 regardless of how good
|
|
2334
|
+
the password is; the key identifies the tenant before the credentials
|
|
2335
|
+
identify the person.
|
|
2336
|
+
|
|
2337
|
+
The role is the other half. A session that came back with an admin role would
|
|
2338
|
+
mean the invitation's `role: 'member'` was ignored, and everything downstream
|
|
2339
|
+
that trusts the platform to enforce a role would be resting on nothing.
|
|
2340
|
+
|
|
2341
|
+
```typescript
|
|
2342
|
+
const teammateScoped = await createSession({
|
|
2343
|
+
baseUrl: process.env.ALVERA_BASE_URL!,
|
|
2344
|
+
email: ctx.teammateEmail,
|
|
2345
|
+
password: ctx.teammatePassword,
|
|
2346
|
+
tenantSlug,
|
|
2347
|
+
apiKey: ctx.tenantApiKey,
|
|
2348
|
+
})
|
|
2349
|
+
if (teammateScoped.tenant?.slug !== tenantSlug) {
|
|
2350
|
+
throw new Error(`expected a tenant-scoped session on ${tenantSlug}, got ${teammateScoped.tenant?.slug}`)
|
|
2351
|
+
}
|
|
2352
|
+
if (!/member/i.test(teammateScoped.role?.name ?? '')) {
|
|
2353
|
+
throw new Error(`expected a member role, got ${teammateScoped.role?.name}`)
|
|
2354
|
+
}
|
|
2355
|
+
```
|
|
2356
|
+
|
|
2357
|
+
# Outcome
|
|
2358
|
+
|
|
2359
|
+
After all forty-one steps run green, on one tenant and **one raw datalake**:
|
|
2360
|
+
|
|
2361
|
+
- **One customers table** holds eleven rows that arrived three different
|
|
2362
|
+
ways — four inline through an ingest call, four from a CSV through a
|
|
2363
|
+
presigned upload, three pulled from a REST API — and every one of them
|
|
2364
|
+
carries a `legal_entity_id` stamp written by a contract that never
|
|
2365
|
+
declares the column.
|
|
2366
|
+
- **Three workflows** run over that one table and disagree about which rows
|
|
2367
|
+
matter. The welcome workflow gates on reachability, the dunning workflow on
|
|
2368
|
+
reachability *and* KYC, and the triage workflow hands the question to a
|
|
2369
|
+
model and fans out on its answer.
|
|
2370
|
+
- **Three outbound messages** are readable from the `message` dataset, each
|
|
2371
|
+
under its own idempotency key, two carrying deep links that resolve against
|
|
2372
|
+
two separate connected apps on two separate routes.
|
|
2373
|
+
- **A second person** is on the tenant with a `member` role, invited and
|
|
2374
|
+
accepted entirely through the SDK.
|
|
2375
|
+
|
|
2376
|
+
Four things this walk teaches that are easy to learn the expensive way:
|
|
2377
|
+
|
|
2378
|
+
1. **Never bind a contract pair to one activation client.** The fan-out
|
|
2379
|
+
staggers by row, so a pair on one client resolves the same subject twice
|
|
2380
|
+
at the same instant and the unique index refuses the loser. Split them and
|
|
2381
|
+
sequence them — §012, and again at §030 and §033.
|
|
2382
|
+
2. **Gate on `dataset_updated`, not `rows_ingested`.** Received is not
|
|
2383
|
+
written. A refused row reports the first and not the second, and a walk
|
|
2384
|
+
gated on the wrong one goes green over an empty table.
|
|
2385
|
+
3. **Search the idempotency key, never the workflow id.** The key is built
|
|
2386
|
+
from a stable business id and names the same message forever; a workflow
|
|
2387
|
+
id changes the moment the resource is recreated.
|
|
2388
|
+
4. **Assert the end state, not the transition.** An idempotency key fires
|
|
2389
|
+
once ever, so every assertion here accepts both the first run's shape and
|
|
2390
|
+
the repeat run's — §027 spells that out, and §040 defers its proof to
|
|
2391
|
+
§041 for the same reason.
|
|
2392
|
+
|
|
2393
|
+
# See also
|
|
2394
|
+
|
|
2395
|
+
- `payments-compliance.md` — the same contract-pair shape on a tokenized
|
|
2396
|
+
surface, with the derived lakes provisioned
|
|
2397
|
+
- `organic-marketing.md` — three scenarios on one raw lake, the sibling of
|
|
2398
|
+
this walk
|
|
2399
|
+
- `../workflows.md` — the standard workflow primitive, `context_datasets`,
|
|
2400
|
+
and `workflows.run`
|
|
2401
|
+
- `../data_activation_clients.md` — data source → tool → contract → client
|
|
2402
|
+
- `../interoperability_contracts.md` — `template_config` vs `mdm_input_config`
|
|
2403
|
+
- `../connected_apps.md` — registration, `resolvePage`, message tracking
|