@alvera-ai/platform-sdk 0.10.0-rc.2 → 0.10.0-rc.21

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/.agent/AGENTS.md +440 -0
  2. package/.agent/account_management.md +455 -0
  3. package/.agent/action_status_updaters.md +262 -0
  4. package/.agent/ai_agents.md +423 -0
  5. package/.agent/ai_sandbox.md +265 -0
  6. package/.agent/async.md +111 -0
  7. package/.agent/connected_apps.md +407 -0
  8. package/.agent/cookbook/_fixtures/README.md +99 -0
  9. package/.agent/cookbook/_fixtures/accounts_receivable/_customers_accounts_receivable_customer.liquid +32 -0
  10. package/.agent/cookbook/_fixtures/accounts_receivable/_customers_accounts_receivable_mdm.liquid +20 -0
  11. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_generic_table.liquid +33 -0
  12. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_legal_entity.liquid +88 -0
  13. package/.agent/cookbook/_fixtures/foundation/_lead_submissions_foundation_mdm.liquid +48 -0
  14. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_appointment.liquid +47 -0
  15. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_mdm.liquid +24 -0
  16. package/.agent/cookbook/_fixtures/healthcare/_cahps_appointments_healthcare_patient.liquid +38 -0
  17. package/.agent/cookbook/_fixtures/payment_risk/_compliance_screenings_payment_risk_compliance_screening.liquid +59 -0
  18. package/.agent/cookbook/_fixtures/payment_risk/_compliance_screenings_payment_risk_mdm.liquid +36 -0
  19. package/.agent/cookbook/_fixtures/payment_risk/_payment_accounts_payment_risk_mdm.liquid +30 -0
  20. package/.agent/cookbook/_fixtures/payment_risk/_payment_accounts_payment_risk_payment_account.liquid +55 -0
  21. package/.agent/cookbook/_setup/accounts_receivable.md +282 -0
  22. package/.agent/cookbook/_setup/foundation.md +277 -0
  23. package/.agent/cookbook/_setup/healthcare.md +279 -0
  24. package/.agent/cookbook/_setup/payment_risk.md +283 -0
  25. package/.agent/cookbook/appointment-review-sms-workflow.md +761 -0
  26. package/.agent/cookbook/birthday-greeting-sms-trigger.md +656 -0
  27. package/.agent/cookbook/contact-us-triage-with-llm.md +603 -0
  28. package/.agent/cookbook/dunning-sms-for-delinquent.md +619 -0
  29. package/.agent/cookbook/kyc-notification-on-account-activation.md +619 -0
  30. package/.agent/cookbook/sanctions-screening-with-agent-review.md +711 -0
  31. package/.agent/cookbook/score-leads-with-llm-categorization.md +602 -0
  32. package/.agent/cookbook/welcome-sms-for-customers.md +607 -0
  33. package/.agent/data_activation_clients.md +557 -0
  34. package/.agent/data_sources.md +234 -0
  35. package/.agent/datalakes.md +712 -0
  36. package/.agent/debugging.md +137 -0
  37. package/.agent/errors.md +196 -0
  38. package/.agent/generic_tables.md +351 -0
  39. package/.agent/interoperability_contracts.md +351 -0
  40. package/.agent/mdm.md +293 -0
  41. package/.agent/mutations.md +152 -0
  42. package/.agent/templates.md +98 -0
  43. package/.agent/tool-call-configs.md +90 -0
  44. package/.agent/tools.md +546 -0
  45. package/.agent/type_naming.md +131 -0
  46. package/.agent/workflows.md +601 -0
  47. package/README.md +46 -0
  48. package/dist/bin/platform-sdk.d.mts +1 -0
  49. package/dist/bin/platform-sdk.mjs +106 -0
  50. package/dist/bin/platform-sdk.mjs.map +1 -0
  51. package/dist/index.d.mts +1200 -43201
  52. package/dist/index.d.mts.map +1 -1
  53. package/dist/index.mjs +1859 -7319
  54. package/dist/index.mjs.map +1 -1
  55. package/package.json +19 -10
@@ -0,0 +1,351 @@
1
+ # Interoperability contracts
2
+
3
+ An **interoperability contract** is the transformation step between
4
+ a raw inbound row (CSV, JSON, etc.) and a typed datalake resource.
5
+ Each contract carries:
6
+
7
+ - a `resource_type` (the target shape — `patient`, `customer`,
8
+ `payment_account`, `legal_entity`, `generic_table`, etc.)
9
+ - a `template_config` (the Liquid template that renders the inbound
10
+ row into the target shape)
11
+ - an `mdm_input_config` (the Liquid template that renders the
12
+ matching subject identity for entity resolution — see `mdm.md`
13
+ for the full MDM surface, including the `api.mdm.verify` verb
14
+ invoked at workflow action time)
15
+ - an optional `filter_template` (a Liquid pre-filter that decides
16
+ whether a row enters the pipeline at all)
17
+
18
+ SDK namespace: `api.interoperabilityContracts`.
19
+
20
+ The contract is **Datalake-DB-resident** — POSTing requires the
21
+ parent datalake at `status: 'ready'` (see `datalakes.md` §5).
22
+
23
+ ## 1. Wire shape
24
+
25
+ ```typescript
26
+ import type {
27
+ InteroperabilityContractRequestWritable,
28
+ InteroperabilityContractResponse,
29
+ } from '@alvera-ai/platform-sdk'
30
+
31
+ const { data: created } = await api.interoperabilityContracts.create(
32
+ tenantSlug,
33
+ datalakeSlug,
34
+ {
35
+ name: 'Inbound CRM → Customer',
36
+ description: 'Vendor CRM rows → AccountsReceivable Customer',
37
+ resource_type: 'customer',
38
+ filter_template: undefined, // optional
39
+ template_config: { type: 'custom', body: liquidBody },
40
+ mdm_input_config: { type: 'custom', body: mdmBody },
41
+ generic_table_id: null, // see §2
42
+ },
43
+ )
44
+ // created.id, created.slug — server-derived
45
+ // created.type — mirrors template_config.type
46
+ // created.resource_type === body.resource_type
47
+ ```
48
+
49
+ Both polymorphic embeds use the SAME shared `TemplateConfig`
50
+ shape with discriminator `type`:
51
+
52
+ ```
53
+ template_config.type 'system' | 'custom' | 'identity' | 'null'
54
+ mdm_input_config.type 'system' | 'custom' | 'identity' | 'null'
55
+ ```
56
+
57
+ Default for both is `'null'`. The wire accepts all four values
58
+ on each surface; semantic narrowing is a runtime concern, not
59
+ a schema one:
60
+
61
+ - `custom` — caller-supplied `body` string (Liquid template).
62
+ Pair with `body`; leave `path` null.
63
+ - `system` — references a platform-shipped template by
64
+ filesystem path. Pair with `path` (required); leave `body`
65
+ null. See `templates.md`.
66
+ - `identity` — pass-through. The inbound row is forwarded
67
+ unchanged. Used by the auto-created identity contracts for
68
+ generic tables. Leave both `body` and `path` null.
69
+ - `null` — disabled / "row IS the subject" on the mdm surface.
70
+ Leave both `body` and `path` null.
71
+
72
+ ### TemplateConfig embed fields
73
+
74
+ Both `template_config` and `mdm_input_config` are instances of
75
+ the shared `TemplateConfig` embed:
76
+
77
+ ```
78
+ type required enum — 'system' | 'custom' | 'identity' | 'null' (default 'null')
79
+ body optional string — Liquid template body; REQUIRED iff type === 'custom'
80
+ path optional string — filesystem path; REQUIRED iff type === 'system'
81
+ output_schema server-injected — JSON Schema { type: 'object' } for this surface; not caller-set
82
+ ```
83
+
84
+ Server validation enforces the body/path pairing — supplying
85
+ `body` with `type: 'system'`, or `path` with `type: 'custom'`,
86
+ returns 422.
87
+
88
+ ## 2. Rules the type cannot encode
89
+
90
+ ### `generic_table_id` is conditionally required
91
+
92
+ When `resource_type === 'generic_table'`, the body MUST include a
93
+ non-null `generic_table_id` identifying the target table. For
94
+ every other `resource_type`, the field must be `null`. The
95
+ TypeScript type permits either value; the validator and server
96
+ enforce the conditional rule.
97
+
98
+ ### `resource_type` is industry-scoped
99
+
100
+ The valid set of `resource_type` values is determined by the
101
+ parent datalake's `data_domain`. Generic concepts like
102
+ `'generic_table'` are accepted everywhere; domain-anchored
103
+ concepts (`'patient'`, `'customer'`, `'payment_account'`,
104
+ `'legal_entity'`, ...) only resolve within their owning industry.
105
+ Don't hardcode a fixed enum across industries — discover the
106
+ valid set at runtime via the platform's domain catalog (see
107
+ `datalakes.md` §5 `.systemDatasets`).
108
+
109
+ ### `filter_template` has INVERTED semantics
110
+
111
+ This is the single highest-value gotcha on this surface. A
112
+ `filter_template` is a Liquid template that decides whether a
113
+ row enters the pipeline:
114
+
115
+ - **Empty render → PASS** (row enters the pipeline).
116
+ - **Non-empty render → SKIP** (row is dropped).
117
+
118
+ This is the opposite of most filter DSLs ("truthy = keep"). A
119
+ correctly-shaped filter looks like:
120
+
121
+ ```liquid
122
+ {% if msg.row.should_drop == 'Y' %}drop{% endif %}
123
+ ```
124
+
125
+ Renders `"drop"` when the condition fires (SKIP); renders empty
126
+ otherwise (PASS). Inverting `{% unless %}...{% endunless %}` is
127
+ equivalent and stylistically common — but always anchor on the
128
+ empty-render-passes rule, not on the Liquid keyword.
129
+
130
+ If you omit `filter_template`, every row passes through.
131
+
132
+ ### `template_config` body is Liquid, with row data on `msg`
133
+
134
+ Templates receive the inbound row on the `msg` variable. The
135
+ exact accessor depends on the source shape — CSV-shaped sources
136
+ typically use `msg.<column>` (with `msg.row.<column>` available
137
+ for nested cases); JSON-shaped sources mirror the inbound JSON
138
+ under `msg.<field>`. The render output is JSON — the platform
139
+ parses the rendered string as JSON and treats it as the target
140
+ resource body.
141
+
142
+ A render that produces invalid JSON fails at sandbox-run with
143
+ `stage: 'failed'` (see §5). Always sandbox-run a contract before
144
+ binding it to a live data activation client.
145
+
146
+ ## 3. Field ownership
147
+
148
+ **Server-derived (Response-only).** Universal set from
149
+ `type_naming.md`, plus:
150
+
151
+ ```
152
+ type string — mirrors template_config.type at create time
153
+ ```
154
+
155
+ **Caller-supplied (round-trip).**
156
+
157
+ ```
158
+ name required string — agent-facing label
159
+ description optional string — narrative
160
+ resource_type required string — target shape (industry-scoped)
161
+ filter_template optional string — Liquid pre-filter (inverted semantics)
162
+ template_config required embed — { type, body?, path? }
163
+ mdm_input_config optional embed — { type, body?, path? }; omit equals { type: 'null' }
164
+ generic_table_id required UUID|null — required iff resource_type='generic_table'
165
+ ```
166
+
167
+ **Write-only (Request-only).** None — every caller-supplied field
168
+ round-trips on the response.
169
+
170
+ ## 4. Error envelopes
171
+
172
+ Standard JSON:API envelopes per `errors.md`. Common rejections:
173
+
174
+ | `source.pointer` | Cause |
175
+ |------------------------------|--------------------------------------------------|
176
+ | `/resource_type` | value not valid for the parent datalake's domain |
177
+ | `/template_config/type` | enum mismatch |
178
+ | `/template_config/body` | empty when `type === 'custom'` |
179
+ | `/mdm_input_config/type` | enum mismatch |
180
+ | `/generic_table_id` | missing when resource_type='generic_table', |
181
+ | | or non-null when not |
182
+ | `/name` | uniqueness within datalake |
183
+ | `errors.base` | "System-created contracts cannot be edited" (PUT) or "...deleted" (DELETE) — see §5 |
184
+
185
+ ## 5. Lifecycle
186
+
187
+ ### Create
188
+
189
+ Synchronous. The create response is immediately usable — no async
190
+ deployment phase. Templates are compiled lazily at sandbox-run
191
+ time and at activation-client fan-out time, so a `body`
192
+ containing invalid Liquid syntax will surface at first use, not
193
+ at create. Always sandbox-run before binding to a live data
194
+ activation client.
195
+
196
+ ### Sandbox run — `.run(slug, msg)`
197
+
198
+ The contract's templates can be exercised against a sample row
199
+ WITHOUT actually ingesting:
200
+
201
+ ```typescript
202
+ const { data: result } = await api.interoperabilityContracts.run(
203
+ tenantSlug, datalakeSlug, contractSlug,
204
+ { /* sample row, same shape as inbound CSV/JSON */ },
205
+ )
206
+ // result.stage === 'completed' | 'failed'
207
+ // result.filter_result === 'pass' | 'skip'
208
+ // result.transformed object | null — the rendered resource body
209
+ ```
210
+
211
+ Use this verb as a CI / pre-bind verification step: produce a
212
+ representative sample row, run it through `.run`, and assert
213
+ the `transformed` output shape matches the consumer's
214
+ expectations. If `filter_result === 'skip'`, the row was filtered
215
+ out before rendering — useful for verifying the filter semantics
216
+ on a known-skip example.
217
+
218
+ ### Generic-table contracts: auto-created identity contracts
219
+
220
+ When you create a generic table (see `generic_tables.md`), the
221
+ platform automatically creates an **identity contract** bound
222
+ to that table:
223
+
224
+ ```
225
+ new generic table (T) → auto-created contract with
226
+ resource_type: 'generic_table'
227
+ generic_table_id: T.id
228
+ template_config.type: 'identity'
229
+ ```
230
+
231
+ Discover this contract via `.list()` + filter on
232
+ `generic_table_id`:
233
+
234
+ ```typescript
235
+ const { data: all } = await api.interoperabilityContracts.list(
236
+ tenantSlug, datalakeSlug,
237
+ )
238
+ const identityContract = all.data?.find(
239
+ (c) => c.generic_table_id === tableId && c.type === 'identity',
240
+ )
241
+ ```
242
+
243
+ Multiple contracts MAY bind to the same generic table — one
244
+ identity (auto) plus any number of caller-created `custom`
245
+ contracts. Distinguish by `type` when iterating. The platform
246
+ enforces at most ONE identity contract per generic table;
247
+ attempts to create a second `type: 'identity'` contract for
248
+ the same table are rejected.
249
+
250
+ ### Read shapes
251
+
252
+ ```
253
+ .list(tenantSlug, datalakeSlug) Paged: { data, meta }
254
+ .get(tenantSlug, datalakeSlug, slug) One row, by slug
255
+ .metadata(tenantSlug, datalakeSlug, slug) Markdown string — one contract
256
+ .run(tenantSlug, datalakeSlug, slug, msg) Sandbox runner (above)
257
+ ```
258
+
259
+ ### Delete
260
+
261
+ `DELETE` removes the contract. Contracts referenced by active
262
+ data activation clients reject with a 422 naming the dependency.
263
+
264
+ **System-created contracts cannot be edited or deleted.** The
265
+ contract record carries a `system_created` flag (server-derived).
266
+ When `system_created === true`, both PUT and DELETE reject with
267
+ 422 on `/base` and a literal message:
268
+
269
+ ```
270
+ PUT → "System-created contracts cannot be edited"
271
+ DELETE → "System-created contracts cannot be deleted"
272
+ ```
273
+
274
+ System-created contracts include the auto-generated identity
275
+ contracts the platform provisions when a generic table is
276
+ created (see §5 "Generic-table contracts: auto-created identity
277
+ contracts") and any industry-built templates the platform ships
278
+ with the datalake's data_domain. User-created contracts of any
279
+ `type` (including caller-created `type: 'identity'` rows bound
280
+ to a generic table) ARE editable and deletable — the gate is
281
+ `system_created`, not `type`.
282
+
283
+ The flag IS exposed on the `InteroperabilityContractResponse`,
284
+ so consumer agents can pre-filter before attempting a mutation:
285
+
286
+ ```typescript
287
+ const { data: list } = await api.interoperabilityContracts.list(
288
+ tenantSlug, datalakeSlug,
289
+ )
290
+ const editable = list.data.filter((c) => !c.system_created)
291
+ ```
292
+
293
+ ## 6. Gotchas
294
+
295
+ 1. **`filter_template` is empty-passes, non-empty-skips.**
296
+ Already documented in §2 — repeated here because it's the
297
+ #1 source of "ingests nothing / ingests too much" silent
298
+ failures. Always anchor on this rule.
299
+
300
+ 2. **`mdm_input_config` is optional; omitting equals `{ type:
301
+ 'null' }`.** The shared TemplateConfig embed defaults `type`
302
+ to `'null'`, so a contract that omits `mdm_input_config`
303
+ behaves as "the row IS the subject — no entity resolution
304
+ runs". Declaring `{ type: 'null' }` explicitly is allowed
305
+ for readability but not required.
306
+
307
+ 3. **`resource_type` is industry-scoped.** A datalake with
308
+ `data_domain: 'foundation'` accepts `'legal_entity'` but
309
+ rejects `'patient'`. The error surfaces as
310
+ `/resource_type` 422 at create time. Don't ship a fixed
311
+ enum across industries.
312
+
313
+ 4. **Sandbox-run is not optional in practice.** A template that
314
+ compiles but renders structurally-invalid JSON, or that
315
+ silently drops fields, won't surface until live ingestion —
316
+ and at that point the failures are batched and harder to
317
+ debug. Run `.run(slug, sampleRow)` against a representative
318
+ row for every new contract before binding it to a data
319
+ activation client.
320
+
321
+ 5. **Generic-table identity contracts are auto-created; don't
322
+ create them yourself.** Creating a generic table fires a
323
+ side-effect that creates a `type: 'identity'` contract bound
324
+ to the new table. Posting a second identity contract for the
325
+ same table is rejected. To add custom transformation on top
326
+ of a generic table, create an additional `type: 'custom'`
327
+ contract bound to the same `generic_table_id`.
328
+
329
+ 6. **Templates use `msg` to access the row.** Liquid receives
330
+ the row under the `msg` variable. CSV columns are typically
331
+ `msg.<column>`; nested cases use `msg.row.<column>`. JSON
332
+ sources expose the inbound fields under `msg.<field>`. Read
333
+ any platform-shipped template (see `templates.md`) to learn
334
+ the exact convention for your row shape.
335
+
336
+ 7. **No partial updates.** PUT replays the full body per
337
+ `mutations.md`; there's no PATCH affordance. Re-supply
338
+ `template_config` / `mdm_input_config` verbatim if not
339
+ changing them.
340
+
341
+ 8. **Custom Liquid filters are a closed allowlist of eleven.**
342
+ The platform exposes exactly eleven custom filters beyond
343
+ stock Liquid (`e164`, `age`, `date`, `to_json`,
344
+ `json_escape`, `now`, `uuid`, `parse_date`, `convert_time`,
345
+ `random`, `tz_offset`). Other filter names — including the
346
+ Shopify-style `| json` — fall through to stock Liquid and
347
+ produce malformed output silently. See `ai_sandbox.md` §2
348
+ for the full per-filter reference and the safety rationale
349
+ for the closed set. The allowlist is the same in
350
+ `template_config.body`, `mdm_input_config.body`, and
351
+ `filter_template`.
package/.agent/mdm.md ADDED
@@ -0,0 +1,293 @@
1
+ # MDM — entity resolution + identity verification
2
+
3
+ Master Data Management is the platform's entity-resolution layer.
4
+ It binds every inbound row to a canonical **subject** record (the
5
+ domain's primary entity — e.g. `patient` for healthcare-domain
6
+ datalakes, `legal_entity` for foundation-domain datalakes) so
7
+ that downstream datasets, workflows, and agent prompts reference
8
+ a single, deduplicated identity rather than the raw inbound
9
+ identifiers each upstream system carries.
10
+
11
+ SDK namespace: `api.mdm`. Today the wire exposes a single verb:
12
+
13
+ ```
14
+ api.mdm.verify(tenantSlug, datalakeSlug, body)
15
+ ```
16
+
17
+ Other MDM operations (subject resolution during ingestion,
18
+ record merging, conflict reconciliation) run inside the platform
19
+ on consumer-supplied inputs but are not directly SDK-callable —
20
+ see §5 Lifecycle.
21
+
22
+ MDM is **Datalake-DB-resident** — the parent datalake must be
23
+ `status: 'ready'` before any MDM verb is called, since
24
+ resolution / verification both look up against the datalake's
25
+ regulated master records.
26
+
27
+ ## 1. Wire shape — `.verify`
28
+
29
+ Verify that a person on the other end of a workflow action or a
30
+ connected-app form submission matches an existing subject in the
31
+ datalake. Used at action time (workflows confirming an SMS
32
+ recipient) and at runtime (connected apps confirming the form
33
+ submitter before posting tracking updates).
34
+
35
+ ```typescript
36
+ import type {
37
+ MDMVerifyRequest,
38
+ MDMVerifyResponse,
39
+ } from '@alvera-ai/platform-sdk'
40
+
41
+ const { data: result } = await api.mdm.verify(
42
+ tenantSlug,
43
+ datalakeSlug,
44
+ {
45
+ subject_id: subjectUuid, // required
46
+ family_name: 'Garcia', // optional
47
+ given_name: 'Maria', // optional
48
+ birth_date: '1978-11-03', // optional; ISO 8601 date
49
+ identifiers: [ // optional
50
+ { system: 'http://hospital.org/mrn', value: 'MRN12345' },
51
+ ],
52
+ },
53
+ )
54
+ // result.status === 'verified' | 'not_verified'
55
+ // result.subject_id === body.subject_id
56
+ // result.verified_at — ISO 8601 timestamp
57
+ // when status === 'not_verified':
58
+ // result.errors — { <field_name>: [reason, ...] } per-field failure detail
59
+ ```
60
+
61
+ The response is **polymorphic on `status`**:
62
+
63
+ ```
64
+ status === 'verified' → { subject_id, verified_at }
65
+ status === 'not_verified' → { subject_id, verified_at, errors }
66
+ ```
67
+
68
+ A `not_verified` response is NOT an error — it's a successful
69
+ identity check that returned a negative result. Distinguish from
70
+ 402-level rejections (404 for missing subject, 401 for missing
71
+ auth, 422 for malformed body).
72
+
73
+ ### Subject per data domain
74
+
75
+ Every datalake has a single canonical **subject entity** — the
76
+ domain's primary identity that all other rows attribute to. The
77
+ verify call resolves against that subject. The subject's wire
78
+ name follows the dataset's `resource_type` (see
79
+ `interoperability_contracts.md` §2 `resource_type`):
80
+
81
+ ```
82
+ data_domain subject resource_type
83
+ ─────────── ─────────────────────
84
+ healthcare patient
85
+ core_banking party
86
+ payment_risk legal_entity
87
+ accounts_receivable customer
88
+ service_commerce consumer
89
+ trading trading_account
90
+ foundation legal_entity
91
+ ```
92
+
93
+ (`legal_entity` is the subject for both `payment_risk` and
94
+ `foundation` — different domain context, same canonical entity
95
+ shape.)
96
+
97
+ Generic tables in a datalake automatically carry a foreign-key
98
+ column to their subject (e.g. `patient_id` in a healthcare-
99
+ domain generic table). Querying that column scopes any custom
100
+ dataset to subjects the verify call has confirmed.
101
+
102
+ ## 2. Rules the type cannot encode
103
+
104
+ ### `subject_id` is required; the identity fields are not
105
+
106
+ `subject_id` (the master record to verify against) is the only
107
+ required field. The identity fields (`family_name`, `given_name`,
108
+ `birth_date`, `identifiers`) are individually optional, but the
109
+ platform requires **at least one** identity field to be supplied
110
+ at runtime — a body carrying only `subject_id` returns a 422.
111
+
112
+ The runtime check is a "must have something to match against"
113
+ rule that the schema can't express in TypeScript / JSON Schema.
114
+
115
+ ### Verification is per-field, with field-specific match rules
116
+
117
+ Each supplied identity field is checked independently against
118
+ the subject record. The platform uses fuzzy matching for the
119
+ string fields and exact matching for the structured ones:
120
+
121
+ ```
122
+ family_name Jaro distance ≥ 0.85
123
+ given_name Jaro distance ≥ 0.80
124
+ birth_date component-wise fuzzy match ≥ 0.80
125
+ identifiers exact (system, value) match
126
+ ```
127
+
128
+ All supplied fields must pass for the overall result to be
129
+ `'verified'`. A single failing field flips the result to
130
+ `'not_verified'` and surfaces the field name in `errors`.
131
+
132
+ The exact thresholds and field set are **domain-scoped** —
133
+ healthcare-domain datalakes use the FHIR-shaped demographics
134
+ listed above; future domains may extend the field set with
135
+ additional verification predicates. The verify endpoint accepts
136
+ the union of fields the parent datalake's domain knows how to
137
+ match against; unknown fields are rejected as a 422.
138
+
139
+ ### `subject_id` must exist in the datalake's regulated schema
140
+
141
+ If the supplied `subject_id` doesn't resolve to a master record
142
+ in the datalake's regulated schema, the response is a 404, not
143
+ a `'not_verified'`. The distinction matters:
144
+
145
+ - `'not_verified'` → "subject exists, identity fields didn't match"
146
+ - 404 → "subject doesn't exist at all"
147
+
148
+ A consumer that conflates the two will mis-route an identity
149
+ challenge as a missing-record error.
150
+
151
+ ## 3. Field ownership
152
+
153
+ **Server-derived (Response-only).**
154
+
155
+ ```
156
+ status enum — 'verified' | 'not_verified'
157
+ verified_at string — ISO 8601 timestamp of the verification attempt
158
+ errors object — present only when status === 'not_verified';
159
+ maps field name to an array of failure reasons
160
+ ```
161
+
162
+ **Caller-supplied (round-trip on the verified branch).**
163
+
164
+ ```
165
+ subject_id required UUID — the master record to verify against
166
+ family_name optional string — fuzzy-matched (string distance)
167
+ given_name optional string — fuzzy-matched (string distance)
168
+ birth_date optional date — fuzzy-matched component-wise
169
+ identifiers optional array — [{ system, value }] exact-matched
170
+ ```
171
+
172
+ **Write-only (Request-only).** None.
173
+
174
+ ## 4. Error envelopes
175
+
176
+ Standard JSON:API envelopes per `errors.md`. Common rejections:
177
+
178
+ | `source.pointer` | Cause |
179
+ |-------------------------------|--------------------------------------------------------|
180
+ | `/subject_id` | missing, or the subject doesn't exist in the datalake |
181
+ | `errors.base` | no identity fields supplied (subject_id alone) — 422 response with `status: "not_verified"` in the body |
182
+ | `/family_name` etc. | field structurally invalid (wrong type, malformed) |
183
+ | `/identifiers/0/system` | missing system on an identifier entry |
184
+ | `/identifiers/0/value` | missing value on an identifier entry |
185
+
186
+ The `not_verified` outcome surfaces in two response shapes:
187
+
188
+ - **Identity mismatch (most common)** — 200 response with
189
+ `status: 'not_verified'` and a field-keyed `errors` map naming
190
+ the demographics that failed (e.g. `errors.family_name`). The
191
+ subject exists; the supplied identity didn't match.
192
+ - **No identity fields supplied** — 422 response with
193
+ `status: 'not_verified'` in the body and `errors.base` set to
194
+ `["at least one verification field is required"]`. Verification
195
+ was structurally impossible.
196
+
197
+ Pattern-match on `status` first, then on the HTTP status code to
198
+ distinguish the two not_verified shapes. The 404 case (subject
199
+ doesn't exist) is a separate envelope — see §6 gotcha 2.
200
+
201
+ ## 5. Lifecycle
202
+
203
+ ### `.verify` is synchronous
204
+
205
+ The verify call returns the resolution outcome in the response —
206
+ no async polling, no job id. The platform performs the fuzzy
207
+ match against the regulated master record inline.
208
+
209
+ ### Other MDM operations are platform-internal
210
+
211
+ Beyond `.verify`, MDM also runs at two points without a direct
212
+ SDK surface:
213
+
214
+ - **During data activation ingestion.** Each inbound row that
215
+ carries an `mdm_input_config` on its bound interoperability
216
+ contract gets routed through the platform's resolution
217
+ pipeline before the dataset upsert. The pipeline matches the
218
+ rendered MDM input against existing subjects (or creates one
219
+ if no match) and produces an `mdm_output` carried forward to
220
+ the regulated upsert. Consumers configure this implicitly via
221
+ the contract's `mdm_input_config` — see
222
+ `interoperability_contracts.md` §2 and `data_activation_clients.md`
223
+ §6.
224
+
225
+ - **Conflict reconciliation.** When two master records turn out
226
+ to represent the same subject (e.g. an MRN-only record and a
227
+ demographics-only record that operators later confirm are the
228
+ same person), the platform supports merge / unmerge operations.
229
+ These are administrative actions surfaced in the admin console,
230
+ not SDK-exposed verbs.
231
+
232
+ ### Read shapes
233
+
234
+ ```
235
+ .verify(tenantSlug, datalakeSlug, body) Single round-trip; see §1
236
+ ```
237
+
238
+ Today there is no `.list`, `.get`, or `.metadata` on the MDM
239
+ namespace — subjects themselves are queried through the dataset
240
+ search surface (see `data_activation_clients.md` §6.5), not
241
+ through `api.mdm`.
242
+
243
+ ## 6. Gotchas
244
+
245
+ 1. **`'not_verified'` is usually a successful 200 response, not
246
+ an error — but the "no fields supplied" case is a 422 with a
247
+ `not_verified` body.** Pattern-match on `result.status === 'verified'`
248
+ to gate the downstream "act on this person" path. Throwing on
249
+ `'not_verified'` collapses two distinct outcomes (subject
250
+ verified vs identity mismatch) into one branch and loses the
251
+ `errors` payload that explains which fields failed. The 422
252
+ sub-case (no identity fields, just `subject_id`) needs HTTP-status
253
+ handling on top of body-status handling — see §4 above.
254
+
255
+ 2. **`subject_id` missing → 404, identity mismatch → 200.** A
256
+ verify call against a deleted or never-existed subject id
257
+ returns a 404 envelope, NOT a `'not_verified'` response.
258
+ Consumers that handle only the polymorphic 200 response will
259
+ surface the 404 as an unexpected exception. Wrap the verify
260
+ call in standard 4xx handling per `errors.md`.
261
+
262
+ 3. **At least one identity field is required.** A body with
263
+ only `subject_id` and no identity fields returns a 422 with
264
+ `errors.base` set to `["at least one verification field is
265
+ required"]` and `status: "not_verified"` in the response
266
+ body. The platform cannot verify identity against zero
267
+ attributes; this is enforced at runtime rather than in the
268
+ request schema, so the TypeScript type permits the
269
+ no-identity-field body but the wire rejects it.
270
+
271
+ 4. **The verification field set is domain-scoped.** Today the
272
+ healthcare-domain verifier accepts FHIR-shaped demographics
273
+ (family_name, given_name, birth_date, identifiers). Other
274
+ domains will expose verifier-appropriate field sets as they
275
+ land. Querying `api.mdm.verify` against a datalake whose
276
+ domain doesn't support a supplied field returns a 422 on
277
+ that field's pointer — don't assume the full healthcare
278
+ field set works on every domain.
279
+
280
+ 5. **MDM also runs at ingestion — and you don't call it
281
+ directly.** The `mdm_input_config` on each interoperability
282
+ contract feeds the platform's resolution pipeline at
283
+ ingestion time, producing the `mdm_output` carried into the
284
+ dataset upsert. There is no `api.mdm.resolve` SDK verb; you
285
+ configure resolution by shaping the contract's
286
+ `mdm_input_config`. See `interoperability_contracts.md` §2.
287
+
288
+ 6. **`identifiers` matches are exact, not fuzzy.** Whitespace,
289
+ case, and trailing slashes in `system` URIs all matter.
290
+ Compare with the master record's identifier system strings
291
+ character-for-character; a `'http://hospital.org/mrn'` vs
292
+ `'http://hospital.org/mrn/'` mismatch will flip the verifier
293
+ to `'not_verified'` even when the value is correct.