@tanstack/ai-persistence 0.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/dist/esm/blob-range.d.ts +51 -0
  2. package/dist/esm/blob-range.js +84 -0
  3. package/dist/esm/blob-range.js.map +1 -0
  4. package/dist/esm/capabilities.d.ts +5 -0
  5. package/dist/esm/capabilities.js +16 -0
  6. package/dist/esm/capabilities.js.map +1 -0
  7. package/dist/esm/index.d.ts +13 -0
  8. package/dist/esm/index.js +9 -0
  9. package/dist/esm/memory.d.ts +19 -0
  10. package/dist/esm/memory.js +319 -0
  11. package/dist/esm/memory.js.map +1 -0
  12. package/dist/esm/middleware.d.ts +252 -0
  13. package/dist/esm/middleware.js +872 -0
  14. package/dist/esm/middleware.js.map +1 -0
  15. package/dist/esm/reconstruct-generation.d.ts +129 -0
  16. package/dist/esm/reconstruct-generation.js +148 -0
  17. package/dist/esm/reconstruct-generation.js.map +1 -0
  18. package/dist/esm/reconstruct.d.ts +79 -0
  19. package/dist/esm/reconstruct.js +75 -0
  20. package/dist/esm/reconstruct.js.map +1 -0
  21. package/dist/esm/retrieve.d.ts +40 -0
  22. package/dist/esm/retrieve.js +54 -0
  23. package/dist/esm/retrieve.js.map +1 -0
  24. package/dist/esm/testkit/conformance.d.ts +33 -0
  25. package/dist/esm/testkit/conformance.js +997 -0
  26. package/dist/esm/testkit/conformance.js.map +1 -0
  27. package/dist/esm/types.d.ts +554 -0
  28. package/dist/esm/types.js +103 -0
  29. package/dist/esm/types.js.map +1 -0
  30. package/package.json +71 -0
  31. package/skills/ai-persistence/SKILL.md +218 -0
  32. package/skills/ai-persistence/build-cloudflare-adapter/SKILL.md +313 -0
  33. package/skills/ai-persistence/build-cloudflare-artifact-store/SKILL.md +693 -0
  34. package/skills/ai-persistence/build-custom-adapter/SKILL.md +328 -0
  35. package/skills/ai-persistence/build-drizzle-adapter/SKILL.md +562 -0
  36. package/skills/ai-persistence/build-prisma-adapter/SKILL.md +518 -0
  37. package/skills/ai-persistence/server/SKILL.md +210 -0
  38. package/skills/ai-persistence/stores/SKILL.md +485 -0
  39. package/src/blob-range.ts +101 -0
  40. package/src/capabilities.ts +18 -0
  41. package/src/index.ts +114 -0
  42. package/src/memory.ts +491 -0
  43. package/src/middleware.ts +1795 -0
  44. package/src/reconstruct-generation.ts +244 -0
  45. package/src/reconstruct.ts +149 -0
  46. package/src/retrieve.ts +77 -0
  47. package/src/testkit/conformance.ts +1288 -0
  48. package/src/types.ts +878 -0
@@ -0,0 +1,693 @@
1
+ ---
2
+ name: ai-persistence/build-cloudflare-artifact-store
3
+ description: Use when a Cloudflare Worker needs durable byte storage for TanStack AI generated media (images, audio, video, transcripts) — writes a BlobStore backed by R2 and an ArtifactStore backed by D1 (or KV), composes them onto the generation persistence so withGenerationPersistence persists artifact bytes, and serves them back from a Worker GET route. Includes one-line sketches for S3, GCS, Vercel Blob, Supabase, and a dev filesystem BlobStore.
4
+ ---
5
+
6
+ # Cloudflare Artifact + Blob Store
7
+
8
+ `withGenerationPersistence(persistence)` needs only `stores.generationRuns` to track a
9
+ generation's lifecycle. Add `stores.artifacts` (metadata) **and** `stores.blobs`
10
+ (the bytes) — both, or neither — and the middleware also persists the generated
11
+ media: image/audio/TTS/video/transcription bytes land at blob key
12
+ `artifacts/<runId>/<artifactId>`, with an `ArtifactRecord` row describing each.
13
+
14
+ The deliverable is **one file in the Worker** — e.g.
15
+ `src/lib/generation-persistence.ts` — exporting a factory that builds an
16
+ `AIPersistence` from the request's R2 + D1 bindings, plus a GET route that serves
17
+ artifact bytes with `retrieveArtifact` / `retrieveBlob`.
18
+
19
+ Read the sibling **ai-persistence/build-cloudflare-adapter** skill for the
20
+ per-request-binding rule, `wrangler` config shape, D1 migration workflow, and the
21
+ chat (generation-run/message) side. This skill covers only the two byte-storage stores and
22
+ how to compose them.
23
+
24
+ ## The two contracts, verbatim
25
+
26
+ Both come from `@tanstack/ai-persistence`. `defineBlobStore` / `defineArtifactStore`
27
+ type an object literal inline (autocomplete + contract checking, no separate
28
+ annotation).
29
+
30
+ ```ts
31
+ // BlobStore — the byte layer. R2 backs it.
32
+ interface BlobStore {
33
+ put(
34
+ key: string,
35
+ body: BlobBody,
36
+ options?: BlobPutOptions,
37
+ ): Promise<BlobRecord>
38
+ // metadata + byte accessors; `options.range` reads one slice (for `206`s)
39
+ get(key: string, options?: BlobGetOptions): Promise<BlobObject | null>
40
+ head(key: string): Promise<BlobRecord | null> // metadata only
41
+ delete(key: string): Promise<void> // no-op if absent
42
+ list(options?: BlobListOptions): Promise<BlobListPage>
43
+ }
44
+
45
+ // ArtifactStore — the metadata layer. D1 (or KV) backs it.
46
+ interface ArtifactStore {
47
+ save(record: ArtifactRecord): Promise<void> // insert or overwrite
48
+ get(artifactId: string): Promise<ArtifactRecord | null>
49
+ list(runId: string): Promise<Array<ArtifactRecord>> // [] when none
50
+ delete(artifactId: string): Promise<void>
51
+ deleteForRun(runId: string): Promise<void>
52
+ }
53
+ ```
54
+
55
+ `BlobBody` is `ReadableStream<Uint8Array> | ArrayBuffer | ArrayBufferView |
56
+ string | Blob`. The non-stream shapes flow straight into `R2Bucket.put`
57
+ unchanged — but a `ReadableStream` body does **not**, in the general case:
58
+ workerd's `put` requires a stream with a known length (a `Response` body or the
59
+ readable half of a `FixedLengthStream`), and the artifact middleware hands you a
60
+ `TransformStream`-wrapped body whenever it had to cap a fetched body as it
61
+ drains. Passing that stream to `bucket.put` throws `TypeError: Provided readable
62
+ stream must have a known length`.
63
+
64
+ **When does that actually happen?** The wrapper only exists to enforce
65
+ `maxArtifactBytes` during the drain, so the middleware applies it only when
66
+ nothing else bounds the transfer:
67
+
68
+ | Provider response | Body handed to `put` | R2 path |
69
+ | --------------------------------------- | --------------------------------- | ------------------- |
70
+ | `content-length`, no `content-encoding` | untouched, declared length intact | `bucket.put` direct |
71
+ | chunked (no declared length) | wrapped, length-less | multipart |
72
+ | `content-encoding: gzip` | wrapped, length-less | multipart |
73
+
74
+ A provider CDN normally sends `content-length`, so the first row is the common
75
+ case and `bucket.put(key, body)` just works. The recipe below is what makes the
76
+ other two rows work: it re-declares the length from
77
+ `BlobPutOptions.expectedLength` when the middleware could vouch for one, and
78
+ otherwise streams through a multipart upload (one 8 MiB part at a time — flat
79
+ memory at any artifact size). Write it once and every response shape is
80
+ covered.
81
+
82
+ `withGenerationPersistence(persistence, { maxArtifactBytes: false })` drops the
83
+ ceiling and the wrapper altogether, so even a chunked reply arrives untouched.
84
+ It buys nothing extra for R2 (a chunked body has no length to preserve), so
85
+ choose it on its own merits: no _application_ limit on what an origin can
86
+ stream into your bucket. R2's own limits still apply — 5 GiB per single-shot
87
+ put, and 10,000 multipart parts (~80 GiB at the 8 MiB part size below). Keep
88
+ the cap when `allowInputUrl` lets callers name the URL.
89
+
90
+ `BlobPutOptions` is
91
+ `{ contentType?, customMetadata?, expectedLength? }`; `BlobGetOptions` is
92
+ `{ range?: { offset: number, length?: number } }` and maps onto R2's own
93
+ `range`; `BlobListOptions` is `{ prefix?, cursor?, limit? }`; `BlobListPage` is
94
+ `{ objects: BlobRecord[], cursor?, truncated? }`.
95
+
96
+ ## 1. BlobStore backed by R2
97
+
98
+ `R2Object` carries `size`, `etag`, `httpMetadata.contentType`, `customMetadata`,
99
+ and `uploaded` (a `Date`). `BlobRecord` wants `createdAt` / `updatedAt` as epoch
100
+ ms — R2 tracks only the single `uploaded` instant, so map it to both. `get` /
101
+ `head` are the byte-body vs metadata-only split; `R2ObjectBody` already exposes
102
+ `body`, `arrayBuffer()`, and `text()`, so a `BlobObject` is essentially the R2
103
+ object plus the mapped metadata.
104
+
105
+ ```ts ignore
106
+ import { defineBlobStore, resolveBlobRange } from '@tanstack/ai-persistence'
107
+ import type { BlobObject, BlobRecord } from '@tanstack/ai-persistence'
108
+
109
+ // R2 multipart parts must be ≥ 5 MiB and — except for the last — all exactly
110
+ // the SAME size, so a part reader has to cut on an exact boundary and carry the
111
+ // remainder. 8 MiB × the 10,000-part ceiling puts the multipart path's limit at
112
+ // ~80 GiB; raise this for larger objects, and check R2's current object-size
113
+ // limits before promising more.
114
+ const MULTIPART_PART_SIZE = 8 * 1024 * 1024
115
+
116
+ /**
117
+ * Cut exactly `limit` bytes off the stream (fewer only at EOF), carrying any
118
+ * overshoot into the next part.
119
+ *
120
+ * `carry` is the leftover from the previous call. Chunks are collected by
121
+ * reference and copied once per part: growing a `Uint8Array` chunk-by-chunk
122
+ * instead would re-copy the whole part on every read — quadratic work for a
123
+ * part built from hundreds of small chunks.
124
+ */
125
+ async function readPart(
126
+ reader: ReadableStreamDefaultReader<Uint8Array>,
127
+ limit: number,
128
+ carry: Uint8Array,
129
+ ): Promise<{ bytes: Uint8Array; carry: Uint8Array; eof: boolean }> {
130
+ const chunks: Array<Uint8Array> = carry.byteLength > 0 ? [carry] : []
131
+ let total = carry.byteLength
132
+ let eof = false
133
+ while (total < limit) {
134
+ const { value, done } = await reader.read()
135
+ if (done) {
136
+ eof = true
137
+ break
138
+ }
139
+ chunks.push(value)
140
+ total += value.byteLength
141
+ }
142
+ const joined = new Uint8Array(total)
143
+ let offset = 0
144
+ for (const chunk of chunks) {
145
+ joined.set(chunk, offset)
146
+ offset += chunk.byteLength
147
+ }
148
+ // Equal-sized parts: hand back exactly `limit` and keep the rest for next
149
+ // time. At EOF the last part is whatever is left, which R2 allows.
150
+ if (!eof && total > limit) {
151
+ return {
152
+ bytes: joined.subarray(0, limit),
153
+ carry: joined.subarray(limit),
154
+ eof,
155
+ }
156
+ }
157
+ return { bytes: joined, carry: new Uint8Array(0), eof }
158
+ }
159
+
160
+ // R2 takes a single-shot put up to 5 GiB. Above that the upload has to be
161
+ // multipart even when the length is known.
162
+ const MAX_SINGLE_SHOT_BYTES = 5 * 1024 * 1024 * 1024
163
+
164
+ /**
165
+ * `R2Bucket.put` requires a stream with a known length; a pass-through
166
+ * TransformStream (what the middleware sends when it had to cap the body) has
167
+ * none. Re-declare the length when `expectedLength` provides it — the
168
+ * middleware only sets it when it is the exact decoded byte count — and
169
+ * otherwise stream through a multipart upload, one buffered part at a time.
170
+ */
171
+ async function putStream(
172
+ bucket: R2Bucket,
173
+ key: string,
174
+ body: ReadableStream<Uint8Array>,
175
+ options: R2PutOptions,
176
+ expectedLength: number | undefined,
177
+ ): Promise<R2Object | null> {
178
+ if (expectedLength !== undefined && expectedLength <= MAX_SINGLE_SHOT_BYTES) {
179
+ return bucket.put(
180
+ key,
181
+ body.pipeThrough(new FixedLengthStream(expectedLength)),
182
+ options,
183
+ )
184
+ }
185
+ // Unknown length, or too big for one shot: multipart, one part at a time.
186
+ const reader = body.getReader()
187
+ const first = await readPart(reader, MULTIPART_PART_SIZE, new Uint8Array(0))
188
+ if (first.eof) {
189
+ // The whole body fit in one part — a plain put is enough.
190
+ return bucket.put(key, first.bytes, options)
191
+ }
192
+ const upload = await bucket.createMultipartUpload(key, options)
193
+ try {
194
+ const parts: Array<R2UploadedPart> = [
195
+ await upload.uploadPart(1, first.bytes),
196
+ ]
197
+ let carry = first.carry
198
+ let partNumber = 2
199
+ for (;;) {
200
+ const part = await readPart(reader, MULTIPART_PART_SIZE, carry)
201
+ carry = part.carry
202
+ if (part.bytes.byteLength > 0) {
203
+ // 10,000 parts is R2's ceiling; failing here beats a confusing error
204
+ // from `complete` after uploading gigabytes.
205
+ if (partNumber > 10_000) {
206
+ throw new Error(
207
+ `Artifact at ${key} exceeds the multipart part limit — raise MULTIPART_PART_SIZE.`,
208
+ )
209
+ }
210
+ parts.push(await upload.uploadPart(partNumber, part.bytes))
211
+ partNumber += 1
212
+ }
213
+ if (part.eof) break
214
+ }
215
+ return await upload.complete(parts)
216
+ } catch (error) {
217
+ await upload.abort().catch(() => undefined)
218
+ throw error
219
+ }
220
+ }
221
+
222
+ function toRecord(obj: R2Object): BlobRecord {
223
+ const uploaded = obj.uploaded.getTime()
224
+ return {
225
+ key: obj.key,
226
+ size: obj.size,
227
+ etag: obj.etag,
228
+ ...(obj.httpMetadata?.contentType
229
+ ? { contentType: obj.httpMetadata.contentType }
230
+ : {}),
231
+ ...(obj.customMetadata ? { customMetadata: obj.customMetadata } : {}),
232
+ createdAt: uploaded,
233
+ updatedAt: uploaded,
234
+ }
235
+ }
236
+
237
+ export function r2BlobStore(bucket: R2Bucket) {
238
+ return defineBlobStore({
239
+ async put(key, body, options) {
240
+ const r2Options = {
241
+ ...(options?.contentType
242
+ ? { httpMetadata: { contentType: options.contentType } }
243
+ : {}),
244
+ ...(options?.customMetadata
245
+ ? { customMetadata: options.customMetadata }
246
+ : {}),
247
+ }
248
+ const obj =
249
+ body instanceof ReadableStream
250
+ ? await putStream(
251
+ bucket,
252
+ key,
253
+ body,
254
+ r2Options,
255
+ options?.expectedLength,
256
+ )
257
+ : await bucket.put(key, body, r2Options)
258
+ // R2.put returns null only when an onlyIf precondition fails — not used here.
259
+ if (!obj) throw new Error(`R2 put failed for ${key}`)
260
+ return toRecord(obj)
261
+ },
262
+
263
+ async get(key, options): Promise<BlobObject | null> {
264
+ if (!options?.range) {
265
+ const whole = await bucket.get(key)
266
+ if (!whole) return null
267
+ return {
268
+ ...toRecord(whole),
269
+ body: whole.body,
270
+ arrayBuffer: () => whole.arrayBuffer(),
271
+ text: () => whole.text(),
272
+ }
273
+ }
274
+ // A ranged read is what a serve route turns into `206` — R2 slices in
275
+ // the bucket, so a seek never streams the whole object. `resolveBlobRange`
276
+ // clamps a too-long length and throws on an offset past the end (the
277
+ // route answers `416` from `record.size` before ever calling in). The
278
+ // `head` buys that clamp deterministically — one class-B op against
279
+ // bytes you did not need to send.
280
+ //
281
+ // Retry a bounded number of times: each attempt measures the object and
282
+ // reads a slice of THAT version, and only an overwrite landing between
283
+ // the two calls costs another lap.
284
+ for (let attempt = 0; attempt < 3; attempt++) {
285
+ const head = await bucket.head(key)
286
+ // A ranged read of a key that is not there is a miss, not a
287
+ // whole-object read: falling through to an un-ranged `get` would
288
+ // answer a `206` carrying the entire file.
289
+ if (!head) return null
290
+ const served = resolveBlobRange(head.size, options.range)
291
+ const obj = await bucket.get(key, {
292
+ range: { offset: served.offset, length: served.length },
293
+ // Tie the slice to the version `head` measured. Without this, a
294
+ // `put` landing between the two calls yields a `Content-Range`
295
+ // computed from one object and bytes from another.
296
+ onlyIf: { etagMatches: head.etag },
297
+ })
298
+ // No body means the precondition failed — the object changed under us.
299
+ if (!obj) return null
300
+ if (!('body' in obj)) continue
301
+ return {
302
+ // `toRecord(obj)` reports the WHOLE object's size even on a ranged
303
+ // read — R2's `R2Object.size` is the object, not the slice — which
304
+ // is exactly the contract, and what `Content-Range`'s total needs.
305
+ ...toRecord(obj),
306
+ range: served,
307
+ body: obj.body,
308
+ arrayBuffer: () => obj.arrayBuffer(),
309
+ text: () => obj.text(),
310
+ }
311
+ }
312
+ throw new Error(`R2 object ${key} changed under every ranged read.`)
313
+ },
314
+
315
+ async head(key) {
316
+ const obj = await bucket.head(key)
317
+ return obj ? toRecord(obj) : null
318
+ },
319
+
320
+ async delete(key) {
321
+ await bucket.delete(key)
322
+ },
323
+
324
+ async list(options) {
325
+ // R2 reads `limit: 0` as "use the default", so short-circuit it.
326
+ if (options?.limit === 0) {
327
+ return { objects: [] }
328
+ }
329
+ const page = await bucket.list({
330
+ ...(options?.prefix !== undefined ? { prefix: options.prefix } : {}),
331
+ ...(options?.cursor !== undefined ? { cursor: options.cursor } : {}),
332
+ ...(options?.limit !== undefined ? { limit: options.limit } : {}),
333
+ // R2 omits httpMetadata/customMetadata from list rows unless asked.
334
+ include: ['httpMetadata', 'customMetadata'],
335
+ })
336
+ return {
337
+ objects: page.objects.map(toRecord),
338
+ ...(page.truncated ? { cursor: page.cursor, truncated: true } : {}),
339
+ }
340
+ },
341
+ })
342
+ }
343
+ ```
344
+
345
+ Invariants that matter (asserted by the conformance testkit):
346
+
347
+ - `get` / `head` return `null` for a missing key; `delete` is a silent no-op.
348
+ - `put` **overwrites** an existing key.
349
+ - `put` accepts a `ReadableStream` body with **no declared length** — the
350
+ middleware streams URL-fetched artifacts as exactly that. This is where the
351
+ naive "pass the body straight to `bucket.put`" recipe fails at runtime
352
+ (workerd requires a known length), which is what `putStream` above handles.
353
+ - `get` honours `options.range`: it returns **only** that slice, reports it as
354
+ `range`, and keeps `size` on the whole object. That is the `206` a video
355
+ player's seeking depends on, and R2 slices in the bucket so the bytes never
356
+ cross the Worker.
357
+ - `list` filters by `prefix` literally (R2 prefix is a literal byte prefix — no
358
+ glob), returns keys in ascending order, and pages via the opaque `cursor` when
359
+ `truncated`. R2's own cursor is opaque and satisfies this directly. `limit: 0`
360
+ must yield an empty, untruncated page — R2 treats `limit: 0` as "use the
361
+ default", so special-case it: `if (options?.limit === 0) return { objects: [] }`.
362
+
363
+ ## 2. ArtifactStore backed by D1
364
+
365
+ `ArtifactRecord` is `{ artifactId, runId, threadId, name, mimeType, size,
366
+ sourceUrl?, createdAt }` (`createdAt` epoch ms). One flat table, keyed by
367
+ `artifact_id`, indexed by `run_id` for `list`.
368
+
369
+ ```sql
370
+ CREATE TABLE IF NOT EXISTS generation_artifacts (
371
+ artifact_id text PRIMARY KEY NOT NULL,
372
+ run_id text NOT NULL,
373
+ thread_id text NOT NULL,
374
+ blob_key text,
375
+ name text NOT NULL,
376
+ mime_type text NOT NULL,
377
+ size integer NOT NULL,
378
+ source_url text,
379
+ created_at integer NOT NULL
380
+ );
381
+ CREATE INDEX IF NOT EXISTS generation_artifacts_run ON generation_artifacts (run_id);
382
+ ```
383
+
384
+ ```ts ignore
385
+ import { defineArtifactStore } from '@tanstack/ai-persistence'
386
+ import type { ArtifactRecord } from '@tanstack/ai-persistence'
387
+
388
+ interface ArtifactRow {
389
+ artifact_id: string
390
+ run_id: string
391
+ thread_id: string
392
+ blob_key: string | null
393
+ name: string
394
+ mime_type: string
395
+ size: number
396
+ source_url: string | null
397
+ created_at: number
398
+ }
399
+
400
+ function fromRow(row: ArtifactRow): ArtifactRecord {
401
+ return {
402
+ artifactId: row.artifact_id,
403
+ runId: row.run_id,
404
+ threadId: row.thread_id,
405
+ ...(row.blob_key != null ? { blobKey: row.blob_key } : {}),
406
+ name: row.name,
407
+ mimeType: row.mime_type,
408
+ size: row.size,
409
+ ...(row.source_url != null ? { sourceUrl: row.source_url } : {}),
410
+ createdAt: row.created_at,
411
+ }
412
+ }
413
+
414
+ export function d1ArtifactStore(db: D1Database) {
415
+ return defineArtifactStore({
416
+ async save(record) {
417
+ // Insert or overwrite (artifact ids are unique).
418
+ await db
419
+ .prepare(
420
+ `INSERT INTO generation_artifacts
421
+ (artifact_id, run_id, thread_id, blob_key, name, mime_type, size, source_url, created_at)
422
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
423
+ ON CONFLICT(artifact_id) DO UPDATE SET
424
+ run_id = excluded.run_id, thread_id = excluded.thread_id,
425
+ blob_key = excluded.blob_key, name = excluded.name,
426
+ mime_type = excluded.mime_type, size = excluded.size,
427
+ source_url = excluded.source_url, created_at = excluded.created_at`,
428
+ )
429
+ .bind(
430
+ record.artifactId,
431
+ record.runId,
432
+ record.threadId,
433
+ record.blobKey ?? null,
434
+ record.name,
435
+ record.mimeType,
436
+ record.size,
437
+ record.sourceUrl ?? null,
438
+ record.createdAt,
439
+ )
440
+ .run()
441
+ },
442
+
443
+ async get(artifactId) {
444
+ const row = await db
445
+ .prepare(`SELECT * FROM generation_artifacts WHERE artifact_id = ?`)
446
+ .bind(artifactId)
447
+ .first<ArtifactRow>()
448
+ return row ? fromRow(row) : null
449
+ },
450
+
451
+ async list(runId) {
452
+ const { results } = await db
453
+ .prepare(`SELECT * FROM generation_artifacts WHERE run_id = ?`)
454
+ .bind(runId)
455
+ .all<ArtifactRow>()
456
+ return results.map(fromRow)
457
+ },
458
+
459
+ async delete(artifactId) {
460
+ await db
461
+ .prepare(`DELETE FROM generation_artifacts WHERE artifact_id = ?`)
462
+ .bind(artifactId)
463
+ .run()
464
+ },
465
+
466
+ async deleteForRun(runId) {
467
+ await db
468
+ .prepare(`DELETE FROM generation_artifacts WHERE run_id = ?`)
469
+ .bind(runId)
470
+ .run()
471
+ },
472
+ })
473
+ }
474
+ ```
475
+
476
+ Omitting `source_url` / `blob_key` from the record when the column is `NULL`
477
+ keeps records comparing cleanly against the reference in-memory store. Persist
478
+ `blob_key` verbatim: a `storageKey` mapper can put the bytes anywhere, so a
479
+ reader cannot recompute the path — `resolveArtifactBlobKey(record)` falls back
480
+ to the default convention only for rows written before the column existed.
481
+ **KV alternative:** if
482
+ you have no D1, back `save`/`get` with `KV.put(artifactId, JSON.stringify(record))`
483
+ / `KV.get(artifactId, 'json')`, and maintain a `run:<runId>` index key (a JSON
484
+ array of artifact ids) for `list` — KV has no query, so `list` needs that
485
+ secondary index.
486
+
487
+ ## 3. Compose and wire
488
+
489
+ Bindings are per-request on Workers, so export a **factory**. Combine the byte
490
+ stores with a generation-run store (and, if this Worker also does chat, the chat stores).
491
+ Either build the whole `AIPersistence` with `defineAIPersistence`, or layer the
492
+ artifact stores onto an existing chat persistence with `composePersistence`:
493
+
494
+ ```ts ignore
495
+ import {
496
+ defineAIPersistence,
497
+ composePersistence,
498
+ withGenerationPersistence,
499
+ } from '@tanstack/ai-persistence'
500
+ import { r2BlobStore } from './r2-blob-store'
501
+ import { d1ArtifactStore } from './d1-artifact-store'
502
+ import { d1GenerationRunStore } from './d1-job-store' // your GenerationRunStore
503
+
504
+ /** Call inside a request handler — bindings are not available at module scope. */
505
+ export function generationPersistence(env: Env) {
506
+ return defineAIPersistence({
507
+ stores: {
508
+ generationRuns: d1GenerationRunStore(env.DB),
509
+ artifacts: d1ArtifactStore(env.DB),
510
+ blobs: r2BlobStore(env.ARTIFACTS_BUCKET),
511
+ },
512
+ })
513
+ }
514
+
515
+ // …or add bytes to a persistence that already has chat + generation runs:
516
+ // composePersistence(chatAndJobsPersistence(env), {
517
+ // overrides: {
518
+ // artifacts: d1ArtifactStore(env.DB),
519
+ // blobs: r2BlobStore(env.ARTIFACTS_BUCKET),
520
+ // },
521
+ // })
522
+ ```
523
+
524
+ `withGenerationPersistence` throws if exactly one of `artifacts` / `blobs` is
525
+ present — provide both or neither. Wire it as generation middleware:
526
+
527
+ ```ts ignore
528
+ import { generateImage, toServerSentEventsResponse } from '@tanstack/ai'
529
+ import { openaiImage } from '@tanstack/ai-openai'
530
+ import { withGenerationPersistence } from '@tanstack/ai-persistence'
531
+ import { generationPersistence } from './lib/generation-persistence'
532
+
533
+ export default {
534
+ async fetch(request: Request, env: Env) {
535
+ const { prompt, threadId } = await request.json()
536
+ const stream = generateImage({
537
+ adapter: openaiImage('gpt-image-1'),
538
+ prompt,
539
+ threadId, // the slot recorded on the job + artifacts
540
+ stream: true,
541
+ middleware: [
542
+ // Nothing is buffered: the artifact streams from the provider CDN
543
+ // into R2. Add `maxArtifactBytes: false` if you want no ceiling at
544
+ // all on what an origin can stream into your bucket (the default is
545
+ // 1 GiB); it is not needed for the streaming itself.
546
+ withGenerationPersistence(generationPersistence(env), { threadId }),
547
+ ],
548
+ })
549
+ return toServerSentEventsResponse(stream)
550
+ },
551
+ }
552
+ ```
553
+
554
+ ## 4. Serve the bytes back
555
+
556
+ A GET route resolves an `artifactId` to its record and its stored bytes.
557
+ `retrieveArtifact` returns the `ArtifactRecord` (or `null` → 404);
558
+ `retrieveBlob` returns the `BlobObject` (metadata + a streamable `body`). Both
559
+ resolve the blob key from the record internally, so you never build the key
560
+ yourself.
561
+
562
+ Honour `Range` requests: `<video>` seeking is built on `206` / `Content-Range`,
563
+ and Safari refuses to play a source that ignores `Range` entirely. Images never
564
+ notice; a few-hundred-MB clip is unwatchable without it. Pass `range` to
565
+ `retrieveBlob` — the store slices in R2 — rather than reaching into the bucket
566
+ binding from the route, which would tie the route to R2 and bypass the store's
567
+ own key resolution.
568
+
569
+ ```ts ignore
570
+ import {
571
+ parseRangeHeader,
572
+ retrieveArtifact,
573
+ retrieveBlob,
574
+ } from '@tanstack/ai-persistence'
575
+ import { generationPersistence } from './lib/generation-persistence'
576
+
577
+ export async function GET(request: Request, env: Env) {
578
+ const artifactId = new URL(request.url).searchParams.get('id') ?? ''
579
+ const persistence = generationPersistence(env)
580
+
581
+ // Authorize before serving — derive the owner from the session, never trust
582
+ // a client-supplied id. (The record carries runId/threadId to check against.)
583
+ const record = await retrieveArtifact(persistence, artifactId)
584
+ if (!record) return new Response('Not found', { status: 404 })
585
+
586
+ // `parseRangeHeader` resolves the header against the size on the record —
587
+ // suffix ranges included — so an unsatisfiable range is a 416 here and the
588
+ // store only ever sees a range it can serve.
589
+ const range = parseRangeHeader(request.headers.get('range'), record.size)
590
+ if (range === 'unsatisfiable') {
591
+ return new Response('Range not satisfiable', {
592
+ status: 416,
593
+ headers: { 'content-range': `bytes */${record.size}` },
594
+ })
595
+ }
596
+
597
+ // Pass the record, not the id: no second metadata lookup.
598
+ const blob = await retrieveBlob(
599
+ persistence,
600
+ record,
601
+ range ? { range } : undefined,
602
+ )
603
+ if (!blob?.body) return new Response('Not found', { status: 404 })
604
+
605
+ // `accept-ranges` on every response, including the whole-file one: it is how
606
+ // a player learns it may seek at all.
607
+ const headers = {
608
+ 'content-type': record.mimeType,
609
+ 'accept-ranges': 'bytes',
610
+ }
611
+ if (!blob.range) {
612
+ return new Response(blob.body, {
613
+ headers: { ...headers, 'content-length': String(record.size) },
614
+ })
615
+ }
616
+ const { offset, length } = blob.range
617
+ return new Response(blob.body, {
618
+ status: 206,
619
+ headers: {
620
+ ...headers,
621
+ 'content-length': String(length),
622
+ 'content-range': `bytes ${offset}-${offset + length - 1}/${record.size}`,
623
+ },
624
+ })
625
+ }
626
+ ```
627
+
628
+ To hydrate a **server-driven generation client** (`persistence: true` + a stable
629
+ `threadId`) on mount, also expose `reconstructGeneration(persistence, request)`
630
+ on a GET that reads `?threadId=` / `?runId=` — see `ai-core/client-persistence`.
631
+
632
+ ## Other backends — the contract is tiny, here's how each maps
633
+
634
+ `BlobStore` is five methods over an object store. Any of these backs it; swap the
635
+ factory, keep everything else. `put` maps to the SDK's upload, `get` to a
636
+ download that exposes `body`/`arrayBuffer`/`text`, `head` to a metadata fetch,
637
+ `delete` to a delete, `list` to a prefixed, cursor-paged list.
638
+
639
+ | Backend | npm | One-line sketch |
640
+ | ------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
641
+ | **AWS S3** | `@aws-sdk/client-s3` | `put`→`PutObjectCommand`; `get`→`GetObjectCommand` (`Body` is a stream → `body`, `.transformToByteArray()`/`.transformToString()`); `head`→`HeadObjectCommand`; `delete`→`DeleteObjectCommand`; `list`→`ListObjectsV2Command` (`Prefix`, `ContinuationToken`↔`cursor`, `MaxKeys`↔`limit`, `IsTruncated`↔`truncated`). |
642
+ | **Google Cloud Storage** | `@google-cloud/storage` | `bucket.file(key)`: `put`→`.save(body, { contentType, metadata })`; `get`→`.createReadStream()` for `body` + `.download()` for bytes; `head`→`.getMetadata()`; `delete`→`.delete({ ignoreNotFound: true })`; `list`→`bucket.getFiles({ prefix, maxResults, pageToken })`. |
643
+ | **Vercel Blob** | `@vercel/blob` | `put`→`put(key, body, { access: 'public', contentType })`; `get`→`fetch(head(key).url)` (stream `res.body`); `head`→`head(key)` (returns `null`→catch as absent); `delete`→`del(key)`; `list`→`list({ prefix, cursor, limit })` (`hasMore`↔`truncated`). |
644
+ | **Supabase Storage** | `@supabase/supabase-js` | `storage.from(bucket)`: `put`→`.upload(key, body, { contentType, upsert: true })`; `get`→`.download(key)` (returns a `Blob` → `body`/`arrayBuffer`/`text`); `head`→`.info(key)` or list-one; `delete`→`.remove([key])`; `list`→`.list(prefix, { limit })` (offset/limit paging → synthesize a `cursor`). |
645
+ | **Filesystem (dev only)** | `node:fs/promises` | Root each key under a dir: `put`→`mkdir(dirname, { recursive: true })` + `writeFile`; `get`→`createReadStream` for `body` + `readFile`; `head`→`stat` (`size`, `mtimeMs`→`updatedAt`); `delete`→`rm(path, { force: true })`; `list`→recursive `readdir` filtered by prefix, sorted, sliced by `limit`, cursor = last key. Not for production — no concurrency guarantees. |
646
+
647
+ For each: `contentType` and `customMetadata` ride the SDK's own metadata fields;
648
+ `BlobRecord.createdAt`/`updatedAt` come from the object's stored timestamps
649
+ (epoch ms); return `null` from `get`/`head` on a not-found rather than throwing.
650
+
651
+ ## Verify
652
+
653
+ The shared `runPersistenceConformance` testkit covers all seven stores, including
654
+ `generationRuns`, `artifacts`, and `blobs` — point it at your factory rather than
655
+ hand-writing these assertions:
656
+
657
+ ```ts ignore
658
+ import { runPersistenceConformance } from '@tanstack/ai-persistence/testkit'
659
+ import { env } from 'cloudflare:test'
660
+ import { generationPersistence } from '../src/lib/generation-persistence'
661
+
662
+ // A generation-only Worker declares the chat state stores it does not provide.
663
+ runPersistenceConformance('app-r2', () => generationPersistence(env), {
664
+ skip: ['messages', 'runs', 'interrupts', 'metadata'],
665
+ })
666
+ ```
667
+
668
+ Run it against a Miniflare R2 + D1 binding with the migration applied, reset
669
+ between runs (see **ai-persistence/build-cloudflare-adapter** for the
670
+ `cloudflare:test` harness pattern). It exercises, among the rest:
671
+
672
+ - `put` then `get` round-trips bytes and metadata; `get`/`head` return `null` for
673
+ a missing key; `delete` is a silent no-op on an absent key.
674
+ - `put` accepts a `ReadableStream` body with no declared length (a
675
+ `TransformStream`-wrapped stream) and records the real drained size — the
676
+ shape every URL-fetched artifact arrives in.
677
+ - `get` with a `range` returns just that slice, reports it as `range`, and
678
+ still reports the whole object's `size` — what a `206` / `Content-Range`
679
+ response is built from.
680
+ - `put` overwrites an existing key (and its `contentType`/`customMetadata`).
681
+ - `list` filters by `prefix` **literally and case-sensitively**, returns
682
+ ascending keys, pages through the `cursor` when `truncated` without gaps or
683
+ repeats, and returns an empty untruncated page for `limit: 0`.
684
+ - The `ArtifactStore`: `save` is insert-or-overwrite, `get` returns `null` when
685
+ absent, `list(runId)` returns `[]` for an unknown run, and `delete` /
686
+ `deleteForRun` remove exactly the expected rows.
687
+ - The `GenerationRunStore`: `createOrResume` idempotency, no-op `update` on an
688
+ unknown id, and `findLatestForThread` returning the most recently started
689
+ linked run (terminal ones included).
690
+
691
+ An end-to-end check is the strongest signal: run `generateImage` through
692
+ `withGenerationPersistence(generationPersistence(env), { threadId })`, then confirm the blob
693
+ exists at `artifacts/<runId>/<artifactId>` and `retrieveBlob` streams it back.