@proveanything/smartlinks 2.0.34 → 2.0.35

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/docs/ai.md CHANGED
@@ -1,1779 +1,1805 @@
1
- # SmartLinks AI
2
-
3
- Build AI-powered SmartLinks experiences with a practical SDK guide for responses, chat, product assistants, streaming, voice, and real-world integration patterns.
4
-
5
- ---
6
-
7
- ## Table of Contents
8
-
9
- - [Overview](#overview)
10
- - [Quick Start](#quick-start)
11
- - [Authentication](#authentication)
12
- - [Responses API](#responses-api)
13
- - [Chat Completions](#chat-completions)
14
- - [RAG: Product Assistants](#rag-product-assistants)
15
- - [Voice Integration](#voice-integration)
16
- - [Podcast Generation](#podcast-generation)
17
- - [Types & API reference](#types--api-reference)
18
- - [Usage Examples](#usage-examples)
19
- - [Error Handling](#error-handling)
20
- - [Rate Limiting](#rate-limiting)
21
- - [Best Practices](#best-practices)
22
- - [Providing Content to the AI Assistant](#providing-content-to-the-ai-assistant)
23
-
24
- ---
25
-
26
- ## Overview
27
-
28
- This guide is written for SDK users building real products, not backend operators. It focuses on the public SDK surface, recommended starting points, and examples you can adapt directly.
29
-
30
- ### Start with the path that matches your job
31
-
32
- | If you want to... | Start here |
33
- |---|---|
34
- | Build a new AI workflow | Use [Responses API](#responses-api) |
35
- | Add compatibility with existing chat clients | Use [Chat Completions](#chat-completions) |
36
- | Build a product/manual assistant | Use [RAG: Product Assistants](#rag-product-assistants) |
37
- | Add spoken input/output | Use [Voice Integration](#voice-integration) |
38
- | Add progressive rendering | Use [Streaming Responses](#streaming-responses) or [Streaming Chat](#streaming-chat) |
39
-
40
- > **Which AI chat API? Three surfaces, three jobs — don't confuse them:**
41
- > - **`ai.chat.responses`** — the **default for new work**. Structured input, tool use, multi-agent, streaming; you manage any history. Reach for this first.
42
- > - **`ai.chat.completions`** — **OpenAI-compatible** (`messages[]`). Use it *only* to drop into existing OpenAI-style code; prefer Responses for anything new.
43
- > - **`ai.public.chat`** (+ `getSession` / `clearSession`) — the **public product assistant** with **server-managed conversation sessions** (RAG-grounded). This is the "ongoing conversation" surface, and it's public-facing.
44
- >
45
- > Responses and Completions are *generation* APIs (not single-shot vs conversation — both take history); `ai.public.chat` is the one that keeps a conversation for you.
46
-
47
- ### Recommended starting points
48
-
49
- - New AI features: start with `ai.chat.responses.create(...)`
50
- - Product/manual assistants: start with `ai.public.chat(...)`
51
- - Existing OpenAI-style clients: use `ai.chat.completions.create(...)`
52
- - Real-time voice: use `ai.public.getToken(...)` and your provider's live client
53
-
54
- SmartLinks AI provides five main capabilities:
55
-
56
- 1. **Responses API** - Preferred API for agentic workflows, multimodal inputs, and tool-driven responses
57
- 2. **Chat Completions** - OpenAI-compatible text generation with streaming and tool calling
58
- 3. **RAG (Retrieval-Augmented Generation)** - Document-grounded Q&A for product assistants
59
- 4. **Voice Integration** - Voice-to-text and text-to-voice for hands-free interaction
60
- 5. **Podcast Generation** - NotebookLM-style multi-voice conversational podcasts from documents
61
-
62
- ### Key Features
63
-
64
- - ✅ Full TypeScript support with type safety
65
- - ✅ Streaming responses with async iterators
66
- - ✅ Automatic rate limit handling
67
- - ✅ Session management for conversations
68
- - ✅ Voice input/output helpers
69
- - ✅ Tool/function calling support
70
- - ✅ OpenAI-style Responses API support
71
- - ✅ Document indexing and retrieval
72
- - ✅ Customizable assistant behavior
73
-
74
- ---
75
-
76
- ## Quick Start
77
-
78
- If you're only reading one section, start here. The three snippets below cover the most common public SDK use cases.
79
-
80
- ### 1. Generate a response
81
-
82
- ```typescript
83
- import { initializeApi, ai } from '@proveanything/smartlinks';
84
-
85
- // Initialize the SDK
86
- initializeApi({
87
- baseURL: 'https://smartlinks.app/api/v1',
88
- apiKey: process.env.SMARTLINKS_API_KEY // Required for admin endpoints
89
- });
90
-
91
- // Preferred: create a response
92
- const response = await ai.chat.responses.create('my-collection', {
93
- model: 'google/gemini-2.5-flash',
94
- input: 'Summarize the key safety steps for descaling a coffee maker.'
95
- });
96
-
97
- console.log(response.output_text);
98
- ```
99
-
100
- ### 2. Build a product assistant
101
-
102
- ```typescript
103
- import { initializeApi, ai } from '@proveanything/smartlinks';
104
-
105
- initializeApi({ baseURL: 'https://smartlinks.app/api/v1' });
106
-
107
- const answer = await ai.public.chat('my-collection', {
108
- productId: 'coffee-maker-deluxe',
109
- userId: 'user-123',
110
- message: 'How do I descale this machine?'
111
- });
112
-
113
- console.log(answer.message);
114
- ```
115
-
116
- ### 3. Stream output into your UI
117
-
118
- ```typescript
119
- const stream = await ai.chat.responses.create('my-collection', {
120
- input: 'Write a launch checklist for a new product page.',
121
- stream: true
122
- });
123
-
124
- for await (const event of stream) {
125
- if (event.type === 'response.output_text.delta') {
126
- updateUi(event.delta);
127
- }
128
- }
129
- ```
130
-
131
- ---
132
-
133
- ## Authentication
134
-
135
- ### Admin Endpoints
136
-
137
- Admin endpoints require an API key passed during initialization:
138
-
139
- ```typescript
140
- initializeApi({
141
- baseURL: 'https://smartlinks.app/api/v1',
142
- apiKey: process.env.SMARTLINKS_API_KEY
143
- });
144
- ```
145
-
146
- The SDK automatically includes the API key in the `Authorization: Bearer <token>` header.
147
-
148
- ### Public Endpoints
149
-
150
- Public endpoints don't require an API key but are rate-limited by `userId`:
151
-
152
- ```typescript
153
- // No API key needed
154
- const response = await ai.public.chat('my-collection', {
155
- productId: 'coffee-maker',
156
- userId: 'user-123',
157
- message: 'How do I clean this?'
158
- });
159
- ```
160
-
161
- ---
162
-
163
- ## Responses API
164
-
165
- The Responses API is the recommended starting point for new integrations. Use it when you want a single endpoint for structured input, tool use, and streaming output.
166
-
167
- ### Basic Response
168
-
169
- ```typescript
170
- const response = await ai.chat.responses.create('my-collection', {
171
- model: 'google/gemini-2.5-flash',
172
- input: 'Write a friendly two-sentence welcome for a product assistant.'
173
- });
174
-
175
- console.log(response.output_text);
176
- ```
177
-
178
- ### Multimessage Input
179
-
180
- ```typescript
181
- const response = await ai.chat.responses.create('my-collection', {
182
- model: 'google/gemini-2.5-flash',
183
- input: [
184
- {
185
- role: 'system',
186
- content: [
187
- { type: 'input_text', text: 'You are a concise support assistant.' }
188
- ]
189
- },
190
- {
191
- role: 'user',
192
- content: [
193
- { type: 'input_text', text: 'Give me three troubleshooting steps for a grinder that will not start.' }
194
- ]
195
- }
196
- ]
197
- });
198
-
199
- console.log(response.output_text);
200
- ```
201
-
202
- ### Streaming Responses
203
-
204
- When you pass `stream: true`, the SDK returns an `AsyncIterable` of SSE events instead of a final JSON object. You do not need to parse raw SSE frames yourself — just iterate with `for await...of`.
205
-
206
- ```typescript
207
- const result = await ai.chat.responses.create('my-collection', {
208
- input: 'Summarize the manual',
209
- stream: true
210
- });
211
-
212
- for await (const event of result) {
213
- if (event.type === 'response.output_text.delta') {
214
- process.stdout.write(event.delta);
215
- }
216
- }
217
- ```
218
-
219
- If you omit `stream: true`, the same method returns the final `ResponsesResult` object instead.
220
-
221
- ```typescript
222
- const stream = await ai.chat.responses.create('my-collection', {
223
- model: 'google/gemini-2.5-flash',
224
- input: 'Explain how to descale an espresso machine step by step.',
225
- stream: true
226
- });
227
-
228
- for await (const event of stream) {
229
- if (event.type === 'response.output_text.delta') {
230
- process.stdout.write(event.delta);
231
- }
232
- }
233
- ```
234
-
235
- ### Tool Calling
236
-
237
- ```typescript
238
- const response = await ai.chat.responses.create('my-collection', {
239
- model: 'google/gemini-2.5-flash',
240
- input: 'What is the weather in Paris?',
241
- tools: [
242
- {
243
- type: 'function',
244
- name: 'get_weather',
245
- description: 'Get the current weather for a city',
246
- parameters: {
247
- type: 'object',
248
- properties: {
249
- location: { type: 'string' }
250
- },
251
- required: ['location']
252
- }
253
- }
254
- ]
255
- });
256
-
257
- console.log(response.output);
258
- ```
259
-
260
- ### Server-side tools (built-in agent loop)
261
-
262
- The example above is **client-relayed** tool calling: you define the tools, the model returns
263
- `tool_call` requests, and *your app* executes them and sends results back. For the common tools —
264
- reading and searching the web, vision, reading documents, generating images — the platform ships a
265
- curated, tested **built-in toolset it runs itself**. Opt in with `server_tools` and the server
266
- executes each tool and feeds the result back automatically, looping until the model has its answer.
267
- You get one final response; no relay code.
268
-
269
- ```typescript
270
- // Enable the whole built-in toolset:
271
- const res = await ai.chat.responses.create('my-collection', {
272
- model: 'balanced',
273
- input: 'Research acme.com and summarise what they sell, with their brand colours.',
274
- server_tools: true
275
- });
276
- console.log(res.output_text);
277
- console.log(res._agent.toolResults); // trace: which tools ran, with what result
278
- ```
279
-
280
- Scope it to specific tools (recommended — smaller blast radius, faster), by name or capability:
281
-
282
- ```typescript
283
- import { AI_TOOL_NAMES } from '@proveanything/smartlinks';
284
-
285
- const res = await ai.chat.responses.create('my-collection', {
286
- input: 'Find the current price of this product and return it as JSON.',
287
- server_tools: [AI_TOOL_NAMES.WEB_SEARCH, AI_TOOL_NAMES.DATA_EXTRACT],
288
- // or: allowCapabilities: ['web:read'], exclude: ['image.generate'],
289
- maxSteps: 6 // cap model round-trips (1–12, default 8)
290
- });
291
- ```
292
-
293
- **Streaming** surfaces tool progress as it happens — ideal for a "thinking…" UI. You get
294
- `agent.tool_call` / `agent.tool_result` events, then a final `response.completed`:
295
-
296
- ```typescript
297
- const stream = await ai.chat.responses.create('my-collection', {
298
- input: 'Research acme.com', server_tools: true, stream: true
299
- });
300
- for await (const ev of stream) {
301
- if (ev.type === 'agent.tool_call') showStep(`Running ${ev.name}…`);
302
- if (ev.type === 'agent.tool_result') showStep(`${ev.name} done`);
303
- if (ev.type === 'response.completed') render(ev.response.output_text);
304
- }
305
- ```
306
-
307
- `server_tools` can't be combined with `previous_response_id`/`conversation` yet — pass prior turns
308
- in `input`.
309
-
310
- #### Built-in tools
311
-
312
- | Tool | Does |
313
- |------|------|
314
- | `web.search` | Live web search → candidate results (url/title/description). |
315
- | `web.fetchPage` | Fetch a page → clean markdown + metadata + schema.org JSON-LD. |
316
- | `web.extractSchema` | Return only a page's schema.org data of a given `@type` (deterministic). |
317
- | `document.read` | Read a document at a URL — **PDF, deck, doc**, or article — into markdown. |
318
- | `data.extract` | Page + JSON-schema/prompt → **typed JSON** (turn a page into UI data). |
319
- | `brand.assets` | Extract a site's logo, colours, and design. |
320
- | `web.screenshot` | Screenshot a page → hosted image URL (feed to `image.describe`). |
321
- | `image.describe` | Vision: describe an image / read its text. |
322
- | `image.generate` | Generate an image from a prompt → hosted URL. |
323
- | `image.fromReference` | Image-to-image: generate guided by reference image(s). |
324
- | `image.searchStock` | Search real stock photos (Unsplash). |
325
- | `image.transform` | Resize / crop / rotate / grayscale / format-convert / compress → hosted URL. |
326
- | `pdf.create` | Render HTML → PDF → hosted URL. |
327
- | `pdf.fill` | Fill an AcroForm PDF's fields (`{ field: value }`) → hosted URL. |
328
- | `pdf.merge` | Merge several PDFs into one, in order → hosted URL. |
329
- | `pdf.inspect` | Cheap, no-AI introspection: page count/sizes, which pages have a real text layer, raster present, and a routing hint (`text` vs `vision`). |
330
- | `pdf.render` | Rasterize one page — or just a region (`clip: { x0, y0, x1, y1 }`, 0–1 from top-left) — to a PNG at a chosen DPI → hosted image URL (permanent). Max 40 MP per call; larger requests return `code: "too_large"` + `suggestedDpi`. |
331
- | `image.ocr` | **Deterministic OCR** (Google Vision, no generative model): exact characters with per-word `confidence` (0–1) and pixel `bbox`, plus lines with spacing as printed. Input: `imageUrl`, or a PDF `url` + `page` + `dpi` + optional `clip`. Optional `languages` hints. Use for small print and text outlined to curves. |
332
- | `pdf.extract` | PDF → typed JSON in one call (schema and/or prompt). Auto-routes text vs vision. Optional per-field `confidence`/`source` (`includeConfidence`) and `bbox` (`includeBoxes`) in `fieldsMeta`. |
333
- | `pdf.decodeBarcodes` | Deterministically decode barcodes/QR on a page (WASM, no AI) → value + symbology + page + bbox + confidence. Use this for barcode digits, never vision. |
334
- | `pdf.inspectGraphics` | Prepress inspection: per-page path/image/outlined-text counts + colour spaces, plus named SPOT colours (e.g. "PANTONE 871 C"). No AI. |
335
- | `http.request` | SSRF-guarded outbound HTTP(S) to a public URL (call a REST API). |
336
- | `translate` | Translate text into one or more languages (generic, model-based). |
337
-
338
- **Working with PDFs.** The PDF tools split into *read/analyse* and *produce*:
339
-
340
- - **Read a whole PDF as text** → `document.read` (Firecrawl; good for prose/decks).
341
- - **Turn a PDF into structured fields** → `pdf.extract` (give it a JSON schema and/or a prompt; it
342
- returns typed JSON). It routes itself, but you can drive the route yourself: call `pdf.inspect`
343
- first (deterministic, no AI) to see whether each page has a real text layer, then `pdf.extract`
344
- (cheap text path) or `pdf.render` → `image.describe`/vision for curve-only or raster artwork.
345
- - **Read small print exactly** (INCI/allergen lists, net weight, text outlined to curves) → `image.ocr`
346
- with a PDF `url`, `page`, `dpi: 300–400` and a `clip` around the panel. It returns the characters it
347
- sees with a per-word confidence — it never "corrects" a misspelling the way a vision model can, so
348
- flag words below ~0.9 for human review rather than re-reading them with a model.
349
- - **Zoom in visually** (layout, artwork) → `pdf.render` with `clip` to rasterise just a region at a high
350
- DPI, instead of the whole sheet. Renders over 40 MP return `code: "too_large"` with `suggestedDpi`
351
- — retry at that dpi or with a smaller clip, don't blind-retry.
352
- - **Barcodes / QR** → `pdf.decodeBarcodes` (deterministic WASM decode) — never trust vision for barcode
353
- digits; it hallucinates them.
354
- - **Prepress / print QA** (spot colours, colour spaces, vector vs raster) → `pdf.inspectGraphics`.
355
- - **A review UI that flags guessed fields** → `pdf.extract` with `includeConfidence` (per-field
356
- confidence + source) and `includeBoxes` (per-field bbox on text-native pages) in `fieldsMeta`.
357
- - **Produce a PDF** → `pdf.create` (HTML → PDF), `pdf.fill` (populate an AcroForm's fields),
358
- `pdf.merge` (combine several). These return a hosted `hostedUrl` (`pdf.create` also returns it as
359
- `url`). `pdf.create` loads remote `<img>` URLs before rendering and honours `page-break-inside`; pass
360
- `margin: "0"` when your HTML sets its own body margin (default page margin is 18mm/14mm).
361
-
362
- Every tool is also directly callable without the model loop via `ai.tools.run(collectionId, name,
363
- args)` — e.g. render a page or extract fields straight from a UI, no agent round-trip:
364
-
365
- ```ts
366
- // Read the ingredients panel of a sleeve exactly (deterministic OCR of a clipped region)
367
- const { result } = await SL.ai.tools.run(collectionId, 'image.ocr', {
368
- url: sleevePdfUrl, page: 1, dpi: 400,
369
- clip: { x0: 0.40, y0: 0.20, x1: 0.60, y1: 0.40 },
370
- languages: ['en', 'fr'],
371
- })
372
- // result.lines → [{ text: 'INGREDIENTS: Aqua, Glycerin, …', confidence: 0.97, bbox: {…} }]
373
- const toReview = result.words.filter((w) => w.confidence < 0.9)
374
- ```
375
-
376
- Discover tools two ways:
377
- - **Design time (typed):** import `BUILTIN_AI_TOOLS`, `AI_TOOL_NAMES`, and the per-tool arg types
378
- (`WebSearchArgs`, `DataExtractArgs`, `PdfExtractArgs`, …) from the SDK. This is the core set —
379
- stable, versioned, documented here.
380
- - **Runtime (live):** `await ai.catalog(collectionId)` returns the registry as the server sees it,
381
- including any future app-contributed tools. The built-in set above is always present.
382
-
383
- ### Recommended Models
384
-
385
- For agentic workflows on `v1/responses`, GPT-5.6 ships in three tiers. Pass either the full model
386
- id (`openai/gpt-5.6-sol`) or the shorthand tier alias (`'cheap'`, `'balanced'`, `'premium'`) as
387
- `model` — both resolve through the same server-side model registry.
388
-
389
- | Tier | Model | Alias | Use for |
390
- |------|-------|-------|---------|
391
- | Cheapest (default) | `openai/gpt-5.6-luna` | `'cheap'` / omit `model` | Everyday chat, high-volume/simple turns |
392
- | Balanced | `openai/gpt-5.6-terra` | `'balanced'` | Admin agent / tool-heavy setup workflows (route default) |
393
- | Flagship | `openai/gpt-5.6-sol` | `'premium'` | Most demanding reasoning, coding, and multi-step tool use |
394
-
395
- Every tier supports function/tool calling, Programmatic Tool Calling, and Multi-agent (see below), and all
396
- three accept `service_tier: 'flex'` (≈50% cheaper, Batch-API rates, slower) or `'priority'` (premium, guaranteed
397
- throughput) alongside the default `'standard'` tier. Pass it straight through in the request body:
398
-
399
- ```typescript
400
- const response = await ai.chat.responses.create('my-collection', {
401
- model: 'balanced',
402
- input: 'Summarize this week\'s ingested documents.',
403
- service_tier: 'flex' // non-interactive/background work: cheaper, but can take minutes
404
- });
405
- ```
406
-
407
- Flex requests run noticeably slower — reserve `flex` for non-production/background work (batch enrichment,
408
- evaluations, async jobs), not user-facing chat turns. The server already raises its own timeout to 15
409
- minutes for any request with `service_tier: 'flex'`, matching OpenAI's guidance — nothing to configure on
410
- the client. On a `429` ("resource unavailable") the request wasn't charged; retry with backoff or drop
411
- `service_tier` (or set it to `'auto'`) to fall back to standard processing.
412
-
413
- ### Programmatic Tool Calling
414
-
415
- Lets the model write a short in-memory program that orchestrates several tool calls (filtering,
416
- looping, aggregating) before handing back one result, instead of a full message round-trip per
417
- call. Add the hosted `programmatic_tool_calling` tool, and mark each function tool the program is
418
- allowed to invoke with `allowed_callers: ['programmatic']`:
419
-
420
- ```typescript
421
- const response = await ai.chat.responses.create('my-collection', {
422
- model: 'balanced',
423
- input: 'Compare inventory with demand for every SKU in this collection.',
424
- tools: [
425
- {
426
- type: 'function',
427
- name: 'get_inventory',
428
- description: 'Return available_units for a sku',
429
- parameters: { type: 'object', properties: { sku: { type: 'string' } }, required: ['sku'] },
430
- output_schema: { type: 'object', properties: { sku: { type: 'string' }, available_units: { type: 'number' } } },
431
- allowed_callers: ['programmatic']
432
- },
433
- {
434
- type: 'function',
435
- name: 'get_demand',
436
- description: 'Return requested_units for a sku',
437
- parameters: { type: 'object', properties: { sku: { type: 'string' } }, required: ['sku'] },
438
- output_schema: { type: 'object', properties: { sku: { type: 'string' }, requested_units: { type: 'number' } } },
439
- allowed_callers: ['programmatic']
440
- },
441
- { type: 'programmatic_tool_calling' }
442
- ]
443
- });
444
- ```
445
-
446
- The response's `output` array can include a `program` item (the generated code), one
447
- `function_call` item per program-issued call (tagged with `caller: { type: 'program', caller_id }`),
448
- and a `program_output` item with the program's final `result`/`status`. Execute the `function_call`
449
- items as normal and send results back keyed by their `call_id` on the next turn.
450
-
451
- ### Multi-agent (subagents)
452
-
453
- Lets the root agent spawn a tree of subagents that run in parallel and get synthesized back into
454
- one response — useful for tasks that split cleanly into independent workstreams (e.g. research +
455
- draft + review). Enable it by setting `multi_agent.enabled: true` — that's the entire contract on
456
- your end; nothing else to configure or pass.
457
-
458
- This is upstream beta functionality (OpenAI may still be limiting it to certain accounts/rollout),
459
- so treat behavior and availability as subject to change until it's GA.
460
-
461
- ```typescript
462
- const response = await ai.chat.responses.create('my-collection', {
463
- model: 'premium', // Sol recommended for the root agent when spawning subagents
464
- input: 'Research three competitor loyalty programs, then draft a comparison summary.',
465
- multi_agent: { enabled: true, max_concurrent_subagents: 3 }
466
- });
467
- ```
468
-
469
- Subagent output arrives as additional items in `output`, each tagged `agent: { agent_name: '/root/<name>' }`
470
- (root agent output uses `/root`). Notes/limits carried over from OpenAI's beta:
471
-
472
- - `max_concurrent_subagents` defaults to `3`
473
- - `reasoning.summary` and `max_tool_calls` are not supported while multi-agent is enabled
474
- - Streaming is not supported while multi-agent is enabled — the SDK throws if you pass both
475
- `stream: true` and `multi_agent.enabled: true`
476
-
477
- ---
478
-
479
- ## Chat Completions
480
-
481
- OpenAI-compatible chat completions with streaming and tool calling support. Use this for compatibility with existing Chat Completions integrations; prefer the Responses API for new agentic features.
482
-
483
- ### Basic Chat
484
-
485
- ```typescript
486
- const response = await ai.chat.completions.create('my-collection', {
487
- model: 'google/gemini-2.5-flash',
488
- messages: [
489
- { role: 'system', content: 'You are a helpful assistant.' },
490
- { role: 'user', content: 'What is the capital of France?' }
491
- ]
492
- });
493
-
494
- console.log(response.choices[0].message.content);
495
- // Output: "The capital of France is Paris."
496
- ```
497
-
498
- ### Streaming Chat
499
-
500
- Stream responses in real-time for better UX:
501
-
502
- - Set `stream: true`
503
- - The SDK returns an `AsyncIterable<ChatCompletionChunk>`
504
- - Iterate over chunks with `for await...of`
505
- - Read incremental text from `chunk.choices[0]?.delta?.content`
506
- - If `stream` is omitted or `false`, the method returns the normal `ChatCompletionResponse`
507
-
508
- ```typescript
509
- const stream = await ai.chat.completions.create('my-collection', {
510
- model: 'google/gemini-2.5-flash',
511
- messages: [
512
- { role: 'user', content: 'Write a short poem about coding' }
513
- ],
514
- stream: true
515
- });
516
-
517
- for await (const chunk of stream) {
518
- const content = chunk.choices[0]?.delta?.content || '';
519
- process.stdout.write(content);
520
- }
521
- ```
522
-
523
- ### Tool/Function Calling
524
-
525
- Define tools (functions) that the AI can call:
526
-
527
- ```typescript
528
- const tools = [
529
- {
530
- type: 'function',
531
- function: {
532
- name: 'get_weather',
533
- description: 'Get the current weather for a location',
534
- parameters: {
535
- type: 'object',
536
- properties: {
537
- location: {
538
- type: 'string',
539
- description: 'City name'
540
- },
541
- unit: {
542
- type: 'string',
543
- enum: ['celsius', 'fahrenheit']
544
- }
545
- },
546
- required: ['location']
547
- }
548
- }
549
- }
550
- ];
551
-
552
- const response = await ai.chat.completions.create('my-collection', {
553
- model: 'google/gemini-2.5-flash',
554
- messages: [
555
- { role: 'user', content: 'What\'s the weather in Paris?' }
556
- ],
557
- tools
558
- });
559
-
560
- const toolCall = response.choices[0].message.tool_calls?.[0];
561
- if (toolCall) {
562
- console.log('Function:', toolCall.function.name);
563
- console.log('Arguments:', JSON.parse(toolCall.function.arguments));
564
- // { location: "Paris", unit: "celsius" }
565
- }
566
- ```
567
-
568
- ### Available Models
569
-
570
- ```typescript
571
- // List all available models
572
- const models = await ai.models.list('my-collection');
573
-
574
- // Or filter by provider / capability
575
- const openAiModels = await ai.models.list('my-collection', {
576
- provider: 'openai'
577
- });
578
-
579
- const visionModels = await ai.models.list('my-collection', {
580
- capability: 'vision'
581
- });
582
-
583
- models.data.forEach(model => {
584
- console.log(`${model.name}`);
585
- console.log(` Provider: ${model.provider}`);
586
- console.log(` Context: ${model.contextWindow} tokens`);
587
- console.log(` Pricing: $${model.pricing.input}/1M input tokens`);
588
- });
589
-
590
- // Get specific model info
591
- const model = await ai.models.get('my-collection', 'google/gemini-2.5-flash');
592
- console.log(model.capabilities); // ['text', 'vision', 'audio', 'code']
593
- ```
594
-
595
- Use `ai.models.list(collectionId)` as the source of truth for what your collection can use at runtime. The public docs provide recommendations, but actual availability depends on the SmartLinks model catalog exposed to that collection.
596
-
597
- **Recommended Models:**
598
-
599
- | Model | Use Case | Speed | Cost |
600
- |-------|----------|-------|------|
601
- | `openai/gpt-5.4` | Default for new agentic and structured-output workflows | Balanced | Medium |
602
- | `openai/gpt-5-mini` | Lower-cost general purpose and JSON tasks | Fast | Low |
603
- | `google/gemini-2.5-flash` | Fast multimodal and cost-sensitive general use | Fast | Low |
604
- | `google/gemini-2.5-pro` | Complex reasoning and heavier multimodal tasks | Slower | Higher |
605
-
606
- If you want a safe default for most new work, start with `openai/gpt-5.4`. If you want a lower-cost fallback, use `openai/gpt-5-mini` or `google/gemini-2.5-flash` depending on your latency and pricing goals.
607
-
608
- ---
609
-
610
- ## RAG: Product Assistants
611
-
612
- Create intelligent product assistants that answer questions based on product documentation.
613
-
614
- ### Setup: Index Documents
615
-
616
- First, index your product documentation:
617
-
618
- ```typescript
619
- // Index a product manual from URL
620
- const result = await ai.rag.indexDocument('my-collection', {
621
- productId: 'coffee-maker-deluxe',
622
- documentUrl: 'https://example.com/manuals/coffee-maker.pdf',
623
- chunkSize: 500, // Tokens per chunk
624
- overlap: 50, // Token overlap between chunks
625
- provider: 'openai' // Embedding provider
626
- });
627
-
628
- console.log(`Indexed ${result.chunks} chunks`);
629
- console.log(`Dimensions: ${result.metadata.embeddingDimensions}`);
630
-
631
- // Or index from text directly
632
- await ai.rag.indexDocument('my-collection', {
633
- productId: 'coffee-maker-deluxe',
634
- text: 'Your product manual content here...',
635
- metadata: {
636
- source: 'manual',
637
- version: '2.0'
638
- }
639
- });
640
- ```
641
-
642
- ### Configure Assistant
643
-
644
- Customize the assistant's behavior:
645
-
646
- ```typescript
647
- await ai.rag.configureAssistant('my-collection', {
648
- productId: 'coffee-maker-deluxe',
649
- systemPrompt: 'You are a helpful coffee maker assistant. Be concise and friendly.',
650
- model: 'google/gemini-2.5-flash',
651
- temperature: 0.7,
652
- maxTokensPerResponse: 500,
653
- rateLimitPerUser: 20,
654
- allowedTopics: ['usage', 'cleaning', 'troubleshooting'],
655
- customInstructions: {
656
- tone: 'friendly',
657
- additionalRules: 'Always include safety warnings when relevant.'
658
- }
659
- });
660
- ```
661
-
662
- ### Public Chat
663
-
664
- Users can chat with the product assistant without authentication:
665
-
666
- ```typescript
667
- // First question
668
- const response = await ai.public.chat('my-collection', {
669
- productId: 'coffee-maker-deluxe',
670
- userId: 'user-123',
671
- message: 'How do I descale my coffee maker?'
672
- });
673
-
674
- console.log('Answer:', response.message);
675
- console.log('Used', response.context.chunksUsed, 'document sections');
676
- console.log('Top similarity:', response.context.topSimilarity);
677
- ```
678
-
679
- ### Conversation History
680
-
681
- Maintain conversation context with sessions:
682
-
683
- ```typescript
684
- const sessionId = `session-${Date.now()}`;
685
-
686
- // First question
687
- const q1 = await ai.public.chat('my-collection', {
688
- productId: 'coffee-maker-deluxe',
689
- userId: 'user-123',
690
- message: 'How do I clean it?',
691
- sessionId
692
- });
693
-
694
- // Follow-up question (uses history)
695
- const q2 = await ai.public.chat('my-collection', {
696
- productId: 'coffee-maker-deluxe',
697
- userId: 'user-123',
698
- message: 'How often should I do that?',
699
- sessionId
700
- });
701
-
702
- // Get full conversation history
703
- const session = await ai.public.getSession('my-collection', sessionId);
704
- console.log('Messages:', session.messages);
705
- console.log('Total messages:', session.messageCount);
706
-
707
- // Clear session when done
708
- await ai.public.clearSession('my-collection', sessionId);
709
- ```
710
-
711
- ### Session Management
712
-
713
- ```typescript
714
- // Get session statistics (admin)
715
- const stats = await ai.sessions.stats('my-collection');
716
- console.log('Total sessions:', stats.totalSessions);
717
- console.log('Active sessions:', stats.activeSessions);
718
- console.log('Total messages:', stats.totalMessages);
719
- console.log('Rate-limited users:', stats.rateLimitedUsers);
720
- ```
721
-
722
- ---
723
-
724
- ## Voice Integration
725
-
726
- Enable voice input and output for hands-free interaction.
727
-
728
- ### Voice Patterns
729
-
730
- The SDK supports three practical voice patterns:
731
-
732
- | Pattern | Best For | SDK Building Blocks |
733
- |---------|----------|---------------------|
734
- | Voice → Text → AI → Text | Manual helper Q&A, troubleshooting steps | `ai.voice.listen()` + `ai.public.chat()` |
735
- | Voice → Text → AI → Voice | Hands-free assistants, accessibility | `ai.voice.listen()` + `ai.public.chat()` + `ai.voice.speak()` or `ai.tts.generate()` |
736
- | Real-time Voice | Low-latency spoken conversation | `ai.public.getToken()` + Gemini Live client |
737
-
738
- ### Current SDK Support
739
-
740
- - `ai.voice.listen()` and `ai.voice.speak()` are browser helpers built on the Web Speech APIs.
741
- - `ai.public.getToken()` generates ephemeral tokens for Gemini Live sessions.
742
- - `ai.tts.generate()` supports server-side text-to-speech generation.
743
- - The SDK does not currently expose a first-class transcription endpoint like Whisper; if you need that flow, implement it as your own backend endpoint and feed the transcribed text into `ai.public.chat()` or `ai.chat.responses.create()`.
744
-
745
- ### Recommended Approach
746
-
747
- For product assistants and RAG-backed support, start with Voice → Text → AI → Text/Voice. It gives you the best control over retrieval, session history, and cost. Use Gemini Live when low-latency spoken conversation matters more than deep document grounding.
748
-
749
- ### Browser Voice Helpers
750
-
751
- These helpers are browser-only and rely on native speech recognition / speech synthesis support.
752
-
753
- ```typescript
754
- // Check if voice is supported
755
- if (ai.voice.isSupported()) {
756
- // Listen for voice input
757
- const question = await ai.voice.listen('en-US');
758
- console.log('User said:', question);
759
-
760
- // Get answer from AI
761
- const response = await ai.public.chat('my-collection', {
762
- productId: 'coffee-maker-deluxe',
763
- userId: 'user-123',
764
- message: question
765
- });
766
-
767
- // Speak the answer
768
- await ai.voice.speak(response.message, {
769
- voice: 'alloy',
770
- rate: 1.0
771
- });
772
- }
773
- ```
774
-
775
- ### Voice Assistant Class
776
-
777
- Create a complete voice assistant:
778
-
779
- ```typescript
780
- class ProductVoiceAssistant {
781
- private collectionId: string;
782
- private productId: string;
783
- private userId: string;
784
- private sessionId: string;
785
-
786
- constructor(config: {
787
- collectionId: string;
788
- productId: string;
789
- userId: string;
790
- }) {
791
- this.collectionId = config.collectionId;
792
- this.productId = config.productId;
793
- this.userId = config.userId;
794
- this.sessionId = `voice-${Date.now()}`;
795
- }
796
-
797
- async ask(): Promise<string> {
798
- // Listen for question
799
- console.log('Listening...');
800
- const question = await ai.voice.listen();
801
-
802
- // Get answer
803
- console.log('Processing...');
804
- const response = await ai.public.chat(this.collectionId, {
805
- productId: this.productId,
806
- userId: this.userId,
807
- message: question,
808
- sessionId: this.sessionId
809
- });
810
-
811
- // Speak answer
812
- console.log('Speaking...');
813
- await ai.voice.speak(response.message);
814
-
815
- return response.message;
816
- }
817
-
818
- async getRemainingQuestions(): Promise<number> {
819
- const status = await ai.public.getRateLimit(this.collectionId, this.userId);
820
- return status.remaining;
821
- }
822
- }
823
-
824
- // Usage
825
- const assistant = new ProductVoiceAssistant({
826
- collectionId: 'my-collection',
827
- productId: 'coffee-maker-deluxe',
828
- userId: 'user-123'
829
- });
830
-
831
- await assistant.ask(); // Voice question → Voice answer
832
- const remaining = await assistant.getRemainingQuestions();
833
- console.log(`${remaining} questions remaining`);
834
- ```
835
-
836
- ### Gemini Live Integration
837
-
838
- Generate ephemeral tokens for Gemini Live (multimodal voice):
839
-
840
- Use this path for real-time voice sessions. The SDK only issues the short-lived token; the actual live connection is made with the provider client.
841
-
842
- ```typescript
843
- // Generate token for voice session
844
- const token = await ai.public.getToken('my-collection', {
845
- settings: {
846
- ttl: 3600, // 1 hour
847
- voice: 'alloy',
848
- language: 'en-US'
849
- }
850
- });
851
-
852
- console.log('Token:', token.token);
853
- console.log('Expires at:', new Date(token.expiresAt));
854
-
855
- // Use token with Gemini Live API
856
- // (See Google's Gemini documentation)
857
- ```
858
-
859
- ### Voice + RAG Guidance
860
-
861
- For document-grounded assistants, prefer this pattern:
862
-
863
- 1. Capture voice with `ai.voice.listen()` or your own transcription flow.
864
- 2. Send the transcribed text to `ai.public.chat()`.
865
- 3. Render the text response for readability.
866
- 4. Optionally speak the answer with `ai.voice.speak()` or `ai.tts.generate()`.
867
-
868
- This is usually a better fit for manuals and procedural guidance than trying to use a live voice session as the primary retrieval layer.
869
-
870
- ---
871
-
872
- ## Podcast Generation
873
-
874
- Generate NotebookLM-style multi-voice conversational podcasts from product documentation.
875
-
876
- ### Generate a Podcast
877
-
878
- ```typescript
879
- const podcast = await ai.podcast.generate('my-collection', {
880
- productId: 'coffee-maker-deluxe',
881
- duration: 5, // Target 5 minutes
882
- style: 'casual', // 'casual' | 'professional' | 'educational' | 'entertaining'
883
- voices: {
884
- host1: 'nova', // Female voice
885
- host2: 'onyx' // Male voice
886
- },
887
- includeAudio: true // Generate audio files
888
- });
889
-
890
- console.log('Podcast Title:', podcast.script.title);
891
- console.log('Duration:', podcast.metadata.duration, 'seconds');
892
- console.log('Download:', podcast.audio?.mixedUrl);
893
- ```
894
-
895
- ### Available Voices
896
-
897
- | Voice | Gender | Personality | Best For |
898
- |-------|--------|-------------|----------|
899
- | `alloy` | Neutral | Balanced, neutral | Professional podcasts |
900
- | `echo` | Male | Clear, authoritative | Expert/teacher role |
901
- | `fable` | Neutral | Warm, storytelling | Narrative content |
902
- | `onyx` | Male | Deep, engaging | Main host, discussions |
903
- | `nova` | Female | Friendly, enthusiastic | Co-host, questions |
904
- | `shimmer` | Female | Bright, energetic | Entertaining content |
905
-
906
- **Recommended Combinations:**
907
- - **Casual**: Nova + Onyx - Friendly and engaging
908
- - **Professional**: Alloy + Echo - Authoritative and clear
909
- - **Educational**: Fable + Echo - Teaching style
910
- - **Entertaining**: Shimmer + Onyx - High energy
911
-
912
- ### Access the Script
913
-
914
- ```typescript
915
- // View the generated script
916
- podcast.script.segments.forEach((segment, i) => {
917
- const speaker = segment.speaker === 'host1' ? 'Host 1' : 'Host 2';
918
- console.log(`${speaker}: ${segment.text}`);
919
- });
920
- ```
921
-
922
- ### Check Generation Status
923
-
924
- For long-running podcast generation, poll for status:
925
-
926
- ```typescript
927
- // Start generation
928
- const podcast = await ai.podcast.generate('my-collection', {
929
- productId: 'coffee-maker-deluxe',
930
- duration: 10,
931
- includeAudio: true
932
- });
933
-
934
- // Poll for status
935
- const checkStatus = async () => {
936
- const status = await ai.podcast.getStatus('my-collection', podcast.podcastId);
937
-
938
- console.log(`Status: ${status.status} (${status.progress}%)`);
939
-
940
- if (status.status === 'completed' && status.result) {
941
- console.log('Podcast ready!');
942
- console.log('Listen:', status.result.audio?.mixedUrl);
943
- return true;
944
- } else if (status.status === 'failed') {
945
- console.error('Generation failed:', status.error);
946
- return true;
947
- }
948
-
949
- return false;
950
- };
951
-
952
- // Check every 5 seconds
953
- const interval = setInterval(async () => {
954
- const done = await checkStatus();
955
- if (done) clearInterval(interval);
956
- }, 5000);
957
- ```
958
-
959
- ### Text-to-Speech (TTS)
960
-
961
- Generate custom audio from text:
962
-
963
- ```typescript
964
- const audioBlob = await ai.tts.generate('my-collection', {
965
- text: 'Welcome to our podcast about coffee makers!',
966
- voice: 'nova',
967
- speed: 1.0,
968
- format: 'mp3'
969
- });
970
-
971
- // Create audio URL for playback
972
- const audioUrl = URL.createObjectURL(audioBlob);
973
- ```
974
-
975
- ---
976
-
977
-
978
- ## Types & API reference
979
-
980
- Full TypeScript types and the per-endpoint HTTP reference for every `SL.ai.*` method live in the
981
- **generated** references — they're not duplicated here so they can't drift:
982
-
983
- - **[`API_SUMMARY.md`](API_SUMMARY.md)** — every function signature + type (search `ai.`).
984
- - **[`openapi.yaml`](../openapi.yaml)** — the raw HTTP endpoints.
985
-
986
- For inline types, use LSP hover / go-to-definition on `SL.ai.*`.
987
-
988
- ## Usage Examples
989
-
990
- ### Example 1: Product FAQ Bot
991
-
992
- ```typescript
993
- async function createProductFAQ() {
994
- const collectionId = 'my-collection';
995
- const productId = 'coffee-maker-deluxe';
996
-
997
- // 1. Index product documentation
998
- await ai.rag.indexDocument(collectionId, {
999
- productId,
1000
- documentUrl: 'https://example.com/manual.pdf'
1001
- });
1002
-
1003
- // 2. Configure assistant
1004
- await ai.rag.configureAssistant(collectionId, {
1005
- productId,
1006
- systemPrompt: 'You are a coffee maker expert. Provide clear, step-by-step instructions.',
1007
- rateLimitPerUser: 30
1008
- });
1009
-
1010
- // 3. Answer user questions
1011
- const answer = await ai.public.chat(collectionId, {
1012
- productId,
1013
- userId: 'user-123',
1014
- message: 'How do I make espresso?'
1015
- });
1016
-
1017
- console.log(answer.message);
1018
- }
1019
- ```
1020
-
1021
- ### Example 2: Streaming Chatbot UI
1022
-
1023
- ```typescript
1024
- async function streamingChatbot(userMessage: string) {
1025
- const stream = await ai.chat.completions.create('my-collection', {
1026
- model: 'google/gemini-2.5-flash',
1027
- messages: [
1028
- { role: 'system', content: 'You are a helpful assistant.' },
1029
- { role: 'user', content: userMessage }
1030
- ],
1031
- stream: true
1032
- });
1033
-
1034
- let fullResponse = '';
1035
-
1036
- for await (const chunk of stream) {
1037
- const content = chunk.choices[0]?.delta?.content || '';
1038
- fullResponse += content;
1039
-
1040
- // Update UI in real-time
1041
- updateChatUI(content);
1042
- }
1043
-
1044
- return fullResponse;
1045
- }
1046
- ```
1047
-
1048
- ### Example 3: Multi-Turn Conversation
1049
-
1050
- ```typescript
1051
- async function chatConversation() {
1052
- const collectionId = 'my-collection';
1053
- const sessionId = `chat-${Date.now()}`;
1054
- const userId = 'user-123';
1055
- const productId = 'coffee-maker-deluxe';
1056
-
1057
- // Question 1
1058
- const a1 = await ai.public.chat(collectionId, {
1059
- productId,
1060
- userId,
1061
- message: 'How do I clean the machine?',
1062
- sessionId
1063
- });
1064
- console.log('A1:', a1.message);
1065
-
1066
- // Question 2 (references previous context)
1067
- const a2 = await ai.public.chat(collectionId, {
1068
- productId,
1069
- userId,
1070
- message: 'How often should I do that?',
1071
- sessionId
1072
- });
1073
- console.log('A2:', a2.message);
1074
-
1075
- // Get full history
1076
- const session = await ai.public.getSession(collectionId, sessionId);
1077
- console.log('Full conversation:', session.messages);
1078
- }
1079
- ```
1080
-
1081
- ### Example 4: React Hook for Product Assistant
1082
-
1083
- ```typescript
1084
- import { useState, useCallback } from 'react';
1085
- import { ai } from '@proveanything/smartlinks';
1086
-
1087
- export function useProductAssistant(
1088
- collectionId: string,
1089
- productId: string,
1090
- userId: string
1091
- ) {
1092
- const [loading, setLoading] = useState(false);
1093
- const [error, setError] = useState<string | null>(null);
1094
- const [rateLimit, setRateLimit] = useState({ remaining: 20, limit: 20 });
1095
-
1096
- const ask = useCallback(async (message: string) => {
1097
- setLoading(true);
1098
- setError(null);
1099
-
1100
- try {
1101
- const response = await ai.public.chat(collectionId, {
1102
- productId,
1103
- userId,
1104
- message
1105
- });
1106
-
1107
- setRateLimit(prev => ({
1108
- ...prev,
1109
- remaining: prev.remaining - 1
1110
- }));
1111
-
1112
- return response.message;
1113
- } catch (err: any) {
1114
- setError(err.message);
1115
- throw err;
1116
- } finally {
1117
- setLoading(false);
1118
- }
1119
- }, [collectionId, productId, userId]);
1120
-
1121
- return { ask, loading, error, rateLimit };
1122
- }
1123
-
1124
- // Usage in component
1125
- function ProductHelp() {
1126
- const { ask, loading, rateLimit } = useProductAssistant(
1127
- 'my-collection',
1128
- 'coffee-maker',
1129
- 'user-123'
1130
- );
1131
- const [answer, setAnswer] = useState('');
1132
-
1133
- const handleAsk = async () => {
1134
- const response = await ask('How do I clean this?');
1135
- setAnswer(response);
1136
- };
1137
-
1138
- return (
1139
- <div>
1140
- <button onClick={handleAsk} disabled={loading}>
1141
- {loading ? 'Asking...' : 'Ask Question'}
1142
- </button>
1143
- {answer && <p>{answer}</p>}
1144
- <p>{rateLimit.remaining} questions remaining</p>
1145
- </div>
1146
- );
1147
- }
1148
- ```
1149
-
1150
-
1151
- ---
1152
-
1153
- ## Error Handling
1154
-
1155
- ### Error Codes
1156
-
1157
- | Code | Type | HTTP Status | Description |
1158
- |------|------|-------------|-------------|
1159
- | `rate_limit_exceeded` | `rate_limit_error` | 429 | User exceeded rate limit |
1160
- | `invalid_request` | `invalid_request_error` | 400 | Invalid parameters |
1161
- | `authentication_error` | `authentication_error` | 401 | Invalid/missing API key |
1162
- | `permission_denied` | `permission_error` | 403 | Insufficient permissions |
1163
- | `not_found` | `not_found_error` | 404 | Resource not found |
1164
- | `document_not_found` | `not_found_error` | 404 | Product document not indexed |
1165
- | `server_error` | `server_error` | 500 | Internal server error |
1166
- | `service_unavailable` | `server_error` | 503 | Service temporarily unavailable |
1167
-
1168
- ### Error Handling Pattern
1169
-
1170
- ```typescript
1171
- import { SmartLinksAIError } from '@proveanything/smartlinks';
1172
-
1173
- async function robustChat() {
1174
- try {
1175
- const response = await ai.public.chat('my-collection', {
1176
- productId: 'coffee-maker',
1177
- userId: 'user-123',
1178
- message: 'Help!'
1179
- });
1180
- return response.message;
1181
- } catch (error) {
1182
- if (error instanceof SmartLinksAIError) {
1183
- switch (error.code) {
1184
- case 'rate_limit_exceeded':
1185
- console.error('Rate limit exceeded');
1186
- console.log('Try again at:', new Date(error.resetAt!));
1187
- break;
1188
- case 'document_not_found':
1189
- console.error('Product manual not indexed yet');
1190
- break;
1191
- case 'authentication_error':
1192
- console.error('Invalid API key');
1193
- break;
1194
- case 'invalid_request':
1195
- console.error('Invalid request:', error.message);
1196
- break;
1197
- default:
1198
- console.error('API error:', error.message);
1199
- }
1200
- } else {
1201
- console.error('Unexpected error:', error);
1202
- }
1203
- throw error;
1204
- }
1205
- }
1206
- ```
1207
-
1208
- ### Rate Limit Retry
1209
-
1210
- ```typescript
1211
- async function chatWithRetry(
1212
- request: PublicChatRequest,
1213
- maxRetries = 3
1214
- ) {
1215
- let retries = 0;
1216
-
1217
- while (retries < maxRetries) {
1218
- try {
1219
- return await ai.public.chat('my-collection', request);
1220
- } catch (error) {
1221
- if (error instanceof SmartLinksAIError && error.isRateLimitError()) {
1222
- if (retries === maxRetries - 1) throw error;
1223
-
1224
- const resetTime = new Date(error.resetAt!).getTime();
1225
- const waitTime = resetTime - Date.now();
1226
-
1227
- console.log(`Rate limited. Waiting ${waitTime}ms...`);
1228
- await new Promise(resolve => setTimeout(resolve, waitTime));
1229
-
1230
- retries++;
1231
- } else {
1232
- throw error;
1233
- }
1234
- }
1235
- }
1236
- }
1237
- ```
1238
-
1239
- ---
1240
-
1241
- ## Rate Limiting
1242
-
1243
- ### Rate Limit Overview
1244
-
1245
- Public endpoints are rate-limited per `userId`:
1246
-
1247
- | Endpoint Type | Default Limit | Window |
1248
- |--------------|---------------|--------|
1249
- | Public Chat | 20 requests | 1 hour |
1250
- | Token Generation | 10 requests | 1 hour |
1251
- | Admin Endpoints | Unlimited* | - |
1252
-
1253
- *Admin endpoints use API key authentication and are not rate-limited by default.
1254
-
1255
- ### Rate Limit Headers
1256
-
1257
- All API responses include rate limit information:
1258
-
1259
- ```
1260
- X-RateLimit-Limit: 20
1261
- X-RateLimit-Remaining: 15
1262
- X-RateLimit-Reset: 1707300000000
1263
- ```
1264
-
1265
- ### Checking Rate Limit
1266
-
1267
- ```typescript
1268
- // Check rate limit status before making requests
1269
- const status = await ai.public.getRateLimit('my-collection', 'user-123');
1270
-
1271
- console.log('Used:', status.used);
1272
- console.log('Remaining:', status.remaining);
1273
- console.log('Resets at:', new Date(status.resetAt));
1274
-
1275
- if (status.remaining > 0) {
1276
- // Safe to make request
1277
- await ai.public.chat(/* ... */);
1278
- } else {
1279
- // Show user when they can ask again
1280
- console.log('Rate limit reached. Try again at:', status.resetAt);
1281
- }
1282
- ```
1283
-
1284
- ### Resetting Rate Limits (Admin)
1285
-
1286
- ```typescript
1287
- // Reset rate limit for a specific user
1288
- await ai.rateLimit.reset('my-collection', 'user-123');
1289
- console.log('Rate limit reset for user-123');
1290
- ```
1291
-
1292
- ---
1293
-
1294
- ## Best Practices
1295
-
1296
- ### 1. Choose the Right Model
1297
-
1298
- ```typescript
1299
- // For most new workflows (recommended default)
1300
- await ai.chat.responses.create('my-collection', {
1301
- model: 'openai/gpt-5.4',
1302
- input: 'Create a concise onboarding checklist.'
1303
- });
1304
-
1305
- // For lower-cost structured or JSON-oriented work
1306
- await ai.chat.responses.create('my-collection', {
1307
- model: 'openai/gpt-5-mini',
1308
- input: 'Return a color palette as JSON.'
1309
- });
1310
-
1311
- // For fast multimodal or cost-sensitive general use
1312
- await ai.chat.completions.create('my-collection', {
1313
- model: 'google/gemini-2.5-flash',
1314
- messages: [...]
1315
- });
1316
-
1317
- // When you need the actual available catalog for this collection
1318
- const available = await ai.models.list('my-collection');
1319
- console.log(available.data.map(model => model.id));
1320
- ```
1321
-
1322
- ### 2. Use Streaming for Long Responses
1323
-
1324
- Improve perceived performance with streaming:
1325
-
1326
- ```typescript
1327
- // Non-streaming: User waits for full response
1328
- const response = await ai.chat.completions.create('my-collection', {
1329
- messages: [...]
1330
- });
1331
-
1332
- // Streaming: User sees progress immediately
1333
- const stream = await ai.chat.completions.create('my-collection', {
1334
- stream: true,
1335
- messages: [...]
1336
- });
1337
-
1338
- for await (const chunk of stream) {
1339
- updateUI(chunk.choices[0]?.delta?.content);
1340
- }
1341
- ```
1342
-
1343
- ### 3. Maintain Session Context
1344
-
1345
- Keep conversations coherent with session IDs:
1346
-
1347
- ```typescript
1348
- // Generate unique session ID per conversation
1349
- const sessionId = `user-${userId}-product-${productId}`;
1350
-
1351
- // All questions in same conversation use same sessionId
1352
- await ai.public.chat('my-collection', {
1353
- sessionId,
1354
- message: 'First question',
1355
- ...
1356
- });
1357
-
1358
- await ai.public.chat('my-collection', {
1359
- sessionId,
1360
- message: 'Follow-up question',
1361
- ...
1362
- });
1363
- ```
1364
-
1365
- ### 4. Handle Rate Limits Gracefully
1366
-
1367
- Show clear feedback to users:
1368
-
1369
- ```typescript
1370
- try {
1371
- await ai.public.chat('my-collection', {...});
1372
- } catch (error) {
1373
- if (error instanceof SmartLinksAIError && error.isRateLimitError()) {
1374
- const resetTime = new Date(error.resetAt!);
1375
- showNotification(
1376
- `You've reached your question limit. ` +
1377
- `Try again at ${resetTime.toLocaleTimeString()}`
1378
- );
1379
- }
1380
- }
1381
- ```
1382
-
1383
- ### 5. Optimize Voice UX
1384
-
1385
- Provide clear status updates:
1386
-
1387
- ```typescript
1388
- async function voiceAssistant() {
1389
- try {
1390
- showStatus('Listening...');
1391
- const question = await ai.voice.listen();
1392
-
1393
- showStatus('Processing...');
1394
- const answer = await ai.public.chat('my-collection', {
1395
- message: question,
1396
- ...
1397
- });
1398
-
1399
- showStatus('Speaking...');
1400
- await ai.voice.speak(answer.message);
1401
-
1402
- showStatus('Ready');
1403
- } catch (error) {
1404
- showStatus('Error', error.message);
1405
- }
1406
- }
1407
- ```
1408
-
1409
- ### 6. Chunk Large Documents
1410
-
1411
- For better RAG performance, chunk documents appropriately:
1412
-
1413
- ```typescript
1414
- // For technical manuals
1415
- await ai.rag.indexDocument('my-collection', {
1416
- productId: 'coffee-maker',
1417
- documentUrl: '...',
1418
- chunkSize: 500, // Smaller chunks for precise answers
1419
- overlap: 50 // Overlap maintains context
1420
- });
1421
-
1422
- // For narrative content
1423
- await ai.rag.indexDocument('my-collection', {
1424
- productId: 'coffee-maker',
1425
- documentUrl: '...',
1426
- chunkSize: 1000, // Larger chunks for coherent responses
1427
- overlap: 100
1428
- });
1429
- ```
1430
-
1431
- ### 7. Use System Prompts Effectively
1432
-
1433
- Provide clear instructions:
1434
-
1435
- ```typescript
1436
- await ai.chat.completions.create('my-collection', {
1437
- messages: [
1438
- {
1439
- role: 'system',
1440
- content: `You are a coffee maker expert assistant.
1441
- - Be concise and clear
1442
- - Use numbered lists for steps
1443
- - Always mention safety precautions
1444
- - If unsure, ask for clarification`
1445
- },
1446
- { role: 'user', content: 'How do I descale?' }
1447
- ]
1448
- });
1449
- ```
1450
-
1451
- ### 8. Monitor Usage and Costs
1452
-
1453
- Track usage for cost management:
1454
-
1455
- ```typescript
1456
- const response = await ai.chat.completions.create('my-collection', {
1457
- messages: [...]
1458
- });
1459
-
1460
- // Log token usage
1461
- console.log('Usage:', response.usage);
1462
- console.log('Prompt tokens:', response.usage.prompt_tokens);
1463
- console.log('Completion tokens:', response.usage.completion_tokens);
1464
- console.log('Total tokens:', response.usage.total_tokens);
1465
-
1466
- // Calculate estimated cost
1467
- const model = await ai.models.get('my-collection', response.model);
1468
- const cost =
1469
- (response.usage.prompt_tokens * model.pricing.input / 1_000_000) +
1470
- (response.usage.completion_tokens * model.pricing.output / 1_000_000);
1471
- console.log('Estimated cost: $', cost.toFixed(4));
1472
- ```
1473
-
1474
- ---
1475
-
1476
- ## Providing Content to the AI Assistant
1477
-
1478
- The portal's built-in AI assistant can discuss the content currently visible to the user. To make this work, your app must supply contextual content when the assistant requests it. There are three extraction methods, tried in priority order:
1479
-
1480
- | Priority | Method | When Used |
1481
- |----------|--------|-----------|
1482
- | 1 | **Direct prop callback** | Container/widget rendered in the parent React context |
1483
- | 2 | **PostMessage protocol** | App rendered in an iframe |
1484
- | 3 | **DOM text fallback** | Neither of the above responded |
1485
-
1486
- This section covers **methods 1 and 2** — the ones you implement in your app.
1487
-
1488
- ### Method 1: Direct Prop (`onRequestAIContent`)
1489
-
1490
- When your app is rendered as a **container** (a direct component in the parent's React tree), the framework passes an `onRequestAIContent` prop to your exported component. You do **not** call this prop yourself — the framework calls it when the AI assistant needs context.
1491
-
1492
- Structure your component to make current state accessible when the callback fires:
1493
-
1494
- ```tsx
1495
- import { useEffect, useRef } from 'react';
1496
-
1497
- export function PublicContainer(props) {
1498
- const { onRequestAIContent, appId, ...rest } = props;
1499
-
1500
- // Keep a ref to the latest content so the callback always returns fresh data
1501
- const currentContentRef = useRef(null);
1502
-
1503
- useEffect(() => {
1504
- currentContentRef.current = buildCurrentContent();
1505
- }, [relevantState]);
1506
-
1507
- return <div data-app-container={appId}>{/* your UI */}</div>;
1508
- }
1509
- ```
1510
-
1511
- The framework registers the content provider internally via `useAIContentExtraction.registerContentProvider()`. Your component just needs to respond when the callback is invoked.
1512
-
1513
- ### Method 2: PostMessage Protocol (Iframes)
1514
-
1515
- For iframe-embedded apps the framework sends a `postMessage` request and expects a response within **500 ms**. If your app doesn't reply in time the framework falls back to DOM text extraction.
1516
-
1517
- **Request sent by the framework to your iframe:**
1518
-
1519
- ```typescript
1520
- {
1521
- type: 'smartlinks:request-ai-content',
1522
- requestId: 'ai-content-1709834567890-abc123' // unique per request
1523
- }
1524
- ```
1525
-
1526
- **Response your app must send back:**
1527
-
1528
- ```typescript
1529
- {
1530
- type: 'smartlinks:ai-content-response',
1531
- requestId: 'ai-content-1709834567890-abc123', // echo back the requestId
1532
- content: AIContentResponse
1533
- }
1534
- ```
1535
-
1536
- **Minimal implementation:**
1537
-
1538
- ```typescript
1539
- window.addEventListener('message', async (event) => {
1540
- if (event.data?.type === 'smartlinks:request-ai-content') {
1541
- const content = await gatherAIContent();
1542
-
1543
- window.parent.postMessage({
1544
- type: 'smartlinks:ai-content-response',
1545
- requestId: event.data.requestId,
1546
- content,
1547
- }, '*');
1548
- }
1549
- });
1550
-
1551
- async function gatherAIContent(): Promise<AIContentResponse> {
1552
- return {
1553
- text: 'The user is viewing the warranty registration form...',
1554
- contentLabel: 'Warranty Registration',
1555
- };
1556
- }
1557
- ```
1558
-
1559
- ### The `AIContentResponse` Interface
1560
-
1561
- ```typescript
1562
- interface AIContentResponse {
1563
- /**
1564
- * Plain text or markdown for the AI to use as context.
1565
- * Injected into the system prompt. Max recommended: ~4000 characters.
1566
- */
1567
- text: string;
1568
-
1569
- /**
1570
- * Optional structured metadata (key-value pairs).
1571
- * Not used directly in prompts but available for custom providers.
1572
- */
1573
- metadata?: Record<string, unknown>;
1574
-
1575
- /**
1576
- * Pre-built RAG configuration. When provided, the assistant can use
1577
- * SL.ai.public.chat() to ground answers in indexed product documents.
1578
- */
1579
- ragHint?: {
1580
- /** The product ID whose indexed documents should be queried */
1581
- productId: string;
1582
- /** Optional session ID for multi-turn RAG conversations */
1583
- sessionId?: string;
1584
- /** Optional context hint to help scope the RAG query */
1585
- context?: string;
1586
- };
1587
-
1588
- /**
1589
- * Human-readable label shown in context-update messages
1590
- * (e.g. "Product Manual", "FAQ").
1591
- */
1592
- contentLabel?: string;
1593
-
1594
- /**
1595
- * How the assistant should use this content:
1596
- * - 'context' (default): inject text into the system prompt
1597
- * - 'rag': use ragHint to query indexed docs via SL.ai.public.chat()
1598
- * - 'hybrid': inject text as context AND ground answers via RAG
1599
- */
1600
- strategy?: 'context' | 'rag' | 'hybrid';
1601
- }
1602
- ```
1603
-
1604
- ### Response Strategies
1605
-
1606
- #### `context` (Default)
1607
-
1608
- Return readable text. The assistant injects it into the system prompt as background knowledge. Good for descriptions, summaries, structured data, and FAQs.
1609
-
1610
- ```typescript
1611
- return {
1612
- text: `
1613
- ## Wine Details
1614
- - **Name**: 2023 Château Margaux
1615
- - **Region**: Bordeaux, France
1616
- - **Tasting Notes**: Dark fruit, cedar, tobacco
1617
- - **Food Pairing**: Lamb, aged cheese
1618
- `,
1619
- contentLabel: 'Wine Information',
1620
- strategy: 'context',
1621
- };
1622
- ```
1623
-
1624
- #### `rag`
1625
-
1626
- Tell the assistant to query pre-indexed documents via SmartLinks RAG. The `text` field is minimal — the real knowledge comes from the indexed docs. Good for product manuals, large document sets, and technical specs.
1627
-
1628
- ```typescript
1629
- return {
1630
- text: 'The user is viewing the espresso machine product page.',
1631
- contentLabel: 'Product Assistant',
1632
- strategy: 'rag',
1633
- ragHint: {
1634
- productId: 'espresso-machine-pro',
1635
- sessionId: `rag-session-${userId}`,
1636
- context: 'User is on the troubleshooting section',
1637
- },
1638
- };
1639
- ```
1640
-
1641
- When the assistant receives a `rag` strategy it routes the question through `SL.ai.public.chat()`:
1642
-
1643
- ```typescript
1644
- const response = await SL.ai.public.chat(collectionId, {
1645
- productId: ragHint.productId,
1646
- userId: currentUserId,
1647
- message: userQuestion,
1648
- sessionId: ragHint.sessionId,
1649
- });
1650
- ```
1651
-
1652
- #### `hybrid`
1653
-
1654
- Combines both — the `text` is injected as additional context **and** the user's questions are also grounded via RAG. Good for museum exhibits, guided experiences, or any scenario with rich metadata alongside large document sets.
1655
-
1656
- ```typescript
1657
- return {
1658
- text: `
1659
- ## Exhibit: The Starry Night
1660
- - **Artist**: Vincent van Gogh
1661
- - **Year**: 1889
1662
- - **Current Location**: Gallery 3, East Wing
1663
- - **Audio Guide**: Available in 12 languages
1664
- `,
1665
- contentLabel: 'Museum Exhibit',
1666
- strategy: 'hybrid',
1667
- ragHint: {
1668
- productId: 'starry-night-exhibit',
1669
- context: 'Art history and technique questions',
1670
- },
1671
- };
1672
- ```
1673
-
1674
- ### Content Extraction Examples
1675
-
1676
- #### Museum Guide App
1677
-
1678
- ```typescript
1679
- async function gatherAIContent(): Promise<AIContentResponse> {
1680
- const exhibit = getCurrentExhibit();
1681
-
1682
- return {
1683
- text: `
1684
- Exhibit: "${exhibit.title}" by ${exhibit.artist}
1685
- Period: ${exhibit.period}
1686
- Medium: ${exhibit.medium}
1687
- Description: ${exhibit.curatorNotes}
1688
- Related works in this gallery: ${exhibit.relatedWorks.join(', ')}
1689
- `.trim(),
1690
- contentLabel: `Exhibit: ${exhibit.title}`,
1691
- strategy: 'hybrid',
1692
- ragHint: {
1693
- productId: exhibit.smartlinksProductId,
1694
- context: `Art history, technique, and visitor information for ${exhibit.title}`,
1695
- },
1696
- metadata: {
1697
- exhibitId: exhibit.id,
1698
- gallery: exhibit.gallery,
1699
- audioGuideAvailable: exhibit.hasAudioGuide,
1700
- },
1701
- };
1702
- }
1703
- ```
1704
-
1705
- #### Wine Product App
1706
-
1707
- ```typescript
1708
- async function gatherAIContent(): Promise<AIContentResponse> {
1709
- const wine = getCurrentWine();
1710
- const reviews = await fetchRecentReviews(wine.id, 5);
1711
-
1712
- return {
1713
- text: `
1714
- Wine: ${wine.name} (${wine.vintage})
1715
- Winery: ${wine.winery}
1716
- Region: ${wine.region}, ${wine.country}
1717
- Grape: ${wine.grape}
1718
- ABV: ${wine.abv}%
1719
- Price: ${wine.price}
1720
- Tasting Notes: ${wine.tastingNotes}
1721
-
1722
- Recent Reviews:
1723
- ${reviews.map(r => `- "${r.text}" (${r.rating}/5)`).join('\n')}
1724
- `.trim(),
1725
- contentLabel: 'Wine Details',
1726
- strategy: 'context',
1727
- };
1728
- }
1729
- ```
1730
-
1731
- #### Equipment Manual App (RAG-only)
1732
-
1733
- ```typescript
1734
- async function gatherAIContent(): Promise<AIContentResponse> {
1735
- const equipment = getCurrentEquipment();
1736
-
1737
- return {
1738
- text: `User is viewing: ${equipment.name} (Model: ${equipment.modelNumber})`,
1739
- contentLabel: `${equipment.name} Assistant`,
1740
- strategy: 'rag',
1741
- ragHint: {
1742
- productId: equipment.smartlinksProductId,
1743
- sessionId: `manual-${equipment.id}-${Date.now()}`,
1744
- context: 'Technical manual, troubleshooting, and maintenance',
1745
- },
1746
- };
1747
- }
1748
- ```
1749
-
1750
- ### Timing & Lifecycle
1751
-
1752
- - **On navigation**: When the user navigates to a new product/proof/app, the assistant automatically requests fresh content and injects a context-update system message.
1753
- - **On first message**: If no content has been gathered yet, the assistant requests it before the first AI call.
1754
- - **On manual refresh**: The assistant can re-request content at any time (e.g. if the user's view within the app has changed).
1755
-
1756
- Content is requested **lazily** — your handler is only called when the AI assistant is active and needs context. If the user never opens the assistant, your handler is never called.
1757
-
1758
- ### Method 3: DOM Fallback (Automatic)
1759
-
1760
- If your app doesn't implement either of the above, the framework extracts `innerText` from the DOM element with `data-app-container="{appId}"`, truncated to ~4000 characters. This is a **last resort** — content quality is much lower than a structured response. Implementing Method 1 or 2 is strongly recommended.
1761
-
1762
- ---
1763
-
1764
- ## Related Documentation
1765
-
1766
- - [API Summary](./API_SUMMARY.md) - Complete API reference
1767
- - [Widgets](./widgets.md) - Embedding SmartLinks components
1768
- - [Realtime](./realtime.md) - Realtime data updates
1769
- - [iframe Responder](./iframe-responder.md) - iframe integration
1770
-
1771
- ---
1772
-
1773
- ## Support
1774
-
1775
- For questions or issues:
1776
-
1777
- - **Documentation:** https://smartlinks.app/docs
1778
- - **GitHub:** https://github.com/Prove-Anything/smartlinks
1779
- - **Email:** support@smartlinks.app
1
+ # SmartLinks AI
2
+
3
+ Build AI-powered SmartLinks experiences with a practical SDK guide for responses, chat, product assistants, streaming, voice, and real-world integration patterns.
4
+
5
+ ---
6
+
7
+ ## Table of Contents
8
+
9
+ - [Overview](#overview)
10
+ - [Quick Start](#quick-start)
11
+ - [Authentication](#authentication)
12
+ - [Responses API](#responses-api)
13
+ - [Chat Completions](#chat-completions)
14
+ - [RAG: Product Assistants](#rag-product-assistants)
15
+ - [Voice Integration](#voice-integration)
16
+ - [Podcast Generation](#podcast-generation)
17
+ - [Types & API reference](#types--api-reference)
18
+ - [Usage Examples](#usage-examples)
19
+ - [Error Handling](#error-handling)
20
+ - [Rate Limiting](#rate-limiting)
21
+ - [Best Practices](#best-practices)
22
+ - [Providing Content to the AI Assistant](#providing-content-to-the-ai-assistant)
23
+
24
+ ---
25
+
26
+ ## Overview
27
+
28
+ This guide is written for SDK users building real products, not backend operators. It focuses on the public SDK surface, recommended starting points, and examples you can adapt directly.
29
+
30
+ ### Start with the path that matches your job
31
+
32
+ | If you want to... | Start here |
33
+ |---|---|
34
+ | Build a new AI workflow | Use [Responses API](#responses-api) |
35
+ | Add compatibility with existing chat clients | Use [Chat Completions](#chat-completions) |
36
+ | Build a product/manual assistant | Use [RAG: Product Assistants](#rag-product-assistants) |
37
+ | Add spoken input/output | Use [Voice Integration](#voice-integration) |
38
+ | Add progressive rendering | Use [Streaming Responses](#streaming-responses) or [Streaming Chat](#streaming-chat) |
39
+
40
+ > **Which AI chat API? Three surfaces, three jobs — don't confuse them:**
41
+ > - **`ai.chat.responses`** — the **default for new work**. Structured input, tool use, multi-agent, streaming; you manage any history. Reach for this first.
42
+ > - **`ai.chat.completions`** — **OpenAI-compatible** (`messages[]`). Use it *only* to drop into existing OpenAI-style code; prefer Responses for anything new.
43
+ > - **`ai.public.chat`** (+ `getSession` / `clearSession`) — the **public product assistant** with **server-managed conversation sessions** (RAG-grounded). This is the "ongoing conversation" surface, and it's public-facing.
44
+ >
45
+ > Responses and Completions are *generation* APIs (not single-shot vs conversation — both take history); `ai.public.chat` is the one that keeps a conversation for you.
46
+
47
+ ### Recommended starting points
48
+
49
+ - New AI features: start with `ai.chat.responses.create(...)`
50
+ - Product/manual assistants: start with `ai.public.chat(...)`
51
+ - Existing OpenAI-style clients: use `ai.chat.completions.create(...)`
52
+ - Real-time voice: use `ai.public.getToken(...)` and your provider's live client
53
+
54
+ SmartLinks AI provides five main capabilities:
55
+
56
+ 1. **Responses API** - Preferred API for agentic workflows, multimodal inputs, and tool-driven responses
57
+ 2. **Chat Completions** - OpenAI-compatible text generation with streaming and tool calling
58
+ 3. **RAG (Retrieval-Augmented Generation)** - Document-grounded Q&A for product assistants
59
+ 4. **Voice Integration** - Voice-to-text and text-to-voice for hands-free interaction
60
+ 5. **Podcast Generation** - NotebookLM-style multi-voice conversational podcasts from documents
61
+
62
+ ### Key Features
63
+
64
+ - ✅ Full TypeScript support with type safety
65
+ - ✅ Streaming responses with async iterators
66
+ - ✅ Automatic rate limit handling
67
+ - ✅ Session management for conversations
68
+ - ✅ Voice input/output helpers
69
+ - ✅ Tool/function calling support
70
+ - ✅ OpenAI-style Responses API support
71
+ - ✅ Document indexing and retrieval
72
+ - ✅ Customizable assistant behavior
73
+
74
+ ---
75
+
76
+ ## Quick Start
77
+
78
+ If you're only reading one section, start here. The three snippets below cover the most common public SDK use cases.
79
+
80
+ ### 1. Generate a response
81
+
82
+ ```typescript
83
+ import { initializeApi, ai } from '@proveanything/smartlinks';
84
+
85
+ // Initialize the SDK
86
+ initializeApi({
87
+ baseURL: 'https://smartlinks.app/api/v1',
88
+ apiKey: process.env.SMARTLINKS_API_KEY // Required for admin endpoints
89
+ });
90
+
91
+ // Preferred: create a response
92
+ const response = await ai.chat.responses.create('my-collection', {
93
+ model: 'google/gemini-2.5-flash',
94
+ input: 'Summarize the key safety steps for descaling a coffee maker.'
95
+ });
96
+
97
+ console.log(response.output_text);
98
+ ```
99
+
100
+ ### 2. Build a product assistant
101
+
102
+ ```typescript
103
+ import { initializeApi, ai } from '@proveanything/smartlinks';
104
+
105
+ initializeApi({ baseURL: 'https://smartlinks.app/api/v1' });
106
+
107
+ const answer = await ai.public.chat('my-collection', {
108
+ productId: 'coffee-maker-deluxe',
109
+ userId: 'user-123',
110
+ message: 'How do I descale this machine?'
111
+ });
112
+
113
+ console.log(answer.message);
114
+ ```
115
+
116
+ ### 3. Stream output into your UI
117
+
118
+ ```typescript
119
+ const stream = await ai.chat.responses.create('my-collection', {
120
+ input: 'Write a launch checklist for a new product page.',
121
+ stream: true
122
+ });
123
+
124
+ for await (const event of stream) {
125
+ if (event.type === 'response.output_text.delta') {
126
+ updateUi(event.delta);
127
+ }
128
+ }
129
+ ```
130
+
131
+ ---
132
+
133
+ ## Authentication
134
+
135
+ ### Admin Endpoints
136
+
137
+ Admin endpoints require an API key passed during initialization:
138
+
139
+ ```typescript
140
+ initializeApi({
141
+ baseURL: 'https://smartlinks.app/api/v1',
142
+ apiKey: process.env.SMARTLINKS_API_KEY
143
+ });
144
+ ```
145
+
146
+ The SDK automatically includes the API key in the `Authorization: Bearer <token>` header.
147
+
148
+ ### Public Endpoints
149
+
150
+ Public endpoints don't require an API key but are rate-limited by `userId`:
151
+
152
+ ```typescript
153
+ // No API key needed
154
+ const response = await ai.public.chat('my-collection', {
155
+ productId: 'coffee-maker',
156
+ userId: 'user-123',
157
+ message: 'How do I clean this?'
158
+ });
159
+ ```
160
+
161
+ ---
162
+
163
+ ## Responses API
164
+
165
+ The Responses API is the recommended starting point for new integrations. Use it when you want a single endpoint for structured input, tool use, and streaming output.
166
+
167
+ ### Basic Response
168
+
169
+ ```typescript
170
+ const response = await ai.chat.responses.create('my-collection', {
171
+ model: 'google/gemini-2.5-flash',
172
+ input: 'Write a friendly two-sentence welcome for a product assistant.'
173
+ });
174
+
175
+ console.log(response.output_text);
176
+ ```
177
+
178
+ ### Multimessage Input
179
+
180
+ ```typescript
181
+ const response = await ai.chat.responses.create('my-collection', {
182
+ model: 'google/gemini-2.5-flash',
183
+ input: [
184
+ {
185
+ role: 'system',
186
+ content: [
187
+ { type: 'input_text', text: 'You are a concise support assistant.' }
188
+ ]
189
+ },
190
+ {
191
+ role: 'user',
192
+ content: [
193
+ { type: 'input_text', text: 'Give me three troubleshooting steps for a grinder that will not start.' }
194
+ ]
195
+ }
196
+ ]
197
+ });
198
+
199
+ console.log(response.output_text);
200
+ ```
201
+
202
+ ### Streaming Responses
203
+
204
+ When you pass `stream: true`, the SDK returns an `AsyncIterable` of SSE events instead of a final JSON object. You do not need to parse raw SSE frames yourself — just iterate with `for await...of`.
205
+
206
+ ```typescript
207
+ const result = await ai.chat.responses.create('my-collection', {
208
+ input: 'Summarize the manual',
209
+ stream: true
210
+ });
211
+
212
+ for await (const event of result) {
213
+ if (event.type === 'response.output_text.delta') {
214
+ process.stdout.write(event.delta);
215
+ }
216
+ }
217
+ ```
218
+
219
+ If you omit `stream: true`, the same method returns the final `ResponsesResult` object instead.
220
+
221
+ ```typescript
222
+ const stream = await ai.chat.responses.create('my-collection', {
223
+ model: 'google/gemini-2.5-flash',
224
+ input: 'Explain how to descale an espresso machine step by step.',
225
+ stream: true
226
+ });
227
+
228
+ for await (const event of stream) {
229
+ if (event.type === 'response.output_text.delta') {
230
+ process.stdout.write(event.delta);
231
+ }
232
+ }
233
+ ```
234
+
235
+ ### Tool Calling
236
+
237
+ ```typescript
238
+ const response = await ai.chat.responses.create('my-collection', {
239
+ model: 'google/gemini-2.5-flash',
240
+ input: 'What is the weather in Paris?',
241
+ tools: [
242
+ {
243
+ type: 'function',
244
+ name: 'get_weather',
245
+ description: 'Get the current weather for a city',
246
+ parameters: {
247
+ type: 'object',
248
+ properties: {
249
+ location: { type: 'string' }
250
+ },
251
+ required: ['location']
252
+ }
253
+ }
254
+ ]
255
+ });
256
+
257
+ console.log(response.output);
258
+ ```
259
+
260
+ ### Server-side tools (built-in agent loop)
261
+
262
+ The example above is **client-relayed** tool calling: you define the tools, the model returns
263
+ `tool_call` requests, and *your app* executes them and sends results back. For the common tools —
264
+ reading and searching the web, vision, reading documents, generating images — the platform ships a
265
+ curated, tested **built-in toolset it runs itself**. Opt in with `server_tools` and the server
266
+ executes each tool and feeds the result back automatically, looping until the model has its answer.
267
+ You get one final response; no relay code.
268
+
269
+ ```typescript
270
+ // Enable the whole built-in toolset:
271
+ const res = await ai.chat.responses.create('my-collection', {
272
+ model: 'balanced',
273
+ input: 'Research acme.com and summarise what they sell, with their brand colours.',
274
+ server_tools: true
275
+ });
276
+ console.log(res.output_text);
277
+ console.log(res._agent.toolResults); // trace: which tools ran, with what result
278
+ ```
279
+
280
+ Scope it to specific tools (recommended — smaller blast radius, faster), by name or capability:
281
+
282
+ ```typescript
283
+ import { AI_TOOL_NAMES } from '@proveanything/smartlinks';
284
+
285
+ const res = await ai.chat.responses.create('my-collection', {
286
+ input: 'Find the current price of this product and return it as JSON.',
287
+ server_tools: [AI_TOOL_NAMES.WEB_SEARCH, AI_TOOL_NAMES.DATA_EXTRACT],
288
+ // or: allowCapabilities: ['web:read'], exclude: ['image.generate'],
289
+ maxSteps: 6 // cap model round-trips (1–12, default 8)
290
+ });
291
+ ```
292
+
293
+ **Streaming** surfaces tool progress as it happens — ideal for a "thinking…" UI. You get
294
+ `agent.tool_call` / `agent.tool_result` events, then a final `response.completed`:
295
+
296
+ ```typescript
297
+ const stream = await ai.chat.responses.create('my-collection', {
298
+ input: 'Research acme.com', server_tools: true, stream: true
299
+ });
300
+ for await (const ev of stream) {
301
+ if (ev.type === 'agent.tool_call') showStep(`Running ${ev.name}…`);
302
+ if (ev.type === 'agent.tool_result') showStep(`${ev.name} done`);
303
+ if (ev.type === 'response.completed') render(ev.response.output_text);
304
+ }
305
+ ```
306
+
307
+ `server_tools` can't be combined with `previous_response_id`/`conversation` yet — pass prior turns
308
+ in `input`.
309
+
310
+ #### Built-in tools
311
+
312
+ | Tool | Does |
313
+ |------|------|
314
+ | `web.search` | Live web search → candidate results (url/title/description). |
315
+ | `web.fetchPage` | Fetch a page → clean markdown + metadata + schema.org JSON-LD. |
316
+ | `web.extractSchema` | Return only a page's schema.org data of a given `@type` (deterministic). |
317
+ | `document.read` | Read a document at a URL — **PDF, deck, doc**, or article — into markdown. |
318
+ | `data.extract` | Page + JSON-schema/prompt → **typed JSON** (turn a page into UI data). |
319
+ | `brand.assets` | Extract a site's logo, colours, and design. |
320
+ | `web.screenshot` | Screenshot a page → hosted image URL (feed to `image.describe`). |
321
+ | `image.describe` | Vision: describe an image / read its text. |
322
+ | `image.generate` | Generate an image from a prompt → hosted URL. |
323
+ | `image.fromReference` | Image-to-image: generate guided by reference image(s). |
324
+ | `image.searchStock` | Search real stock photos (Unsplash). |
325
+ | `image.transform` | Resize / crop / rotate / grayscale / format-convert / compress → hosted URL. |
326
+ | `pdf.create` | Render HTML → PDF → hosted URL. |
327
+ | `pdf.fill` | Fill an AcroForm PDF's fields (`{ field: value }`) → hosted URL. |
328
+ | `pdf.merge` | Merge several PDFs into one, in order → hosted URL. |
329
+ | `pdf.inspect` | Cheap, no-AI introspection: page count/sizes, which pages have a real text layer, raster present, and a routing hint (`text` vs `vision`). |
330
+ | `pdf.render` | Rasterize one page — or just a region (`clip: { x0, y0, x1, y1 }`, 0–1 from top-left) — to a PNG at a chosen DPI → hosted image URL (permanent). Max 40 MP per call; larger requests return `code: "too_large"` + `suggestedDpi`. |
331
+ | `image.ocr` | **Deterministic OCR** (Google Vision, no generative model): exact characters with per-word `confidence` (0–1) and pixel `bbox`, plus lines with spacing as printed. Input: `imageUrl`, or a PDF `url` + `page` + `dpi` + optional `clip`. Optional `languages` hints. Use for small print and text outlined to curves. |
332
+ | `pdf.extract` | PDF → typed JSON in one call (schema and/or prompt). Auto-routes text vs vision. Optional per-field `confidence`/`source` (`includeConfidence`) and `bbox` (`includeBoxes`) in `fieldsMeta`. |
333
+ | `pdf.decodeBarcodes` | Deterministically decode barcodes/QR on a page (WASM, no AI) → value + symbology + page + bbox + confidence. Use this for barcode digits, never vision. |
334
+ | `pdf.inspectGraphics` | Prepress inspection: per-page path/image/outlined-text counts + colour spaces, plus named SPOT colours (e.g. "PANTONE 871 C"). No AI. |
335
+ | `pdf.edit` | Edit an existing PDF: replace **live** text (glyph-exact removal, neighbours never move; original font when it has the glyphs, else an embedded fallback; CMYK/spot colour kept; matches may span lines), add text / images in boxes (`colorSpace: "cmyk"` converts images for print). Per-item `results` with `verified`. |
336
+ | `pdf.preflight` | Prepress preflight, no changes: fonts, RGB/Lab, TrimBox/BleedBox + bleed, effective image dpi, transparency, white overprint, output intent. No AI. |
337
+ | `pdf.printReady` | RGB→CMYK with a press profile (spot colours kept), embed fonts, optional flattening, mark PDF/X-1a or PDF/X-4 with the output intent embedded, then preflight the result → `{ url, report }`. |
338
+ | `http.request` | SSRF-guarded outbound HTTP(S) to a public URL (call a REST API). |
339
+ | `translate` | Translate text into one or more languages (generic, model-based). |
340
+
341
+ **Working with PDFs.** The PDF tools split into *read/analyse* and *produce*:
342
+
343
+ - **Read a whole PDF as text** → `document.read` (Firecrawl; good for prose/decks).
344
+ - **Turn a PDF into structured fields** → `pdf.extract` (give it a JSON schema and/or a prompt; it
345
+ returns typed JSON). It routes itself, but you can drive the route yourself: call `pdf.inspect`
346
+ first (deterministic, no AI) to see whether each page has a real text layer, then `pdf.extract`
347
+ (cheap text path) or `pdf.render` → `image.describe`/vision for curve-only or raster artwork.
348
+ - **Read small print exactly** (INCI/allergen lists, net weight, text outlined to curves) → `image.ocr`
349
+ with a PDF `url`, `page`, `dpi: 300–400` and a `clip` around the panel. It returns the characters it
350
+ sees with a per-word confidence — it never "corrects" a misspelling the way a vision model can, so
351
+ flag words below ~0.9 for human review rather than re-reading them with a model.
352
+ - **Zoom in visually** (layout, artwork) → `pdf.render` with `clip` to rasterise just a region at a high
353
+ DPI, instead of the whole sheet. Renders over 40 MP return `code: "too_large"` with `suggestedDpi`
354
+ — retry at that dpi or with a smaller clip, don't blind-retry.
355
+ - **Barcodes / QR** → `pdf.decodeBarcodes` (deterministic WASM decode) — never trust vision for barcode
356
+ digits; it hallucinates them.
357
+ - **Prepress / print QA** (spot colours, colour spaces, vector vs raster) → `pdf.inspectGraphics`.
358
+ - **Change the artwork itself** → `pdf.edit`. `replaceText` edits only LIVE text: the matched glyphs are
359
+ cut out with an exact-width gap (nothing else moves) and the new text is drawn at the same spot, in the
360
+ same colour, in the original font if its embedded glyphs cover the new text (otherwise an embedded
361
+ fallback, reported as `usedFallbackFont`). A replacement longer than the original is flagged in `note`
362
+ — it may run into neighbouring text. Text outlined to curves can't be edited: the item fails with
363
+ "not found as live text" — put it on the designer change list. `verified: true` means the new text
364
+ reads back from the output.
365
+ - **Is it print-ready?** → `pdf.preflight` (report only). **Make it print-ready** → `pdf.printReady`
366
+ with `profile` (`pdfx-4` default; `pdfx-1a` for CMYK-only workflows), `outputIntent` (e.g. `FOGRA39`)
367
+ and `bleed: { mm: 3 }`. `report.compliant` comes from preflighting the OUTPUT, so remaining issues
368
+ (low-res images, missing bleed, white overprint) are reported rather than hidden. Flattening for
369
+ PDF/X-1a turns affected pages into images — opt in with `flattenTransparency: true`.
370
+
371
+ ```ts
372
+ const edited = await SL.ai.tools.run(collectionId, 'pdf.edit', {
373
+ url: artworkUrl,
374
+ replaceText: [{ id: 'f1', page: 1, find: 'Best before: see lid', replace: 'Best before: see base' }],
375
+ })
376
+ const ready = await SL.ai.tools.run(collectionId, 'pdf.printReady', {
377
+ url: edited.result.url, profile: 'pdfx-4', outputIntent: 'FOGRA39', bleed: { mm: 3 },
378
+ })
379
+ if (!ready.result.report.compliant) showFindings(ready.result.report.findings)
380
+ ```
381
+ - **A review UI that flags guessed fields** → `pdf.extract` with `includeConfidence` (per-field
382
+ confidence + source) and `includeBoxes` (per-field bbox on text-native pages) in `fieldsMeta`.
383
+ - **Produce a PDF** → `pdf.create` (HTML → PDF), `pdf.fill` (populate an AcroForm's fields),
384
+ `pdf.merge` (combine several). These return a hosted `hostedUrl` (`pdf.create` also returns it as
385
+ `url`). `pdf.create` loads remote `<img>` URLs before rendering and honours `page-break-inside`; pass
386
+ `margin: "0"` when your HTML sets its own body margin (default page margin is 18mm/14mm).
387
+
388
+ Every tool is also directly callable without the model loop via `ai.tools.run(collectionId, name,
389
+ args)` — e.g. render a page or extract fields straight from a UI, no agent round-trip:
390
+
391
+ ```ts
392
+ // Read the ingredients panel of a sleeve exactly (deterministic OCR of a clipped region)
393
+ const { result } = await SL.ai.tools.run(collectionId, 'image.ocr', {
394
+ url: sleevePdfUrl, page: 1, dpi: 400,
395
+ clip: { x0: 0.40, y0: 0.20, x1: 0.60, y1: 0.40 },
396
+ languages: ['en', 'fr'],
397
+ })
398
+ // result.lines → [{ text: 'INGREDIENTS: Aqua, Glycerin, …', confidence: 0.97, bbox: {…} }]
399
+ const toReview = result.words.filter((w) => w.confidence < 0.9)
400
+ ```
401
+
402
+ Discover tools two ways:
403
+ - **Design time (typed):** import `BUILTIN_AI_TOOLS`, `AI_TOOL_NAMES`, and the per-tool arg types
404
+ (`WebSearchArgs`, `DataExtractArgs`, `PdfExtractArgs`, …) from the SDK. This is the core set —
405
+ stable, versioned, documented here.
406
+ - **Runtime (live):** `await ai.catalog(collectionId)` returns the registry as the server sees it,
407
+ including any future app-contributed tools. The built-in set above is always present.
408
+
409
+ ### Recommended Models
410
+
411
+ For agentic workflows on `v1/responses`, GPT-5.6 ships in three tiers. Pass either the full model
412
+ id (`openai/gpt-5.6-sol`) or the shorthand tier alias (`'cheap'`, `'balanced'`, `'premium'`) as
413
+ `model` — both resolve through the same server-side model registry.
414
+
415
+ | Tier | Model | Alias | Use for |
416
+ |------|-------|-------|---------|
417
+ | Cheapest (default) | `openai/gpt-5.6-luna` | `'cheap'` / omit `model` | Everyday chat, high-volume/simple turns |
418
+ | Balanced | `openai/gpt-5.6-terra` | `'balanced'` | Admin agent / tool-heavy setup workflows (route default) |
419
+ | Flagship | `openai/gpt-5.6-sol` | `'premium'` | Most demanding reasoning, coding, and multi-step tool use |
420
+
421
+ Every tier supports function/tool calling, Programmatic Tool Calling, and Multi-agent (see below), and all
422
+ three accept `service_tier: 'flex'` (≈50% cheaper, Batch-API rates, slower) or `'priority'` (premium, guaranteed
423
+ throughput) alongside the default `'standard'` tier. Pass it straight through in the request body:
424
+
425
+ ```typescript
426
+ const response = await ai.chat.responses.create('my-collection', {
427
+ model: 'balanced',
428
+ input: 'Summarize this week\'s ingested documents.',
429
+ service_tier: 'flex' // non-interactive/background work: cheaper, but can take minutes
430
+ });
431
+ ```
432
+
433
+ Flex requests run noticeably slower — reserve `flex` for non-production/background work (batch enrichment,
434
+ evaluations, async jobs), not user-facing chat turns. The server already raises its own timeout to 15
435
+ minutes for any request with `service_tier: 'flex'`, matching OpenAI's guidance — nothing to configure on
436
+ the client. On a `429` ("resource unavailable") the request wasn't charged; retry with backoff or drop
437
+ `service_tier` (or set it to `'auto'`) to fall back to standard processing.
438
+
439
+ ### Programmatic Tool Calling
440
+
441
+ Lets the model write a short in-memory program that orchestrates several tool calls (filtering,
442
+ looping, aggregating) before handing back one result, instead of a full message round-trip per
443
+ call. Add the hosted `programmatic_tool_calling` tool, and mark each function tool the program is
444
+ allowed to invoke with `allowed_callers: ['programmatic']`:
445
+
446
+ ```typescript
447
+ const response = await ai.chat.responses.create('my-collection', {
448
+ model: 'balanced',
449
+ input: 'Compare inventory with demand for every SKU in this collection.',
450
+ tools: [
451
+ {
452
+ type: 'function',
453
+ name: 'get_inventory',
454
+ description: 'Return available_units for a sku',
455
+ parameters: { type: 'object', properties: { sku: { type: 'string' } }, required: ['sku'] },
456
+ output_schema: { type: 'object', properties: { sku: { type: 'string' }, available_units: { type: 'number' } } },
457
+ allowed_callers: ['programmatic']
458
+ },
459
+ {
460
+ type: 'function',
461
+ name: 'get_demand',
462
+ description: 'Return requested_units for a sku',
463
+ parameters: { type: 'object', properties: { sku: { type: 'string' } }, required: ['sku'] },
464
+ output_schema: { type: 'object', properties: { sku: { type: 'string' }, requested_units: { type: 'number' } } },
465
+ allowed_callers: ['programmatic']
466
+ },
467
+ { type: 'programmatic_tool_calling' }
468
+ ]
469
+ });
470
+ ```
471
+
472
+ The response's `output` array can include a `program` item (the generated code), one
473
+ `function_call` item per program-issued call (tagged with `caller: { type: 'program', caller_id }`),
474
+ and a `program_output` item with the program's final `result`/`status`. Execute the `function_call`
475
+ items as normal and send results back keyed by their `call_id` on the next turn.
476
+
477
+ ### Multi-agent (subagents)
478
+
479
+ Lets the root agent spawn a tree of subagents that run in parallel and get synthesized back into
480
+ one response — useful for tasks that split cleanly into independent workstreams (e.g. research +
481
+ draft + review). Enable it by setting `multi_agent.enabled: true` — that's the entire contract on
482
+ your end; nothing else to configure or pass.
483
+
484
+ This is upstream beta functionality (OpenAI may still be limiting it to certain accounts/rollout),
485
+ so treat behavior and availability as subject to change until it's GA.
486
+
487
+ ```typescript
488
+ const response = await ai.chat.responses.create('my-collection', {
489
+ model: 'premium', // Sol recommended for the root agent when spawning subagents
490
+ input: 'Research three competitor loyalty programs, then draft a comparison summary.',
491
+ multi_agent: { enabled: true, max_concurrent_subagents: 3 }
492
+ });
493
+ ```
494
+
495
+ Subagent output arrives as additional items in `output`, each tagged `agent: { agent_name: '/root/<name>' }`
496
+ (root agent output uses `/root`). Notes/limits carried over from OpenAI's beta:
497
+
498
+ - `max_concurrent_subagents` defaults to `3`
499
+ - `reasoning.summary` and `max_tool_calls` are not supported while multi-agent is enabled
500
+ - Streaming is not supported while multi-agent is enabled — the SDK throws if you pass both
501
+ `stream: true` and `multi_agent.enabled: true`
502
+
503
+ ---
504
+
505
+ ## Chat Completions
506
+
507
+ OpenAI-compatible chat completions with streaming and tool calling support. Use this for compatibility with existing Chat Completions integrations; prefer the Responses API for new agentic features.
508
+
509
+ ### Basic Chat
510
+
511
+ ```typescript
512
+ const response = await ai.chat.completions.create('my-collection', {
513
+ model: 'google/gemini-2.5-flash',
514
+ messages: [
515
+ { role: 'system', content: 'You are a helpful assistant.' },
516
+ { role: 'user', content: 'What is the capital of France?' }
517
+ ]
518
+ });
519
+
520
+ console.log(response.choices[0].message.content);
521
+ // Output: "The capital of France is Paris."
522
+ ```
523
+
524
+ ### Streaming Chat
525
+
526
+ Stream responses in real-time for better UX:
527
+
528
+ - Set `stream: true`
529
+ - The SDK returns an `AsyncIterable<ChatCompletionChunk>`
530
+ - Iterate over chunks with `for await...of`
531
+ - Read incremental text from `chunk.choices[0]?.delta?.content`
532
+ - If `stream` is omitted or `false`, the method returns the normal `ChatCompletionResponse`
533
+
534
+ ```typescript
535
+ const stream = await ai.chat.completions.create('my-collection', {
536
+ model: 'google/gemini-2.5-flash',
537
+ messages: [
538
+ { role: 'user', content: 'Write a short poem about coding' }
539
+ ],
540
+ stream: true
541
+ });
542
+
543
+ for await (const chunk of stream) {
544
+ const content = chunk.choices[0]?.delta?.content || '';
545
+ process.stdout.write(content);
546
+ }
547
+ ```
548
+
549
+ ### Tool/Function Calling
550
+
551
+ Define tools (functions) that the AI can call:
552
+
553
+ ```typescript
554
+ const tools = [
555
+ {
556
+ type: 'function',
557
+ function: {
558
+ name: 'get_weather',
559
+ description: 'Get the current weather for a location',
560
+ parameters: {
561
+ type: 'object',
562
+ properties: {
563
+ location: {
564
+ type: 'string',
565
+ description: 'City name'
566
+ },
567
+ unit: {
568
+ type: 'string',
569
+ enum: ['celsius', 'fahrenheit']
570
+ }
571
+ },
572
+ required: ['location']
573
+ }
574
+ }
575
+ }
576
+ ];
577
+
578
+ const response = await ai.chat.completions.create('my-collection', {
579
+ model: 'google/gemini-2.5-flash',
580
+ messages: [
581
+ { role: 'user', content: 'What\'s the weather in Paris?' }
582
+ ],
583
+ tools
584
+ });
585
+
586
+ const toolCall = response.choices[0].message.tool_calls?.[0];
587
+ if (toolCall) {
588
+ console.log('Function:', toolCall.function.name);
589
+ console.log('Arguments:', JSON.parse(toolCall.function.arguments));
590
+ // { location: "Paris", unit: "celsius" }
591
+ }
592
+ ```
593
+
594
+ ### Available Models
595
+
596
+ ```typescript
597
+ // List all available models
598
+ const models = await ai.models.list('my-collection');
599
+
600
+ // Or filter by provider / capability
601
+ const openAiModels = await ai.models.list('my-collection', {
602
+ provider: 'openai'
603
+ });
604
+
605
+ const visionModels = await ai.models.list('my-collection', {
606
+ capability: 'vision'
607
+ });
608
+
609
+ models.data.forEach(model => {
610
+ console.log(`${model.name}`);
611
+ console.log(` Provider: ${model.provider}`);
612
+ console.log(` Context: ${model.contextWindow} tokens`);
613
+ console.log(` Pricing: $${model.pricing.input}/1M input tokens`);
614
+ });
615
+
616
+ // Get specific model info
617
+ const model = await ai.models.get('my-collection', 'google/gemini-2.5-flash');
618
+ console.log(model.capabilities); // ['text', 'vision', 'audio', 'code']
619
+ ```
620
+
621
+ Use `ai.models.list(collectionId)` as the source of truth for what your collection can use at runtime. The public docs provide recommendations, but actual availability depends on the SmartLinks model catalog exposed to that collection.
622
+
623
+ **Recommended Models:**
624
+
625
+ | Model | Use Case | Speed | Cost |
626
+ |-------|----------|-------|------|
627
+ | `openai/gpt-5.4` | Default for new agentic and structured-output workflows | Balanced | Medium |
628
+ | `openai/gpt-5-mini` | Lower-cost general purpose and JSON tasks | Fast | Low |
629
+ | `google/gemini-2.5-flash` | Fast multimodal and cost-sensitive general use | Fast | Low |
630
+ | `google/gemini-2.5-pro` | Complex reasoning and heavier multimodal tasks | Slower | Higher |
631
+
632
+ If you want a safe default for most new work, start with `openai/gpt-5.4`. If you want a lower-cost fallback, use `openai/gpt-5-mini` or `google/gemini-2.5-flash` depending on your latency and pricing goals.
633
+
634
+ ---
635
+
636
+ ## RAG: Product Assistants
637
+
638
+ Create intelligent product assistants that answer questions based on product documentation.
639
+
640
+ ### Setup: Index Documents
641
+
642
+ First, index your product documentation:
643
+
644
+ ```typescript
645
+ // Index a product manual from URL
646
+ const result = await ai.rag.indexDocument('my-collection', {
647
+ productId: 'coffee-maker-deluxe',
648
+ documentUrl: 'https://example.com/manuals/coffee-maker.pdf',
649
+ chunkSize: 500, // Tokens per chunk
650
+ overlap: 50, // Token overlap between chunks
651
+ provider: 'openai' // Embedding provider
652
+ });
653
+
654
+ console.log(`Indexed ${result.chunks} chunks`);
655
+ console.log(`Dimensions: ${result.metadata.embeddingDimensions}`);
656
+
657
+ // Or index from text directly
658
+ await ai.rag.indexDocument('my-collection', {
659
+ productId: 'coffee-maker-deluxe',
660
+ text: 'Your product manual content here...',
661
+ metadata: {
662
+ source: 'manual',
663
+ version: '2.0'
664
+ }
665
+ });
666
+ ```
667
+
668
+ ### Configure Assistant
669
+
670
+ Customize the assistant's behavior:
671
+
672
+ ```typescript
673
+ await ai.rag.configureAssistant('my-collection', {
674
+ productId: 'coffee-maker-deluxe',
675
+ systemPrompt: 'You are a helpful coffee maker assistant. Be concise and friendly.',
676
+ model: 'google/gemini-2.5-flash',
677
+ temperature: 0.7,
678
+ maxTokensPerResponse: 500,
679
+ rateLimitPerUser: 20,
680
+ allowedTopics: ['usage', 'cleaning', 'troubleshooting'],
681
+ customInstructions: {
682
+ tone: 'friendly',
683
+ additionalRules: 'Always include safety warnings when relevant.'
684
+ }
685
+ });
686
+ ```
687
+
688
+ ### Public Chat
689
+
690
+ Users can chat with the product assistant without authentication:
691
+
692
+ ```typescript
693
+ // First question
694
+ const response = await ai.public.chat('my-collection', {
695
+ productId: 'coffee-maker-deluxe',
696
+ userId: 'user-123',
697
+ message: 'How do I descale my coffee maker?'
698
+ });
699
+
700
+ console.log('Answer:', response.message);
701
+ console.log('Used', response.context.chunksUsed, 'document sections');
702
+ console.log('Top similarity:', response.context.topSimilarity);
703
+ ```
704
+
705
+ ### Conversation History
706
+
707
+ Maintain conversation context with sessions:
708
+
709
+ ```typescript
710
+ const sessionId = `session-${Date.now()}`;
711
+
712
+ // First question
713
+ const q1 = await ai.public.chat('my-collection', {
714
+ productId: 'coffee-maker-deluxe',
715
+ userId: 'user-123',
716
+ message: 'How do I clean it?',
717
+ sessionId
718
+ });
719
+
720
+ // Follow-up question (uses history)
721
+ const q2 = await ai.public.chat('my-collection', {
722
+ productId: 'coffee-maker-deluxe',
723
+ userId: 'user-123',
724
+ message: 'How often should I do that?',
725
+ sessionId
726
+ });
727
+
728
+ // Get full conversation history
729
+ const session = await ai.public.getSession('my-collection', sessionId);
730
+ console.log('Messages:', session.messages);
731
+ console.log('Total messages:', session.messageCount);
732
+
733
+ // Clear session when done
734
+ await ai.public.clearSession('my-collection', sessionId);
735
+ ```
736
+
737
+ ### Session Management
738
+
739
+ ```typescript
740
+ // Get session statistics (admin)
741
+ const stats = await ai.sessions.stats('my-collection');
742
+ console.log('Total sessions:', stats.totalSessions);
743
+ console.log('Active sessions:', stats.activeSessions);
744
+ console.log('Total messages:', stats.totalMessages);
745
+ console.log('Rate-limited users:', stats.rateLimitedUsers);
746
+ ```
747
+
748
+ ---
749
+
750
+ ## Voice Integration
751
+
752
+ Enable voice input and output for hands-free interaction.
753
+
754
+ ### Voice Patterns
755
+
756
+ The SDK supports three practical voice patterns:
757
+
758
+ | Pattern | Best For | SDK Building Blocks |
759
+ |---------|----------|---------------------|
760
+ | Voice → Text → AI → Text | Manual helper Q&A, troubleshooting steps | `ai.voice.listen()` + `ai.public.chat()` |
761
+ | Voice → Text → AI → Voice | Hands-free assistants, accessibility | `ai.voice.listen()` + `ai.public.chat()` + `ai.voice.speak()` or `ai.tts.generate()` |
762
+ | Real-time Voice | Low-latency spoken conversation | `ai.public.getToken()` + Gemini Live client |
763
+
764
+ ### Current SDK Support
765
+
766
+ - `ai.voice.listen()` and `ai.voice.speak()` are browser helpers built on the Web Speech APIs.
767
+ - `ai.public.getToken()` generates ephemeral tokens for Gemini Live sessions.
768
+ - `ai.tts.generate()` supports server-side text-to-speech generation.
769
+ - The SDK does not currently expose a first-class transcription endpoint like Whisper; if you need that flow, implement it as your own backend endpoint and feed the transcribed text into `ai.public.chat()` or `ai.chat.responses.create()`.
770
+
771
+ ### Recommended Approach
772
+
773
+ For product assistants and RAG-backed support, start with Voice → Text → AI → Text/Voice. It gives you the best control over retrieval, session history, and cost. Use Gemini Live when low-latency spoken conversation matters more than deep document grounding.
774
+
775
+ ### Browser Voice Helpers
776
+
777
+ These helpers are browser-only and rely on native speech recognition / speech synthesis support.
778
+
779
+ ```typescript
780
+ // Check if voice is supported
781
+ if (ai.voice.isSupported()) {
782
+ // Listen for voice input
783
+ const question = await ai.voice.listen('en-US');
784
+ console.log('User said:', question);
785
+
786
+ // Get answer from AI
787
+ const response = await ai.public.chat('my-collection', {
788
+ productId: 'coffee-maker-deluxe',
789
+ userId: 'user-123',
790
+ message: question
791
+ });
792
+
793
+ // Speak the answer
794
+ await ai.voice.speak(response.message, {
795
+ voice: 'alloy',
796
+ rate: 1.0
797
+ });
798
+ }
799
+ ```
800
+
801
+ ### Voice Assistant Class
802
+
803
+ Create a complete voice assistant:
804
+
805
+ ```typescript
806
+ class ProductVoiceAssistant {
807
+ private collectionId: string;
808
+ private productId: string;
809
+ private userId: string;
810
+ private sessionId: string;
811
+
812
+ constructor(config: {
813
+ collectionId: string;
814
+ productId: string;
815
+ userId: string;
816
+ }) {
817
+ this.collectionId = config.collectionId;
818
+ this.productId = config.productId;
819
+ this.userId = config.userId;
820
+ this.sessionId = `voice-${Date.now()}`;
821
+ }
822
+
823
+ async ask(): Promise<string> {
824
+ // Listen for question
825
+ console.log('Listening...');
826
+ const question = await ai.voice.listen();
827
+
828
+ // Get answer
829
+ console.log('Processing...');
830
+ const response = await ai.public.chat(this.collectionId, {
831
+ productId: this.productId,
832
+ userId: this.userId,
833
+ message: question,
834
+ sessionId: this.sessionId
835
+ });
836
+
837
+ // Speak answer
838
+ console.log('Speaking...');
839
+ await ai.voice.speak(response.message);
840
+
841
+ return response.message;
842
+ }
843
+
844
+ async getRemainingQuestions(): Promise<number> {
845
+ const status = await ai.public.getRateLimit(this.collectionId, this.userId);
846
+ return status.remaining;
847
+ }
848
+ }
849
+
850
+ // Usage
851
+ const assistant = new ProductVoiceAssistant({
852
+ collectionId: 'my-collection',
853
+ productId: 'coffee-maker-deluxe',
854
+ userId: 'user-123'
855
+ });
856
+
857
+ await assistant.ask(); // Voice question → Voice answer
858
+ const remaining = await assistant.getRemainingQuestions();
859
+ console.log(`${remaining} questions remaining`);
860
+ ```
861
+
862
+ ### Gemini Live Integration
863
+
864
+ Generate ephemeral tokens for Gemini Live (multimodal voice):
865
+
866
+ Use this path for real-time voice sessions. The SDK only issues the short-lived token; the actual live connection is made with the provider client.
867
+
868
+ ```typescript
869
+ // Generate token for voice session
870
+ const token = await ai.public.getToken('my-collection', {
871
+ settings: {
872
+ ttl: 3600, // 1 hour
873
+ voice: 'alloy',
874
+ language: 'en-US'
875
+ }
876
+ });
877
+
878
+ console.log('Token:', token.token);
879
+ console.log('Expires at:', new Date(token.expiresAt));
880
+
881
+ // Use token with Gemini Live API
882
+ // (See Google's Gemini documentation)
883
+ ```
884
+
885
+ ### Voice + RAG Guidance
886
+
887
+ For document-grounded assistants, prefer this pattern:
888
+
889
+ 1. Capture voice with `ai.voice.listen()` or your own transcription flow.
890
+ 2. Send the transcribed text to `ai.public.chat()`.
891
+ 3. Render the text response for readability.
892
+ 4. Optionally speak the answer with `ai.voice.speak()` or `ai.tts.generate()`.
893
+
894
+ This is usually a better fit for manuals and procedural guidance than trying to use a live voice session as the primary retrieval layer.
895
+
896
+ ---
897
+
898
+ ## Podcast Generation
899
+
900
+ Generate NotebookLM-style multi-voice conversational podcasts from product documentation.
901
+
902
+ ### Generate a Podcast
903
+
904
+ ```typescript
905
+ const podcast = await ai.podcast.generate('my-collection', {
906
+ productId: 'coffee-maker-deluxe',
907
+ duration: 5, // Target 5 minutes
908
+ style: 'casual', // 'casual' | 'professional' | 'educational' | 'entertaining'
909
+ voices: {
910
+ host1: 'nova', // Female voice
911
+ host2: 'onyx' // Male voice
912
+ },
913
+ includeAudio: true // Generate audio files
914
+ });
915
+
916
+ console.log('Podcast Title:', podcast.script.title);
917
+ console.log('Duration:', podcast.metadata.duration, 'seconds');
918
+ console.log('Download:', podcast.audio?.mixedUrl);
919
+ ```
920
+
921
+ ### Available Voices
922
+
923
+ | Voice | Gender | Personality | Best For |
924
+ |-------|--------|-------------|----------|
925
+ | `alloy` | Neutral | Balanced, neutral | Professional podcasts |
926
+ | `echo` | Male | Clear, authoritative | Expert/teacher role |
927
+ | `fable` | Neutral | Warm, storytelling | Narrative content |
928
+ | `onyx` | Male | Deep, engaging | Main host, discussions |
929
+ | `nova` | Female | Friendly, enthusiastic | Co-host, questions |
930
+ | `shimmer` | Female | Bright, energetic | Entertaining content |
931
+
932
+ **Recommended Combinations:**
933
+ - **Casual**: Nova + Onyx - Friendly and engaging
934
+ - **Professional**: Alloy + Echo - Authoritative and clear
935
+ - **Educational**: Fable + Echo - Teaching style
936
+ - **Entertaining**: Shimmer + Onyx - High energy
937
+
938
+ ### Access the Script
939
+
940
+ ```typescript
941
+ // View the generated script
942
+ podcast.script.segments.forEach((segment, i) => {
943
+ const speaker = segment.speaker === 'host1' ? 'Host 1' : 'Host 2';
944
+ console.log(`${speaker}: ${segment.text}`);
945
+ });
946
+ ```
947
+
948
+ ### Check Generation Status
949
+
950
+ For long-running podcast generation, poll for status:
951
+
952
+ ```typescript
953
+ // Start generation
954
+ const podcast = await ai.podcast.generate('my-collection', {
955
+ productId: 'coffee-maker-deluxe',
956
+ duration: 10,
957
+ includeAudio: true
958
+ });
959
+
960
+ // Poll for status
961
+ const checkStatus = async () => {
962
+ const status = await ai.podcast.getStatus('my-collection', podcast.podcastId);
963
+
964
+ console.log(`Status: ${status.status} (${status.progress}%)`);
965
+
966
+ if (status.status === 'completed' && status.result) {
967
+ console.log('Podcast ready!');
968
+ console.log('Listen:', status.result.audio?.mixedUrl);
969
+ return true;
970
+ } else if (status.status === 'failed') {
971
+ console.error('Generation failed:', status.error);
972
+ return true;
973
+ }
974
+
975
+ return false;
976
+ };
977
+
978
+ // Check every 5 seconds
979
+ const interval = setInterval(async () => {
980
+ const done = await checkStatus();
981
+ if (done) clearInterval(interval);
982
+ }, 5000);
983
+ ```
984
+
985
+ ### Text-to-Speech (TTS)
986
+
987
+ Generate custom audio from text:
988
+
989
+ ```typescript
990
+ const audioBlob = await ai.tts.generate('my-collection', {
991
+ text: 'Welcome to our podcast about coffee makers!',
992
+ voice: 'nova',
993
+ speed: 1.0,
994
+ format: 'mp3'
995
+ });
996
+
997
+ // Create audio URL for playback
998
+ const audioUrl = URL.createObjectURL(audioBlob);
999
+ ```
1000
+
1001
+ ---
1002
+
1003
+
1004
+ ## Types & API reference
1005
+
1006
+ Full TypeScript types and the per-endpoint HTTP reference for every `SL.ai.*` method live in the
1007
+ **generated** references — they're not duplicated here so they can't drift:
1008
+
1009
+ - **[`API_SUMMARY.md`](API_SUMMARY.md)** — every function signature + type (search `ai.`).
1010
+ - **[`openapi.yaml`](../openapi.yaml)** — the raw HTTP endpoints.
1011
+
1012
+ For inline types, use LSP hover / go-to-definition on `SL.ai.*`.
1013
+
1014
+ ## Usage Examples
1015
+
1016
+ ### Example 1: Product FAQ Bot
1017
+
1018
+ ```typescript
1019
+ async function createProductFAQ() {
1020
+ const collectionId = 'my-collection';
1021
+ const productId = 'coffee-maker-deluxe';
1022
+
1023
+ // 1. Index product documentation
1024
+ await ai.rag.indexDocument(collectionId, {
1025
+ productId,
1026
+ documentUrl: 'https://example.com/manual.pdf'
1027
+ });
1028
+
1029
+ // 2. Configure assistant
1030
+ await ai.rag.configureAssistant(collectionId, {
1031
+ productId,
1032
+ systemPrompt: 'You are a coffee maker expert. Provide clear, step-by-step instructions.',
1033
+ rateLimitPerUser: 30
1034
+ });
1035
+
1036
+ // 3. Answer user questions
1037
+ const answer = await ai.public.chat(collectionId, {
1038
+ productId,
1039
+ userId: 'user-123',
1040
+ message: 'How do I make espresso?'
1041
+ });
1042
+
1043
+ console.log(answer.message);
1044
+ }
1045
+ ```
1046
+
1047
+ ### Example 2: Streaming Chatbot UI
1048
+
1049
+ ```typescript
1050
+ async function streamingChatbot(userMessage: string) {
1051
+ const stream = await ai.chat.completions.create('my-collection', {
1052
+ model: 'google/gemini-2.5-flash',
1053
+ messages: [
1054
+ { role: 'system', content: 'You are a helpful assistant.' },
1055
+ { role: 'user', content: userMessage }
1056
+ ],
1057
+ stream: true
1058
+ });
1059
+
1060
+ let fullResponse = '';
1061
+
1062
+ for await (const chunk of stream) {
1063
+ const content = chunk.choices[0]?.delta?.content || '';
1064
+ fullResponse += content;
1065
+
1066
+ // Update UI in real-time
1067
+ updateChatUI(content);
1068
+ }
1069
+
1070
+ return fullResponse;
1071
+ }
1072
+ ```
1073
+
1074
+ ### Example 3: Multi-Turn Conversation
1075
+
1076
+ ```typescript
1077
+ async function chatConversation() {
1078
+ const collectionId = 'my-collection';
1079
+ const sessionId = `chat-${Date.now()}`;
1080
+ const userId = 'user-123';
1081
+ const productId = 'coffee-maker-deluxe';
1082
+
1083
+ // Question 1
1084
+ const a1 = await ai.public.chat(collectionId, {
1085
+ productId,
1086
+ userId,
1087
+ message: 'How do I clean the machine?',
1088
+ sessionId
1089
+ });
1090
+ console.log('A1:', a1.message);
1091
+
1092
+ // Question 2 (references previous context)
1093
+ const a2 = await ai.public.chat(collectionId, {
1094
+ productId,
1095
+ userId,
1096
+ message: 'How often should I do that?',
1097
+ sessionId
1098
+ });
1099
+ console.log('A2:', a2.message);
1100
+
1101
+ // Get full history
1102
+ const session = await ai.public.getSession(collectionId, sessionId);
1103
+ console.log('Full conversation:', session.messages);
1104
+ }
1105
+ ```
1106
+
1107
+ ### Example 4: React Hook for Product Assistant
1108
+
1109
+ ```typescript
1110
+ import { useState, useCallback } from 'react';
1111
+ import { ai } from '@proveanything/smartlinks';
1112
+
1113
+ export function useProductAssistant(
1114
+ collectionId: string,
1115
+ productId: string,
1116
+ userId: string
1117
+ ) {
1118
+ const [loading, setLoading] = useState(false);
1119
+ const [error, setError] = useState<string | null>(null);
1120
+ const [rateLimit, setRateLimit] = useState({ remaining: 20, limit: 20 });
1121
+
1122
+ const ask = useCallback(async (message: string) => {
1123
+ setLoading(true);
1124
+ setError(null);
1125
+
1126
+ try {
1127
+ const response = await ai.public.chat(collectionId, {
1128
+ productId,
1129
+ userId,
1130
+ message
1131
+ });
1132
+
1133
+ setRateLimit(prev => ({
1134
+ ...prev,
1135
+ remaining: prev.remaining - 1
1136
+ }));
1137
+
1138
+ return response.message;
1139
+ } catch (err: any) {
1140
+ setError(err.message);
1141
+ throw err;
1142
+ } finally {
1143
+ setLoading(false);
1144
+ }
1145
+ }, [collectionId, productId, userId]);
1146
+
1147
+ return { ask, loading, error, rateLimit };
1148
+ }
1149
+
1150
+ // Usage in component
1151
+ function ProductHelp() {
1152
+ const { ask, loading, rateLimit } = useProductAssistant(
1153
+ 'my-collection',
1154
+ 'coffee-maker',
1155
+ 'user-123'
1156
+ );
1157
+ const [answer, setAnswer] = useState('');
1158
+
1159
+ const handleAsk = async () => {
1160
+ const response = await ask('How do I clean this?');
1161
+ setAnswer(response);
1162
+ };
1163
+
1164
+ return (
1165
+ <div>
1166
+ <button onClick={handleAsk} disabled={loading}>
1167
+ {loading ? 'Asking...' : 'Ask Question'}
1168
+ </button>
1169
+ {answer && <p>{answer}</p>}
1170
+ <p>{rateLimit.remaining} questions remaining</p>
1171
+ </div>
1172
+ );
1173
+ }
1174
+ ```
1175
+
1176
+
1177
+ ---
1178
+
1179
+ ## Error Handling
1180
+
1181
+ ### Error Codes
1182
+
1183
+ | Code | Type | HTTP Status | Description |
1184
+ |------|------|-------------|-------------|
1185
+ | `rate_limit_exceeded` | `rate_limit_error` | 429 | User exceeded rate limit |
1186
+ | `invalid_request` | `invalid_request_error` | 400 | Invalid parameters |
1187
+ | `authentication_error` | `authentication_error` | 401 | Invalid/missing API key |
1188
+ | `permission_denied` | `permission_error` | 403 | Insufficient permissions |
1189
+ | `not_found` | `not_found_error` | 404 | Resource not found |
1190
+ | `document_not_found` | `not_found_error` | 404 | Product document not indexed |
1191
+ | `server_error` | `server_error` | 500 | Internal server error |
1192
+ | `service_unavailable` | `server_error` | 503 | Service temporarily unavailable |
1193
+
1194
+ ### Error Handling Pattern
1195
+
1196
+ ```typescript
1197
+ import { SmartLinksAIError } from '@proveanything/smartlinks';
1198
+
1199
+ async function robustChat() {
1200
+ try {
1201
+ const response = await ai.public.chat('my-collection', {
1202
+ productId: 'coffee-maker',
1203
+ userId: 'user-123',
1204
+ message: 'Help!'
1205
+ });
1206
+ return response.message;
1207
+ } catch (error) {
1208
+ if (error instanceof SmartLinksAIError) {
1209
+ switch (error.code) {
1210
+ case 'rate_limit_exceeded':
1211
+ console.error('Rate limit exceeded');
1212
+ console.log('Try again at:', new Date(error.resetAt!));
1213
+ break;
1214
+ case 'document_not_found':
1215
+ console.error('Product manual not indexed yet');
1216
+ break;
1217
+ case 'authentication_error':
1218
+ console.error('Invalid API key');
1219
+ break;
1220
+ case 'invalid_request':
1221
+ console.error('Invalid request:', error.message);
1222
+ break;
1223
+ default:
1224
+ console.error('API error:', error.message);
1225
+ }
1226
+ } else {
1227
+ console.error('Unexpected error:', error);
1228
+ }
1229
+ throw error;
1230
+ }
1231
+ }
1232
+ ```
1233
+
1234
+ ### Rate Limit Retry
1235
+
1236
+ ```typescript
1237
+ async function chatWithRetry(
1238
+ request: PublicChatRequest,
1239
+ maxRetries = 3
1240
+ ) {
1241
+ let retries = 0;
1242
+
1243
+ while (retries < maxRetries) {
1244
+ try {
1245
+ return await ai.public.chat('my-collection', request);
1246
+ } catch (error) {
1247
+ if (error instanceof SmartLinksAIError && error.isRateLimitError()) {
1248
+ if (retries === maxRetries - 1) throw error;
1249
+
1250
+ const resetTime = new Date(error.resetAt!).getTime();
1251
+ const waitTime = resetTime - Date.now();
1252
+
1253
+ console.log(`Rate limited. Waiting ${waitTime}ms...`);
1254
+ await new Promise(resolve => setTimeout(resolve, waitTime));
1255
+
1256
+ retries++;
1257
+ } else {
1258
+ throw error;
1259
+ }
1260
+ }
1261
+ }
1262
+ }
1263
+ ```
1264
+
1265
+ ---
1266
+
1267
+ ## Rate Limiting
1268
+
1269
+ ### Rate Limit Overview
1270
+
1271
+ Public endpoints are rate-limited per `userId`:
1272
+
1273
+ | Endpoint Type | Default Limit | Window |
1274
+ |--------------|---------------|--------|
1275
+ | Public Chat | 20 requests | 1 hour |
1276
+ | Token Generation | 10 requests | 1 hour |
1277
+ | Admin Endpoints | Unlimited* | - |
1278
+
1279
+ *Admin endpoints use API key authentication and are not rate-limited by default.
1280
+
1281
+ ### Rate Limit Headers
1282
+
1283
+ All API responses include rate limit information:
1284
+
1285
+ ```
1286
+ X-RateLimit-Limit: 20
1287
+ X-RateLimit-Remaining: 15
1288
+ X-RateLimit-Reset: 1707300000000
1289
+ ```
1290
+
1291
+ ### Checking Rate Limit
1292
+
1293
+ ```typescript
1294
+ // Check rate limit status before making requests
1295
+ const status = await ai.public.getRateLimit('my-collection', 'user-123');
1296
+
1297
+ console.log('Used:', status.used);
1298
+ console.log('Remaining:', status.remaining);
1299
+ console.log('Resets at:', new Date(status.resetAt));
1300
+
1301
+ if (status.remaining > 0) {
1302
+ // Safe to make request
1303
+ await ai.public.chat(/* ... */);
1304
+ } else {
1305
+ // Show user when they can ask again
1306
+ console.log('Rate limit reached. Try again at:', status.resetAt);
1307
+ }
1308
+ ```
1309
+
1310
+ ### Resetting Rate Limits (Admin)
1311
+
1312
+ ```typescript
1313
+ // Reset rate limit for a specific user
1314
+ await ai.rateLimit.reset('my-collection', 'user-123');
1315
+ console.log('Rate limit reset for user-123');
1316
+ ```
1317
+
1318
+ ---
1319
+
1320
+ ## Best Practices
1321
+
1322
+ ### 1. Choose the Right Model
1323
+
1324
+ ```typescript
1325
+ // For most new workflows (recommended default)
1326
+ await ai.chat.responses.create('my-collection', {
1327
+ model: 'openai/gpt-5.4',
1328
+ input: 'Create a concise onboarding checklist.'
1329
+ });
1330
+
1331
+ // For lower-cost structured or JSON-oriented work
1332
+ await ai.chat.responses.create('my-collection', {
1333
+ model: 'openai/gpt-5-mini',
1334
+ input: 'Return a color palette as JSON.'
1335
+ });
1336
+
1337
+ // For fast multimodal or cost-sensitive general use
1338
+ await ai.chat.completions.create('my-collection', {
1339
+ model: 'google/gemini-2.5-flash',
1340
+ messages: [...]
1341
+ });
1342
+
1343
+ // When you need the actual available catalog for this collection
1344
+ const available = await ai.models.list('my-collection');
1345
+ console.log(available.data.map(model => model.id));
1346
+ ```
1347
+
1348
+ ### 2. Use Streaming for Long Responses
1349
+
1350
+ Improve perceived performance with streaming:
1351
+
1352
+ ```typescript
1353
+ // Non-streaming: User waits for full response
1354
+ const response = await ai.chat.completions.create('my-collection', {
1355
+ messages: [...]
1356
+ });
1357
+
1358
+ // Streaming: User sees progress immediately
1359
+ const stream = await ai.chat.completions.create('my-collection', {
1360
+ stream: true,
1361
+ messages: [...]
1362
+ });
1363
+
1364
+ for await (const chunk of stream) {
1365
+ updateUI(chunk.choices[0]?.delta?.content);
1366
+ }
1367
+ ```
1368
+
1369
+ ### 3. Maintain Session Context
1370
+
1371
+ Keep conversations coherent with session IDs:
1372
+
1373
+ ```typescript
1374
+ // Generate unique session ID per conversation
1375
+ const sessionId = `user-${userId}-product-${productId}`;
1376
+
1377
+ // All questions in same conversation use same sessionId
1378
+ await ai.public.chat('my-collection', {
1379
+ sessionId,
1380
+ message: 'First question',
1381
+ ...
1382
+ });
1383
+
1384
+ await ai.public.chat('my-collection', {
1385
+ sessionId,
1386
+ message: 'Follow-up question',
1387
+ ...
1388
+ });
1389
+ ```
1390
+
1391
+ ### 4. Handle Rate Limits Gracefully
1392
+
1393
+ Show clear feedback to users:
1394
+
1395
+ ```typescript
1396
+ try {
1397
+ await ai.public.chat('my-collection', {...});
1398
+ } catch (error) {
1399
+ if (error instanceof SmartLinksAIError && error.isRateLimitError()) {
1400
+ const resetTime = new Date(error.resetAt!);
1401
+ showNotification(
1402
+ `You've reached your question limit. ` +
1403
+ `Try again at ${resetTime.toLocaleTimeString()}`
1404
+ );
1405
+ }
1406
+ }
1407
+ ```
1408
+
1409
+ ### 5. Optimize Voice UX
1410
+
1411
+ Provide clear status updates:
1412
+
1413
+ ```typescript
1414
+ async function voiceAssistant() {
1415
+ try {
1416
+ showStatus('Listening...');
1417
+ const question = await ai.voice.listen();
1418
+
1419
+ showStatus('Processing...');
1420
+ const answer = await ai.public.chat('my-collection', {
1421
+ message: question,
1422
+ ...
1423
+ });
1424
+
1425
+ showStatus('Speaking...');
1426
+ await ai.voice.speak(answer.message);
1427
+
1428
+ showStatus('Ready');
1429
+ } catch (error) {
1430
+ showStatus('Error', error.message);
1431
+ }
1432
+ }
1433
+ ```
1434
+
1435
+ ### 6. Chunk Large Documents
1436
+
1437
+ For better RAG performance, chunk documents appropriately:
1438
+
1439
+ ```typescript
1440
+ // For technical manuals
1441
+ await ai.rag.indexDocument('my-collection', {
1442
+ productId: 'coffee-maker',
1443
+ documentUrl: '...',
1444
+ chunkSize: 500, // Smaller chunks for precise answers
1445
+ overlap: 50 // Overlap maintains context
1446
+ });
1447
+
1448
+ // For narrative content
1449
+ await ai.rag.indexDocument('my-collection', {
1450
+ productId: 'coffee-maker',
1451
+ documentUrl: '...',
1452
+ chunkSize: 1000, // Larger chunks for coherent responses
1453
+ overlap: 100
1454
+ });
1455
+ ```
1456
+
1457
+ ### 7. Use System Prompts Effectively
1458
+
1459
+ Provide clear instructions:
1460
+
1461
+ ```typescript
1462
+ await ai.chat.completions.create('my-collection', {
1463
+ messages: [
1464
+ {
1465
+ role: 'system',
1466
+ content: `You are a coffee maker expert assistant.
1467
+ - Be concise and clear
1468
+ - Use numbered lists for steps
1469
+ - Always mention safety precautions
1470
+ - If unsure, ask for clarification`
1471
+ },
1472
+ { role: 'user', content: 'How do I descale?' }
1473
+ ]
1474
+ });
1475
+ ```
1476
+
1477
+ ### 8. Monitor Usage and Costs
1478
+
1479
+ Track usage for cost management:
1480
+
1481
+ ```typescript
1482
+ const response = await ai.chat.completions.create('my-collection', {
1483
+ messages: [...]
1484
+ });
1485
+
1486
+ // Log token usage
1487
+ console.log('Usage:', response.usage);
1488
+ console.log('Prompt tokens:', response.usage.prompt_tokens);
1489
+ console.log('Completion tokens:', response.usage.completion_tokens);
1490
+ console.log('Total tokens:', response.usage.total_tokens);
1491
+
1492
+ // Calculate estimated cost
1493
+ const model = await ai.models.get('my-collection', response.model);
1494
+ const cost =
1495
+ (response.usage.prompt_tokens * model.pricing.input / 1_000_000) +
1496
+ (response.usage.completion_tokens * model.pricing.output / 1_000_000);
1497
+ console.log('Estimated cost: $', cost.toFixed(4));
1498
+ ```
1499
+
1500
+ ---
1501
+
1502
+ ## Providing Content to the AI Assistant
1503
+
1504
+ The portal's built-in AI assistant can discuss the content currently visible to the user. To make this work, your app must supply contextual content when the assistant requests it. There are three extraction methods, tried in priority order:
1505
+
1506
+ | Priority | Method | When Used |
1507
+ |----------|--------|-----------|
1508
+ | 1 | **Direct prop callback** | Container/widget rendered in the parent React context |
1509
+ | 2 | **PostMessage protocol** | App rendered in an iframe |
1510
+ | 3 | **DOM text fallback** | Neither of the above responded |
1511
+
1512
+ This section covers **methods 1 and 2** — the ones you implement in your app.
1513
+
1514
+ ### Method 1: Direct Prop (`onRequestAIContent`)
1515
+
1516
+ When your app is rendered as a **container** (a direct component in the parent's React tree), the framework passes an `onRequestAIContent` prop to your exported component. You do **not** call this prop yourself — the framework calls it when the AI assistant needs context.
1517
+
1518
+ Structure your component to make current state accessible when the callback fires:
1519
+
1520
+ ```tsx
1521
+ import { useEffect, useRef } from 'react';
1522
+
1523
+ export function PublicContainer(props) {
1524
+ const { onRequestAIContent, appId, ...rest } = props;
1525
+
1526
+ // Keep a ref to the latest content so the callback always returns fresh data
1527
+ const currentContentRef = useRef(null);
1528
+
1529
+ useEffect(() => {
1530
+ currentContentRef.current = buildCurrentContent();
1531
+ }, [relevantState]);
1532
+
1533
+ return <div data-app-container={appId}>{/* your UI */}</div>;
1534
+ }
1535
+ ```
1536
+
1537
+ The framework registers the content provider internally via `useAIContentExtraction.registerContentProvider()`. Your component just needs to respond when the callback is invoked.
1538
+
1539
+ ### Method 2: PostMessage Protocol (Iframes)
1540
+
1541
+ For iframe-embedded apps the framework sends a `postMessage` request and expects a response within **500 ms**. If your app doesn't reply in time the framework falls back to DOM text extraction.
1542
+
1543
+ **Request sent by the framework to your iframe:**
1544
+
1545
+ ```typescript
1546
+ {
1547
+ type: 'smartlinks:request-ai-content',
1548
+ requestId: 'ai-content-1709834567890-abc123' // unique per request
1549
+ }
1550
+ ```
1551
+
1552
+ **Response your app must send back:**
1553
+
1554
+ ```typescript
1555
+ {
1556
+ type: 'smartlinks:ai-content-response',
1557
+ requestId: 'ai-content-1709834567890-abc123', // echo back the requestId
1558
+ content: AIContentResponse
1559
+ }
1560
+ ```
1561
+
1562
+ **Minimal implementation:**
1563
+
1564
+ ```typescript
1565
+ window.addEventListener('message', async (event) => {
1566
+ if (event.data?.type === 'smartlinks:request-ai-content') {
1567
+ const content = await gatherAIContent();
1568
+
1569
+ window.parent.postMessage({
1570
+ type: 'smartlinks:ai-content-response',
1571
+ requestId: event.data.requestId,
1572
+ content,
1573
+ }, '*');
1574
+ }
1575
+ });
1576
+
1577
+ async function gatherAIContent(): Promise<AIContentResponse> {
1578
+ return {
1579
+ text: 'The user is viewing the warranty registration form...',
1580
+ contentLabel: 'Warranty Registration',
1581
+ };
1582
+ }
1583
+ ```
1584
+
1585
+ ### The `AIContentResponse` Interface
1586
+
1587
+ ```typescript
1588
+ interface AIContentResponse {
1589
+ /**
1590
+ * Plain text or markdown for the AI to use as context.
1591
+ * Injected into the system prompt. Max recommended: ~4000 characters.
1592
+ */
1593
+ text: string;
1594
+
1595
+ /**
1596
+ * Optional structured metadata (key-value pairs).
1597
+ * Not used directly in prompts but available for custom providers.
1598
+ */
1599
+ metadata?: Record<string, unknown>;
1600
+
1601
+ /**
1602
+ * Pre-built RAG configuration. When provided, the assistant can use
1603
+ * SL.ai.public.chat() to ground answers in indexed product documents.
1604
+ */
1605
+ ragHint?: {
1606
+ /** The product ID whose indexed documents should be queried */
1607
+ productId: string;
1608
+ /** Optional session ID for multi-turn RAG conversations */
1609
+ sessionId?: string;
1610
+ /** Optional context hint to help scope the RAG query */
1611
+ context?: string;
1612
+ };
1613
+
1614
+ /**
1615
+ * Human-readable label shown in context-update messages
1616
+ * (e.g. "Product Manual", "FAQ").
1617
+ */
1618
+ contentLabel?: string;
1619
+
1620
+ /**
1621
+ * How the assistant should use this content:
1622
+ * - 'context' (default): inject text into the system prompt
1623
+ * - 'rag': use ragHint to query indexed docs via SL.ai.public.chat()
1624
+ * - 'hybrid': inject text as context AND ground answers via RAG
1625
+ */
1626
+ strategy?: 'context' | 'rag' | 'hybrid';
1627
+ }
1628
+ ```
1629
+
1630
+ ### Response Strategies
1631
+
1632
+ #### `context` (Default)
1633
+
1634
+ Return readable text. The assistant injects it into the system prompt as background knowledge. Good for descriptions, summaries, structured data, and FAQs.
1635
+
1636
+ ```typescript
1637
+ return {
1638
+ text: `
1639
+ ## Wine Details
1640
+ - **Name**: 2023 Château Margaux
1641
+ - **Region**: Bordeaux, France
1642
+ - **Tasting Notes**: Dark fruit, cedar, tobacco
1643
+ - **Food Pairing**: Lamb, aged cheese
1644
+ `,
1645
+ contentLabel: 'Wine Information',
1646
+ strategy: 'context',
1647
+ };
1648
+ ```
1649
+
1650
+ #### `rag`
1651
+
1652
+ Tell the assistant to query pre-indexed documents via SmartLinks RAG. The `text` field is minimal — the real knowledge comes from the indexed docs. Good for product manuals, large document sets, and technical specs.
1653
+
1654
+ ```typescript
1655
+ return {
1656
+ text: 'The user is viewing the espresso machine product page.',
1657
+ contentLabel: 'Product Assistant',
1658
+ strategy: 'rag',
1659
+ ragHint: {
1660
+ productId: 'espresso-machine-pro',
1661
+ sessionId: `rag-session-${userId}`,
1662
+ context: 'User is on the troubleshooting section',
1663
+ },
1664
+ };
1665
+ ```
1666
+
1667
+ When the assistant receives a `rag` strategy it routes the question through `SL.ai.public.chat()`:
1668
+
1669
+ ```typescript
1670
+ const response = await SL.ai.public.chat(collectionId, {
1671
+ productId: ragHint.productId,
1672
+ userId: currentUserId,
1673
+ message: userQuestion,
1674
+ sessionId: ragHint.sessionId,
1675
+ });
1676
+ ```
1677
+
1678
+ #### `hybrid`
1679
+
1680
+ Combines both — the `text` is injected as additional context **and** the user's questions are also grounded via RAG. Good for museum exhibits, guided experiences, or any scenario with rich metadata alongside large document sets.
1681
+
1682
+ ```typescript
1683
+ return {
1684
+ text: `
1685
+ ## Exhibit: The Starry Night
1686
+ - **Artist**: Vincent van Gogh
1687
+ - **Year**: 1889
1688
+ - **Current Location**: Gallery 3, East Wing
1689
+ - **Audio Guide**: Available in 12 languages
1690
+ `,
1691
+ contentLabel: 'Museum Exhibit',
1692
+ strategy: 'hybrid',
1693
+ ragHint: {
1694
+ productId: 'starry-night-exhibit',
1695
+ context: 'Art history and technique questions',
1696
+ },
1697
+ };
1698
+ ```
1699
+
1700
+ ### Content Extraction Examples
1701
+
1702
+ #### Museum Guide App
1703
+
1704
+ ```typescript
1705
+ async function gatherAIContent(): Promise<AIContentResponse> {
1706
+ const exhibit = getCurrentExhibit();
1707
+
1708
+ return {
1709
+ text: `
1710
+ Exhibit: "${exhibit.title}" by ${exhibit.artist}
1711
+ Period: ${exhibit.period}
1712
+ Medium: ${exhibit.medium}
1713
+ Description: ${exhibit.curatorNotes}
1714
+ Related works in this gallery: ${exhibit.relatedWorks.join(', ')}
1715
+ `.trim(),
1716
+ contentLabel: `Exhibit: ${exhibit.title}`,
1717
+ strategy: 'hybrid',
1718
+ ragHint: {
1719
+ productId: exhibit.smartlinksProductId,
1720
+ context: `Art history, technique, and visitor information for ${exhibit.title}`,
1721
+ },
1722
+ metadata: {
1723
+ exhibitId: exhibit.id,
1724
+ gallery: exhibit.gallery,
1725
+ audioGuideAvailable: exhibit.hasAudioGuide,
1726
+ },
1727
+ };
1728
+ }
1729
+ ```
1730
+
1731
+ #### Wine Product App
1732
+
1733
+ ```typescript
1734
+ async function gatherAIContent(): Promise<AIContentResponse> {
1735
+ const wine = getCurrentWine();
1736
+ const reviews = await fetchRecentReviews(wine.id, 5);
1737
+
1738
+ return {
1739
+ text: `
1740
+ Wine: ${wine.name} (${wine.vintage})
1741
+ Winery: ${wine.winery}
1742
+ Region: ${wine.region}, ${wine.country}
1743
+ Grape: ${wine.grape}
1744
+ ABV: ${wine.abv}%
1745
+ Price: ${wine.price}
1746
+ Tasting Notes: ${wine.tastingNotes}
1747
+
1748
+ Recent Reviews:
1749
+ ${reviews.map(r => `- "${r.text}" (${r.rating}/5)`).join('\n')}
1750
+ `.trim(),
1751
+ contentLabel: 'Wine Details',
1752
+ strategy: 'context',
1753
+ };
1754
+ }
1755
+ ```
1756
+
1757
+ #### Equipment Manual App (RAG-only)
1758
+
1759
+ ```typescript
1760
+ async function gatherAIContent(): Promise<AIContentResponse> {
1761
+ const equipment = getCurrentEquipment();
1762
+
1763
+ return {
1764
+ text: `User is viewing: ${equipment.name} (Model: ${equipment.modelNumber})`,
1765
+ contentLabel: `${equipment.name} Assistant`,
1766
+ strategy: 'rag',
1767
+ ragHint: {
1768
+ productId: equipment.smartlinksProductId,
1769
+ sessionId: `manual-${equipment.id}-${Date.now()}`,
1770
+ context: 'Technical manual, troubleshooting, and maintenance',
1771
+ },
1772
+ };
1773
+ }
1774
+ ```
1775
+
1776
+ ### Timing & Lifecycle
1777
+
1778
+ - **On navigation**: When the user navigates to a new product/proof/app, the assistant automatically requests fresh content and injects a context-update system message.
1779
+ - **On first message**: If no content has been gathered yet, the assistant requests it before the first AI call.
1780
+ - **On manual refresh**: The assistant can re-request content at any time (e.g. if the user's view within the app has changed).
1781
+
1782
+ Content is requested **lazily** — your handler is only called when the AI assistant is active and needs context. If the user never opens the assistant, your handler is never called.
1783
+
1784
+ ### Method 3: DOM Fallback (Automatic)
1785
+
1786
+ If your app doesn't implement either of the above, the framework extracts `innerText` from the DOM element with `data-app-container="{appId}"`, truncated to ~4000 characters. This is a **last resort** — content quality is much lower than a structured response. Implementing Method 1 or 2 is strongly recommended.
1787
+
1788
+ ---
1789
+
1790
+ ## Related Documentation
1791
+
1792
+ - [API Summary](./API_SUMMARY.md) - Complete API reference
1793
+ - [Widgets](./widgets.md) - Embedding SmartLinks components
1794
+ - [Realtime](./realtime.md) - Realtime data updates
1795
+ - [iframe Responder](./iframe-responder.md) - iframe integration
1796
+
1797
+ ---
1798
+
1799
+ ## Support
1800
+
1801
+ For questions or issues:
1802
+
1803
+ - **Documentation:** https://smartlinks.app/docs
1804
+ - **GitHub:** https://github.com/Prove-Anything/smartlinks
1805
+ - **Email:** support@smartlinks.app