@cogitator-ai/openai-compat 21.0.4 → 21.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,26 +1,35 @@
1
1
  # @cogitator-ai/openai-compat
2
2
 
3
- OpenAI Assistants API compatibility layer for Cogitator. Use OpenAI SDK clients with Cogitator backend, or integrate Cogitator with existing OpenAI-based applications.
3
+ OpenAI Assistants API compatibility layer for Cogitator. Point the official OpenAI SDK (or any Assistants API client) at a Cogitator server, or drive the same assistants/threads/runs model in-process.
4
+
5
+ Full guide: [cogitator.app/docs/integrations/openai-compat](https://cogitator.app/docs/integrations/openai-compat)
4
6
 
5
7
  ## Installation
6
8
 
7
9
  ```bash
8
- pnpm add @cogitator-ai/openai-compat
10
+ pnpm add @cogitator-ai/openai-compat @cogitator-ai/core
11
+
12
+ # optional: the OpenAI SDK for client code
13
+ pnpm add openai
9
14
  ```
10
15
 
16
+ `ioredis` and `pg` are optional peer dependencies, needed only for `RedisThreadStorage` / `PostgresThreadStorage`.
17
+
11
18
  ## Features
12
19
 
13
- - **OpenAI Server** - Expose Cogitator as OpenAI-compatible REST API
14
- - **OpenAI Adapter** - In-process adapter for programmatic access
15
- - **Thread Manager** - Manage conversations, messages, and assistants
16
- - **Persistent Storage** - Pluggable backends: In-memory, Redis, PostgreSQL
17
- - **SSE Streaming** - Real-time token streaming with OpenAI-compatible events
18
- - **File Operations** - Upload and manage files for assistants
19
- - **Full Assistants API** - Create, update, delete assistants
20
- - **Run Management** - Execute, stream, cancel and list runs
20
+ - **OpenAI Server** - Expose Cogitator as an OpenAI Assistants API (Fastify)
21
+ - **OpenAI Adapter** - In-process access to the same assistants/threads/runs API
22
+ - **Thread Manager** - Threads, messages, assistants and files over pluggable storage
23
+ - **Persistent Storage** - In-memory, Redis or PostgreSQL backends
24
+ - **SSE Streaming** - Token streaming with OpenAI-compatible stream events
25
+ - **Files** - Upload, list, download and delete files (`multipart/form-data`)
21
26
  - **Function Calling** - Assistant `function` tools pause the run with `requires_action` until the client submits outputs
22
- - **Authentication** - Optional API key authentication
23
- - **CORS Support** - Configurable cross-origin requests
27
+ - **Server-side tools** - Cogitator tools passed to the server run inside Cogitator for every run
28
+ - **Vision & JSON output** - `image_url` / image `image_file` parts reach the model; `response_format` supports `json_object` and `json_schema`
29
+ - **Authentication** - Optional API keys (constant-time check, `/health` stays public)
30
+ - **CORS** - Configurable cross-origin requests
31
+
32
+ The package implements the Assistants API surface (models, assistants, threads, messages, runs, files). There is no `/v1/chat/completions`, run steps, or vector stores endpoint; `code_interpreter` and `file_search` assistant tools are stored but not executed.
24
33
 
25
34
  ---
26
35
 
@@ -92,7 +101,7 @@ const stream = openai.beta.threads.runs
92
101
  await stream.finalRun();
93
102
  ```
94
103
 
95
- Assistant `model` values are Cogitator model strings (`openai/gpt-6.1-sol`, `ollama/llama3.2:latest`, ...). The id `cogitator` listed by `GET /v1/models` maps to the server's `defaultModel`.
104
+ Assistant `model` values are Cogitator model strings (`openai/gpt-6.1-sol`, `ollama/llama3.2:latest`, ...). `GET /v1/models` lists a single model, `cogitator`, which maps to the server's `defaultModel`; runs that use it fail when no `defaultModel` is configured.
96
105
 
97
106
  ---
98
107
 
@@ -103,7 +112,7 @@ The `OpenAIServer` exposes Cogitator as an OpenAI-compatible REST API.
103
112
  ### Configuration
104
113
 
105
114
  ```typescript
106
- import { OpenAIServer, createOpenAIServer } from '@cogitator-ai/openai-compat';
115
+ import { OpenAIServer } from '@cogitator-ai/openai-compat';
107
116
 
108
117
  const server = new OpenAIServer(cogitator, {
109
118
  port: 8080,
@@ -126,13 +135,13 @@ const server = new OpenAIServer(cogitator, {
126
135
 
127
136
  | Option | Type | Default | Description |
128
137
  | -------------- | ------------------------------- | -------------------------------------- | ------------------------------------------------ |
129
- | `port` | `number` | `8080` | Port to listen on |
138
+ | `port` | `number` | `8080` | Port to listen on (`0` picks a free port) |
130
139
  | `host` | `string` | `'0.0.0.0'` | Host to bind to |
131
140
  | `apiKeys` | `string[]` | `[]` | API keys for authentication. Empty disables auth |
132
141
  | `tools` | `Tool[]` | `[]` | Server-side tools available to every run |
133
142
  | `defaultModel` | `string` | — | Model used for the `cogitator` model id |
134
143
  | `storage` | `ThreadStorage` | in-memory | Persistence backend (connect before passing) |
135
- | `maxFileSize` | `number` | `512 MB` | Upload limit for `POST /v1/files` |
144
+ | `maxFileSize` | `number` | `512 MB` | Upload limit for `POST /v1/files`, in bytes |
136
145
  | `logging` | `boolean` | `false` | Enable Fastify request logging (JSON logs) |
137
146
  | `cors.origin` | `string \| string[] \| boolean` | `true` | CORS origin configuration |
138
147
  | `cors.methods` | `string[]` | `['GET', 'POST', 'DELETE', 'OPTIONS']` | Allowed HTTP methods |
@@ -140,21 +149,26 @@ const server = new OpenAIServer(cogitator, {
140
149
  ### Server Lifecycle
141
150
 
142
151
  ```typescript
143
- await server.start();
152
+ await server.start(); // throws if already started
144
153
 
145
154
  console.log(server.getUrl()); // uses the bound port, so `port: 0` works
146
- console.log(server.getBaseUrl());
155
+ console.log(server.getBaseUrl()); // getUrl() + '/v1'
147
156
 
148
157
  console.log(server.isRunning());
149
158
 
150
- const adapter = server.getAdapter();
159
+ const adapter = server.getAdapter(); // the OpenAIAdapter behind the routes
160
+ const fastify = server.getFastify(); // underlying Fastify instance
151
161
 
152
162
  await server.stop();
153
163
  ```
154
164
 
155
- ### Health Check
165
+ For tests without a listening socket, `await server.waitUntilReady()` and then call `server.getFastify().inject(...)`.
166
+
167
+ ### Authentication
156
168
 
157
- The server provides a health endpoint (public even when `apiKeys` are configured):
169
+ With `apiKeys` set, every request except `GET /health` needs `Authorization: Bearer <key>`. A missing header answers `401` with code `missing_api_key`; a malformed header or unknown key answers `401` with code `invalid_api_key`.
170
+
171
+ ### Health Check
158
172
 
159
173
  ```bash
160
174
  curl http://localhost:8080/health
@@ -165,13 +179,16 @@ curl http://localhost:8080/health
165
179
 
166
180
  ## OpenAI Adapter
167
181
 
168
- The `OpenAIAdapter` provides in-process access without running a server.
182
+ The `OpenAIAdapter` provides in-process access without running a server. The server uses one internally (`server.getAdapter()`).
169
183
 
170
184
  ```typescript
171
- import { OpenAIAdapter, createOpenAIAdapter } from '@cogitator-ai/openai-compat';
185
+ import { createOpenAIAdapter } from '@cogitator-ai/openai-compat';
172
186
 
173
187
  const adapter = createOpenAIAdapter(cogitator, {
174
- tools: [calculator],
188
+ tools: [calculator], // server-side Cogitator tools
189
+ defaultModel: 'openai/gpt-6-luna', // resolves the `cogitator` model id
190
+ storage, // ThreadStorage, default: in-memory
191
+ maxStoredRuns: 10_000, // finished runs kept in memory (default 10 000)
175
192
  });
176
193
  ```
177
194
 
@@ -181,13 +198,14 @@ const adapter = createOpenAIAdapter(cogitator, {
181
198
  const assistant = await adapter.createAssistant({
182
199
  model: 'openai/gpt-6.1-sol',
183
200
  name: 'Code Helper',
201
+ description: 'Writes TypeScript',
184
202
  instructions: 'You help write code',
185
203
  temperature: 0.7,
186
- tools: [{ type: 'code_interpreter' }],
204
+ response_format: { type: 'json_object' },
187
205
  metadata: { category: 'development' },
188
206
  });
189
207
 
190
- const fetched = await adapter.getAssistant(assistant.id);
208
+ const fetched = await adapter.getAssistant(assistant.id); // undefined if missing
191
209
 
192
210
  const updated = await adapter.updateAssistant(assistant.id, {
193
211
  name: 'Code Expert',
@@ -196,16 +214,18 @@ const updated = await adapter.updateAssistant(assistant.id, {
196
214
 
197
215
  const all = await adapter.listAssistants();
198
216
 
199
- const deleted = await adapter.deleteAssistant(assistant.id);
217
+ const deleted = await adapter.deleteAssistant(assistant.id); // boolean
200
218
  ```
201
219
 
202
220
  ### Thread Operations
203
221
 
204
222
  ```typescript
205
- const thread = await adapter.createThread({ project: 'demo' });
223
+ const thread = await adapter.createThread({ project: 'demo' }); // metadata
206
224
 
207
225
  const fetched = await adapter.getThread(thread.id);
208
226
 
227
+ await adapter.updateThread(thread.id, { metadata: { stage: 'review' } }); // merged into existing metadata
228
+
209
229
  const message = await adapter.addMessage(thread.id, {
210
230
  role: 'user',
211
231
  content: 'Hello, how are you?',
@@ -225,6 +245,8 @@ const msg = await adapter.getMessage(thread.id, 'msg_abc123');
225
245
  await adapter.deleteThread(thread.id);
226
246
  ```
227
247
 
248
+ Message `content` is a string or an array of parts: `{ type: 'text', text }`, `{ type: 'image_url', image_url: { url } }` or `{ type: 'image_file', image_file: { file_id } }`.
249
+
228
250
  ### Run Execution
229
251
 
230
252
  ```typescript
@@ -237,9 +259,9 @@ const run = await adapter.createRun(thread.id, {
237
259
  metadata: { source: 'api' },
238
260
  });
239
261
 
240
- const status = adapter.getRun(thread.id, run.id);
262
+ const status = adapter.getRun(thread.id, run.id); // synchronous
241
263
 
242
- const runs = adapter.listRuns(thread.id);
264
+ const runs = adapter.listRuns(thread.id); // newest first
243
265
 
244
266
  const cancelled = adapter.cancelRun(thread.id, run.id); // throws if the run already finished
245
267
 
@@ -248,20 +270,26 @@ for await (const { event, data } of adapter.streamRunEvents(run.id)) {
248
270
  }
249
271
  ```
250
272
 
251
- Runs use the whole thread as context (earlier messages are replayed to the agent), receive an abort signal on cancel, and honour `additional_instructions`, `response_format` (`json_object` / `json_schema`), `max_completion_tokens`, `top_p`, `tool_choice` (`none` or a specific function), `parallel_tool_calls` and `truncation_strategy` (`last_messages`). A thread can have only one active run at a time. Image parts (`image_url`, and `image_file` uploads with an image extension) are passed to the model.
273
+ `createRun` returns the `queued` run immediately and executes it in the background. It rejects with an `InvalidRequestError` when the assistant or thread is missing, the thread already has an active run (one active run per thread) or `max_prompt_tokens` is not a positive integer.
274
+
275
+ Runs use the whole thread as context (earlier messages are replayed to the agent), receive an abort signal on cancel, and honour `additional_instructions`, `response_format` (`json_object` / `json_schema`), `max_completion_tokens`, `top_p`, `tool_choice` (`none` or a specific function), `parallel_tool_calls` and `truncation_strategy` (`last_messages`). `max_prompt_tokens` is a best-effort budget (estimated at about 4 characters per token, instructions included): the oldest replayed messages are dropped until the prompt fits, and when the last user message alone does not fit the run ends `incomplete` with `incomplete_details: { reason: 'max_prompt_tokens' }`. Image parts (`image_url`, and `image_file` uploads with a `png`/`jpg`/`jpeg`/`gif`/`webp` extension) are passed to the model.
276
+
277
+ `streamRunEvents(runId, fromIndex = 0)` replays the run's event log from `fromIndex` and ends after the next `done` event. `getRunEventCursor(runId)` returns the current log position (use it before `submitToolOutputs` to stream only the continuation), and `getStreamEmitter(runId)` exposes the raw `EventEmitter`. Token deltas are emitted only for runs created with `stream: true`.
278
+
279
+ Runs live in the adapter's memory, unlike assistants, threads, messages and files, which go to `storage`. Poll, cancel or submit tool outputs to a run on the same process that created it: with several processes, run a single instance or route each thread to the same one (sticky routing).
252
280
 
253
281
  ### Tool Outputs
254
282
 
255
- Server-side `tools` run inside Cogitator. Assistant tools of type `function` are executed by the API client: when the model calls one, the run moves to `requires_action` with the pending calls, and continues after all outputs are submitted (runs waiting longer than 10 minutes expire). With the OpenAI SDK use `submitToolOutputsAndPoll` / `submitToolOutputsStream`.
283
+ Server-side `tools` run inside Cogitator. Assistant tools of type `function` are executed by the API client: when the model calls one, the run moves to `requires_action` with the pending calls, and continues after all outputs are submitted (runs waiting longer than 10 minutes expire). Outputs must cover every pending call; missing or unknown `tool_call_id`s are rejected. With the OpenAI SDK use `submitToolOutputsAndPoll` / `submitToolOutputsStream`.
256
284
 
257
285
  ```typescript
258
286
  const run = adapter.getRun(thread.id, runId);
259
287
 
260
288
  if (run?.status === 'requires_action') {
261
- const toolCalls = run.required_action?.submit_tool_outputs.tool_calls;
289
+ const toolCalls = run.required_action?.submit_tool_outputs.tool_calls ?? [];
262
290
 
263
291
  const outputs = await Promise.all(
264
- toolCalls!.map(async (call) => ({
292
+ toolCalls.map(async (call) => ({
265
293
  tool_call_id: call.id,
266
294
  output: await executeMyTool(call.function.name, call.function.arguments),
267
295
  }))
@@ -277,12 +305,12 @@ if (run?.status === 'requires_action') {
277
305
 
278
306
  ## Thread Manager
279
307
 
280
- The `ThreadManager` handles storage for threads, messages, assistants, and files.
308
+ The `ThreadManager` handles storage for threads, messages, assistants, and files. Get the adapter's instance with `adapter.getThreadManager()`, or create one directly.
281
309
 
282
310
  ```typescript
283
311
  import { ThreadManager } from '@cogitator-ai/openai-compat';
284
312
 
285
- const manager = new ThreadManager();
313
+ const manager = new ThreadManager(); // or new ThreadManager(storage)
286
314
  ```
287
315
 
288
316
  ### Assistant Storage
@@ -291,16 +319,19 @@ const manager = new ThreadManager();
291
319
  interface StoredAssistant {
292
320
  id: string;
293
321
  name: string | null;
322
+ description?: string | null;
294
323
  model: string;
295
324
  instructions: string | null;
296
325
  tools: AssistantTool[];
297
326
  metadata: Record<string, string>;
298
327
  temperature?: number;
328
+ top_p?: number;
329
+ response_format?: ResponseFormat;
299
330
  created_at: number;
300
331
  }
301
332
 
302
333
  const assistant = await manager.createAssistant({
303
- model: 'gpt-6.1-sol',
334
+ model: 'openai/gpt-6.1-sol',
304
335
  name: 'Helper',
305
336
  instructions: 'Be helpful',
306
337
  });
@@ -316,6 +347,7 @@ await manager.deleteAssistant(assistant.id);
316
347
  ```typescript
317
348
  const thread = await manager.createThread({ key: 'value' });
318
349
  const fetched = await manager.getThread(thread.id);
350
+ await manager.updateThread(thread.id, { metadata: { key: 'other' } });
319
351
  await manager.deleteThread(thread.id);
320
352
  ```
321
353
 
@@ -332,6 +364,7 @@ const assistantMsg = await manager.addAssistantMessage(
332
364
  'Hi there!',
333
365
  assistant.id,
334
366
  run.id
367
+ // optional 5th argument: message id to reuse
335
368
  );
336
369
 
337
370
  const messages = await manager.listMessages(thread.id, {
@@ -339,15 +372,19 @@ const messages = await manager.listMessages(thread.id, {
339
372
  order: 'desc',
340
373
  });
341
374
 
342
- const llmMessages = await manager.getMessagesForLLM(thread.id);
375
+ const one = await manager.getMessage(thread.id, message!.id);
376
+
377
+ const llmMessages = await manager.getMessagesForLLM(thread.id); // { role, content, images? }[]
343
378
  ```
344
379
 
380
+ `addMessage` / `addAssistantMessage` resolve to `undefined` when the thread does not exist.
381
+
345
382
  ### File Management
346
383
 
347
384
  ```typescript
348
- const file = await manager.addFile(Buffer.from('file content'), 'document.txt');
385
+ const file = await manager.addFile(Buffer.from('file content'), 'document.txt', 'assistants');
349
386
 
350
- const fetched = await manager.getFile(file.id);
387
+ const fetched = await manager.getFile(file.id); // StoredFile with `content: Buffer`
351
388
 
352
389
  const all = await manager.listFiles();
353
390
 
@@ -376,13 +413,13 @@ import {
376
413
 
377
414
  ```typescript
378
415
  const manager = new ThreadManager();
379
- // or explicitly:
380
- const manager = new ThreadManager(new InMemoryThreadStorage());
416
+ // equivalent to:
417
+ const explicit = new ThreadManager(new InMemoryThreadStorage());
381
418
  ```
382
419
 
383
420
  ### Redis Storage
384
421
 
385
- Requires `ioredis` peer dependency:
422
+ Requires the `ioredis` peer dependency:
386
423
 
387
424
  ```bash
388
425
  pnpm add ioredis
@@ -390,10 +427,10 @@ pnpm add ioredis
390
427
 
391
428
  ```typescript
392
429
  const storage = new RedisThreadStorage({
393
- host: 'localhost',
394
- port: 6379,
395
- keyPrefix: 'cogitator:openai:', // optional, default: 'cogitator:openai:'
396
- ttl: 86400, // optional, default: 86400 (24h)
430
+ host: 'localhost', // default: 'localhost'
431
+ port: 6379, // default: 6379
432
+ keyPrefix: 'cogitator:openai:', // default: 'cogitator:openai:'
433
+ ttl: 86400, // seconds, default: 86400 (24h); 0 disables expiry
397
434
  });
398
435
  await storage.connect();
399
436
 
@@ -403,7 +440,7 @@ const manager = new ThreadManager(storage);
403
440
  await storage.disconnect();
404
441
  ```
405
442
 
406
- With connection URL:
443
+ With a connection URL (takes precedence over `host` / `port`):
407
444
 
408
445
  ```typescript
409
446
  const storage = new RedisThreadStorage({
@@ -413,7 +450,7 @@ const storage = new RedisThreadStorage({
413
450
 
414
451
  ### PostgreSQL Storage
415
452
 
416
- Requires `pg` peer dependency:
453
+ Requires the `pg` peer dependency:
417
454
 
418
455
  ```bash
419
456
  pnpm add pg
@@ -422,10 +459,10 @@ pnpm add pg
422
459
  ```typescript
423
460
  const storage = new PostgresThreadStorage({
424
461
  connectionString: 'postgresql://user:pass@localhost:5432/db',
425
- schema: 'public', // optional, default: 'public'
426
- tableName: 'openai_compat_data', // optional, default: 'openai_compat_data'
462
+ schema: 'public', // default: 'public'
463
+ tableName: 'openai_compat_data', // default: 'openai_compat_data'
427
464
  });
428
- await storage.connect();
465
+ await storage.connect(); // creates the table and index if missing
429
466
 
430
467
  const manager = new ThreadManager(storage);
431
468
 
@@ -433,33 +470,33 @@ const manager = new ThreadManager(storage);
433
470
  await storage.disconnect();
434
471
  ```
435
472
 
473
+ `schema` and `tableName` must be plain SQL identifiers (`/^[a-zA-Z_][a-zA-Z0-9_]*$/`); the constructor throws otherwise.
474
+
436
475
  ### Factory Function
437
476
 
438
477
  ```typescript
439
- // In-memory (default)
440
- const storage = createThreadStorage();
441
- // or: createThreadStorage({ type: 'memory' })
478
+ const memory = createThreadStorage(); // or createThreadStorage({ type: 'memory' })
442
479
 
443
- // Redis
444
- const storage = createThreadStorage({
480
+ const redis = createThreadStorage({
445
481
  type: 'redis',
446
482
  host: 'localhost',
447
483
  port: 6379,
448
484
  });
449
- await storage.connect!();
485
+ await redis.connect?.();
450
486
 
451
- // PostgreSQL
452
- const storage = createThreadStorage({
487
+ const postgres = createThreadStorage({
453
488
  type: 'postgres',
454
489
  connectionString: 'postgresql://localhost/db',
455
490
  });
456
- await storage.connect!();
491
+ await postgres.connect?.();
457
492
  ```
458
493
 
459
- ### Using with OpenAI Adapter
494
+ The factory returns an unconnected `ThreadStorage`; call `connect()` before use.
495
+
496
+ ### Using with the Adapter or Server
460
497
 
461
498
  ```typescript
462
- import { OpenAIAdapter, ThreadManager, RedisThreadStorage } from '@cogitator-ai/openai-compat';
499
+ import { OpenAIAdapter, RedisThreadStorage, createOpenAIServer } from '@cogitator-ai/openai-compat';
463
500
  import { Cogitator } from '@cogitator-ai/core';
464
501
 
465
502
  const storage = new RedisThreadStorage({ host: 'localhost' });
@@ -474,7 +511,7 @@ const adapter = new OpenAIAdapter(cogitator, { tools: [], storage });
474
511
  const server = createOpenAIServer(cogitator, { storage });
475
512
  ```
476
513
 
477
- Storage is the single source of truth (no in-process cache), so several server instances can share one Redis or PostgreSQL backend. Redis listings use `SCAN` + `MGET`; `ttl: 0` disables expiry. `ioredis` and `pg` are optional peer dependencies loaded on `connect()`.
514
+ Storage is the single source of truth for assistants, threads, messages and files (no in-process cache), so several server instances can share one Redis or PostgreSQL backend. Runs are not persisted: they stay in the memory of the instance that executes them. Redis listings use `SCAN` + `MGET`. `ioredis` and `pg` are loaded on `connect()`, with an install hint if missing.
478
515
 
479
516
  ### ThreadStorage Interface
480
517
 
@@ -506,10 +543,12 @@ interface ThreadStorage {
506
543
  }
507
544
  ```
508
545
 
546
+ `StoredThread` is `{ thread: Thread; messages: Message[] }`; `StoredFile` is `{ id, content: Buffer, filename, created_at, purpose? }`.
547
+
509
548
  ### Custom Storage Implementation
510
549
 
511
550
  ```typescript
512
- import type { ThreadStorage } from '@cogitator-ai/openai-compat';
551
+ import type { ThreadStorage, StoredThread } from '@cogitator-ai/openai-compat';
513
552
 
514
553
  class MyCustomStorage implements ThreadStorage {
515
554
  async saveThread(id: string, thread: StoredThread): Promise<void> {
@@ -525,11 +564,14 @@ const manager = new ThreadManager(new MyCustomStorage());
525
564
 
526
565
  ## Supported Endpoints
527
566
 
567
+ All endpoints except `/health` live under `/v1`. List endpoints for assistants, messages and runs accept `limit` (1-100, default 20), `order` (`asc` / `desc`, default `desc`), `after` and `before`, and return `{ object: 'list', data, first_id, last_id, has_more }`. `GET /v1/files` accepts `limit` (1-10 000, default 10 000), `order` (default `desc`) and `after`, plus a `purpose` filter.
568
+
528
569
  ### Models
529
570
 
530
- | Method | Endpoint | Description |
531
- | ------ | ------------ | --------------------- |
532
- | GET | `/v1/models` | List available models |
571
+ | Method | Endpoint | Description |
572
+ | ------ | ------------ | ---------------------------------- |
573
+ | GET | `/v1/models` | Lists the single `cogitator` model |
574
+ | GET | `/health` | Health check (public, no auth) |
533
575
 
534
576
  ### Assistants
535
577
 
@@ -543,20 +585,20 @@ const manager = new ThreadManager(new MyCustomStorage());
543
585
 
544
586
  ### Threads
545
587
 
546
- | Method | Endpoint | Description |
547
- | ------ | ----------------- | ---------------------- |
548
- | POST | `/v1/threads` | Create thread |
549
- | GET | `/v1/threads/:id` | Get thread |
550
- | POST | `/v1/threads/:id` | Update thread metadata |
551
- | DELETE | `/v1/threads/:id` | Delete thread |
588
+ | Method | Endpoint | Description |
589
+ | ------ | ----------------- | ------------------------------------------- |
590
+ | POST | `/v1/threads` | Create thread (optional initial `messages`) |
591
+ | GET | `/v1/threads/:id` | Get thread |
592
+ | POST | `/v1/threads/:id` | Update thread metadata (merged) |
593
+ | DELETE | `/v1/threads/:id` | Delete thread |
552
594
 
553
595
  ### Messages
554
596
 
555
- | Method | Endpoint | Description |
556
- | ------ | ---------------------------------- | ------------- |
557
- | POST | `/v1/threads/:id/messages` | Add message |
558
- | GET | `/v1/threads/:id/messages` | List messages |
559
- | GET | `/v1/threads/:id/messages/:msg_id` | Get message |
597
+ | Method | Endpoint | Description |
598
+ | ------ | ---------------------------------- | ------------------------------------- |
599
+ | POST | `/v1/threads/:id/messages` | Add message (`user` or `assistant`) |
600
+ | GET | `/v1/threads/:id/messages` | List messages (also filters `run_id`) |
601
+ | GET | `/v1/threads/:id/messages/:msg_id` | Get message |
560
602
 
561
603
  ### Runs
562
604
 
@@ -569,15 +611,19 @@ const manager = new ThreadManager(new MyCustomStorage());
569
611
  | POST | `/v1/threads/:id/runs/:run_id/cancel` | Cancel run |
570
612
  | POST | `/v1/threads/:id/runs/:run_id/submit_tool_outputs` | Submit tool outputs |
571
613
 
614
+ The create and `submit_tool_outputs` endpoints stream SSE when the body has `stream: true`.
615
+
572
616
  ### Files
573
617
 
574
- | Method | Endpoint | Description |
575
- | ------ | ----------------------- | --------------------- |
576
- | POST | `/v1/files` | Upload file |
577
- | GET | `/v1/files` | List files |
578
- | GET | `/v1/files/:id` | Get file metadata |
579
- | GET | `/v1/files/:id/content` | Download file content |
580
- | DELETE | `/v1/files/:id` | Delete file |
618
+ | Method | Endpoint | Description |
619
+ | ------ | ----------------------- | ------------------------------------------------- |
620
+ | POST | `/v1/files` | Upload file (`multipart/form-data`) |
621
+ | GET | `/v1/files` | List files (`purpose`, `limit`, `order`, `after`) |
622
+ | GET | `/v1/files/:id` | Get file metadata |
623
+ | GET | `/v1/files/:id/content` | Download file content |
624
+ | DELETE | `/v1/files/:id` | Delete file |
625
+
626
+ Uploads take a `file` part and an optional `purpose` field (`assistants`, `assistants_output`, `batch`, `batch_output`, `fine-tune`, `fine-tune-results`, `vision`; default `assistants`).
581
627
 
582
628
  ---
583
629
 
@@ -596,14 +642,21 @@ interface OpenAIError {
596
642
  }
597
643
  ```
598
644
 
645
+ `formatOpenAIError(code, message, type?, param?)` builds this shape if you add your own routes via `server.getFastify()`.
646
+
599
647
  ### Error Types
600
648
 
601
- | HTTP Status | Type | Description |
602
- | ----------- | ----------------------- | -------------------------- |
603
- | 400 | `invalid_request_error` | Invalid request parameters |
604
- | 401 | `authentication_error` | Invalid or missing API key |
605
- | 404 | `invalid_request_error` | Resource not found |
606
- | 500 | `server_error` | Internal server error |
649
+ | HTTP Status | Type | Code | When |
650
+ | ----------- | ----------------------- | ------------------------------------ | ------------------------------------------------------ |
651
+ | 400 | `invalid_request_error` | `invalid_request`, `missing_file` | Invalid parameters, unknown assistant, thread busy |
652
+ | 401 | `invalid_request_error` | `missing_api_key`, `invalid_api_key` | Missing, malformed or unknown API key |
653
+ | 404 | `invalid_request_error` | `not_found` | Unknown thread, message, run, assistant, file or route |
654
+ | 429 | `rate_limit_error` | `rate_limit_exceeded` | Errors thrown with status 429 |
655
+ | 500 | `server_error` | `internal_error` | Unhandled errors |
656
+
657
+ Every 5xx answer says `Internal server error`; the details are logged on the server, never sent to the client. The adapter refuses invalid requests (unknown assistant or thread, a busy thread, a run in the wrong state, missing or unknown tool outputs, an invalid `max_prompt_tokens`) by throwing `InvalidRequestError` (exported, with an optional `param`), which the server answers as `400 invalid_request`; any other error from the adapter is a server failure.
658
+
659
+ A run that fails during execution does not produce an HTTP error: it ends with `status: 'failed'` and `last_error: { code: 'server_error', message }`. `message` is `Internal server error` unless the cause is a `CogitatorError` or an `InvalidRequestError`, whose message is kept; the original error is logged.
607
660
 
608
661
  ### Client-Side Error Handling
609
662
 
@@ -614,9 +667,9 @@ try {
614
667
  });
615
668
  } catch (error) {
616
669
  if (error instanceof OpenAI.APIError) {
617
- console.log(error.status);
670
+ console.log(error.status); // 400
618
671
  console.log(error.message);
619
- console.log(error.code);
672
+ console.log(error.code); // 'invalid_request'
620
673
  }
621
674
  }
622
675
  ```
@@ -647,17 +700,24 @@ queued → in_progress → completed
647
700
  → failed
648
701
  → requires_action → in_progress → ...
649
702
  → expired (no outputs within 10 minutes)
703
+ → incomplete (last message exceeds max_prompt_tokens)
650
704
 
651
705
  queued / in_progress / requires_action → cancelling → cancelled
652
706
  ```
653
707
 
654
- Streaming runs (`stream: true`) emit `thread.run.created`, `thread.run.queued`, `thread.run.in_progress`, `thread.message.created`, `thread.message.delta`, `thread.message.completed`, `thread.run.completed` (or `failed` / `cancelled` / `expired` / `requires_action`) and a final `done` event.
708
+ `incomplete` is produced only by `max_prompt_tokens` (see [Run Execution](#run-execution)).
655
709
 
656
710
  ### Polling for Completion
657
711
 
658
712
  ```typescript
659
- async function waitForRun(openai: OpenAI, threadId: string, runId: string): Promise<Run> {
660
- const terminalStates = ['completed', 'failed', 'cancelled', 'expired'];
713
+ import OpenAI from 'openai';
714
+
715
+ async function waitForRun(
716
+ openai: OpenAI,
717
+ threadId: string,
718
+ runId: string
719
+ ): Promise<OpenAI.Beta.Threads.Run> {
720
+ const terminalStates = ['completed', 'failed', 'cancelled', 'expired', 'requires_action'];
661
721
 
662
722
  while (true) {
663
723
  const run = await openai.beta.threads.runs.retrieve(runId, { thread_id: threadId });
@@ -666,38 +726,32 @@ async function waitForRun(openai: OpenAI, threadId: string, runId: string): Prom
666
726
  return run;
667
727
  }
668
728
 
669
- if (run.status === 'requires_action') {
670
- return run;
671
- }
672
-
673
729
  await new Promise((r) => setTimeout(r, 1000));
674
730
  }
675
731
  }
676
732
  ```
677
733
 
734
+ The SDK's `createAndPoll` / `submitToolOutputsAndPoll` do the same.
735
+
678
736
  ---
679
737
 
680
738
  ## SSE Streaming
681
739
 
682
- Real-time streaming support with Server-Sent Events (SSE). Streams tokens as they are generated for immediate feedback.
740
+ Runs created with `stream: true` answer with Server-Sent Events and stream tokens as they are generated.
683
741
 
684
742
  ### Streaming with OpenAI SDK
685
743
 
686
744
  ```typescript
687
745
  const run = await openai.beta.threads.runs.create(threadId, {
688
746
  assistant_id: assistant.id,
689
- stream: true, // Enable streaming
747
+ stream: true,
690
748
  });
691
749
 
692
- // Handle streaming events
693
750
  for await (const event of run) {
694
751
  if (event.event === 'thread.message.delta') {
695
- const delta = event.data.delta;
696
- if (delta.content) {
697
- for (const content of delta.content) {
698
- if (content.type === 'text' && content.text?.value) {
699
- process.stdout.write(content.text.value);
700
- }
752
+ for (const content of event.data.delta.content ?? []) {
753
+ if (content.type === 'text' && content.text?.value) {
754
+ process.stdout.write(content.text.value);
701
755
  }
702
756
  }
703
757
  }
@@ -715,7 +769,6 @@ const run = await openai.beta.threads.createAndRun({
715
769
  stream: true,
716
770
  });
717
771
 
718
- // Process stream
719
772
  for await (const event of run) {
720
773
  console.log(event.event, event.data);
721
774
  }
@@ -723,8 +776,10 @@ for await (const event of run) {
723
776
 
724
777
  ### Stream Events
725
778
 
779
+ The server emits these events (the exported `StreamEventType` union also lists run-step events, which are not emitted):
780
+
726
781
  ```typescript
727
- type StreamEvent =
782
+ type EmittedEvent =
728
783
  | { event: 'thread.run.created'; data: Run }
729
784
  | { event: 'thread.run.queued'; data: Run }
730
785
  | { event: 'thread.run.in_progress'; data: Run }
@@ -734,6 +789,7 @@ type StreamEvent =
734
789
  | { event: 'thread.run.cancelling'; data: Run }
735
790
  | { event: 'thread.run.cancelled'; data: Run }
736
791
  | { event: 'thread.run.expired'; data: Run }
792
+ | { event: 'thread.run.incomplete'; data: Run }
737
793
  | { event: 'thread.message.created'; data: Message }
738
794
  | { event: 'thread.message.in_progress'; data: Message }
739
795
  | { event: 'thread.message.delta'; data: MessageDelta }
@@ -741,7 +797,7 @@ type StreamEvent =
741
797
  | { event: 'done'; data: '[DONE]' };
742
798
  ```
743
799
 
744
- A stream ends with `done` after the run finishes or pauses on `requires_action`; continue a paused run with `submit_tool_outputs` and `stream: true`.
800
+ A stream ends with `done` after the run finishes or pauses on `requires_action`; continue a paused run with `submit_tool_outputs` and `stream: true`, which streams only the events after the submission.
745
801
 
746
802
  ### Message Delta Format
747
803
 
@@ -781,17 +837,23 @@ Response format (Server-Sent Events):
781
837
  event: thread.run.created
782
838
  data: {"id":"run_xxx","status":"queued",...}
783
839
 
840
+ event: thread.run.queued
841
+ data: {"id":"run_xxx","status":"queued",...}
842
+
784
843
  event: thread.run.in_progress
785
844
  data: {"id":"run_xxx","status":"in_progress",...}
786
845
 
787
846
  event: thread.message.created
788
847
  data: {"id":"msg_xxx","status":"in_progress",...}
789
848
 
849
+ event: thread.message.in_progress
850
+ data: {"id":"msg_xxx","status":"in_progress",...}
851
+
790
852
  event: thread.message.delta
791
- data: {"id":"msg_xxx","delta":{"content":[{"index":0,"type":"text","text":{"value":"Hello"}}]}}
853
+ data: {"id":"msg_xxx","object":"thread.message.delta","delta":{"content":[{"index":0,"type":"text","text":{"value":"Hello"}}]}}
792
854
 
793
855
  event: thread.message.delta
794
- data: {"id":"msg_xxx","delta":{"content":[{"index":1,"type":"text","text":{"value":" world"}}]}}
856
+ data: {"id":"msg_xxx","object":"thread.message.delta","delta":{"content":[{"index":0,"type":"text","text":{"value":" world"}}]}}
795
857
 
796
858
  event: thread.message.completed
797
859
  data: {"id":"msg_xxx","status":"completed",...}
@@ -807,6 +869,29 @@ data: [DONE]
807
869
 
808
870
  ## Type Reference
809
871
 
872
+ ### Server & Adapter
873
+
874
+ ```typescript
875
+ import type {
876
+ OpenAIServerConfig,
877
+ AuthConfig,
878
+ OpenAIAdapterOptions,
879
+ StreamEventType,
880
+ StreamEventData,
881
+ StreamEmitterEvents,
882
+ RunStreamEvent,
883
+ LLMThreadMessage,
884
+ CreateAssistantParams,
885
+ UpdateAssistantParams,
886
+ } from '@cogitator-ai/openai-compat';
887
+
888
+ import {
889
+ COGITATOR_MODEL_ID,
890
+ formatOpenAIError,
891
+ InvalidRequestError,
892
+ } from '@cogitator-ai/openai-compat';
893
+ ```
894
+
810
895
  ### Core Types
811
896
 
812
897
  ```typescript
@@ -817,6 +902,7 @@ import type {
817
902
  AssistantTool,
818
903
  FunctionDefinition,
819
904
  ResponseFormat,
905
+ JsonSchema,
820
906
  CreateAssistantRequest,
821
907
  UpdateAssistantRequest,
822
908
  } from '@cogitator-ai/openai-compat';
@@ -836,10 +922,11 @@ import type {
836
922
  MessageContent,
837
923
  TextContent,
838
924
  TextAnnotation,
925
+ ImageFileContent,
926
+ ImageUrlContent,
839
927
  Attachment,
840
928
  CreateMessageRequest,
841
929
  MessageContentPart,
842
- MessageDelta,
843
930
  } from '@cogitator-ai/openai-compat';
844
931
  ```
845
932
 
@@ -854,6 +941,8 @@ import type {
854
941
  RunError,
855
942
  Usage,
856
943
  ToolChoice,
944
+ IncompleteDetails,
945
+ TruncationStrategy,
857
946
  CreateRunRequest,
858
947
  SubmitToolOutputsRequest,
859
948
  ToolOutput,
@@ -875,7 +964,12 @@ import type { FileObject, FilePurpose, UploadFileRequest } from '@cogitator-ai/o
875
964
  ### Stream Types
876
965
 
877
966
  ```typescript
878
- import type { StreamEvent, MessageDelta, RunStepDelta } from '@cogitator-ai/openai-compat';
967
+ import type {
968
+ StreamEvent,
969
+ MessageDelta,
970
+ MessageContentDelta,
971
+ RunStepDelta,
972
+ } from '@cogitator-ai/openai-compat';
879
973
  ```
880
974
 
881
975
  ### Storage Types
@@ -953,7 +1047,7 @@ console.log(await chat('Hi, my name is Alex'));
953
1047
  console.log(await chat('What is my name?'));
954
1048
  ```
955
1049
 
956
- ### Code Assistant with Tools
1050
+ ### Server-Side Tools
957
1051
 
958
1052
  ```typescript
959
1053
  import { createOpenAIServer } from '@cogitator-ai/openai-compat';
@@ -961,15 +1055,13 @@ import { Cogitator, tool } from '@cogitator-ai/core';
961
1055
  import { z } from 'zod';
962
1056
  import OpenAI from 'openai';
963
1057
 
964
- const runCode = tool({
965
- name: 'run_code',
966
- description: 'Execute Python code',
1058
+ const lookupOrder = tool({
1059
+ name: 'lookup_order',
1060
+ description: 'Look up an order by id',
967
1061
  parameters: z.object({
968
- code: z.string().describe('Python code to execute'),
1062
+ orderId: z.string().describe('Order id'),
969
1063
  }),
970
- execute: async ({ code }) => {
971
- return `Output: ${code.length} characters`;
972
- },
1064
+ execute: async ({ orderId }) => ({ orderId, status: 'shipped' }),
973
1065
  });
974
1066
 
975
1067
  const cogitator = new Cogitator({
@@ -978,19 +1070,19 @@ const cogitator = new Cogitator({
978
1070
 
979
1071
  const server = createOpenAIServer(cogitator, {
980
1072
  port: 8080,
981
- tools: [runCode],
1073
+ tools: [lookupOrder],
982
1074
  });
983
1075
 
984
1076
  await server.start();
985
1077
 
986
1078
  const openai = new OpenAI({
987
1079
  baseURL: server.getBaseUrl(),
988
- apiKey: process.env.OPENAI_API_KEY,
1080
+ apiKey: 'not-needed',
989
1081
  });
990
1082
 
991
1083
  const assistant = await openai.beta.assistants.create({
992
- name: 'Code Runner',
993
- instructions: 'You can run Python code using the run_code tool.',
1084
+ name: 'Support Bot',
1085
+ instructions: 'Use lookup_order to answer questions about orders.',
994
1086
  model: 'openai/gpt-6.1-sol',
995
1087
  });
996
1088
  // Server-side tools (`tools` on the server) are available to every run without
@@ -998,9 +1090,47 @@ const assistant = await openai.beta.assistants.create({
998
1090
  // your client executes (they surface as `requires_action`).
999
1091
  ```
1000
1092
 
1093
+ ### Client-Executed Function Tools
1094
+
1095
+ ```typescript
1096
+ const assistant = await openai.beta.assistants.create({
1097
+ model: 'openai/gpt-6.1-sol',
1098
+ tools: [
1099
+ {
1100
+ type: 'function',
1101
+ function: {
1102
+ name: 'get_weather',
1103
+ description: 'Current weather for a city',
1104
+ parameters: {
1105
+ type: 'object',
1106
+ properties: { city: { type: 'string' } },
1107
+ required: ['city'],
1108
+ },
1109
+ },
1110
+ },
1111
+ ],
1112
+ });
1113
+
1114
+ let run = await openai.beta.threads.runs.createAndPoll(thread.id, {
1115
+ assistant_id: assistant.id,
1116
+ });
1117
+
1118
+ while (run.status === 'requires_action' && run.required_action) {
1119
+ const tool_outputs = run.required_action.submit_tool_outputs.tool_calls.map((call) => ({
1120
+ tool_call_id: call.id,
1121
+ output: JSON.stringify({ city: JSON.parse(call.function.arguments).city, tempC: 21 }),
1122
+ }));
1123
+ run = await openai.beta.threads.runs.submitToolOutputsAndPoll(run.id, {
1124
+ thread_id: thread.id,
1125
+ tool_outputs,
1126
+ });
1127
+ }
1128
+ ```
1129
+
1001
1130
  ### File Upload
1002
1131
 
1003
1132
  ```typescript
1133
+ import fs from 'node:fs';
1004
1134
  import OpenAI from 'openai';
1005
1135
 
1006
1136
  const openai = new OpenAI({
@@ -1018,7 +1148,7 @@ console.log('Uploaded:', file.id);
1018
1148
  const content = await openai.files.content(file.id);
1019
1149
  console.log('Content:', await content.text());
1020
1150
 
1021
- await openai.files.del(file.id);
1151
+ await openai.files.delete(file.id);
1022
1152
  ```
1023
1153
 
1024
1154
  ### Multi-Model Setup